Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

An implementation of a high-order generalized finite difference method for solving the time-harmonic cold plasma wave equation in toroidal geometry

A high-order physics-informed meshless finite difference numerical technique is introduced for solving the time-harmonic cold plasma wave equation in toroidal geometries, presenting a novel application of the generalized finite difference (GFD) method to plasma wave simulations. The algorithm employs an irregular distribution of computational points, with local point density informed by the shortest wavelength derived from the cold plasma dispersion relation. Numerical stability and robustness are addressed using regularization techniques. The algorithm, implemented for two spatial dimensions, solves for the wave electric field and is demonstrated to achieve convergence rates of $\mathcal{O}$($\mathcal{h}$ $\mathcal{P}$ )⁠. Verification tests reproduce plane wave solutions, and example simulations of ion cyclotron resonance heating and electron cyclotron resonance heating demonstrate its capability, approaching realistic tokamak plasma scenarios. This work contributes to laying a foundation for the GFD method to be used in more sophisticated, optimized, and physically realistic full-wave simulations in time-harmonic plasma wave research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)↗

Initializing BSQ with Open-Source ICCING

While it is well known that there is a significant amount of conserved charges in the initial state of nuclear collisions, the production of these due to gluon splitting has yet to be thoroughly investigated. The ICCING (Initial Conserved Charges in Nuclear Geometry) algorithm reconstructs these quark distributions, providing conserved strange, baryon, and electric charges, by sampling a given model for the g → qq¯ splitting function over the initial energy density, which is valid at top collider energies, even when µB = 0. The ICCING algorithm includes fluctuations in the gluon longitudinal momenta, a structure that supports the implementation of dynamical processes, and the c++ version is now open-source. A full analysis of parameter choices on the model has been done to quantify the effect these have on the underlying physics. We find there is a sustained difference across the different charges that indicates sensitivity to hot spot geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An Initial Study of the Convergence Rate of Griffin’s Pebble Bed Reactors Algorithm

This paper presents an initial study of the convergence properties of an iterative algorithm for computing the burnup distribution in a pebble bed reactor (PBR) in its equilibrium core condition. The algorithm is implemented in the Griffin code. Griffin is a reactor multiphysics analysis application jointly developed by Idaho National Laboratory (INL) and Argonne National Laboratory (ANL). Griffin’s PBR algorithm is discussed and simulation data are presented. An alternative matrix formulation of the algorithm is presented that facilitates analysis of the iterative algorithm. The dependence of the spectral radius of the iterative algorithm on operational and discretization parameters is investigated.

97 MATHEMATICS AND COMPUTING↗

DESI mock challenge: Halo and galaxy catalogues with the bias assignment method

We present a novel approach to the construction of mock galaxy catalogues for large-scale structure analysis based on the distribution of dark matter halos obtained with effective bias models at the field level. We aim to produce mock galaxy catalogues capable of generating accurate covariance matrices for a number of cosmological probes that are expected to be measured in current and forthcoming galaxy redshift surveys (e.g. two- and three-point statistics). The construction of the catalogues shown in this paper is part of a mock-comparison project within the Dark Energy Spectroscopic Instrument (DESI) collaboration. We use the bias assignment method ( BAM ) to model the statistics of halo distribution through a learning algorithm using a few detailed N-body simulations, and approximated gravity solvers based on Lagrangian perturbation theory. We introduce cosmic-web-dependent corrections to modelling redshift-space distortions at the N-body level – both in the halo and galaxy distributions –, as well as a multi-scale approach for accurate assignment of halo properties. Using specific models of halo occupation distributions to populate halos, we generate galaxy mocks with the expected number density and central-satellite fraction of emission-line galaxies, which are a key target of the DESI experiment. BAM generates mock catalogues with per cent accuracy in a number of summary statistics, such as the abundance, the two- and three-point statistics of halo distributions, both in real and redshift space. In particular, the mock galaxy catalogues display ~3%-10% accuracy in the multipoles of the power spectrum up to scales of k ~ 0.4 h -1 Mpc. We show that covariance matrices of two- and three-point statistics obtained with BAM display a similar structure to the reference simulation. BAM offers an efficient way to produce mock halo catalogues with accurate two- and three-point statistics and is able to generate a variety of multi-tracer catalogues with precise covariance matrices of several cosmological probes. We discuss future developments of the algorithm towards mock production in DESI and other galaxy-redshift surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Real-Time Hardware-in-the-Loop Testbed to Evaluate FLISR Implemented with OpenFMB

With the increasing complexity of the distribution smart grid architecture, algorithms such as the fault location, isolation, and service restoration (FLISR) scheme rely on robust communications that are resilient to natural and man-made adverse conditions and exhibit robustness. Existing communications infrastructure for information exchange are centralized at the distribution management system, with very little autonomy or intelligence at the grid-edge. As a first step towards achieving grid-edge self-healing, this paper aims to bridge this shortcoming by implementing a centrally coordinated rules-based FLISR scheme and integrating it with Open Field Message Bus (OpenFMB), which is a flexible publish-subscribe architecture with the potential to enable point-to-multipoint communications and is more robust and resilient to natural and man-made adverse conditions. A proof of concept is developed to validate the centrally coordinated FLISR and OpenFMB mounted on an SEL-3360 computer that interacts with a simple feeder network of five SEL-651R relays, an SEL-3530 RTAC, and a hardware-in-the-loop testbed. Results demonstrate the efficacy of this approach in enabling direct, low-latency information exchange. OpenFMB's publish-subscribe data model also opens new ways to enable grid-edge interoperability among devices of different vendors interacting with different protocols.

Sundararajan, Aditya↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Explainable Neural Architecture Search (XNAS)

Code for the paper Learning Interpretable Models Through Multi-Objective Neural Architecture Search by Zachariah Carmichael, Tim Moon, and Sam Ade Jacobs. Monumental advances in deep learning have led to unprecedented achievements across a multitude of domains. While the performance of deep neural networks is indubitable, the architectural design and interpretability of such models are nontrivial. Research has been introduced to automate the design of neural network architectures through neural architecture search (NAS). Recent progress has made these methods more pragmatic by exploiting distributed computation and novel optimization algorithms. However, there is little work in optimizing architectures for interpretability. To this end, we propose a multiobjective distributed NAS framework that optimizes for both task performance and introspection. We leverage the non-dominated sorting genetic algorithm (NSGA-II) and explainable AI (XAI) techniques to reward architectures that can be better comprehended by humans. The framework is evaluated on several image classification datasets. We demonstrate that jointly optimizing for introspection ability and task error leads to more disentangled architectures that perform within tolerable error.

Carmichael, ZachariahJ↗

A Peer-to-Peer Market-Based Control Strategy for a Smart Residential Community with Behind-the-Meter Distributed Energy Resources

This paper presents a distributed peer-to-peer market control strategy to manage and to enable resource sharing of behind-the-meter distributed energy resources in a residential community. In the proposed strategy, each consumer or prosumer determines the flexibility of their point of connection to the power network such that the obtained flexibility is network-feasible. Based on the feasible flexibility, the consumers and the prosumers trade power among each other at each time instance to fulfil their preferred load requirements while maximizing their payoffs and helping to regulate node voltages inside the community. Because the problem to be solved is non-convex, a distributed particle swarm optimization algorithm is used to coordinate the consumers/prosumers in a fully autonomous manner without any centralized or hierarchical coordination. Numerical simulations performed on a community of 48 homes demonstrate the efficacy of the proposed approach.

distributed energy resource↗

Optimal resource allocation for flexible-grid entanglement distribution networks

We use a genetic algorithm (GA) as a design aid for determining the optimal provisioning of entangled photon spectrum in flex-grid quantum networks with arbitrary numbers of channels and users. After introducing a general model for entanglement distribution based on frequency-polarization hyperentangled biphotons, we derive upper bounds on fidelity and entangled bit rate for networks comprising one-to-one user connections. Simple conditions based on user detector quality and link efficiencies are found that determine whether entanglement is possible. We successfully apply a GA to find optimal resource allocations in four different representative network scenarios and validate features of our model experimentally in a quantum local area network in deployed fiber. Our results show promise for the rapid design of large-scale entanglement distribution networks.

97 MATHEMATICS AND COMPUTING↗

A Peer-to-Peer Market-Based Control Strategy for a Smart Residential Community with Behind-the-Meter Distributed Energy Resources

This paper presents a distributed peer-to-peer market control strategy to manage and to enable resource sharing of behind-the-meter distributed energy resources in a residential community. In the proposed strategy, each consumer or prosumer determines the flexibility of their point of connection to the power network such that the obtained flexibility is network-feasible. Based on the feasible flexibility, the consumers and the prosumers trade power among each other at each time instance to fulfill their preferred load requirements while maximizing their payoffs and helping to regulate node voltages inside the community. Because the problem to be solved is non-convex, a distributed particle swarm optimization algorithm is used to coordinate the consumers/prosumers in a fully autonomous manner without any centralized or hierarchical coordination. Numerical simulations performed on a community of 48 homes demonstrate the efficacy of the proposed approach.

behind-the-meter↗

A Peer-to-Peer Market-Based Control Strategy for a Smart Residential Community with Behind-the-Meter Distributed Energy Resources: Preprint

This paper presents a distributed peer-to-peer market control strategy to manage and to enable resource sharing of behind-the-meter distributed energy resources in a residential community. In the proposed strategy, each consumer or prosumer determines the flexibility of their point of connection to the power network such that the obtained flexibility is network-feasible. Based on the feasible flexibility, the consumers and the prosumers trade power among each other at each time instance to fulfill their preferred load requirements while maximizing their payoffs and helping to regulate node voltages inside the community. Because the problem to be solved is non-convex, a distributed particle swarm optimization algorithm is used to coordinate the consumers/prosumers in a fully autonomous manner without any centralized or hierarchical coordination. Numerical simulations performed on a community of 48 homes demonstrate the efficacy of the proposed approach.

behind-the-meter↗

Grid Modernization of Cooperatives and Municipal Utilities via Breakthrough System Monitoring, Control and Optimization (CRADA Final Report)

This project aims at developing and demonstrating successful implementation of breakthrough approaches in real-time data visualization as well as real-time distributed DER control and optimization to provide ample benefits to both utilities and end users. The National Renewable Energy Laboratory (NREL), Holy Cross Energy (HCE), National Rural Electric Cooperative Association (NRECA) and Survalent are collaborating to enable Cooperative and Municipal utilities to fully leverage DERs as part of their strategies for providing safe, reliable, and affordable electric services to their customers and help meet DOE Grid modernization goal of achieving at least 10% active devices to provide grid flexibility by 2035. This project will use novel real-time control algorithms and approaches for distributed control recently developed under DOE-funded projects, using the date from the Survalent’s basic SCADA engine, GIS and AMI engines deployed at HCE combined with NRECA’s globally-used MultiSpeak(R) software interoperability specification for seamless and real-time communications between electric utility enterprise software to embrace DER as part of their strategies for providing safe, reliable and affordable electric service to their customers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Moment preserving constrained resampling with applications to particle-in-cell methods

The Moment Preserving Constrained Resampling (MPCR) algorithm for particle resampling is introduced and applied to particle-in-cell (PIC) methods to increase simulation accuracy, reduce compute cost, and/or avoid numerical instabilities. The general algorithm partitions the system space into smaller subsets and resamples the distribution within each subset. Further, the algorithm is designed to conserve any number of particle and grid moments with a high degree of accuracy (i.e. machine accuracy). The effectiveness of MPCR is demonstrated with several numerical tests, including a use-case study in gyrokinetic fusion plasma simulations. Finally, the computational cost of MPCR is negligible compared to the cost of particle evolution in PIC methods, and the tests demonstrate that periodic particle resampling yields a significant improvement in the accuracy and stability of the results.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

SMALE: Enhancing Scalability of Machine Learning Algorithms on Extreme-Scale Computing Platforms

Deployment and execution of machine learning tasks on extreme-scale computing platforms face several significant technical challenges: 1) High computing cost incurred by dense networks – The computing workload of deep networks with densely-connected topology increases rapidly with the network size, imposing a non-scalable computing model of extreme-scale computing platforms; 2) Non-optimized workload distribution – Many advanced deep learning algorithms, e.g., sparsification and irregular net-work topology, produce very unbalanced workload distribution on extreme-scale computing platforms. The computation efficiency is greatly hindered by the incurred data and computation redundancies as well as long tails of the node with extensive workload; 3) Constraints in data movement and I/O bottle-neck – Inter-node data movement in extreme-scale computing platforms are associated with high energy and latency costs, and subject to the constraints of I/O bandwidth; and 4) Generalization of algorithm realization and acceleration on computing platforms – The large varieties of machine learning algorithms and structures of extreme-scale computing platforms make the derivation of a generalized algorithm realization and acceleration method very challenging, which, however, is the requirement by domain scientists and interested users. We call the above challenges Smale’s Problems in Machine Learning and Understanding for High-Performance Computing Scientific Discovery. The objective of our three-year research project is to develop a holistic innovation set at structure, assembly, and acceleration layers of machine learning algorithms to address the above challenges in algorithm deployment and execution. Three tasks are particularly performed, including: At the algorithm structure level, we investigate the techniques that can structurally sparsify on the topology of deep networks for computing workload reduction. We also study clustering and pruning techniques that can optimize the workload distributions over the extreme-scale computing platforms; At the algorithm assembly level, we derive a unified learning framework for unsupervised transfer learning and dynamic growing capabilities. Novel training methods are also exploited to enhance the training efficiency of the proposed framework; At the algorithm acceleration level, we will develop a series of techniques that can accelerate the computation of sparse matrix operations, which are one of the core executions in deep learning and optimize memory access of the concerned platforms. Our proposed techniques attack the fundamental problems in machine learning algorithms running on extreme-scale computing platforms by vertically integrating the solutions at three closely entangled layers, paving the long-term scaling path of machine learning applications under DOE context. Three tasks corresponding to the above respective research orientations are performed during the three-year project period with our collaborators at ORNL. The outcome of the proposed project is anticipated to form a holistic solution set of novel algorithms and network topologies, efficient training techniques, and fast acceleration methods to promote the computing scalability of the machine learning applications of particular interest to DOE.

97 MATHEMATICS AND COMPUTING↗

Resilient Operation of Networked Community Microgrids with High Solar Penetration

This project, funded by the US Department of Energy’s Solar Energy Technologies Office (SETO), focused on the operation of microgrids as a coordinated network. The primary objective, which was successfully achieved, was to develop both control strategies and hardware solutions to support the resilient and efficient operation of networked microgrids with high solar penetration. The work was structured around the following four main tasks: • Development of distributed and scalable optimization algorithms for AC-coupled networked microgrids. • Design and implementation of a novel DC interconnection hardware to enable precise power exchange between microgrids. • Laboratory operational validation of the developed technologies using 480 V testbeds and commercially available hardware. • Field operational validation of the complete solution in Adjuntas, Puerto Rico, interconnecting two kW-scale, split-phase microgrids of Casa Pueblo’s microgrids. This project addressed multiple technical challenges across the domains of optimization, control, hardware interconnection, and protection. One of its key contributions was delivering tangible, real-world solutions for networking microgrids. In contrast to purely theoretical or simulation-based work, this project included full-scale hardware operational validation both in the lab and in the field. The work conducted as part of this project—in collaboration with the University of Puerto Rico; the University of Tennessee, Knoxville; the University of Central Florida; and Casa Pueblo—has advanced the state of the art in networked microgrids. Key contributions include the development of distributed control strategies, practical solutions for real-world implementation challenges, and the introduction of a novel DC interlink approach for microgrid interconnection. The project featured both laboratory and field validation using commercial off-the-shelf components. The field deployment successfully validated that a group of microgrids can operate in a coordinated manner, enabling precise power flow between systems and mutual support during extreme events. This project resulted in 15 journal publications and 15 conference papers; 5 graduate students and 15 undergraduate students were supported. The codes of distributed optimization and forecasting were made open-source through OSTI.gov for distributed optimization and forecasting. All the publications are available in the ORNL-hosted project landing page. The DC interlink with state-of-charge balancing control was operationally validated in Adjuntas by interconnecting two real-world, 240 V split-phase microgrids. To the best knowledge of the team, this represents the first operational validation of AC microgrids interconnected via DC-interlinks. As a culmination of this project, a follow-on grant was awarded to support the technology transfer of the distributed optimization framework to a commercial microgrid controller, Stellar Edge, developed by the California-based company New Sun Road.

14 SOLAR ENERGY↗

Distributed Inertia Management Under Communication Constraints

This paper presents a framework for distributed inertia management based on consensus algorithm. We propose a methodology to achieve optimal operation of inertia sources such as distributed energy resources (DERs) and synchronous generators in real time. Additionally, we analyze the algorithm under communication constraints and evaluate its robustness under scenarios involving communication time delays and packet losses. The proposed approach is validated via simulations on a 4-node system test feeder, demonstrating its effectiveness and resilience.

Yadav, Ajay [ORNL] (ORCID:000000016111881X)↗