Engineering PapersSearch

SEARCH · Engineering Papers

Results for “distributed algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE

Radiation image reconstruction and uncertainty quantification using a Gaussian process prior

We propose a complete framework for Bayesian image reconstruction and uncertainty quantification based on a Gaussian process prior (GPP) to overcome limitations of maximum likelihood expectation maximization (ML-EM) image reconstruction algorithm. The prior distribution is constructed with a zero-mean Gaussian process (GP) with a choice of a covariance function, and a link function is used to map the Gaussian process to an image. Unlike many other maximum a posteriori approaches, our method offers highly interpretable hyperparamters that are selected automatically with the empirical Bayes method. Furthermore, the GP covariance function can be modified to incorporate a priori structural priors, enabling multi-modality imaging or contextual data fusion. Lastly, we illustrate that our approach lends itself to Bayesian uncertainty quantification techniques, such as the preconditioned Crank–Nicolson method and the Laplace approximation. The proposed framework is general and can be employed in most radiation image reconstruction problems, and we demonstrate it with simulated free-moving single detector radiation source imaging scenarios. We compare the reconstruction results from GPP and ML-EM, and show that the proposed method can significantly improve the image quality over ML-EM, all the while providing greater understanding of the source distribution via the uncertainty quantification capability. Furthermore, significant improvement of the image quality by incorporating a structural prior is illustrated.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

A Comprehensive Strategy for Grid Forming Control in DC Coupled Photovoltaic and Battery Energy Storage Inverters

This paper presents an integrated DC-DC and DCAC grid-forming control strategy for DC-coupled photovoltaic (PV) plus battery energy storage systems, considering the effect of DC link voltage variations caused by direct PV connections. A power reference algorithm determines power distribution between the PV and battery to the grid while observing device power ratings to prevent the over-rating of components and keep the battery's state of charge within an acceptable range. The simulated utility-scale model in MATLAB/Simulink illustrates its ability against extreme phase angle variation contingencies in the grid while controlled through grid-forming control with a fast dynamic on DC link voltage. The simulation results confirm the effectiveness of the proposed control in integrating PV plus battery configurations with grid forming control and maintaining reliable grid operation under severe grid disturbances.

battery, boost, control, energy storage, grid form

Adaptive Linear State Estimation for Unbalanced Distribution System

The inclusion of PMU functionality in distribution relays enables the implementation of a linear state estimator (LSE) in Distribution Systems (DS). However, the unbalanced topology and phase coupling in distribution lines necessitate modifications to the LSE formulation. Additionally, the higher fault frequency in distribution systems requires a state estimation approach that is resilient to contingencies. This work proposes an adaptive linear state estimation algorithm tailored for unbalanced distribution systems with single-phase and two-phase laterals. Furthermore, a modified Optimal PMU Placement (OPP) strategy is introduced to ensure full observability in distribution systems with single-phase and two-phase buses. To maintain adaptability to topology changes, the state estimator incorporates circuit breaker status data provided by PMUs, ensuring robust performance during topology changes triggered by faults. The performance of the algorithm is verified on the IEEE 13-bus, 34-bus, and 123-bus systems.

PMUs

Optimal Operation and Impact Assessment of Distributed Wind for Improving Efficiency and Resilience of Rural Electricity Systems

This project aims to empower rural utilities by developing advanced optimization models and algorithms for effectively integrating distributed wind energy alongside battery storage and other distributed energy resources (DERs). The primary objectives are to reduce peak demand, ensure reliable emergency power supply, and regulate voltage and frequency. To address operational challenges, the project introduces innovative mitigation strategies and ultrafast assessment frameworks to evaluate the impacts of distributed wind and DERs on rural grids, offering actionable solutions to potential issues. Economic viability is assessed through cost-benefit analysis using real rural utility data, ensuring the practical application of the project outcomes.

17 WIND ENERGY

Clustering at Massive Scale

ClaMS provides hierarchical clustering technology for use on massive, high-dimensional datasets that require distributed memory for processing. The algorithm employed is inspired by the popular HDBSCAN algorithm but makes use of computational kernels better suited for distributed computing. ClaMS is built on scalable nearest neighbor graph construction, metric forest completion, and approximate minimum spanning tree techniques.

Stanley, ThomasA [Lawrence Livermore National Labo

Particle Tracking Methods for Battery Precipitation Reactions

Precipitation and deposition reactions at solid–liquid interfaces play a key role in a number of battery chemistries, including Li-ion, so-called “anode free” batteries, zinc-based battery chemistries, and lithium–sulfur, among others. Although models with heterogeneous nucleation and growth phenomena are present in the literature, papers have not to date provided much detail on the numerical algorithms used to track the temporal evolution of the particle size distribution of deposits on electrode surfaces. In this paper we examine several approaches to discretize and track the particle size distribution, demonstrating that common approaches lead to anomalous flattening of the particle size distribution. We conclude by presenting an algorithm that preserves the appropriate particle size distribution during particle growth.

Algorithms

Beam loss modeling and mitigation due to intra-beam stripping

Intra-Beam Stripping (IBS) is a critical beam loss mechanism in high-intensity H- linacs and presents a significant limitation to increasing beam power. This work presents a computational framework to evaluate and mitigate IBS-induced beam loss along the Spallation Neutron Source (SNS) LINAC. Our calculation is based on an analytic theory and involves evaluation of a 9D integral using the Monte-Carlo technique. We first benchmarked our calculations against simplified, analytically solvable cases. We then applied our algorithm to Gaussian bunches with a known probability density function (PDF). We next expanded our algorithm to arbitrary bunch distributions using the Neural Spline Flow (NSF) models trained on PyORBIT tracking data. In the future, we plan to validate our algorithm experimentally and apply it to design IBS mitigation strategies.

Nln, Shivam [ORNL]

An implementation of a high-order generalized finite difference method for solving the time-harmonic cold plasma wave equation in toroidal geometry

A high-order physics-informed meshless finite difference numerical technique is introduced for solving the time-harmonic cold plasma wave equation in toroidal geometries, presenting a novel application of the generalized finite difference (GFD) method to plasma wave simulations. The algorithm employs an irregular distribution of computational points, with local point density informed by the shortest wavelength derived from the cold plasma dispersion relation. Numerical stability and robustness are addressed using regularization techniques. The algorithm, implemented for two spatial dimensions, solves for the wave electric field and is demonstrated to achieve convergence rates of $\mathcal{O}$($\mathcal{h}$ $\mathcal{P}$ )⁠. Verification tests reproduce plane wave solutions, and example simulations of ion cyclotron resonance heating and electron cyclotron resonance heating demonstrate its capability, approaching realistic tokamak plasma scenarios. This work contributes to laying a foundation for the GFD method to be used in more sophisticated, optimized, and physically realistic full-wave simulations in time-harmonic plasma wave research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)

Grid Modernization of Cooperatives and Municipal Utilities via Breakthrough System Monitoring, Control and Optimization (CRADA Final Report)

This project aims at developing and demonstrating successful implementation of breakthrough approaches in real-time data visualization as well as real-time distributed DER control and optimization to provide ample benefits to both utilities and end users. The National Renewable Energy Laboratory (NREL), Holy Cross Energy (HCE), National Rural Electric Cooperative Association (NRECA) and Survalent are collaborating to enable Cooperative and Municipal utilities to fully leverage DERs as part of their strategies for providing safe, reliable, and affordable electric services to their customers and help meet DOE Grid modernization goal of achieving at least 10% active devices to provide grid flexibility by 2035. This project will use novel real-time control algorithms and approaches for distributed control recently developed under DOE-funded projects, using the date from the Survalent’s basic SCADA engine, GIS and AMI engines deployed at HCE combined with NRECA’s globally-used MultiSpeak(R) software interoperability specification for seamless and real-time communications between electric utility enterprise software to embrace DER as part of their strategies for providing safe, reliable and affordable electric service to their customers.

24 POWER TRANSMISSION AND DISTRIBUTION

Resilient Operation of Networked Community Microgrids with High Solar Penetration

This project, funded by the US Department of Energy’s Solar Energy Technologies Office (SETO), focused on the operation of microgrids as a coordinated network. The primary objective, which was successfully achieved, was to develop both control strategies and hardware solutions to support the resilient and efficient operation of networked microgrids with high solar penetration. The work was structured around the following four main tasks: • Development of distributed and scalable optimization algorithms for AC-coupled networked microgrids. • Design and implementation of a novel DC interconnection hardware to enable precise power exchange between microgrids. • Laboratory operational validation of the developed technologies using 480 V testbeds and commercially available hardware. • Field operational validation of the complete solution in Adjuntas, Puerto Rico, interconnecting two kW-scale, split-phase microgrids of Casa Pueblo’s microgrids. This project addressed multiple technical challenges across the domains of optimization, control, hardware interconnection, and protection. One of its key contributions was delivering tangible, real-world solutions for networking microgrids. In contrast to purely theoretical or simulation-based work, this project included full-scale hardware operational validation both in the lab and in the field. The work conducted as part of this project—in collaboration with the University of Puerto Rico; the University of Tennessee, Knoxville; the University of Central Florida; and Casa Pueblo—has advanced the state of the art in networked microgrids. Key contributions include the development of distributed control strategies, practical solutions for real-world implementation challenges, and the introduction of a novel DC interlink approach for microgrid interconnection. The project featured both laboratory and field validation using commercial off-the-shelf components. The field deployment successfully validated that a group of microgrids can operate in a coordinated manner, enabling precise power flow between systems and mutual support during extreme events. This project resulted in 15 journal publications and 15 conference papers; 5 graduate students and 15 undergraduate students were supported. The codes of distributed optimization and forecasting were made open-source through OSTI.gov for distributed optimization and forecasting. All the publications are available in the ORNL-hosted project landing page. The DC interlink with state-of-charge balancing control was operationally validated in Adjuntas by interconnecting two real-world, 240 V split-phase microgrids. To the best knowledge of the team, this represents the first operational validation of AC microgrids interconnected via DC-interlinks. As a culmination of this project, a follow-on grant was awarded to support the technology transfer of the distributed optimization framework to a commercial microgrid controller, Stellar Edge, developed by the California-based company New Sun Road.

14 SOLAR ENERGY

Distributed Inertia Management Under Communication Constraints

This paper presents a framework for distributed inertia management based on consensus algorithm. We propose a methodology to achieve optimal operation of inertia sources such as distributed energy resources (DERs) and synchronous generators in real time. Additionally, we analyze the algorithm under communication constraints and evaluate its robustness under scenarios involving communication time delays and packet losses. The proposed approach is validated via simulations on a 4-node system test feeder, demonstrating its effectiveness and resilience.

Yadav, Ajay [ORNL] (ORCID:000000016111881X)

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms

Fast calculation of diffraction patterns from an ensemble of aligned molecules

We report an algorithm to calculate electron diffraction patterns for molecules with anisotropic angular distribution, which is significantly faster than existing methods. The algorithm uses a transform to convert the molecular orientation distribution, which is a function of three Euler angles, to the atom-pair distribution functions which depend on the polar and azimuthal angles. The diffraction signal can then be calculated from the atom-pair distributions. We demonstrate the computation method numerically by calculating electron diffraction patterns for a symmetric top molecule (trifluoroiodomethane) and an asymmetric top molecule (formaldehyde) and show that it reduces the calculation time by approximately two orders of magnitude compared to the standard brute-force method. Here, the method can also be applied to the calculation of x-ray diffraction patterns.

74 ATOMIC AND MOLECULAR PHYSICS

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]

Understanding the Effect of Sample Geometry on Temperature Distribution during Optical Floating Zone Crystal Growth in Vacuum Environment through Heat Transfer Modeling

Optical floating zone furnaces (OFZ) have had a transformative impact on fundamental science due to their ability to rapidly produce large single crystals of a wide variety of complex materials. However, a quantitative understanding of the OFZ growth environment is generally lacking due to the difficulty of measuring the local sample temperatures during OFZ growth, as well as to the general lack of information about the temperature-dependent physical parameters needed to model heat transfer. To overcome these challenges, we apply a physics-based heat transfer model, parametrized by measurements from synchrotron experiments and a machine-learning (ML) algorithm, to simulate the temperature distributions of samples heated in an OFZ furnace in a vacuum environment. This model is used to quantitatively understand how the sample maximum temperature and temperature gradient (key parameters that influence the success of crystal growth) are affected by the rod size, rod shape, and heat-zone position on the rod. The results of this study can be applied to make informed decisions on how crystal growth parameters can be tuned to modify temperature profiles and to optimize crystal growth outcomes even when data on internal sample temperature profiles (e.g., those obtained through in situ synchrotron experiments) are not accessible.

36 MATERIALS SCIENCE