Parallel Computing for Power System Climate Resiliency: Solving a Large-Scale Stochastic Capacity Expansion Problem with mpi-sppy
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.
The RMF (Rotating Magnetic Field) code is designed to calculate the motion of a charged particle in a given electromagnetic field. It integrates Hamilton’s equations in cylindrical coordinates using an adaptive predictor-corrector double-precision variable-coefficient ordinary differential equation solver for speed and accuracy. RMF has multiple capabilities for the field. Particle motion is initialized by specifying the position and velocity vectors. Here, the six-dimensional state vector and derived quantities are saved as functions of time. A post-processing graphics code, XDRAW, is used on the stored output to plot up to 12 windows of any two quantities using different colors to denote successive time intervals. Multiple cases of RMF may be run in parallel and perform data mining on the results. Recent features are a synthetic diagnostic for simulating the observations of charge-exchange-neutral energy distributions and RF grids to explore a Fermi acceleration parallel to static magnetic fields.
We present a computationally improved heuristic algorithm for transmission switching (TS) to recover load shed. Research from the past showed that changing power system topology may control power flows and remove line congestion. Hence, TS may reduce the required load shed. One of the main challenges is to find a potential TS candidate in a suitable time. Here, we propose a novel heuristic method that is capable of finding the potential TS candidate faster than existing algorithms in literature. The proposed method is compatible with both the AC and DC optimal power flows (OPF). Three metrics are used to compare the proposed algorithm with the state-of-the-art from literature to show the speedup and accuracy achieved. The proposed method is implemented on the IEEE 30-bus system, PEGASE 89-bus system, IEEE 118-bus system, and Polish 2383- bus system. The results on the large-scale Polish 2383-bus system shows that the proposed algorithm is scalable to large real-world systems. Parallel computing is implemented to further improve the computational performance of the proposed algorithm.
A general strategy is presented for constructing and validating permutationally invariant polynomial (PIP) expansions for chemical systems of any stoichiometry. Demonstrations are made for three categories of gas-phase dynamics and kinetics: collisional energy-transfer trajectories for predicting pressure-dependent kinetics, three-body collisions for describing transient van der Waals adducts relevant to atmospheric chemistry, and nonthermal reactivity via quasiclassical trajectories. In total, 30 systems are considered with up to 15 atoms and 39 degrees of freedom. Permutational invariance is enforced in PIP expansions with as many as 13 million terms and 13 permutationally distinct atom types by taking advantage of petascale computational resources. The quality of the PIP expansions is demonstrated through the systematic convergence of in-sample and out-of-sample errors with respect to both the number of training data and the order of the expansion, and these errors are shown to predict errors in the dynamics for both reactive and nonreactive applications. Here, the parallelized code distributed as part of this work enables the automation of PIP generation for complex systems with multiple channels and flexible user-defined symmetry constraints and for automatically removing unphysical unconnected terms from the basis set expansions, all of which are required for simulating complex reactive systems.
Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.
We present Phoenix, a scalable hypergraph analytics framework for data analytics and knowledge discovery that was implemented on the leadership class computing platforms at Oak Ridge National Laboratory (ORNL). Our software framework comprises a distributed implementation of a streaming server architecture which acts as a gateway for various hypergraph generators/external sources to connect. Phoenix has the capability to utilize diverse hypergraph generators, including HyGen, a very large-scale hypergraph generator developed by ORNL. Phoenix incorporates specific algorithms for efficient data representation by exploiting hidden structures of the hypergraphs. Our experimental results demonstrate Phoenix’s scalable and stable performance on massively parallel computing platforms. Phoenix’s superior performance is due to the merging of high-performance computing with data analytic.
Particle accelerators are among the most versatile and important tools of scientific discovery. The Nation's accelerators are responsible for a wealth of advances in materials science, chemistry, the biosciences, particle physics, and nuclear physics. They also have important applications to national security, the environment, energy, medicine, and on the quality of people's lives. LANL has a long history of making pioneering contributions to Accelerator Science including key contributions to the field of Computational Accelerator Physics. These include the development of early beam dynamics codes with space charge (such as PARMILA and PARMELA), the development of rf cavity codes and magnet codes (including Poisson and Superfish), and the development and distribution of codes to the accelerator community through the Los Alamos Accelerator Code Group. LANL researchers also helped pioneer the development of massively parallel space-charge codes. In project t22_accelsim we have moved beyond electrostatic models of collective effects (i.e., solving the Poisson equation in the bunch frame) to fully electromagnetic models based on the Lienard-Wiechert formalism. This approach enables the large-scale simulation of radiation production and collective effects in high brightness electron beams. This is highly relevant to LANL given its future goal of developing an X-ray Free Electron Laser (XFEL). It also directly impacts a LANL LDRD project to develop an undulator-based non-invasive beam profile monitor for beams created in laser-plasma accelerator systems.
In this paper, we discuss our experience in implementing homology computation, in particular Betti number calculation in Chapel hypergraph Library (CHGL). Given a dataset represented as a hypergraph, a Betti number for a particular dimension $k$ indicates how many $k$-dimensional `voids' are present in the dataset. Computing Betti number involves various array-centric and linear algebra operations. We demonstrate that implementing these operations in Chapel is both concise and intuitive. In addition, we show that Chapel provides language constructs for implementing parallel and distributed execution of the linear algebra kernels with minimal effort. Syntactically, Chapel provides succinctness of Python, while delivering comparable and better performance than C++-based and Julia-based packages for calculating Betti numbers respectively.
The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.
The versatility and wide-ranging applicability of the Ising model, originally introduced to study phase transitions in magnetic materials, have made it a cornerstone in statistical physics and a valuable tool for evaluating the performance of emerging computer hardware. Here, we present a novel implementation of the two-dimensional Ising model on Cerebras Wafer-Scale Engine (WSE) – a revolutionary processor that is opening new frontiers in computing. In our deployment of the checkerboard algorithm, we optimized the Ising model to take advantage of the unique WSE architecture. Specifically, we employed a compressed bit representation storing 16 spins on each int16 word, and efficiently distributed the spins over the processing units enabling seamless weak scaling and limiting communications to only immediate neighboring units. Our implementation can handle up to 754 simulations in parallel, achieving an aggregate of over 61.8 trillion flip attempts per second for Ising models with up to 200 million spins. This represents a gain of up to 148 times over previously reported single-devices with a highly optimized implementation on NVIDIA V100 and up to 88 times in productivity compared to NVIDIA H100. Our findings highlight the significant potential of the WSE in scientific computing, particularly in the field of materials modeling.
In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.
Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.
Axom is an open-source library of "building block" software components that provide core infrastructure capabilities for HPC applications. Such capabilities include: input file parsing and verification for various file formats, format specification for material shape input to multi-material simulations and tools for placing those shapes on meshes, parallel distributed coordination of diagnostic messages, geometric primitives, spatial queries and spatial search acceleration data structures, in memory key-value data store for managing simulation data and parallel file I/O. Axom components are used widely across LLNL Advanced Simulation and Computing program applications. It is also part of the LLNL institutionally-supported RADIUSS project, which promotes and funds its adoption by projects across LLNL.
Distributed Acoustic Sensing (DAS) is an emerging sensing technology that records the strain-rate along fiber optic cables at high spatial and temporal resolution. This technique is becoming a popular tool in seismology, hydrology, and other subsurface monitoring applications. However, due to the large coverage (10’s of km) and high density of measurements (1m spacing at 100’s of Hz), a DAS installation could produce terabytes of data records per day. Because many DAS instruments are deployed in remote locations, this large data size poses significant challenges to its transfer and storage. In this paper, we explore lossless compression methods to reduce the storage requirement in both real-time and post-hoc scenarios. Here we propose a two-stage compression method to improve the compression ratio and compression speed. This two-stage compression method could reduce the storage requirement by 40%, which is 20% more than other lossless methods, such as ZSTD. We demonstrate that the compression method could complete its operation well before the DAS instrument needs to output the next file, making it suitable for real-time DAS acquisition. We also implement a parallel compression method for a post-hoc scenario and demonstrate that our method could effectively utilize a parallel computer. With 256 CPU cores, our parallel compression method achieves the speed of 26GB/second.
The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.
We develop an Evolutionary Markov Chain Monte Carlo (EMCMC) algorithm for sampling spatial partitions that lie within a large, complex, and constrained spatial state space. Our algorithm combines the advantages of evolutionary algorithms (EAs) as optimization heuristics for state space traversal and the theoretical convergence properties of Markov Chain Monte Carlo algorithms for sampling from unknown distributions. Local optimality information that is identified via a directed search by our optimization heuristic is used to adaptively update a Markov chain in a promising direction within the framework of a Multiple-Try Metropolis Markov Chain model that incorporates a generalized Metropolis-Hastings ratio. We further expand the reach of our EMCMC algorithm by harnessing the computational power afforded by massively parallel computing architecture through the integration of a parallel EA framework that guides Markov chains running in parallel.
The confinement effect on crosslinking and gelation during controlled/living copolymerization of a monovinyl and a divinyl monomer was studied by Monte Carlo computer simulation using the dynamic lattice liquid method. The simulated reactions took place in plane-parallel slits of various widths for various grafting densities (GDs) and various initial initiator/crosslinker ratios ([X] 0 /[Ini] 0 ). Monomers and crosslinkers were initially randomly distributed in the slit, but initiators were either anchored to one solid plane (polymer brushes) or randomly distributed (free chains). For brushes, the gel point (GP) depended on GD and [X] 0 /[Ini] 0 . The higher these parameters, the lower was the gel point. The gel point was also affected by the slit width. For free, nongrafted chains, the dependence of GP on slit width was negligible. For brushes, a nonuniform density of crosslinks was found, the highest density was close to the wall. This asymmetry was high at the gel point and decreased at higher conversions.