Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Scaling theory of three-dimensional magnetic reconnection spreading

We develop a first-principles scaling theory of the spreading of three-dimensional (3D) magnetic reconnection of finite extent in the out of plane direction. This theory addresses systems with or without an out of plane (guide) magnetic field, and with or without Hall physics. The theory reproduces known spreading speeds and directions with and without guide fields, unifying previous knowledge in a single theory. New results include: (1) Reconnection spreads in a particular direction if an x-line is induced at the interface between reconnecting and non-reconnecting regions, which is controlled by the out of plane gradient of the electric field in the outflow direction. (2) The spreading mechanism for anti-parallel collisionless reconnection is convection, as is known, but for guide field reconnection it is magnetic field bending. We confirm the theory using 3D two-fluid and resistive-magnetohydrodynamics simulations. (3) The theory explains why anti-parallel reconnection in resistive-magnetohydrodynamics does not spread. (4) The simulation domain aspect ratio, associated with the free magnetic energy, influences whether reconnection spreads or convects with a fixed x-line length. (5) We perform a simulation initiating anti-parallel collisionless reconnection with a pressure pulse instead of a magnetic perturbation, finding spreading is unchanged rather than spreading at the magnetosonic speed as previously suggested. The results provide a theoretical framework for understanding spreading beyond systems studied here, and are important for applications including two-ribbon solar flares and reconnection in Earth’s magnetosphere.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Less can be more: Insights on the role of electrode microstructure in redox flow batteries from two-dimensional direct numerical simulations

Understanding how to structure a porous electrode to facilitate fluid, mass, and charge transport is key to enhancing the performance of electrochemical devices, such as fuel cells, electrolyzers, and redox flow batteries (RFBs). Here, using a parallel computational framework, direct numerical simulations are carried out on idealized porous electrode microstructures for RFBs. Strategies to improve an electrode design starting from a regular lattice are explored. By introducing vacancies in the ordered arrangement, it is possible to achieve higher voltage efficiency at a given current density, thanks to improved mixing of reactive species, despite reducing the total reactive surface. Careful engineering of the location of vacancies, resulting in a density gradient, outperforms disordered configurations. Our simulation framework is a new tool to explore transport phenomena in RFBs, and our findings suggest new ways to design performant electrodes.

25 ENERGY STORAGE↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Interagency Agreement No. DE-SC0006988 (Final Technical Report)

The overarching goal of this project was to advance understanding of deep convection updraft microphysics—the source of long-lived stratiform ice—by combining detailed mining of in situ and remote sensing data from the MC3E field campaign with detailed 3D simulations. The project began with parallel work, first strictly on the remote-sensing observation side via dedicated analysis of polarimetric radar signatures, which is a relatively new area, on the one hand. On the other hand, a more traditional but well-formulated preparation and comparison of detailed aerosol-aware simulations with in situ measurements was prepared. In the final phase of this work, these parallel elements were brought together into an integrated analysis of in situ and remote-sensing observations with model results. The project also supported team member involvement in collaborative activities.

54 ENVIRONMENTAL SCIENCES↗

Kinetic study of shock formation and particle acceleration in laser-driven quasi-parallel magnetized collisionless shocks

Quasi-parallel magnetized collisionless shocks are believed to be one of the most efficient accelerators in the universe. Compared to quasi-perpendicular shocks, quasi-parallel shocks are more difficult to form in the laboratory and to simulate because of their large spatial scales and long formation times. Our two-dimensional particle-in-cell simulations show that the early stages of quasi-parallel shock formation are achievable in experiments planned for the National Ignition Facility and that particles accelerated by diffusive shock acceleration (DSA) are expected to be observable in the experiment. Repetitive ion acceleration by crossings of the shock front, a key feature of DSA, is seen in the simulations. Other characteristic features of quasi-parallel shocks such as upstream wave excitation by energetic ions are also observed, and energy partition between the ions and the electrons in the downstream of the shock is briefly discussed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cyber-Power Co-Simulation for End-to-End Synchrophasor Network Analysis and Applications

The resiliency, reliability and security of the next generation cyber-power smart grid depend upon efficiently leveraging advanced communication and computing technologies. Also, developing real-time data-driven applications is critical to enable wide-area monitoring and control of the cyber-power grid given high-resolution data from Phasor Measurement Units (PMUs). North American Synchrophasor Initiative Network (NASPlnet) provides guidance for PMU data exchanges. With the advancement in networking and grid operation, it is necessary to evaluate the performance of different data flow architectures suggested by NASPInet and analyze the impact on applications. Therefore, we need a cyber-power co-simulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents an end-to-end automated and user-driven cyber-power co-simulation using NS3 to model communication networks, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application utilizing PMU data in an IEEE 39 bus test system is presented using this cosimulation testbed.

Mustafa, Hussain M.↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

Performance and Feature Improvements in Parareal-based Power System Dynamic Simulation

In recent years, a novel Parareal-based approach has been developed for fast transient simulations of large power system interconnections. Parareal belongs to the class of Parallel-in-time algorithms for solution of systems of differential-algebraic equations in parallel over an interval of time. The selection of a reasonably fast and accurate coarse solution is crucial to improve the performance of Parareal algorithm. Semi-analytical solution methods are one promising approach to achieve this goal. They have been investigated, and some preliminary results are presented here. In addition, Parareal-based simulator has been expanded to enable co-simulation with OpenDSS, a widely used open-source distribution system simulator. Preserving the parallel nature of the Parareal approach and taking advantage of the parallel capabilities of the latest versions of OpenDSS, each distribution system can be solved in their entirety on different processors in parallel within the main Parareal simulator. This paper also presents the structure of the transmission and distribution co-simulation and some results with different dynamic models of inverter-based resources in the distribution systems.

Park, Byungkwon↗

QTENSOR

QTensor is a quantum circuit simulator designed to run efficiently in parallel mode on large supercomputers. It is based on the tensor network representation of quantum circuits. The simulator is flexible and agnostic to both the connectivity map of the quantum device and the types of gates used in the circuit.

Alekseev, Yury↗

PipeSight: A High-Performance Computing Platform for Pipeline Integrity Management

The Phase I feasibility study completed as part of this project has led to a number of innovative technologies being developed and has laid the foundation for a successful Phase II effort to commercialize a platform for managing the integrity of pipelines for the damage mechanisms of the new, hybrid-energy based economy. To ground the development efforts and direction of the project, an extensive market research and customer discovery effort was undertaken early in Phase I. Through this effort, a number of pipeline owners and operators were interviewed, and the following key findings were discovered about the pipeline industry: • Small pipeline operators do not have the central engineering groups necessary to perform their own independent analysis of inspection data, but instead rely on summarized tally sheets provided to them by inspection service providers. • The time it takes to go from an inspection to a completed engineering assessment, even for small segments of pipeline, can take anywhere from 30-120 days. During this delay, critical threats can (and have been known to) cause failures. • Uncertainty is often not accounted for in the assessment of pipeline integrity. The tally sheets provided by third-party service providers are almost always deterministic in nature, identifying threats that present a concern only to the current (not the future) integrity of the pipeline. • It is uncommon to apply the latest technologies to perform advanced assessments of damaged pipelines. There is a desire to use more advanced analysis capabilities to assess threats. Many pipeline operators indicated that they would often excavate a pipeline to perform an inspection and find that the damage was not as bad as they anticipated, thus using limited resources unnecessarily. Companies are not consistent in their use of inspection data to determine corrosion rates, and those that do only calculate deterministic corrosion rates. • The industry has prominently relied on time-based inspections but has recently started to transition to risk-based inspections. However, there appears to be no uniform guidance on how to do so while properly accounting for all sources of uncertainty. • Companies are not storing inspection data in a manner that allows for the ready determination of temporal trends. • Predictive maintenance principles and practices are beginning to be used by early adopters • Some pipelines are being re-purposed to transport different process fluids than they were designed for, e.g., H 2 and CO 2 rich process streams to serve the new hybrid-energy based economy, which are presenting new integrity concerns for the existing pipeline network that crisscrosses the United States. As a result of these discoveries, we were able to target the development efforts in Phase I to best serve the needs of the industry. In Phase I, we developed a way to correlate multiple large-scale scans of the pipeline to determine a probabilistic corrosion rate that accounts for all sources of error and uncertainty in the inspection process. This probabilistic corrosion rate can be used to predict the future thickness distribution of the pipe wall. We demonstrate how this analysis may be performed in an analytical fashion and has been implemented in such a manner that it can be readily distributed using GPU computing through integration of the Kokkos programming model. We also make a very novel extension of the analytical corrosion rate model to Bayesian Networks (an explainable AI technique) that can account for non-parametric distributions of corrosion rates. With the predictions made above for the probabilistic corrosion rate and corresponding future distribution of the pipe wall thickness, we can assess the integrity of the pipeline through the use of a probabilistic engineering assessment. We developed a novel screening data analysis approach that can rapidly identify ‘hotspots’ (local thin areas) where the integrity of the pipeline is a concern. Once more, we implemented this screening approach in C++ to leverage GPU computing via the Kokkos programming model. After the critical hotspots are identified, we developed a program that can automatically generate an advanced finite element model of the damaged regions. Since the number of damaged regions that require advanced analysis can number in the thousands, we integrated an open-source container-native workflow engine for orchestrating parallel jobs on the cloud. Initially, these advanced numerical models were only designed to account for loading due to internal pressure. However, in a slight pivot from the initial Phase I proposal, we developed a complete pipe stress analysis program (called Simflex) which can simulate the complete pipeline and its response to thermal expansion, pressure, thermal bowing, weight, wind, earthquake, support displacement, support friction and external forces. This pipe stress analysis program was written generically, to handle any piping system, but contains the features needed to model long pipelines (i.e., it incorporates a model for soil mechanics and can account for the nonlinear boundary conditions necessary to simulate long underground pipelines). This pipe stress analysis program can simulate any segment of the pipeline (simple or complex) under any set of conditions and loads, to determine the supplemental loads (axial forces and bending moments) at the location of damage. This enables the most accurate state of stress to be accounted for in the pipeline, which can prove critical when evaluating the integrity of a damaged region. In the process of developing the technologies to perform the integrity assessment of the pipeline, we also extended one of the industry standard approaches for performing the assessment of local thin areas that extend more in the circumferential direction than the longitudinal direction of the pipeline. This approach was presented to the API 579-1/AS ME FFS-1 steering committee in November 2021 for consideration in the next edition of the industry standard for Fitness-For-Service (expected to be released in 2023). To help pipeline operators make decisions with the results on any integrity assessment, we developed a new approach to the life-cycle management of pipelines which uses a Bayesian Decision Network. The network is designed to help pipeline operators plan and prioritize inspection activities and ultimately make smarter, more cost-effective decisions. The Bayesian approach accounts for all sources of uncertainty and carries them through to the final optimal decisions, providing a probabilistic framework for optimizing inspection intervals. The proof-of-concept networks developed in the feasibility study are complete, verified, and are focused on a subset of the pipeline. To expand this novel approach to the scale necessary for an entire network of pipelines in Phase II, we will leverage the DOE-funded Bengi solver for industrial-scale decision making with Bayesian Networks [22]. Once implemented, we will be able to provide the pipeline industry with a much-needed tool for optimal inspection planning using truly explainable artificial intelligence (XAI). To handle all of these advanced capabilities into a cloud-based platform, the architecture of the Equity Engineering Cloud (EEC) was extended to include Argo Workflows, a framework capable of distributing and managing a massive number of jobs that consume their own resources, such that thousands of serial finite element simulations can be run in parallel. As part of this substantial undertaking, we also integrated Argo Continuous Delivery (CD) into the EEC, to aid with the rapid prototyping and iterations that will be imperative to the success of the PipeSight platform’s Agile development process in Phase II. As part of the pipe stress analysis program, we also developed a custom visualizer that leverages the DOE-funded VTK visualization library. We added custom contouring capabilities and a means for interacting visually with both the inputs and outputs of the pipe stress analysis program. We also developed routines for automating the post-processing of the finite element simulations to determine if any failure criteria are met and to visualize the deformations, stresses and strains in ParaView using the exodus II file format (a subset of netCDF).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Electron‐Scale Reconnection in Three‐Dimensional Shock Turbulence

Abstract Magnetic reconnection has been observed in the transition region of quasi‐parallel shocks. In this work, the particle‐in‐cell method is used to simulate three‐dimensional reconnection in a quasi‐parallel shock. The shock transition region is turbulent, leading to the formation of reconnecting current sheets with various orientations. Two reconnection sites with weak and strong guide fields are studied, and it is shown that reconnection is fast and transient. Reconnection sites are characterized using diagnostics including electron flows and magnetic flux transport. In contrast to two‐dimensional simulations, weak guide field reconnection is realized. Furthermore, the current sheets in these events form in a direction almost perpendicular to those found in two‐dimensional simulations, where the reconnection geometry is constrained.

58 GEOSCIENCES↗

Demonstration of Advanced Experimental and Theoretical Characterization of Hydrogen Dynamics and Associated Behavior in Advanced Reactors

Advanced materials development, manufacturing, and modeling capabilities for innovative reactor designs support nuclear security and mission-focused science through enhanced technology for safer and more efficient and secure production of nuclear energy. The research in this project has established: 1) a state-of-the-art neutron-based hydrogen mapping and cross-section measurement capability as well as detailed crystallographic characterization of hydrogen atoms at LANSCE, and 2) a multi-physics framework for simulating behavior of moderator materials and other material performance in advanced nuclear reactors. Through the course of this project, we successfully developed and demonstrated measurement techniques for hydrogen distribution and atomistic-scale behavior of hydrogen atoms using pulsed neutron techniques. In parallel, advanced multi-physics simulation tools to predict the behavior of hydrogen atoms, e.g. in a moderator for a nuclear reactor, through materials performance, neutron transport, and thermal mechanical behavior were enhanced. Multi-discipline areas across the laboratory were involved in the project as the integration of improved experimental capabilities with enhanced modeling and simulation through MST, NEN, SIGMA, and XCP division subject matter experts.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Development of a Single-Phase, Transient, Subchannel Code, within the MOOSE Multi-Physics Computational Framework

Subchannel codes have been widely used for thermal-hydraulics analyses in nuclear reactors. This paper details the development of a novel subchannel code within the Idaho National Laboratory’s (INL) Multi-physics Object Oriented Simulation Environment (MOOSE). MOOSE is a parallel computational framework targeted at the solution of systems of coupled, nonlinear partial differential equations, that often arise in the simulation of nuclear processes. As such, it includes codes/modules able to solve the multiple linear and nonlinear physics that describe a nuclear reactor, under normal operation conditions or accidents. This includes thermal-hydraulics, fuel performance, and neutronics codes, between others. A MOOSE-based subchannel code is a new addition to the fleet of INL-developed codes, based on the MOOSE framework. In this work, we present the derivation of the subchannel equations for a single-phase fluid, we proceed with the description of the algorithm that is used to solve these equations and describe how this algorithm was implemented within MOOSE. We also present how this code can be coupled to the BISON fuel performance code. Next, we verify the friction model and the turbulent mixing model. We calibrate the turbulent modeling parameters for momentum mixing and enthalpy mixing, C T , β. We validate the code using experimental results and last demonstrate the coupling capabilities using a simple example.

42 ENGINEERING↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N. One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an “intelligent” mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗