Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

A phase-shift-periodic parallel boundary condition for low-magnetic-shear scenarios

Abstract We formulate a generalized periodic boundary condition as a limit of the standard twist-and-shift parallel boundary condition that is suitable for simulations of plasmas with low magnetic shear. This is done by applying a phase shift in the binormal direction when crossing the parallel boundary. While this phase shift can be set to zero without loss of generality in the local flux-tube limit when employing the twist-and-shift boundary condition, we show that this is not the most general case when employing periodic parallel boundaries, and may not even be the most desirable. A non-zero phase shift can be used to avoid the convective cells that plague simulations of the three-dimensional Hasegawa–Wakatani system, and is shown to have measurable effects in periodic low-magnetic-shear gyrokinetic simulations. We propose a numerical program where a sampling of periodic simulations at random pseudo-irrational flux surfaces are used to determine physical observables in a statistical sense. This approach can serve as an alternative to applying the twist-and-shift boundary condition to low-magnetic-shear scenarios, which, while more straightforward, can be computationally demanding.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ANS MultiApps tutorial

The Multiphysics Object-Oriented Simulation Environment (MOOSE) [1] is an open-source, parallel finite element framework which provides the foundation formany advanced modeling and simulation tools developed under the Department of Energy’s (DOE’s) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program [2] for the analysis of advanced reactors. The MOOSE framework provides the common foundational capability on which many NEAMS codes for reactor analysis are built. The MOOSE framework also includes several systems to assemble unique workflows and coupling amongMOOSE-based applications. In particular, the MultiApp and Transfer systems are widely used to assemble different MOOSE-based or MOOSE- wrapped physics applications together to perform loosely or tightly coupled multiphysics simulations. The National Reactor Innovation Center’s (NRIC’s) Virtual Test Bed (VTB) [3] hosts publicly available nuclear reactor multiphysics simulation examples which leverage MOOSE’s MultiApp system to meet the modeling needs of different reactor types. The flexibility and robustness of coupling provided by MOOSE permits rapid development of coupled physics models for a wide range of reactor types and events.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Performance Modeling of a Variable-Geometry Oscillating Surge Wave Energy Converter on a Raised Foundation: Preprint

This paper analyzes the power capture potential, structural loadings, and costs associated with an oscillating surge wave energy converter (OSWEC) operating on a raised foundation. The raised OSWEC offers opportunities for reduced installation costs, improved energy production, and greater flexibility of deployment when compared with bottom-fixed models. In this investigation, several different foundation geometries were simulated using WEC-Sim to estimate power capture and structural loads. In an effort to maximize power capture, several cases in which flat plates of varying size were attached to the top of the foundation, under and parallel with the OSWEC, were also simulated. These plates were found to enhance power capture by preventing the wave induced pressure from passing underneath the OSWEC, diverting this pressure towards the OSWEC instead. The OSWEC was simulated in the six Wave Energy Prize sea states, which were chosen as a representative sample of U.S. deployment sites. A first-order estimate of structural costs was calculated using the Wave Energy Prize ACE metric, with the foundation comprised predominantly of steel-reinforced concrete and the OSWEC comprised of A36 steel. Influence of foundation geometry on power capture, structural loadings, and ACE are topics of particular interest. This work has been inspired by advances in large scale additive manufacturing techniques that have the potential to dramatically reduce the cost of subsea foundations. These advancements may enable cost effective WEC systems to be deployed on raised foundations.

50 EE - Wind and Water Power Program - Water (EE-4↗

Performance Modeling of a Variable-Geometry Oscillating Surge Wave Energy Converter on a Raised Foundation

This paper analyzes the power capture potential, structural loadings, and costs associated with an oscillating surge wave energy converter (OSWEC) operating on a raised foundation. The raised OSWEC offers opportunities for reduced installation costs, improved energy production, and greater flexibility of deployment when compared with fixed-bottom models. In this investigation, we simulated several different foundation geometries using WEC-Sim to estimate power capture and structural loads. In an effort to maximize power capture, several cases in which flat plates of varying size were attached to the top of the foundation, under and parallel with the OSWEC, were also simulated. These plates were found to enhance power capture by preventing the wave-induced pressure from passing underneath the OSWEC, diverting this pressure toward the OSWEC instead. The OSWEC was simulated in the six Wave Energy Prize sea states, which were chosen as a representative sample of U.S. deployment sites. A first-order estimate of structural costs was calculated using the Wave Energy Prize ACE metric, with the foundation comprised predominantly of steel-reinforced concrete and the OSWEC comprised of A36 steel. Influence of foundation geometry on power capture, structural loadings, and ACE are topics of particular interest. This work has been inspired by advances in large-scale additive manufacturing techniques that have the potential to dramatically reduce the cost of subsea foundations. These advancements may enable cost-effective WEC systems to be deployed on raised foundations.

cost↗

Assembling Multiphysics Nuclear Reactor Simulations Using the MOOSE Framework

The Multiphysics Object Oriented Simulation Environment (MOOSE) [1] is an open-source, parallel finite element framework which provides the foundation for many advanced modeling and simulation tools developed under the Department of Energy (DOE) Nuclear Energy Advanced Modeling and Simulation (NEAMS) Program [2] for the analysis of advanced reactors. The MOOSE framework provides the common foundational capability on which many NEAMS codes for reactor analysis are built. The MOOSE framework also includes several systems to assemble unique workflows and couplingamong MOOSE-based applications. In particular, the MultiApp and Transfer Systems are widely used to assemble different MOOSE-based or MOOSE-wrapped physics applications together to perform loosely or tightly coupled multiphysics simulations. The National Reactor Innovation Center (NRIC) Virtual Test Bed (VTB) [3] hosts publicly available nuclear reactor multiphysics simulation examples which leverage MOOSE’s MultiApp System to meet the modeling needs of different reactor types. The flexibility and robustness of coupling provided by MOOSE permits rapid development of coupled physics models for a wide range of reactor types and events

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

CyDER: A Cyber Physical Co-simulation Platform for Distributed Energy Resources in Smartgrids

The CyDER project aimed at developing an open-source, modular and scalable co-simulation platform for power grids with large shares of Distributed Energy Resources (DERs). The project partners are the Lawrence Berkeley National Lab (LBNL), Lawrence Livermore National Lab (LLNL), PG&E, SolarCity, and ChargePoint. The prime recipient is LBNL; SolarCity and ChargePoint were partners for the project’s first two years. Increased DER integration introduces a number of challenges in power grid operation including a more dynamic interaction between the transmission grid and distribution grids, and increased modeling complexity. Although specialized software exists to precisely model different components of the power system, it is far from trivial to integrate all various models and perform a holistic simulation. Instead of replicating all models in a common simulation program, a commonly accepted approach to tackle this model diversity is to couple third-party simulators and models through a co-simulation platform that coordinates information exchange among the various components. Following this line of research, this project’s objective was to develop a co-simulation platform based on a widely accepted industrial standard called Functional Mockup Interface (FMI). Within this process, the project developed models compliant with the FMI standard, called Functional Mockup Units (FMUs), and used them to perform various operational and planning power system analyses. Relying and building upon an industrial standard is the main differentiation of this project compared with previous or parallel efforts in the co-simulation area. Particular emphasis was put on delivering software utilities to facilitate setting up and running co-simulations by end-users. Furthermore, a strong aspect of this project is demonstrating that co-simulation techniques can be used to perform Hardware-in-the-Loop (HIL) simulations that couple software components (e.g., simulated models) with hardware components (e.g., real devices such PV systems and batteries). The long-term goal of CyDER project is to help establish FMI as a powerful standard for co-simulation and promote adoption by electric utilities and other interested stakeholders. The main accomplishments of the project include the development of several FMUs including distribution and transmission grid models, PV inverters with Volt/Var/Watt controllers, batteries, and predictive optimal controllers. Additionally, a unique software package was developed, called SimulatorToFMU, which is capable of exporting any Python-driven simulator or Python script as an FMU. This is an important contribution towards establishing FMI as one of the main co-simulation standards, because more and more third-party programs for sub-system modeling and simulation are delivered with Python APIs. The CyDER platform was used to perform PV hosting capacity analyses in real utility feeders with and without smart inverter controls, battery storage, and EV charging. Smart inverter controls include conventional Volt/Var/Watt controls for reactive power support and active power curtailment, but also predictive controls that optimize the charging and discharging profile of the battery connected on the DC side in order to minimize the customer’s economic benefit. Finally, an important result of this project is delivering an experimental setup that consists of residential-scale PV inverters with battery storage, a real-time grid simulator with an ideal voltage source as grid emulator, and micro Phasor Measurement Units (PMUs). All these components and additional software modules are coupled to one another using the FMI standard and can be co-simulated with the CyDER platform.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Inverter-Based Operation of Maui: Electromagnetic Transient Simulations

As larger and larger power systems approach and reach 100% inverter-based resource (IBR) operation during some hours of the year, questions arise regarding the stability of such extremely-high IBR power systems, the potential need for grid-forming (GFM) inverter technology, and the potential need for synchronous condensers. Relatedly, questions also arise about the ability of conventional positive sequence power system modeling tools to capture high-IBR system dynamics. This presentation introduces electromagnetic transient (EMT) simulations in PSCAD of the near-future (year 2023) Maui power system at and near 100% IBR operation and compares those simulations to positive sequence (PSSE) simulations. The Maui PSCAD model is parallelized on 30 cores and includes the entire transmission system (>200 three-phase buses), >170 individual and aggregate IBR models, four wind plants, and three synchronous generators plus six synchronous condensers at two locations. We investigate system stability with varying levels of inertia using conventional grid-following IBR controls, and then we investigate the impact of GFM controls on stability. Results suggest that: 1) positive sequence simulations can miss key dynamics in extremely high IBR cases; 2) EMT simulations can also miss key dynamics if inverter inner control dynamics are not modeled; 3) synchronous condensers can stabilize a system in which 100% of the energy is supplied by IBRs, even conventional grid-following IBRs; 4) GFM controls on just some of the IBRs can stabilize a 100% IBR power system, even if that system has zero inertia (i.e. no synchronous condensers, though synchronous condensers may be needed for other purposes such as protection system operations).

41 EE - Solar Energy Technologies Office (EE-4S)↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Mechanisms of heat flux across the Southern Greenland continental shelf in 1/10° and 1/12° ocean/sea ice simulations

The increased presence of warm Atlantic water on the Greenland continental shelf has been connected to the accelerated melting of the Greenland Ice Sheet, particularly in the southwest and southeast shelf regions. Results from two high-resolution coupled ocean-sea ice simulations that utilized either the 1/10-degree Parallel Ocean Program (POP) or the 1/12-degree HYbrid Coordinate Ocean Model (HYCOM) are used to understand the flux of heat on and off the southern Greenland shelf. The analysis reveals that the region of greatest heat flux onto the shelf is southeast Greenland. On the southwestern shelf, heat is mainly exported from the shelf to the interior basins. We identify differences in the shelf break current structure and on-shelf heat content between the two simulations. Just south of the Denmark strait, there is a seasonally persistent pattern of multi-day variability in the cross-shelf heat flux in both simulations. In the POP simulation, this high-frequency signal results in net on-shore heat flux. In the HYCOM simulation, the signal is weaker and results in net off-shelf heat flux. This variability is consistent with Denmark Strait Overflow eddies traveling along the shelf break.

58 GEOSCIENCES↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Combining machine-learned and empirical force fields with the parareal algorithm: application to the diffusion of atomistic defects

We numerically investigate an adaptive version of the parareal algorithm in the context of molecular dynamics. This adaptive variant has been originally introduced in [1]. We focus here on test cases of physical interest where the dynamics of the system is modelled by the Langevin equation and is simulated using the molecular dynamics software LAMMPS. In this work, the parareal algorithm uses a family of machine-learning spectral neighbor analysis potentials (SNAP) as fine, reference, potentials and embedded-atom method potentials (EAM) as coarse potentials. We consider a self-interstitial atom in a tungsten lattice and compute the average residence time of the system in metastable states. Our numerical results demonstrate significant computational gains using the adaptive parareal algorithm in comparison to a sequential integration of the Langevin dynamics. We also identify a large regime of numerical parameters for which statistical accuracy is reached without being a consequence of trajectorial accuracy.

36 MATERIALS SCIENCE↗

Rendezvous algorithms for large-scale modeling and simulation

Rendezvous algorithms encode a communication pattern that is useful when processors sending data do not know who the receiving processors should be, or vice versa. The idea is to define an intermediate decomposition where datums from different sending processors can ”rendezvous” to perform a computation, in a manner that both the senders and eventual receivers of the results can identify the appropriate rendezvous processor. Though they were originally designed for interpolating between overlaid grids with independent parallel decompositions (Plimpton et al., 2004), we have recently found rendezvous algorithms useful for a variety of operations in particle- or grid-based simulation codes when running large problems on large numbers of processors. In particular, we show they can perform well when a load-balanced intermediate decomposition is randomized and not spatial, requiring all-to-all communication to move data between processors. In this case rendezvous algorithms leverage the large bisection communication bandwidths which parallel machines provide. We describe how rendezvous algorithms work in a scientific computing context and give specific examples for molecular dynamics and Direct Simulation Monte Carlo codes which result in dramatic performance improvements versus simpler algorithms which do not scale as well. We explain how a generic rendezvous algorithm can be implemented, and also point out similarities with the MapReduce paradigm popularized by Google and Hadoop.

97 MATHEMATICS AND COMPUTING↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗