Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantitative proton radiography and shadowgraphy for arbitrary intensities

Charged-particle radiography and shadowgraphy data can be directly inverted to obtain a line-integrated transverse Lorentz force or a line-integrated transverse refractive index gradient if intensity modulations due to scattering and absorption are negligible, and angular deflections are small. We develop a new direct-inversion algorithm based on plasma physics and compare it to a new Monge–Ampère code and an existing power diagram code. The measured or source intensity is represented by electrons subject to drag, and the other intensity by fixed ions. The decrease in kinetic plus electrostatic energy determines convergence. The displacement of the electrons from their initial to their equilibrium positions determines the line-integrated force or refractive index gradient. We have implemented two approaches: PIC (particle in cell) and Lagrangian fluid, in 1-D and 2-D. The PIC code works for arbitrary intensities, can work efficiently in parallel, and can make use of existing codes. The Lagrangian code requires less memory and is faster than the PIC code without massively parallel processing, but fails in 2-D for large intensity modulations. The Monge–Ampère code is by far the fastest in 2-D, without massively parallel processing, but fails for intensities with large voids, high contrast ratios and large deflections across the boundaries, and could not obtain the degree of convergence possible with the PIC code. As a result, the power diagram code was by far the slowest and most memory intensive, and failed for large peaks in the measured intensity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

Performance Evaluation of Heterogeneous GPU Programming Frameworks for Hemodynamic Simulations

Preparing for the deployment of large scientific and engineering codes on upcoming exascale systems with GPU-dense nodes is made challenging by the unprecedented diversity of device architectures and heterogeneous programming models. In this work, we evaluate the process of porting a massively parallel, fluid dynamics code written in CUDA to SYCL, HIP, and Kokkos with a range of backends, using a combination of automated tools and manual tuning. We use a proxy application along with a custom performance model to inform the results and identify additional optimization strategies. At scale performance of the programming model implementations are evaluated on pre-production GPU node architectures for Frontier and Aurora, as well as on current NVIDIA device-based systems Summit and Polaris. Real-world workloads representing 3D blood flow calculations in complex vasculature are assessed. Our analysis highlights critical trade-offs between code performance, portability, and development time.

Martin, Aristotle↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

COMPOFF: A Compiler Cost model using Machine Learning to predict the Cost of OpenMP Offloading

The HPC industry is inexorably moving towards an era of extremely heterogeneous architectures, with more devices configured on any given HPC platform and potentially more kinds of devices, some of them highly specialized. Writing a separate code suitable for each target system for a given HPC application is not practical. The better solution is to use directive-based parallel programming models such as OpenMP. OpenMP provides a number of options for offloading a piece of code to devices like GPUs. To select the best option from such options during compilation, most modern compilers use analytical models to estimate the cost of executing the original code and the different offloading code variants. Building such an analytical model for compilers is a difficult task that necessitates a lot of effort on the part of a compiler engineer. Recently, machine learning techniques have been successfully applied to build cost models for a variety of compiler optimization problems. In this paper, we present COMPOFF, a cost model which uses the multi-layer perceptrons to statically estimates the Cost of OpenMP OFFloading. We used six different transformations on a parallel code of Wilson Dslash Operator to support GPU offloading, and we predicted their cost of execution on different GPUs using COMPOFF during compile time. Our results show that this model can predict offloading costs with a root mean squared error in prediction of less than 0.5 seconds. Our preliminary findings indicate that this work will make it much easier and faster for scientists and compiler developers to port legacy HPC applications that use OpenMP to new heterogeneous computing environment.

97 MATHEMATICS AND COMPUTING↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

Implementation of a Mesh refinement algorithm into the quasi-static PIC code QuickPIC

Plasma-based acceleration (PBA) has emerged as a promising candidate for the accelerator technology used to build a future linear collider and/or an advanced light source. In PBA, a trailing or witness particle beam is accelerated in the plasma wave wakefield (WF) created by a laser or particle beam driver. The WF is often nonlinear and involves the crossing of plasma particle trajectories in real space and thus particle-in-cell methods are used. The distance over which the drive beam evolves is several orders of magnitude larger than the wake wavelength. This large disparity in length scales is amenable to the quasi-static approach. Three-dimensional (3D), quasi-static (QS), particle-in-cell (PIC) codes, e.g., QuickPIC, have been shown to provide high fidelity simulation capability with 2-4 orders of magnitude speedup over 3D fully explicit PIC codes. In PBA, the witness beam needs to be matched to the focusing forces of the WF to reduce the emittance growth. In some linear collider designs, the matched spot size of the witness beam can be 2 to 3 orders of magnitude smaller than the spot size (and wavelength) of the wakefield. Such an additional disparity in length scales is ideal for mesh refinement where the WF within the witness beam is described on a finer mesh than the rest of the WF. A mesh refinement scheme is described that has been implemented into the 3D QS PIC code, QuickPIC. Very fine (high) resolution is used in a small spatial region that includes the witness beam and progressively coarser resolutions in the rest of the simulation domain. A fast multigrid Poisson solver has been implemented for the field solve on the refined meshes and a Fast Fourier Transform (FFT) based Poisson solver is used for the coarse mesh. The code has been parallelized with both MPI and OpenMP, and the parallel scalability has also been improved by using pipelining. A preliminary adaptive mesh refinement technique is described to optimize the computational time for simulations with an evolving witness beam size. Several test problems are used to verify that the mesh refinement algorithm provides accurate results. Additionally, the results are benchmarked against highly resolved simulations exhibiting near-azimuthal symmetry, performed using QPAD—a novel hybrid QS PIC code that uses a PIC description in the coordinates (r, ct – z) and a gridless description in the azimuthal angle, Φ.

Linear collider↗

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Towards improved speed and accuracy of laser powder bed fusion simulations via multiscale spatial representations

Due to the growing popularity of laser powder bed fusion (LPBF) as a metal additive manufacturing technique, there is a strong need to be able to accurately predict build outcomes. Full fidelity simulations of this process are not feasible due to the vast range of length and time scales inherent to it. While part-scale codes for simulating residual stress and distortion have shown reasonable predictive capability, they often neglect many aspects of the process occurring over smaller length/time scales, and thus are unable to capture effects of process parameter adjustments or the behavior of fine features. One way of capturing aspects at more refined length scales is through the use of adaptive mesh refinement (AMR). AMR allows for the process to be simulated at scales approaching the physical spatial dimensions without drastically increasing the total degrees of freedom in the simulation. This manuscript describes the implementation of an AMR algorithm within a multiphysics, parallelized finite element code, and its application to the LPBF problem. In this work, part-scale examples are provided where the use of AMR has allowed for higher fidelity thermal and thermomechanical simulations, as compared to experimental measurements. Results from these higher resolution simulations show that while AMR is a necessary component for increased accuracy in a computationally efficient manner, other improvements are also necessary, including handling of the multiple time scales inherent to the problem and the need for improved AM-specific material models.

42 ENGINEERING↗

Multiphysics analysis system for heat pipe cooled micro-reactors employing PRAGMA-OpenFOAM-ANLHTP

A multiphysics analysis system for neutronics/thermo-mechanical/heat-pipe analysis of heat pipe cooled micro-reactors was developed using the PRAGMA code as the neutronics engine. PRAGMA, which used to be a GPU-based continuous-energy MC code for power reactor applications, now has an extended geometry package to handle geometries with unstructured meshes generated by the ANSYS Design-Modeler and Meshing. The NVIDIA ray tracing engine OptiX was exploited for efficient neutron transport on unstructured geometry. On the multiphysics side, the open-source CFD tool OpenFOAM and one-dimensional heat pipe analysis code ANLHTP were adopted. The manager-worker system based on the MPI dynamic process management (DPM) model enables efficient coupling of codes employing different parallelization schemes. With all the features, the multiphysics analysis of the one-sixth symmetrical MegaPower 2D core was performed and it demonstrated the benefit of the tight integration of the three-way coupling system and one-to-one geometry coupling strategy. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

Self-Consistent Relativistic Electron Scattering using the Sherlock Scattering Model for X-ray Diagnostics

We present on a new, self-consistent, arbitrary-temperature Romberg integration scheme for modeling electron scattering in materials in a LANL Lagrangian Shock Hydro (LSH) code. Electron beam-target interactions are fundamental to a wide range of scientific and technological applications. When high-energy electron beams hit their target, they may scatter, deposit energy, or ionize the source. These processes govern the behavior and outcomes in nanotechnology manufacturing, electron microscopy, and modern X-ray diagnostics. Simulating these interactions is essential for interpreting experimental results, predicting material responses, and designing efficient tools and experiments. At Los Alamos, this is done using a LSH code, which is a multi-dimension, multi-material, massively parallel, multi-physics code used to simulate applications from asteroid impacts to electron beam interactions. By effectively and efficiently modeling the way that electrons scatter from the beam we can bolster these simulations and more accurately predict experimental outcomes. The model currently implemented in the LSH of interest is based on work by Papp and does not self-consistently preserve momentum in the slightly relativistic regime; here we adopt a model proposed by Braams and Karney and implement a Romberg integration scheme to compute the diffusion tensor. In this paper we will provide background on the Braams-Karney diffusion tensor as well as the Romberg integration scheme we employed to numerically solve for it. We will show that our integration scheme is accurate in solving for the set of scalar potentials used to re-express the diffusion tensor in differential form, and in solving for the diffusion coefficients in the larger LSH code. By using this diffusion tensor rather than the existing Papp one, and numerically integrating it with a Romberg method, we produce much more accurate, self-consistent results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modelling of edge plasma dynamics with active wall boundary conditions

A self-consistent 2D model is presented for transport in boundary plasma and plasma-facing material walls. Plasma dynamics in the domain is represented by a 2D collisional plasma fluid model in the edge-plasma code UEDGE, and transport of hydrogen and heat in the wall is represented by a system of reaction–diffusion equations in the 1D wall code FACE. To account for variation of parameters along the wall, in the coupled model multiple instances of the FACE code run in parallel. Here, the coupled model provides a tool for investigating a range of dynamic plasma–material interactions phenomena in 2D. For demonstration of its capability, one application of particular interest is the role of active wall in tokamak strike point sweeping proposed for mitigation of divertor heat loads. In the present study, the coupled calculations are applied to investigation of the impact of heat and hydrogen transport in the material wall on the divertor plasma and target heat load during sweeping of the target strike point for parameters of a high-power tokamak.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dataset of simulated vibrational density of states and X-ray diffraction profiles of mechanically deformed and disordered atomic structures in Gold, Iron, Magnesium, and Silicon

This dataset is comprised of a library of atomistic structure files and corresponding X-ray diffraction (XRD) profiles and vibrational density of states (VDoS) profiles for bulk single crystal silicon (Si), gold (Au), magnesium (Mg), and iron (Fe) with and without disorder introduced into the atomic structure and with and without mechanical loading. Included with the atomistic structure files are descriptor files that measure the stress state, phase fractions, and dislocation content of the microstructures. All data was generated via molecular dynamics or molecular statics simulations using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) code. This dataset can inform the understanding of how local or global changes to a materials microstructure can alter their spectroscopic and diffraction behavior across a variety of initial structure types (cubic diamond, face-centered cubic (FCC), hexagonal close-packed (HCP), and body-centered cubic (BCC) for Si, Au, Mg, and Fe, respectively) and overlapping changes to the microstructure (i.e., both disorder insertion and mechanical loading).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

TurboRVB: A many-body toolkit for ab initio electronic simulations by quantum Monte Carlo

TurboRVB is a computational package for ab initio Quantum Monte Carlo (QMC) simulations of both molecular and bulk electronic systems. The code implements two types of well established QMC algorithms: Variational Monte Carlo (VMC) and diffusion Monte Carlo in its robust and efficient lattice regularized variant. A key feature of the code is the possibility of using strongly correlated many-body wave functions (WFs), capable of describing several materials with very high accuracy, even when standard mean-field approaches [e.g., density functional theory (DFT)] fail. The electronic WF is obtained by applying a Jastrow factor, which takes into account dynamical correlations, to the most general mean-field ground state, written either as an antisymmetrized geminal power with spin-singlet pairing or as a Pfaffian, including both singlet and triplet correlations. This WF can be viewed as an efficient implementation of the so-called resonating valence bond (RVB) Ansatz, first proposed by Pauling and Anderson in quantum chemistry [L. Pauling, The Nature of the Chemical Bond (Cornell University Press, 1960)] and condensed matter physics [P.W. Anderson, Mat. Res. Bull 8, 153 (1973)], respectively. The RVB Ansatz implemented in TurboRVB has a large variational freedom, including the Jastrow correlated Slater determinant as its simplest, but nontrivial case. Moreover, it has the remarkable advantage of remaining with an affordable computational cost, proportional to the one spent for the evaluation of a single Slater determinant. Therefore, its application to large systems is computationally feasible. The WF is expanded in a localized basis set. Several basis set functions are implemented, such as Gaussian, Slater, and mixed types, with no restriction on the choice of their contraction. The code implements the adjoint algorithmic differentiation that enables a very efficient evaluation of energy derivatives, comprising the ionic forces. Thus, one can perform structural optimizations and molecular dynamics in the canonical NVT ensemble at the VMC level. For the electronic part, a full WF optimization (Jastrow and antisymmetric parts together) is made possible, thanks to state-of-the-art stochastic algorithms for energy minimization. In the optimization procedure, the first guess can be obtained at the mean-field level by a built-in DFT driver. The code was efficiently parallelized by using a hybrid MPI-OpenMP protocol, which is also an ideal environment for exploiting the computational power of modern Graphics Processing Unit accelerators.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

kokkosSZ (kSZ) A Portable Cellerator Implementation of SZ using Kokkos Programming Model

kSZ is a Kokkos-based implementation of the world-widely used SZ lossy compressor. We use Kokkos because it provides abstractions for both parallel execution of code and data management, which can be used to support portable implementation across different accelerator technologies. Kokkos can support OpenMP/OpenMPTarget, oneAPI, Pthreads, and CUDA as backend programming models.

ECP↗

ML-PSA

The computer code uses a parallel simulated annealing framework with embedded machine learning components to solve multi-constrained optimization problems. The software automatically balances the execution of low and high fidelity physics models within the optimization procedure. The low fidelity model is used to rapidly explore the design space while the high fidelity physics model is executed sparingly to account for complex design constraints that are not resolved by the quickly executing low fidelity model.

Gurecky, William↗