Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Mahakala: A Python-based Modular Ray-tracing and Radiative Transfer Algorithm for Curved Spacetimes

We introduce Mahakala, a Python-based, modular, radiative ray-tracing code for curved spacetimes. We employ Google's JAX framework for accelerated automatic differentiation, which can efficiently compute Christoffel symbols directly from the metric, allowing the user to easily and quickly simulate photon trajectories through non-Kerr spacetimes. JAX also enables Mahakala to run in parallel on both CPUs and GPUs. Mahakala natively uses the Cartesian Kerr–Schild coordinate system, which avoids numerical issues caused by the pole in spherical coordinate systems. We demonstrate Mahakala's capabilities by simulating 1.3 mm wavelength images (the wavelength of Event Horizon Telescope observations) of general relativistic magnetohydrodynamic simulations of low-accretion rate supermassive black holes. The modular nature of Mahakala allows us to quantitatively explore how different regions of the flow influence different image features. We show that most of the emission seen in 1.3 mm images originates close to the black hole and peaks near the photon orbit. We also quantify the relative contribution of the disk, forward jet, and counterjet to 1.3 mm images.

79 ASTRONOMY AND ASTROPHYSICS↗

Parallelized domain decomposition for multi-dimensional Lagrangian random walk mass-transfer particle tracking schemes

Lagrangian particle tracking schemes allow a wide range of flow and transport processes to be simulated accurately, but a major challenge is numerically implementing the inter-particle interactions in an efficient manner. This article develops a multi-dimensional, parallelized domain decomposition (DDC) strategy for mass-transfer particle tracking (MTPT) methods in which particles exchange mass dynamically. We show that this can be efficiently parallelized by employing large numbers of CPU cores to accelerate run times. In order to validate the approach and our theoretical predictions we focus our efforts on a well-known benchmark problem with pure diffusion, where analytical solutions in any number of dimensions are well established. In this work, we investigate different procedures for “tiling” the domain in two and three dimensions (2-D and 3-D), as this type of formal DDC construction is currently limited to 1-D. An optimal tiling is prescribed based on physical problem parameters and the number of available CPU cores, as each tiling provides distinct results in both accuracy and run time. We further extend the most efficient technique to 3-D for comparison, leading to an analytical discussion of the effect of dimensionality on strategies for implementing DDC schemes. Increasing computational resources (cores) within the DDC method produces a trade-off between inter-node communication and on-node work. For an optimally subdivided diffusion problem, the 2-D parallelized algorithm achieves nearly perfect linear speedup in comparison with the serial run-up to around 2700 cores, reducing a 5 h simulation to 8 s, while the 3-D algorithm maintains appreciable speedup up to 1700 cores.

97 MATHEMATICS AND COMPUTING↗

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195↗

MPACT Verification With Magnox Reactor Neutronics Progression Problems

MPACT is a state-of-the-art core simulator designed to perform high-fidelity analysis using whole-core, three-dimensional, pin-resolved neutron transport calculations on modern parallel computing hardware. MPACT was originally developed to model light water reactors, and its capabilities are being extended to simulate gas-cooled, graphite-moderated cores such as Magnox reactors. To verify MPACT’s performance in this new application, the code is being formally benchmarked using representative problems. Progression problems are a series of example models that increase in complexity designed to test a code’s performance. The progression problems include both beginning-of-cycle and depletion calculations. Reference solutions for each progression problem have been generated using Serpent 2, a continuous-energy Monte Carlo reactor physics burnup calculation code.Using the neutron multiplication eigenvalue ke as a metric, MPACT’s performance is assessed on each of the progression problems. Initial results showed that MPACT’s multi-group cross section libraries, originally developed for pressurized water reactor problems, were not sufficient to accurately solve Magnox problems. MPACT’s improved performance on the progression problems is demonstrated using this new optimized cross section library.

Luciano, Nicholas↗

Development of PFLOTRAN Transport Capability for Use in the Waste Isolation Pilot Plant Performance Assessment - 20545

Waste Isolation Pilot Plant (WIPP) performance assessment (PA) calculations estimate the probability of radionuclide release from the repository to the land surface and across the land withdrawal boundary for a regulatory period of 10,000 years after facility closure. Simulations of flow and transport in the repository and the surrounding Salado Formation are foundational to the PA. Because proposed additional waste emplacement panels would result in an asymmetric repository layout, the US Department of Energy (DOE) is preparing to transition to use of a three-dimensional (3-D) model domain for simulation of flow and transport instead of the two-dimensional (2-D) flared grid domain currently used. DOE has charged Sandia National Laboratories with developing the capability necessary to simulate processes affecting flow and transport in the WIPP in PFLOTRAN, an open-source massively parallel multi-phase flow and reactive transport code. The new flow and transport capabilities developed in PFLOTRAN incorporate WIPP-specific process models and will replace the 2-D simulators (BRAGFLO and NUTS) that are currently utilized for Salado flow and transport calculations in WIPP PA. The focus of this paper is on the development of a new Nuclear Waste Transport (NWT) mode in PFLOTRAN that has all of the capabilities necessary for Salado transport simulations, including the ability to handle complete dry-out (100% gas saturation) of arbitrary cells in the model domain, radionuclide mass conservation at step changes in porosity associated with borehole intrusion, and the ability to calculate fluxes on a flared grid. The new PFLOTRAN transport capability and a suite of verification tests were designed around a list of functional requirements for WIPP PA calculations. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Simulating Magnetic Reconnection in ‘Two Ribbon’ Type Solar Flares

Magnetic reconnection is an astrophysical process where neighboring magnetic field lines, facing anti-parallel, are reconfigured. This reconfiguration results in built up magnetic energy being explosively released as it is being converted to plasma kinetic and thermal energies. There are various kinds of simulations used to simulation reconnection; our work begins with Athena++, a magnetohydrodynamic (MHD) simulation code typically used for astrophysical problems, and a reconnection specific code file. The resistive MHD equations are solved with Riemann solvers. There was an initial test run without modifying the code to understand the dynamics of the simulation. We expand on the original reconnection problem file by implementing a radiative cooling term specific to the corona. The radiative cooling is theorized to have an effect on solar coronal plasma and magnetic reconnection dynamics. The cooling term will be tested with various parameters and compared to the case without cooling to study these dynamics. The condensation found in only the with cooling case emphasizes the importance of implementing this feature and will be later tested with a thermal conduction term. We want to determine the parameter regime where non-equilibrium cooling will be important for the reconnection dynamics.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalability of OpenFOAM Density-Based Solver with Runge–Kutta Temporal Discretization Scheme

Compressible density-based solvers are widely used in OpenFOAM, and the parallel scalability of these solvers is crucial for large-scale simulations. In this paper, we report our experiences with the scalability of OpenFOAM’s native rhoCentralFoam solver, and by making a small number of modifications to it, we show the degree to which the scalability of the solver can be improved. The main modification made is to replace the first-order accurate Euler scheme in rhoCentralFoam with a third-order accurate, four-stage Runge-Kutta or RK4 scheme for the time integration. The scaling test we used is the transonic flow over the ONERA M6 wing. This is a common validation test for compressible flows solvers in aerospace and other engineering applications. Numerical experiments show that our modified solver, referred to as rhoCentralRK4Foam, for the same spatial discretization, achieves as much as a 123.2% improvement in scalability over the rhoCentralFoam solver. As expected, the better time resolution of the Runge–Kutta scheme makes it more suitable for unsteady problems such as the Taylor–Green vortex decay where the new solver showed a 50% decrease in the overall time-to-solution compared to rhoCentralFoam to get to the final solution with the same numerical accuracy. Finally, the improved scalability can be traced to the improvement of the computation to communication ratio obtained by substituting the RK4 scheme in place of the Euler scheme. All numerical tests were conducted on a Cray XC40 parallel system, Theta, at Argonne National Laboratory.

Li, Sibo↗

Simulating energetic ions and enhanced fusion rates from ion-cyclotron resonance heating with a full-wave/Fokker–Planck model

Reproducing fast-ion enhanced fusion rates from ion-cyclotron resonance heating (ICRH) in tokamaks requires the self-consistent coupling of a full-wave solver and a Fokker–Planck solver, which evolves multiple simultaneously resonant ion species. We introduce a new self-consistent model that iterates the TORIC full-wave solver with the CQL3D Fokker–Planck solver using the integrated plasma simulator (IPS). This model evolves the bounce-averaged ion distribution functions in both parallel and perpendicular velocity-space with a quasilinear radio frequency (RF) diffusion operator valid in the ion finite Larmor radius (FLR) limit and the RF electric fields with the resultant non-Maxwellian FLR dielectric tensor. This produces non-Maxwellian ICRH simulations that are fully self-consistent, fast, and interoperable with integrated modeling frameworks, such as TRANSP/GACODE/IPS-FASTRAN. We demonstrate our model's capabilities by validating it against experimental data in Alcator C-Mod. We then perform the first RF heating simulations of SPARC using self-consistent non-Maxwellian ion distributions to investigate the potential to enhance fusion rates using ion cyclotron resonance heating generated fast ions.

Physics↗

Enabling Large-Scale Condensed-Phase Hybrid Density Functional Theory Based Ab Initio Molecular Dynamics. 1. Theory, Algorithm, and Performance

By including a fraction of exact exchange (EXX), hybrid functionals reduce the self-interaction error in semilocal density functional theory (DFT) and thereby furnish a more accurate and reliable description of the underlying electronic structure in systems throughout biology, chemistry, physics, and materials science. However, the high computational cost associated with the evaluation of all required EXX quantities has limited the applicability of hybrid DFT in the treatment of large molecules and complex condensed-phase materials. To overcome this limitation, we describe a linear-scaling approach that utilizes a local representation of the occupied orbitals (e.g., maximally localized Wannier functions (MLWFs)) to exploit the sparsity in the real-space evaluation of the quantum mechanical exchange interaction in finite-gap systems. In this work, we present a detailed description of the theoretical and algorithmic advances required to perform MLWF-based ab initio molecular dynamics (AIMD) simulations of large-scale condensed-phase systems of interest at the hybrid DFT level. We focus our theoretical discussion on the integration of this approach into the framework of Car–Parrinello AIMD, and highlight the central role played by the MLWF-product potential (i.e., the solution of Poisson’s equation for each corresponding MLWF-product density) in the evaluation of the EXX energy and wave function forces. We then provide a comprehensive description of the exx algorithm implemented in the open-source Quantum ESPRESSO program, which employs a hybrid MPI/OpenMP parallelization scheme to efficiently utilize the high-performance computing (HPC) resources available on current- and next-generation supercomputer architectures. Furthermore, this is followed by a critical assessment of the accuracy and parallel performance (e.g., strong and weak scaling) of this approach when AIMD simulations of liquid water are performed in the canonical (NVT) ensemble. With access to HPC resources, we demonstrate that exx enables hybrid DFT-based AIMD simulations of condensed-phase systems containing 500–1000 atoms (e.g., (H₂O)₂₅₆) with a wall time cost that is comparable to that of semilocal DFT. In doing so, exx takes us one step closer to routinely performing AIMD simulations of complex and large-scale condensed-phase systems for sufficiently long time scales at the hybrid DFT level of theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Containers for Massive Ensemble of I/O Bound Hierarchical Coupled Simulations

We present our experience using containers to scale up a massive ensemble of coupled I/O bound workloads on the NERSC Cori supercomputer. We describe the design of a hierarchical simulation structure using the Integrated Plasma Simulator (IPS) that enables the flexible execution of coupled simulations at the system, node, and core level using the same coupling abstraction and API. The hierarchical design allows for the node-level execution to be efficiently executed using containers while not impacting the structure of the simulation at the system level. We demonstrate the viability of the approach by presenting experimental results from applications in coupled fusion plasma simulations that illustrate the performance impact of using containers to deploy the node-level workloads, in conjunction with the user mountable XFS file systems to ameliorate the load on the Lustre parallel file system. We also present results from production runs showing the ability of the ensemble simulations to scale to hundreds of Cori Haswell nodes, with little or no overhead.

Elwasif, Wael↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

Attribute-Aware RBFs: Interactive Visualization of Time Series Particle Volumes Using RT Core Range Queries

Smoothed-particle hydrodynamics (SPH) is a mesh-free method used to simulate volumetric media in fluids, astrophysics, and solid mechanics. Visualizing these simulations is problematic because these datasets often contain millions, if not billions of particles carrying physical attributes and moving over time. Radial basis functions (RBFs) are used to model particles, and overlapping particles are interpolated to reconstruct a high-quality volumetric field; however, this interpolation process is expensive and makes interactive visualization difficult. Existing RBF interpolation schemes do not account for color-mapped attributes and are instead constrained to visualizing just the density field. To address these challenges, we exploit ray tracing cores in modern GPU architectures to accelerate scalar field reconstruction. We use a novel RBF interpolation scheme to integrate per-particle colors and densities, and leverage GPU-parallel tree construction and refitting to quickly update the tree as the simulation animates over time or when the user manipulates particle radii. We also propose a Hilbert reordering scheme to cluster particles together at the leaves of the tree to reduce tree memory consumption. Finally, we reduce the noise of volumetric shadows by adopting a spatially temporal blue noise sampling scheme. Our method can provide a more detailed and interactive view of these large, volumetric, time-series particle datasets than traditional methods, leading to new insights into these physics simulations.

Particle Volumes↗

A weighted state redistribution algorithm for embedded boundary grids

State redistribution is an algorithm that stabilizes cut cells for embedded boundary grid methods. This work extends the earlier algorithm in several important ways. First, state redistribution is extended to three spatial dimensions. Second, we discuss several algorithmic changes and improvements motivated by the more complicated cut cell geometries that can occur in higher dimensions. In particular, we introduce a weighted version with less dissipation in an easily generalizable framework. Third, we demonstrate that state redistribution can also stabilize a solution update that includes both advective and diffusive contributions. Notably, the stabilization algorithm is shown to be effective for incompressible as well as compressible reacting flows. Finally, we discuss the implementation of the algorithm for several exascale-ready simulation codes based on AMReX, demonstrating ease of use in combination with domain decomposition, hybrid parallelism and complex physics.

97 MATHEMATICS AND COMPUTING↗

A resistive MHD model and simulation on plasma flow evolution in the presence of resonant magnetic perturbation in a tokamak

Nonaxisymmetric magnetic fields such as the intrinsic error field and the externally applied resonant magnetic perturbation (RMP) in a tokamak are known to influence the plasma momentum transport and flow evolution through plasma response, which itself strongly depends on the plasma flow as well. The nonlinear interaction between plasma response and flow has been previously modeled in the conventional error field theory with the “no-slip” condition, which has been recently extended to allow the “free-slip” condition. In this work, we further target this specific process and numerically simulate the nonlinear plasma response and flow evolution in the presence of a single-helicity RMP in a circular-shaped model tokamak configuration, based on the full resistive MHD model in the initial-value code NIMROD. Time evolution of the parallel (to k) flow or “slip frequency” profile and its asymptotic steady state obtained from the NIMROD simulations are compared with both conventional and extended nonlinear response theories. Here, k is the wave vector of the propagating island. Good agreement with the extended theory with free-slip condition has been achieved for the parallel flow profile evolution in response to RMP in all resistive regimes, whereas the difference from the conventional theory with the no-slip condition tends to diminish as the plasma resistivity approaches zero.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Massively Parallel Bayesian Model Calibration and Uncertainty Quantification with Applications to Nuclear Fuels and Materials

The U.S. Department of Energy (DOE)’s Nuclear Energy Advanced Modeling and Simulation (NEAMS) program aims to develop predictive capabilities by applying computational methods to the analysis and design of advanced reactor and fuel cycle systems. This program has been providing engineering-scale support for the development of BISON, a high-fidelity and high-resolution fuel performance tool. Fuel behavior in a nuclear reactor is governed by a complex network of mechanisms interacting with various other physics aspects in the reactor system. Any model developed to represent the fuel behavior will likely be idealized resulting in uncertainties in their predictions compared to the observed data. As such, this report was motivated by the need to identify the sources of uncertainties and quantify and propagate them through the fuel model outputs. Such quantification of uncertainties will establish a level of model trustworthiness, identify approaches to improve the model trustworthiness, and even guide optimal experiment design for maximal information gain. To accomplish the uncertainty quantification for computational models, this report has relied on the Bayesian framework which provides probabilistic treatment of models their inputs and outputs. The current state-of-the-art on performing Bayesian Uncertainty Quantification (UQ) for nuclear engineering models using High Performance Computing (HPC) resources have been reviewed. Implementation of capabilities for massively parallel Bayesian UQ in Multiphysics Object-Oriented Simulation Environment (MOOSE) is discussed. Several verification cases are discussed to verify the accuracy of the quantified uncertainties using the developed computational capabilities in MOOSE. Then, the problem of quantifying the uncertainties in TRI-Structural isOtropic (TRISO) fuel silver release is addressed. For the first time, the uncertainties arising from the TRISO Fission Gas Release (FGR) model due to model inadequacy and experimental noise are quantified. Also, the Bayesian capabilities are applied to the calibration of the MATPRO creep model, a widely used model in several fuel assessment cases. The impact of the prediction uncertainties in the MATPRO model on the fuel cladding behavior as part of the TRIBULATION assessment case (which is an integral effects case) is investigated. This report concludes with a discussion on the future work for the UQ for computational models.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Fractional Skyrmion Tubes in Chiral‐Interfaced 3D Magnetic Nanowires

Magnetic skyrmions are chiral spin textures with rich physics and great potential for unconventional computing. Typically, skyrmions form in bulk crystals with reduced symmetry or ultrathin film multilayers involving heavy metals. Here, the formation of fractional Bloch skyrmion tubes at room temperature is demonstrated by 3D printing ferromagnetic double‐helix nanowires with two regions of opposite chirality. Using X‐ray microscopy and micromagnetic simulations, it is shown that the coexistence of vortex and anti‐parallel spin states induces the formation of fractional skyrmion tubes at zero magnetic fields, minimizing the energy cost of breaking the coupling between geometric and magnetic chirality. Control over zero‐field states is also demonstrated, including pure vortex, or mixed skyrmion‐vortex states, highlighting the magnetic reconfigurability of these 3D nanowires. This work shows how interfacing chiral geometries at the nanoscale can enable advanced forms of topological spintronics.

X-ray microscopy↗

Designed Spin‐Texture‐Lattice to Control Anisotropic Magnon Transport in Antiferromagnets

Abstract Spin waves in magnetic materials are promising information carriers for future computing technologies due to their ultra‐low energy dissipation and long coherence length. Antiferromagnets are strong candidate materials due, in part, to their stability to external fields and larger group velocities. Multiferroic antiferromagnets, such as BiFeO 3 (BFO), have an additional degree of freedom stemming from magnetoelectric coupling, allowing for control of the magnetic structure, and thus spin waves, with the electric field. Unfortunately, spin‐wave propagation in BFO is not well understood due to the complexity of the magnetic structure. In this work, long‐range spin transport is explored within an epitaxially engineered, electrically tunable, 1D magnonic crystal. A striking anisotropy is discovered in the spin transport parallel and perpendicular to the 1D crystal axis. Multiscale theory and simulation suggest that this preferential magnon conduction emerges from a combination of a population imbalance in its dispersion, as well as anisotropic structural scattering. This work provides a pathway to electrically reconfigurable magnonic crystals in antiferromagnets.

36 MATERIALS SCIENCE↗

Developing Ultrahigh-Resolution E3SM Land Model for GPU Systems

Designing and refactoring complex scientific code, such as the E3SM land model (ELM), for new computing architectures is challenging. This paper presents design strategies and technical approaches to develop a data-oriented, GPU-ready ELM model using compiler directives (OpenACC/OpenMP). We first analyze the datatypes and processes in the original ELM code. Then we present design considerations for ultrahigh-resolution ELM (uELM) development for massive GPU systems. These techniques include the global data-oriented simulation workflow, domain partition, code porting and data copy, memory reduction, parallel loop restructure and flattening, and race condition detection. We implemented the first version of uELM using OpenACC targeting the NVidia GPUs in the Summit supercomputer at Oak Ridge National Laboratory. During the implementation, we developed a software tool (named SPEL) to facilitate code generation, verification, and performance tuning using these techniques. The first uELM implementation for Nvidia GPUs on Summit delivered promising results: 1) over 98% of the ELM code was automatically generated and tuned by scripts. Most ELM modules had better computational performances than the original ELM code for CPUs. The GPU-ready uELM is more scalable than the CPU code on fully-loaded Summit nodes. Example profiling results from several modules are also presented to illustrate the performance improvements and race condition detection. The lessons learned and toolkit developed in the study are also suitable for further uELM deployment using OpenMP on the first US exascale computer, Frontier, equipped with AMD CPUs and GPUs.

Schwartz, Peter↗