Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ParMOO: A Python library for parallel multiobjective simulation optimization

A multiobjective optimization problem (MOOP) is an optimization problem in which multiple objectives are optimized simultaneously. The goal of a MOOP is to find solutions that describe the tradeoff between these (potentially conflicting) objectives. Such a tradeoff surface is called the Pareto front. Real-world MOOPs may also involve constraints – additional hard rules that every solution must adhere to. In a multiobjective simulation optimization problem, the objectives are derived from the outputs of one or more computationally expensive simulations. Such problems are ubiquitous in science and engineering.

97 MATHEMATICS AND COMPUTING↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Virtual Time III, Part 1: Unified Virtual Time Synchronization for Parallel Discrete Event Simulation

Algorithms for synchronization of parallel discrete event simulation have historically been divided between conservative methods that require lookahead but not rollback, and optimistic methods that require rollback but not lookahead. In this paper we present a new approach in the form of a framework called Unified Virtual Time (UVT) that unifies the two approaches, combining the advantages of both within a single synchronization theory. Whenever timely lookahead information is available, a logical process (LP) executes conservatively using an irreversible event handler. When lookahead information is not available the LP does not block, as it would in a classical conservative execution, but instead executes optimistically using a reversible event handler. The switch from conservative to optimistic synchronization and back is decided on an event-by-event basis by the simulator, transparently to the model code. UVT treats conservative synchronization algorithms as optional accelerators for an underlying optimistic synchronization algorithm, enabling the speed of conservative execution whenever it is applicable, but otherwise falling back on the generality of optimistic execution. We describe UVT in a novel way, based on fundamental invariants, monotonicity requirements, and synchronization rules. UVT permits zero-delay messages and pays careful attention to tie-handling using superposition. We prove that under fairly general conditions a UVT simulation always makes progress in virtual time. This is Part 1 of a trio of papers describing the UVT framework for PDES, mixing conservative and optimistic synchronization and integrating throttling control.

97 MATHEMATICS AND COMPUTING↗

Circumbinary Disk Accretion into Spinning Black Hole Binaries

Supermassive black hole binaries are likely to accrete interstellar gas through a circumbinary disk. Shortly before merger, the inner portions of this circumbinary disk are subject to general relativistic effects. To study this regime, we approximate the spacetime metric of close orbiting black holes by superimposing two boosted Kerr–Schild terms. After demonstrating the quality of this approximation, we carry out very long-term general relativistic magnetohydrodynamic simulations of the circumbinary disk. We consider black holes with spin dimensionless parameters of magnitude 0.9, in one simulation parallel to the orbital angular momentum of the binary, but in another anti-parallel. These are contrasted with spinless simulations. We find that, for a fixed surface mass density in the inner circumbinary disk, aligned spins of this magnitude approximately reduce the mass accretion rate by 14% and counter-aligned spins increase it by 45%, leaving many other disk properties unchanged.

79 ASTRONOMY AND ASTROPHYSICS↗

ORNL_AISD_NiPt

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for three different sizes of the crystal structure: 256 atoms, 864 atoms, and 2,048 atoms, each of which contains 1900 configurations. For each size of the crystal structure, the data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in three compressed zipped folders: atoms256.zip, atoms864.zip, atoms2048.zip Each zipped folder contains the data that describes crystals of size 256 atoms, 864 atoms, and 2,048 atoms respectively. Each one of the three zipped folders contains the data structured in the following way: -Ni_ground_state.cfg --> atomic configuration for the pure nickel -Pt_ground_state.cfg --> atomic configuration for the pure platinum -Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder. Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms256.zip is 65,046. The total number of atomic configurations contained in atoms864.zip is 63,936. The total number of atomic configurations contained in atoms2048.zip is 61,997. The total number of atomic configurations spanned by the entire dataset is 190,979. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büccchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D

36 MATERIALS SCIENCE↗

ORNL_AISD_NiPt_108atoms

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for crystal structures with 108 atoms with 1,900 configurations. The data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system 'MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001' [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in a compressed zipped folders atoms108.zip. The zipped folder contains the data structured in the following way: - Ni_ground_state.cfg --> atomic configuration for the pure nickel - Pt_ground_state.cfg --> atomic configuration for the pure platinum - Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder - Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms108.zip is 66,132. This dataset is an extension to the dataset ORNL_AISD_NiPt [5] that has been previously released with crystal structures of 256 atoms, 864 atoms, and 2,048 atoms, with the same methodology for data collection. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D [5] M. Karabin, M. Lupo Pasini, and M. Eisenbach. ORNL_AISD_NiPt. United States: N. p., 2023. Web. doi:10.13139/OLCF/1958172.

36 MATERIALS SCIENCE↗

Scalable Deep Learning-Based Microarchitecture Simulation on GPUs

Cycle-accurate microarchitecture simulators are essential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the potential of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.

97 MATHEMATICS AND COMPUTING↗

Dataset of simulated vibrational density of states and X-ray diffraction profiles of mechanically deformed and disordered atomic structures in Gold, Iron, Magnesium, and Silicon

This dataset is comprised of a library of atomistic structure files and corresponding X-ray diffraction (XRD) profiles and vibrational density of states (VDoS) profiles for bulk single crystal silicon (Si), gold (Au), magnesium (Mg), and iron (Fe) with and without disorder introduced into the atomic structure and with and without mechanical loading. Included with the atomistic structure files are descriptor files that measure the stress state, phase fractions, and dislocation content of the microstructures. All data was generated via molecular dynamics or molecular statics simulations using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) code. This dataset can inform the understanding of how local or global changes to a materials microstructure can alter their spectroscopic and diffraction behavior across a variety of initial structure types (cubic diamond, face-centered cubic (FCC), hexagonal close-packed (HCP), and body-centered cubic (BCC) for Si, Au, Mg, and Fe, respectively) and overlapping changes to the microstructure (i.e., both disorder insertion and mechanical loading).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Cloud Services Enable Efficient AI-Guided Simulation Workflows across Heterogeneous Resources

Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.

Ward, Logan↗

Extending TOUGH + HYDRATE with a parallel particle transport simulator: numerical investigation of sand production during gas production from hydrate deposits

A new parallel code for simulating particle transport in porous media is integrated with the TOUGH + HYDRATE simulator to investigate sand production associated with gas production from unconsolidated gas hydrate-bearing sediments (HBS). Here, the parallel coupled simulator is named THMPT and uses the integral finite difference method to describe the Darcian and non-Darcian flow of fluids and heat transport, the finite element method to describe the associated geomechanical changes, and the discrete element method to track the trajectory of individual sand particles within the HBS. The THMPT simulator is written in Fortran, incorporates multiple optimized algorithms, and can comprehensively address the coupled flow, thermal, chemical, geomechanical, and particle transport processes that characterize the system behaviors during gas production from HBS. The simulator can capture all processes involved in sand particle transport in porous media, including sand detachment, collision, clogging (i.e., bridging), and migration. A benchmark case study of sand production in the course of depressurization-induced gas production from a representative HBS reveals various distinct microscopic particle migration mechanisms and the adverse impact of sand particle detachment, transport, and clogging. The numerical investigation also examines the effect of bottomhole pressure on mitigating sand production. The simulation results indicate that sand clogging near the wellbore significantly reduces permeability, decreasing gas production by at least 50%. Lastly, the efficiency of gravel packing in mitigating sand production is numerically evaluated, revealing that the structure of the porous media appears to profoundly influence the macroscopic motion behavior of sand particles and sand clogging characteristics.

discrete element method↗

Grain structure and texture selection regimes in metal powder bed fusion

Additive manufacturing (AM) offers opportunities to produce complex part geometries not possible with conventional processing and in some cases even improve part performance. However, adoption has been slowed by difficulties assessing microstructure variability and there is no straightforward approach to relate processing to grain structure characteristics. In this study, datasets from AdditiveFOAM heat transport simulations of laser powder bed fusion (LPBF) are used to drive ExaCA simulations of grain structure. The GPU utilization of ExaCA and an algorithmic update for modeling melt pool overlap region solidification enabled rapid and parallel simulation across a wider range of process conditions than previously explored with cellular automata-based solidification models. A texture selection angle $θ_s$ is defined based on melt pool overlap geometry, and the range of $θ_s$ over which a commonly observed texture transition occurs in characterized AM builds was well-reproduced by ExaCA simulations over a wide range of melt pool shape, hatch spacing, and layer height. ExaCA simulations with 90 degree rotation of the scan direction on every other layer reproduced a number of trends from the AM literature including grain refinement, the dominance of layers with larger melt pools on the final grain structure, and the weakening or strengthening of texture depending on odd and even layer melt pool overlap geometry. EBSD data from a benchmark AM part is used to validate the simulated mechanism of a layer rotation-induced texture strengthening effect. Importantly, these results expand the understanding of the mechanisms for texture selection in alloys with cubic crystal symmetry and offer an approach to easily evaluate processing conditions. With this new understanding, these modeling tools will enable anticipation of previously unexpected variations in grain structure and target specific microstructures and properties.

36 MATERIALS SCIENCE↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers, part II: Regional simulations of building response

The existing observational database of the regional-scale distribution of strong ground motions and measured building response for major earthquakes continues to be quite sparse. As a result, details of the regional variability and spatial distribution of ground motions, and the corresponding distribution of risk to buildings and other infrastructure, are not comprehensively understood. Utilizing high-performance computing platforms, emerging high-resolution, physics-based ground motion simulations can now resolve frequencies of engineering interest and provide detailed synthetic ground motions at high spatial density. This provides an opportunity for new insight into the distribution of infrastructure seismic demands and risk. In the work presented herein, the EQSIM fault-to-structure computational framework described in a companion paper, McCallen et al., is employed to investigate the regional-scale response of buildings to large earthquakes. A representative M = 7.0 strike-slip event is used to explore the distribution and amplitude of building demand, and comparisons are made between building response computed with fault-to-structure simulations and building response computed with existing measured near-fault earthquake records. New information on the distribution and variability of building response from high-performance parallel simulations is described and analyzed, and favorable first comparisons between building response predicted with both fault-to-structure simulations and real ground motions records are presented.

58 GEOSCIENCES↗

SimNet: Accurate and High-Performance Computer Architecture Simulation using Deep Learning

While cycle-accurate simulators are essential tools for architecture research, design, and development, their practicality is limited by an extremely long time-to-solution for realistic applications under investigation. This work describes a concerted effort, where machine learning (ML) is used to accelerate microarchitecture simulation. First, an ML-based instruction latency prediction framework that accounts for both static instruction properties and dynamic processor states is constructed. Then, a GPU-accelerated parallel simulator is implemented based on the proposed instruction latency predictor, and its simulation accuracy and throughput are validated and evaluated against a state-of-the-art simulator. Leveraging modern GPUs, the ML-based simulator outperforms traditional CPU-based simulators significantly.

97 MATHEMATICS AND COMPUTING↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

TReactMech v4.217

TReactMech couples geomechanical processes (poroelasticity, failure, and inelastic strain) with multiphase nonisothermal flow (derived from TOUGH2) and reactive geochemical transport. At its core is the reactive-transport code TOUGHREACT v4.13. TReactMech is an efficient hybrid parallel simulator, solving the geomechanics using finite elements and MPI/PetSc, the multiphase flow using integrated finite difference and MPI/PETSc, and the reactive chemistry using OpenMP. The advantages of TReactMech are in its multiphase flow capabilities (e.g., supercritical CO2, supercritical water, air) and parallel geomechanics including full 3-D stress tensor, shear and tensile failure, coupled to porosity and permeability changes. It is backwardly compatible with TOUGH2 and TOUGHREACT v4.13, allowing for easier transitions between the codes. TReactMech can be used to simulate many natural and engineered subsurface systems, including geothermal reservoirs, borehole heat exchangers, geologic carbon sequestration, geologic storage of nuclear waste, groundwater resources, weathering, sediment diagenesis, seafloor hydrothermal circulation, hydrofracturing in unconventional reservoirs, and injection/production-induced surface deformation.

Sonnenthal, Eric↗

Thermal Scattering Law Data Development for Paraffin Wax

Paraffin wax is often used as a nuclear moderator to slow down the fast neutrons in experimental critical assemblies [1]. It is a colorless and soft solid material that consists primarily of straight-chain alkanes (n-alkanes), which are hydrocarbons with the general formula CnH2n+2 [2-3]. The length of the hydrocarbon chain ranges from C20 to C30 and higher [2]. It is distinguished by its solid state at room temperature and begins to melt above approximately 310 K [4]. Paraffin wax is a commonly employed substance in the manufacture of shielding. One of its noteworthy characteristics is its ability to effectively absorb the neutrons. Also, it possesses a high macroscopic cross section, which enables it to efficiently moderate neutrons. As a result, paraffin wax is extensively utilized in various applications where moderation and shielding of neutrons are needed. For simulations, it is necessary to evaluate its thermal scattering law (TSL) and cross sections. Computationally, classical molecular dynamics (CMD) simulations provide the capability of simulating atomic details. For example, several unary, binary, and few multi component mixtures have been investigated of the paraffin model by using molecular dynamics simulations [5-12]. An assessment of thermal neutron scattering in a heavy paraffinic oil treated both as a solid and a viscous fluid containing 25% linear branched paraffin (C30H62), 35% one ring cycloalkane (C30H60), 15% two rings cycloalkane (C30H58), and 25% aromatic (C30H60) chains has been studied using CMD simulations for producing TSL data [13]. Nevertheless, there is lack of TSL and cross section data for paraffin wax as most of the reported analyses focus on the unary and binary mixture of n-alkanes, which is not consistent with actual paraffin wax [2]. In this work, we applied the equilibrium CMD simulations technique to explore the structure and dynamical properties of wax, which are fundamental input to calculate the TSL. A paraffin wax system was modeled using the CMD code LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) [14-15] with the semi-empirical COMPASS [16] force field. The density of state (DOS) was calculated from the normalized velocity autocorrelation function (VACF), which is the Fourier transform of the normalized VACF. The DOS was used for the calculation of the TSL and thermal scattering cross sections. The paraffin wax atomic system was constructed by using the MedeA material design platform [17], and was benchmarked using available properties (i.e., density, bond lengths, angles, diffusivity, and viscosity).

Nuclear Criticality Safety Program (NCSP)↗