Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

ParMOO: A Python library for parallel multiobjective simulation optimization

A multiobjective optimization problem (MOOP) is an optimization problem in which multiple objectives are optimized simultaneously. The goal of a MOOP is to find solutions that describe the tradeoff between these (potentially conflicting) objectives. Such a tradeoff surface is called the Pareto front. Real-world MOOPs may also involve constraints – additional hard rules that every solution must adhere to. In a multiobjective simulation optimization problem, the objectives are derived from the outputs of one or more computationally expensive simulations. Such problems are ubiquitous in science and engineering.

97 MATHEMATICS AND COMPUTING↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Virtual Time III, Part 1: Unified Virtual Time Synchronization for Parallel Discrete Event Simulation

Algorithms for synchronization of parallel discrete event simulation have historically been divided between conservative methods that require lookahead but not rollback, and optimistic methods that require rollback but not lookahead. In this paper we present a new approach in the form of a framework called Unified Virtual Time (UVT) that unifies the two approaches, combining the advantages of both within a single synchronization theory. Whenever timely lookahead information is available, a logical process (LP) executes conservatively using an irreversible event handler. When lookahead information is not available the LP does not block, as it would in a classical conservative execution, but instead executes optimistically using a reversible event handler. The switch from conservative to optimistic synchronization and back is decided on an event-by-event basis by the simulator, transparently to the model code. UVT treats conservative synchronization algorithms as optional accelerators for an underlying optimistic synchronization algorithm, enabling the speed of conservative execution whenever it is applicable, but otherwise falling back on the generality of optimistic execution. We describe UVT in a novel way, based on fundamental invariants, monotonicity requirements, and synchronization rules. UVT permits zero-delay messages and pays careful attention to tie-handling using superposition. We prove that under fairly general conditions a UVT simulation always makes progress in virtual time. This is Part 1 of a trio of papers describing the UVT framework for PDES, mixing conservative and optimistic synchronization and integrating throttling control.

97 MATHEMATICS AND COMPUTING↗

NAS parallel benchmark results

The NAS (Numerical Aerodynamic Simulation) parallel benchmarks have been developed at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a 'pencil and paper' fashion. The performance results of various systems using the NAS parallel benchmarks are presented. These results represent the best results that have been reported to the authors for the specific systems listed. They represent implementation efforts performed by personnel in both the NAS Applied Research Branch of NASA Ames Research Center and in other organizations.

Bailey, D. H.↗

Scalability study of parallel spatial direct numerical simulation code on IBM SP1 parallel supercomputer

The implementation and the performance of a parallel spatial direct numerical simulation (PSDNS) code are reported for the IBM SP1 supercomputer. The spatially evolving disturbances that are associated with laminar-to-turbulent in three-dimensional boundary-layer flows are computed with the PS-DNS code. By remapping the distributed data structure during the course of the calculation, optimized serial library routines can be utilized that substantially increase the computational performance. Although the remapping incurs a high communication penalty, the parallel efficiency of the code remains above 40% for all performed calculations. By using appropriate compile options and optimized library routines, the serial code achieves 52-56 Mflops on a single node of the SP1 (45% of theoretical peak performance). The actual performance of the PSDNS code on the SP1 is evaluated with a 'real world' simulation that consists of 1.7 million grid points. One time step of this simulation is calculated on eight nodes of the SP1 in the same time as required by a Cray Y/MP for the same simulation. The scalability information provides estimated computational costs that match the actual costs relative to changes in the number of grid points.

Hanebutte, Ulf R.↗

Implementation of a blade element UH-60 helicopter simulation on a parallel computer architecture in real-time

A high-performance platform for development of real-time helicopter flight simulations based on a simulation development and analysis platform combining a parallel simulation development and analysis environment with a scalable multiprocessor computer system is described. Simulation functional decomposition is covered, including the sequencing and data dependency of simulation modules and simulation functional mapping to multiple processors. The multiprocessor-based implementation of a blade-element simulation of the UH-60 helicopter is presented, and a prototype developed for a TC2000 computer is generalized in order to arrive at a portable multiprocessor software architecture. It is pointed out that the proposed approach coupled with a pilot's station creates a setting in which simulation engineers, computer scientists, and pilots can work together in the design and evaluation of advanced real-time helicopter simulations.

Moxon, Bruce C.↗

Circumbinary Disk Accretion into Spinning Black Hole Binaries

Supermassive black hole binaries are likely to accrete interstellar gas through a circumbinary disk. Shortly before merger, the inner portions of this circumbinary disk are subject to general relativistic effects. To study this regime, we approximate the spacetime metric of close orbiting black holes by superimposing two boosted Kerr–Schild terms. After demonstrating the quality of this approximation, we carry out very long-term general relativistic magnetohydrodynamic simulations of the circumbinary disk. We consider black holes with spin dimensionless parameters of magnitude 0.9, in one simulation parallel to the orbital angular momentum of the binary, but in another anti-parallel. These are contrasted with spinless simulations. We find that, for a fixed surface mass density in the inner circumbinary disk, aligned spins of this magnitude approximately reduce the mass accretion rate by 14% and counter-aligned spins increase it by 45%, leaving many other disk properties unchanged.

79 ASTRONOMY AND ASTROPHYSICS↗

Circumbinary Disk Accretion into Spinning Black Hole Binaries

Supermassive black hole binaries are likely to accrete interstellar gas through a circumbinary disk. Shortly before merger, the inner portions of this circumbinary disk are subject to general relativistic effects. To study this regime, we approximate the spacetime metric of close orbiting black holes by superimposing two boosted Kerr–Schild terms. After demonstrating the quality of this approximation, we carry out very long-term general relativistic magnetohydrodynamic simulations of the circumbinary disk. We consider black holes with spin dimensionless parameters of magnitude 0.9, in one simulation parallel to the orbital angular momentum of the binary, but in another anti-parallel. These are contrasted with spinless simulations. We find that, for a fixed surface mass density in the inner circumbinary disk, aligned spins of this magnitude approximately reduce the mass accretion rate by 14% and counter-aligned spins increase it by 45%, leaving many other disk properties unchanged.

AGN↗

ORNL_AISD_NiPt

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for three different sizes of the crystal structure: 256 atoms, 864 atoms, and 2,048 atoms, each of which contains 1900 configurations. For each size of the crystal structure, the data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in three compressed zipped folders: atoms256.zip, atoms864.zip, atoms2048.zip Each zipped folder contains the data that describes crystals of size 256 atoms, 864 atoms, and 2,048 atoms respectively. Each one of the three zipped folders contains the data structured in the following way: -Ni_ground_state.cfg --> atomic configuration for the pure nickel -Pt_ground_state.cfg --> atomic configuration for the pure platinum -Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder. Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms256.zip is 65,046. The total number of atomic configurations contained in atoms864.zip is 63,936. The total number of atomic configurations contained in atoms2048.zip is 61,997. The total number of atomic configurations spanned by the entire dataset is 190,979. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büccchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D

36 MATERIALS SCIENCE↗

ORNL_AISD_NiPt_108atoms

This dataset describes the nickel-platinum (NiPt) solid solution binary alloy, where the two constituent elements nickel (Ni) and platinum (Pt) are randomly placed on the face centered cubic (FCC) crystal structure, with the lattice constant of 3.840 angstroms. The dataset comprises data for crystal structures with 108 atoms with 1,900 configurations. The data set was generated for concentrations ranging from 0at% of Pt to 100at% of Pt in the NiPt binary system, with increasing the concentration of Pt in the system every 5at%. For each one of the chemical compositions, 100 random configurations were generated, each with a different random seed. Each of the output files contains the mass, type, atomic coordinates, energy per atom, and forces in x, y, and z directions respectively. For each atomic configuration, the output was collected every 150 steps during the minimization stage and every 1000 steps during the replica exchange stage. Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) [1], which is a molecular dynamics code, was used to generate data for NiPt alloy. The simulation used the interatomic potential for NiPt binary system 'MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001' [3] from the OpenKIM library (Open Knowledgebase of Interatomic Models) [2]. This potential was developed based on the second nearest-neighbor modified embedded-atom method (2NN MEAM). The simulation process begins with the generation of the random NiPt structure and follows with the short minimization and replica exchange simulation. The minimization procedure adjusts atomic coordinates and performs energy minimization, which typically leads to a local potential energy minimum. The method used for the minimization was the conjugate gradient algorithm. A short replica exchange (parallel tempering) simulation involves four replicas (ensembles) of a system and follows the minimization stage. Multiple snapshots of the configuration were collected during the minimization and replica exchange stages. NiPt alloy is interesting due to its magnetic and charge transfer properties [4]. The data is provided in a compressed zipped folders atoms108.zip. The zipped folder contains the data structured in the following way: - Ni_ground_state.cfg --> atomic configuration for the pure nickel - Pt_ground_state.cfg --> atomic configuration for the pure platinum - Pt#_filtered --> folders containing atomic configurations for #at% concentration of platinum. The folder contains 100 atomic configurations, each saved in a subfolder - Each subfolder named config* is associated with a specific atomic configuration. Each of these subfolders contains files with .cfg format, corresponding to outputs for each atomic configuration The total number of atomic configurations contained in atoms108.zip is 66,132. This dataset is an extension to the dataset ORNL_AISD_NiPt [5] that has been previously released with crystal structures of 256 atoms, 864 atoms, and 2,048 atoms, with the same methodology for data collection. References [1] https://www.lammps.org/ [2] https://openkim.org/ [3] https://openkim.org/id/MEAM_LAMMPS_KimSeolJi_2017_PtNi__MO_020840179467_001 [4] El-Gendy, Ahmed A. and Hampel, Silke and Büchner, Bernd and Klingeler, Rüdiger, Tuneable magnetic properties of carbon-shielded NiPt-nanoalloys, RSC Adv., volume 6, issue 57, pages 52427-52433, 2016, The Royal Society of Chemistry, doi:10.1039/C6RA05910D [5] M. Karabin, M. Lupo Pasini, and M. Eisenbach. ORNL_AISD_NiPt. United States: N. p., 2023. Web. doi:10.13139/OLCF/1958172.

36 MATERIALS SCIENCE↗

Scalable Deep Learning-Based Microarchitecture Simulation on GPUs

Cycle-accurate microarchitecture simulators are essential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the potential of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.

97 MATHEMATICS AND COMPUTING↗

Parallel algorithms for simulating continuous time Markov chains

We have previously shown that the mathematical technique of uniformization can serve as the basis of synchronization for the parallel simulation of continuous-time Markov chains. This paper reviews the basic method and compares five different methods based on uniformization, evaluating their strengths and weaknesses as a function of problem characteristics. The methods vary in their use of optimism, logical aggregation, communication management, and adaptivity. Performance evaluation is conducted on the Intel Touchstone Delta multiprocessor, using up to 256 processors.

Nicol, David M.↗

Unsteady turbomachinery flow simulations on massively parallel architectures

The accurate numerical simulation of unsteady, three-dimensional viscous flow in turbomachines is computationally very intensive, requiring prohibitively large amounts of computer time on current vector supercomputers. In recent years, computer systems based on massively parallel architectures have been developed that offer the promise of meeting the computational power requirements of such large-scale simulations. However, a rethinking of existing algorithms and methodology is required in order to fully harness the computational power of such architectures. In this paper the capabilities of the Connection Machine (CM-2) in predicting unsteady flows in turbomachines are evaluated. The implementation on the CM-2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original algorithm (developed for vector, pipelined supercomputers) in order to improve performance on the CM-2 are outlined. Algorithm performance is evaluated and compared with a functionally equivalent code for the CRAY-YMP.

Madavan, N. K.↗

Dataset of simulated vibrational density of states and X-ray diffraction profiles of mechanically deformed and disordered atomic structures in Gold, Iron, Magnesium, and Silicon

This dataset is comprised of a library of atomistic structure files and corresponding X-ray diffraction (XRD) profiles and vibrational density of states (VDoS) profiles for bulk single crystal silicon (Si), gold (Au), magnesium (Mg), and iron (Fe) with and without disorder introduced into the atomic structure and with and without mechanical loading. Included with the atomistic structure files are descriptor files that measure the stress state, phase fractions, and dislocation content of the microstructures. All data was generated via molecular dynamics or molecular statics simulations using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) code. This dataset can inform the understanding of how local or global changes to a materials microstructure can alter their spectroscopic and diffraction behavior across a variety of initial structure types (cubic diamond, face-centered cubic (FCC), hexagonal close-packed (HCP), and body-centered cubic (BCC) for Si, Au, Mg, and Fe, respectively) and overlapping changes to the microstructure (i.e., both disorder insertion and mechanical loading).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗