Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Real time identification of large space structures

Identification of frequencies, damping ratios, and mode shapes of large space structures (LSSs) are examined in real time. Real time processing allows for quick updates of model processing after a reconfiguration of structural failure. Recursive lattice least squares (RLLS) was selected as the baseline algorithm for the identification. Simulation results on a one dimensional LSS demonstrated that it provides good estimates, was not ill-conditioned in the presence of under-excited modes, allowed activity by a supervisory control system which prevented damage to the LSS or excessive drift, and was capable of real-time processing for typical LSS models. A suboptimal version of RLLS, which is equivalent to simulated parallel processing, was derived. A NASTRAN model of the dual keel U.S. space station was used to demonstrate the input/identification algorithm package in a more realistic simulation. Because the first eight flexible modes were very close together, the identification was much more difficult than in the simple examples. Even so, the model was accurately identified in real time.

Voss, Janice E.↗

Performance of a parallel algorithm for standard cell placement on the Intel Hypercube

A parallel simulated annealing algorithm for standard cell placement that is targeted to run on the Intel Hypercube is presented. A tree broadcasting strategy that is used extensively in our algorithm for updating cell locations in the parallel environment is presented. Studies on the performance of our algorithm on example industrial circuits show that it is faster and gives better final placement results than the uniprocessor simulated annealing algorithms.

Jones, Mark↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers, part II: Regional simulations of building response

The existing observational database of the regional-scale distribution of strong ground motions and measured building response for major earthquakes continues to be quite sparse. As a result, details of the regional variability and spatial distribution of ground motions, and the corresponding distribution of risk to buildings and other infrastructure, are not comprehensively understood. Utilizing high-performance computing platforms, emerging high-resolution, physics-based ground motion simulations can now resolve frequencies of engineering interest and provide detailed synthetic ground motions at high spatial density. This provides an opportunity for new insight into the distribution of infrastructure seismic demands and risk. In the work presented herein, the EQSIM fault-to-structure computational framework described in a companion paper, McCallen et al., is employed to investigate the regional-scale response of buildings to large earthquakes. A representative M = 7.0 strike-slip event is used to explore the distribution and amplitude of building demand, and comparisons are made between building response computed with fault-to-structure simulations and building response computed with existing measured near-fault earthquake records. New information on the distribution and variability of building response from high-performance parallel simulations is described and analyzed, and favorable first comparisons between building response predicted with both fault-to-structure simulations and real ground motions records are presented.

58 GEOSCIENCES↗

Comparative energetics of the observed and simulated global circulation during the special observing periods of FGGE

Energetics of the observed and simulated global circulation are evaluated in the zonal spectral domain for the special observing periods of FGGE. The study utilizes GLA analyses of FGGE observational data and parallel simulation experiments. There are noticeable differences in energy transformations between the observation and simulation during SOP-1. These include the baroclinic conversion C(n) by the zonal mean motion and short-wave disturbances, and the nonlinear wave-wave interaction L(n) at the long and short waves. The energy transformations of the short-wave disturbances are much more intense in the simulated circulation than in the observation. However, good agreement is noted in the conversion and dissipation of kinetic energy in the large- and cyclone-wave range n = 1-10. Spectral distributions of global energy transformations at the long- and cyclone-wave range indicate that the SOP-2 simulation agrees more closely with the observed fields than the SOP-1 simulation. Other pertinent points of energetics diagnosis are also included in the discussion.

Kung, E. C.↗

SimNet: Accurate and High-Performance Computer Architecture Simulation using Deep Learning

While cycle-accurate simulators are essential tools for architecture research, design, and development, their practicality is limited by an extremely long time-to-solution for realistic applications under investigation. This work describes a concerted effort, where machine learning (ML) is used to accelerate microarchitecture simulation. First, an ML-based instruction latency prediction framework that accounts for both static instruction properties and dynamic processor states is constructed. Then, a GPU-accelerated parallel simulator is implemented based on the proposed instruction latency predictor, and its simulation accuracy and throughput are validated and evaluated against a state-of-the-art simulator. Leveraging modern GPUs, the ML-based simulator outperforms traditional CPU-based simulators significantly.

97 MATHEMATICS AND COMPUTING↗

A parallel algorithm for channel routing on a hypercube

A new parallel simulated annealing algorithm for channel routing on a P processor hypercube is presented. The basic idea used is to partition a set of tracks equally among processors in the hypercube. In parallel, P/2 pairs of processors perform displacements and exchanges of nets between tracks, compute the changes in cost functions, and accept moves using a parallel annealing criteria. Through the use of a unique distributed data structure, it is possible to minimize message traffic and add versatility and efficiency in a parallel routing tool. The algorithm has been implemented and is being tested on some of the popular channel problems from the literature.

Brouwer, Randall↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

TReactMech v4.217

TReactMech couples geomechanical processes (poroelasticity, failure, and inelastic strain) with multiphase nonisothermal flow (derived from TOUGH2) and reactive geochemical transport. At its core is the reactive-transport code TOUGHREACT v4.13. TReactMech is an efficient hybrid parallel simulator, solving the geomechanics using finite elements and MPI/PetSc, the multiphase flow using integrated finite difference and MPI/PETSc, and the reactive chemistry using OpenMP. The advantages of TReactMech are in its multiphase flow capabilities (e.g., supercritical CO2, supercritical water, air) and parallel geomechanics including full 3-D stress tensor, shear and tensile failure, coupled to porosity and permeability changes. It is backwardly compatible with TOUGH2 and TOUGHREACT v4.13, allowing for easier transitions between the codes. TReactMech can be used to simulate many natural and engineered subsurface systems, including geothermal reservoirs, borehole heat exchangers, geologic carbon sequestration, geologic storage of nuclear waste, groundwater resources, weathering, sediment diagenesis, seafloor hydrothermal circulation, hydrofracturing in unconventional reservoirs, and injection/production-induced surface deformation.

Sonnenthal, Eric↗

Thermal Scattering Law Data Development for Paraffin Wax

Paraffin wax is often used as a nuclear moderator to slow down the fast neutrons in experimental critical assemblies [1]. It is a colorless and soft solid material that consists primarily of straight-chain alkanes (n-alkanes), which are hydrocarbons with the general formula CnH2n+2 [2-3]. The length of the hydrocarbon chain ranges from C20 to C30 and higher [2]. It is distinguished by its solid state at room temperature and begins to melt above approximately 310 K [4]. Paraffin wax is a commonly employed substance in the manufacture of shielding. One of its noteworthy characteristics is its ability to effectively absorb the neutrons. Also, it possesses a high macroscopic cross section, which enables it to efficiently moderate neutrons. As a result, paraffin wax is extensively utilized in various applications where moderation and shielding of neutrons are needed. For simulations, it is necessary to evaluate its thermal scattering law (TSL) and cross sections. Computationally, classical molecular dynamics (CMD) simulations provide the capability of simulating atomic details. For example, several unary, binary, and few multi component mixtures have been investigated of the paraffin model by using molecular dynamics simulations [5-12]. An assessment of thermal neutron scattering in a heavy paraffinic oil treated both as a solid and a viscous fluid containing 25% linear branched paraffin (C30H62), 35% one ring cycloalkane (C30H60), 15% two rings cycloalkane (C30H58), and 25% aromatic (C30H60) chains has been studied using CMD simulations for producing TSL data [13]. Nevertheless, there is lack of TSL and cross section data for paraffin wax as most of the reported analyses focus on the unary and binary mixture of n-alkanes, which is not consistent with actual paraffin wax [2]. In this work, we applied the equilibrium CMD simulations technique to explore the structure and dynamical properties of wax, which are fundamental input to calculate the TSL. A paraffin wax system was modeled using the CMD code LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) [14-15] with the semi-empirical COMPASS [16] force field. The density of state (DOS) was calculated from the normalized velocity autocorrelation function (VACF), which is the Fourier transform of the normalized VACF. The DOS was used for the calculation of the TSL and thermal scattering cross sections. The paraffin wax atomic system was constructed by using the MedeA material design platform [17], and was benchmarked using available properties (i.e., density, bond lengths, angles, diffusivity, and viscosity).

Nuclear Criticality Safety Program (NCSP)↗

Devastator Parallel Discrete Event Simulation Runtime (Devastator) v1.0

The Devastator runtime is a modern C++ implementation of optimistic parallel discrete event simulation methods. Devastator allows simulation application code to productively specify their component and event functionality with C++14 constructs. It utilizes GASNet-EX for distributed memory communication and includes parallel performance optimizations such as light-weight thread message queues and asynchronous GVT. Furthermore, it supports efficient event broadcasts and pause-rewind-resume functionality to support periodic load balancing and outer loop optimization algorithms.

Chan, Cy↗

Scalable High Performance Computing: Direct and Large-Eddy Turbulent Flow Simulations Using Massively Parallel Computers

This final report contains reports of research related to the tasks "Scalable High Performance Computing: Direct and Lark-Eddy Turbulent FLow Simulations Using Massively Parallel Computers" and "Devleop High-Performance Time-Domain Computational Electromagnetics Capability for RCS Prediction, Wave Propagation in Dispersive Media, and Dual-Use Applications. The discussion of Scalable High Performance Computing reports on three objectives: validate, access scalability, and apply two parallel flow solvers for three-dimensional Navier-Stokes flows; develop and validate a high-order parallel solver for Direct Numerical Simulations (DNS) and Large Eddy Simulation (LES) problems; and Investigate and develop a high-order Reynolds averaged Navier-Stokes turbulence model. The discussion of High-Performance Time-Domain Computational Electromagnetics reports on five objectives: enhancement of an electromagnetics code (CHARGE) to be able to effectively model antenna problems; utilize lessons learned in high-order/spectral solution of swirling 3D jets to apply to solving electromagnetics project; transition a high-order fluids code, FDL3DI, to be able to solve Maxwell's Equations using compact-differencing; develop and demonstrate improved radiation absorbing boundary conditions for high-order CEM; and extend high-order CEM solver to address variable material properties. The report also contains a review of work done by the systems engineer.

Morgan, Philip E.↗

Photochemically Induced Acousto-optics Fluid Simulations

PIAFS is a finite-difference code to solve the compressible Navier-Stokes equations with chemical heating on Cartesian grids. It models chemical reactions of air (oxygen and carbon dioxide) with ozone subject to radiation. It uses a high-order WENO spatial discretization and explicit Runge-Kutta time integration. It is capable of parallel simulations using MPI. The code is written in C/C++.

Oudin, AlbertineN [Lawrence Livermore National Lab↗

Molecular Dynamics Simulations of Silicon Carbide, Boron Nitride and Silicon for Ceramic Matrix Composite Applications

A comprehensive computational molecular dynamics study is presented for crystalline α-SiC (6H, 4H, and 2H SiC), β-SiC (3C SiC), layered boron nitride, amorphous boron nitride and silicon, the constituent materials for high-temperature SiC/SiC compositions. Large-scale Atomic/Molecular Parallel Simulator software package was used. The Tersoff Potential force field was utilized to evaluate their mechanical characteristics of most of the materials, and the Reax force field was used to model silicon when the Tersoff Potential did not provide accurate results. Their mechanical behaviors were evaluated at a strain rate of 10(exp 7)/s and the results agree with the experimental data in the literature. The results are foundational for linking constituent behavior to composite performance, particularly when test data is unavailable or suspect.

Aluko, Olanrewaju↗

Development of a two-dimensional zonally averaged statistical-dynamical model. III - The parameterization of the eddy fluxes of heat and moisture

A number of perpetual January simulations are carried out with a two-dimensional zonally averaged model employing various parameterizations of the eddy fluxes of heat (potential temperature) and moisture. The parameterizations are evaluated by comparing these results with the eddy fluxes calculated in a parallel simulation using a three-dimensional general circulation model with zonally symmetric forcing. The three-dimensional model's performance in turn is evaluated by comparing its results using realistic (nonsymmetric) boundary conditions with observations. Branscome's parameterization of the meridional eddy flux of heat and Leovy's parameterization of the meridional eddy flux of moisture simulate the seasonal and latitudinal variations of these fluxes reasonably well, while somewhat underestimating their magnitudes. New parameterizations of the vertical eddy fluxes are developed that take into account the enhancement of the eddy mixing slope in a growing baroclinic wave due to condensation, and also the effect of eddy fluctuations in relative humidity. The new parameterizations, when tested in the two-dimensional model, simulate the seasonal, latitudinal, and vertical variations of the vertical eddy fluxes quite well, when compared with the three-dimensional model, and only underestimate the magnitude of the fluxes by 10 to 20 percent.

Stone, Peter H.↗

Minimum Free-Energy Shapes of Ag Nanocrystals: Vacuum vs Solution

Here, we use two variants of replica-exchange molecular dynamics (MD) simulations, parallel tempering MD and partial replica exchange MD, to probe the minimum free-energy shapes of Ag nanocrystals containing 100–200 atoms in a vacuum, ethylene glycol (EG) solvent, and EG solvent with a PVP polymer containing 100 repeat units. Our simulations reveal a shape intermediate between a Dh and an Ih, a Dh-Ih, that has distinct structural signatures and magic sizes. We find several prominent features associated with entropy: pure FCC nanocrystals are less common than FCC crystals containing stacking faults, and crystals with the minimum potential energy are not always preferred over the range of relevant temperatures. The shapes of the nanocrystals in solution are influenced by the chemical identities of the solution-phase molecules. Comparing Ag nanocrystal shapes in EG to those in an EG+PVP solution, we find more icosahedra in EG and more decahedra in EG+PVP across all of the nanocrystal sizes probed in this study. At certain critical sizes, nanocrystal shapes can change dramatically with the addition and removal of a single atom or with a change in temperature at a fixed size. The information in our study could be useful in efforts to devise processing routes to achieve selective nanocrystal shapes.

36 MATERIALS SCIENCE↗

Data parallel sorting for particle simulation

Sorting on a parallel architecture is a communications intensive event which can incur a high penalty in applications where it is required. In the case of particle simulation, only integer sorting is necessary, and sequential implementations easily attain the minimum performance bound of O (N) for N particles. Parallel implementations, however, have to cope with the parallel sorting problem which, in addition to incurring a heavy communications cost, can make the minimun performance bound difficult to attain. This paper demonstrates how the sorting problem in a particle simulation can be reduced to a merging problem, and describes an efficient data parallel algorithm to solve this merging problem in a particle simulation. The new algorithm is shown to be optimal under conditions usual for particle simulation, and its fieldwise implementation on the Connection Machine is analyzed in detail. The new algorithm is about four times faster than a fieldwise implementation of radix sort on the Connection Machine.

Dagum, Leonardo↗

Ion equation of state in quasi-parallel shocks - A simulation result

Ion equation of state in the quasi-parallel collisionless shock is deduced from simulation results. The simulations were performed for theta(bn) = 10 deg, beta = 0.5 and M sub A in the range from 1.2 to 8, where M sub A is the Alfven Mach number, beta is the upstream ratio of plasma pressure to magnetic pressure, and theta(bn) is the angle between the shock normal and the upstream magnetic field. The equation of state can be approximated by a power law with different exponents in the upstream and downstream sides of the shock transition region. The exponent in the upstream side of the transition region is much greater than the adiabatic value of 5/3 and increases with M sub A. The exponent in the downstream side of the transition region is slightly less than 5/3. The results show that ion heating in the quasi-parallel shock is highly nonadiabatic with a large increase in entropy and in temperature ratio in the upstream side of the transition region, while the heating is highly isentropic with a large increase in temperature difference across the principal density jump in the downstream side of the transition region.

Mandt, M. E.↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗