Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Implementation of Detailed Polyethylene Pyrolysis Kinetics into CFD Simulations using Machine Learning

Municipal solid waste (MSW) and waste plastics have received significant attention due to the issues of waste generation and storage, as well as their potential as an energy resource. High-density polyethylene (HDPE) makes up a large portion of plastic waste and has been the subject of several conversion studies. However, the mechanisms associated with converting HDPE through pyrolysis and gasification are extensive and complex making them difficult to implement into high-fidelity computational fluid dynamic (CFD) simulations. For this project, a primary pyrolysis mechanism containing 42 unique species and 737 heterogeneous reactions was used to generate kinetic data over a range of operating conditions. A machine learning (ML) model was developed to replicate the results of the detailed pyrolysis mechanism while significantly increasing the computational efficiency. A deep operator network (DeepONet) architecture was adopted to train the model using time steps relevant to CFD simulations. The ML used physics-based loss functions to ensure mass conservation. The ML model has been deployed in simple MFiX CFD simulations, single particle, and an experimental drop tube reactor, and has shown promising performance compared to the original scheme.

Houston, Ross↗

Multigrid Reduction in Time for Chaotic Dynamical Systems

As CPU clock speeds have stagnated and high performance computers continue to have ever higher core counts, increased parallelism is needed to take advantage of these new architectures. Traditional serial time-marching schemes can be a significant bottleneck, as many types of simulations require large numbers of time-steps which must be computed sequentially. Parallel-in-time schemes, such as the Multigrid Reduction in Time (MGRIT) method, remedy this by parallelizing across time-steps and have shown promising results for parabolic problems. However, chaotic problems have proved more difficult, since chaotic initial value problems (IVPs) are inherently ill-conditioned. MGRIT relies on a hierarchy of successively coarser time-grids to iteratively correct the solution on the finest time-grid, but due to the nature of chaotic systems, small inaccuracies on the coarser levels can be greatly magnified and lead to poor coarse-grid corrections. Here we introduce a modified MGRIT algorithm based on an existing quadratically converging nonlinear extension to the multigrid Full Approximation Scheme (FAS), as well as a novel time-coarsening scheme. Together, these approaches better capture long-term chaotic behavior on coarse-grids and greatly improve convergence of MGRIT for chaotic IVPs. Further, we introduce a novel low-memory variant of the algorithm for solving chaotic PDEs with MGRIT which not only solves the IVP, but also provides estimates for the unstable Lyapunov vectors of the system. Finally, we provide supporting numerical results for the Lorenz system and demonstrate parallel speedup for the chaotic Kuramoto–Sivashinsky PDE over a significantly longer time-domain than in previous works.

97 MATHEMATICS AND COMPUTING↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗

A PIC Bootstrapping Strategy for Exascale CFD-DEM Simulation

The exascale computing era comes with the release of MFIX-Exa, a new code for studying gas-particle fluidization using CFD-DEM and PIC models on massively parallel, heterogeneous high-performance computing architectures. However, as compute resources continue to grow in scale and complexity, so too does the impetus to use them efficiently. Here, we propose a novel bootstrapping method to minimize neglected simulation time and cast the approach more broadly among a growing field of multi-fidelity methods.

Porcu, Roberto↗

Studying performance portability of LAMMPS across diverse GPU-based platforms

The molecular dynamics simulation software, LAMMPS, utilizes the Kokkos acceleration library to port computation to a diverse set of architectures including those based on GPU accelerators. In addition to Kokkos, LAMMPS contains a vast code base that leverages the CUDA application programming interface using library functions such as cuFFT, CUDA's fast-fourier transform (FFT) library, and, more recently, also support for AMD's Heterogeneous Interface for Portability (HIP) that is rapidly growing. While preparing LAMMPS tests for the AMD GPU-based test system precursors to Frontier, we investigated several strategies for accelerating LAMMPS on AMD GPUs, using the AMD Instinct MI100 and MI250X. In this work, we integrated the HIP FFT library, hipFFT, into the particle-particle particle-mesh (PPPM) long-range solver, which allowed the porting of PPPM calculations to the GPUs. Kokkos behavior on the MI100 and MI250X was also investigated through the package kokkos command of LAMMPS, targeting communication, memory usage, and particle grid decomposition. The Tersoff, Reax, Lennard-Jones (LJ), EAM, Granular, and PPPM potentials were investigated in this effort, and results from these experiments are provided. In conclusion, the selected potentials were run on Spock (AMD Instinct MI100), Crusher (AMD Instinct MI250X), AFW HPC11 (NVIDIA A100) and Summit (NVIDIA V100), for comparison. Operational roofline models were constructed and analyzed for the Tersoff, Reax, and Lennard–Jones potentials on Crusher and Summit.

97 MATHEMATICS AND COMPUTING↗

Transfer learning for probabilistic localization of hidden cracks in concrete structures

Abstract The utility of discriminative supervised learning models built using multiple training-data sources is investigated for hidden crack localization in concrete. Feed-forward neural network (FFNN) is chosen as the model architecture, and transfer learning is used to assimilate the information obtained from different sources (computational physics simulations and laboratory experiments). The labeled training data consists of values of a damage index and the known locations of hidden cracks. The classification models need to learn how the presence of damage (hidden cracks) affects the damage index at different sensors for different test conditions. To this end, diagnostic FFNN models are built by sequentially adding and training new hidden layers to assimilate labeled information from computer models (different model geometries, test conditions, crack lengths, crack locations) and laboratory experiments on a plain cement slab. These transfer learning-based models are then used to localize damage in concrete specimens that reflect real-world conditions (i.e., specimens with steel reinforcement and randomly distributed aggregate). The actual damage state in these specimens is determined by extracting cores and performing petrographic studies on the extracted cores. The damage probability estimated by transfer learning-based models is compared with the petrographic damage rating index (DRI) to identify the most suitable approach to train the diagnostic models. The transfer learning-based diagnostic methodology shows promise and could be used in various structural health monitoring applications, where sufficient labeled data are typically not available from a single data source.

Miele, S.↗

ERF: Energy Research and Forecasting

The Energy Research and Forecasting (ERF) code is a new model that simulates the mesoscale and microscale dynamics of the atmosphere using the latest high-performance computing architectures. It employs hierarchical parallelism using an MPI+X model, where X may be OpenMP on multicore CPU-only systems, or CUDA, HIP, or SYCL on GPU-accelerated systems. ERF is built on AMReX (Zhang et al., 2019, 2021), a block-structured adaptive mesh refinement (AMR) software framework that provides the underlying performance-portable software infrastructure for block-structured mesh operations. The "energy" aspect of ERF indicates that the software has been developed with renewable energy applications in mind. In addition to being a numerical weather prediction model, ERF is designed to provide a flexible computational framework for the exploration and investigation of different physics parameterizations and numerical strategies, and to characterize the flow field that impacts the ability of wind turbines to extract wind energy. The ERF development is part of a broader effort led by the US Department of Energy's Wind Energy Technologies Office.

17 WIND ENERGY↗

Extending SST vanadis to Add SIMT Functional Units

Sandia National Laboratories is currently investigating scalable architectural simulation capabilities, with a focus on simulating and evaluating highly scalable supercomputers for high-performance computing applications. This exploration is driven by the shift toward more specialized forms of compute and the need for a more diverse set of accurate models. This project will explore the use of General-Purpose Graphical Processing Units (GPGPUs) in high-performance computing using both physical systems and new simulator models – traditional GPUs as well as tightly-coupled SIMT accelerators.

97 MATHEMATICS AND COMPUTING↗

Studying Performance Portability of LAMMPS across Diverse GPU-based Platforms

The molecular dynamics simulation software, LAMMPS, utilizes the Kokkos acceleration library to port computation to a diverse set of architectures including those based on GPU accelerators. In addition to Kokkos, LAMMPS contains a vast code base that leverages the CUDA application programming interface using library functions such as cuFFT, CUDA’s fast-fourier transform (FFT) library, and, more recently, also support for AMD’s Heterogeneous Interface for Portability (HIP) that is rapidly growing. While preparing LAMMPS tests for the AMD GPU-based test system precursors to Frontier, we investigated several strategies for accelerating LAMMPS on AMD GPUs, using the AMD Instinct MI100 and MI250X. In this work, we integrated the HIP FFT library, hipFFT, into the particle-particle particle-mesh (PPPM) long-range solver, which allowed the porting of PPPM calculations to the GPUs. Kokkos behavior on the MI100 and MI250X was also investigated through the package kokkos command of LAMMPS, targeting com- munication, memory usage, and particle grid decomposition. The Tersoff, Reax, Lennard-Jones (LJ), EAM, Granular, and PPPM potentials were investigated in this effort, and results from these experiments are provided. The selected potentials were run on Spock (AMD Instinct MI100), Crusher (AMD Instinct MI250X), AFW HPC11 (NVIDIA A100) and Summit (NVIDIA V100), for comparison. Operational roofline models were constructed and analyzed for the Tersoff, Reax, and Lennard-Jones potentials on Crusher and Summit.

Hagerty, Nick↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

Artificial Intelligence for Multiphysics Nuclear Design Optimization with Additive Manufacturing

The geometric flexibility of additively manufactured metals and ceramics generates a very large and open design space that requires advanced modeling and simulation tools for physics simulations and the rigorous definition of design problems. This effort deploys artificial intelligence (AI) and machine learning (ML) algorithms to understand the design space, evaluate potential designs, and more efficiently generate optimized results. The Transformational Challenge Reactor (TCR) program is leveraging advances in several scientific areas—including materials, manufacturing, sensors and control systems, data analytics, and high-fidelity modeling and simulation—to accelerate the design, manufacturing, qualification, and deployment of advanced nuclear energy systems. Through a manufacturing-informed design approach, the TCR program seeks to integrate digital data for rapid nuclear innovation; accelerate the adoption of advances in manufacturing, materials, and computational sciences for nuclear applications; and dramatically reduce deployment costs and timelines for new nuclear reactor technologies. This report documents efforts under the TCR program to leverage advanced modeling and simulation techniques driven by AI/ML algorithms on high-performance computing (HPC) systems to yield more optimized TCR core designs. A multiphysics ML surrogate model was developed to run on the HPC architectures. The surrogate model is trained on high-fidelity simulation data of coupled neutronics and thermofluidics and is used to quickly evaluate thousands of candidate core designs in parallel, which drives the evolution of the cooling channel shapes to minimize temperature peaking and material stress. Outcomes from these activities provide design information and feedback into the core design efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory

The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Portable Programming Models to Accelerate LArTPC Detector Simulations

The Liquid Argon Time Projection Chamber (LArTPC) technology is widely used in high energy physics experiments, including the upcoming Deep Underground Neutrino Experiment (DUNE). Accurately simulating LArTPC detector responses is essential for analysis algorithm development and physics model interpretations. Accurate LArTPC detector response simulations are computationally demanding, and can become a bottleneck in the analysis workflow. Compute devices such as General-Purpose Graphics Processing Units (GPGPUs) have the potential to substantially accelerate simulations compared to traditional CPU-only processing. The software development for these compute accelerators often carries the cost of specialized code refactorization and porting to match the target hardware architecture. With the rapid evolution and increased diversity of the computer architecture landscape, it is highly desirable to have a portable solution that also maintains reasonable performance. We report our ongoing effort in evaluating Kokkos as a basis for this portable programming model using LArTPC simulations in the context of the Wire-Cell Toolkit, a C++ library for LArTPC simulations, data analysis, reconstruction and visualization.

47 OTHER INSTRUMENTATION↗

Real-Time Lifetime Prediction of Semiconductor Devices Using Hardware-in-the-Loop

This paper presents a unique approach to enable real-time lifespan prediction of semiconductor power modules using a Hardware-in-the-Loop (HIL) system. By integrating the module's overall loss characteristics-specifically switching and conduction losses-with a thermoelectric model of the thermal management system, this research demonstrates that the model can dynamically estimates the junction temperature profile of the semiconductor devices in response to a changing torque demand profile for the motor drive system. This capability enables continuous monitoring of the module's operational time and cumulative stress induced on the devices to compute accumulated remaining lifetime or time-to-failure (TTF). This study provides an architectural framework for the HIL system with high-fidelity component models of multiple physical domains, allowing simulation of dynamic behaviors of a closely-coupled motor drive system. The advanced real-time computation and measurement functionalities of the HIL system allow for both dynamic lifetime calculations based on simulated data and aggregate lifetime predictions utilizing historical data. Moreover, this paper details an algorithm that not only computes cumulative damage but also synthesizes these data into a comprehensive aggregated lifetime metric. This methodology can enhance the maintenance scheduling strategies and operational reliability of semiconductor devices in critical applications, ultimately extending their service life while optimizing performance.

hardware-in-the-loop (HIL)↗

Automatic Differentiation of C++ Codes on Emerging Manycore Architectures with Sacado

Automatic differentiation (AD) is a well-known technique for evaluating analytic derivatives of calculations implemented on a computer, with numerous software tools available for incorporating AD technology into complex applications. However, a growing challenge for AD is the efficient differentiation of parallel computations implemented on emerging manycore computing architectures such as multicore CPUs, GPUs, and accelerators as these devices become more pervasive. In this work, we explore forward mode, operator overloading-based differentiation of C++ codes on these architectures using the widely available Sacado AD software package. In particular, we leverage Kokkos, a C++ tool providing APIs for implementing parallel computations that is portable to a wide variety of emerging architectures. Here we describe the challenges that arise when differentiating code for these architectures using Kokkos, and two approaches for overcoming them that ensure optimal memory access patterns as well as expose additional dimensions of fine-grained parallelism in the derivative calculation. We describe the results of several computational experiments that demonstrate the performance of the approach on a few contemporary CPU and GPU architectures. We then conclude with applications of these techniques to the simulation of discretized systems of partial differential equations.

97 MATHEMATICS AND COMPUTING↗

The Multiphysics on Advanced Platforms Project

In 2015, the Lawrence Livermore National Laboratory started development of next-generation multiphysics simulation capabilities for the National Nuclear Security Administration under the Advanced Technologies Development and Mitigation (ATDM) element of the Advanced Simulation and Computing program in collaboration with the Exascale Computing Project (ECP). A key driver for this effort across the NNSA tri-lab was the emergence of advanced high performance computing (HPC) architectures based on heterogeneous compute capabilities, including GPU based systems, as part of the national drive toward exascale computing platforms at multiple Department of Energy (DOE) facilities. Developing a multiphysics code capable of meeting the various simulation needs of the NNSA as defined by the current generation of integrated codes (or ICs), initially developed as part of the Accelerated Strategic Computing Initiative (ASCI) program beginning in 1996, and able to scale to the current 100 petaflop class pre-exascale systems, as well the forthcoming exaflop class computers, is a daunting challenge. To accomplish this ambitious goal, LLNL has embraced two key themes: use of high-order numerical methods and a modular approach to code development. The LLNL next generation effort is organized under the Multi-Physics on Advanced Platforms Project (MAPP). A foundational component of MAPP is the Axom computer science (CS) toolkit which provides infrastructure for the development of modular, performance portable, multi-physics application codes. MARBL is a next-generation application code built on the Axom base to address the modeling needs of the high energy density physics (HEDP) community for simulating high-explosive, magnetic or laser driven experiments such as inertial confinement fusion (ICF), pulsed-power magneto-hydrodynamics (MHD), equation of state (EOS) and material strength studies as part of the NNSA’s stockpile stewardship program (SSP).

97 MATHEMATICS AND COMPUTING↗

Multi-GPU porting of a phase-change cascaded lattice Boltzmann method for three-dimensional pool boiling simulations

The Lattice Boltzmann method (LBM) has proven effective in simulating phase-change phenomena, such as melting, solidification, evaporation, and boiling. In this work, we develop a highly parallelized multi-GPU implementation of LBM for three-dimensional pool boiling simulations. The code is based on the OpenACC programming model, which enables the code to be deployed efficiently on multi-core CPUs, GPUs, and potentially other accelerators, without the need for architecture-specific rewrites. To support large-scale simulations, the domain is decomposed and distributed across multiple compute nodes using MPI. We demonstrate that the code exhibits excellent scaling properties, with ideal strong-scaling running with up to 256 GPUs on the MareNostrum5 cluster.

97 MATHEMATICS AND COMPUTING↗

Scattering phase shift in quantum mechanics on quantum computers

Here, we investigate the feasibility of extracting infinite volume scattering phase shift on quantum computers in a simple one-dimensional quantum mechanical model, using the formalism established in the work by Guo and Gasparian [Phys. Rev. D 108, 074504 (2023)] that relates the integrated correlation functions for a trapped system to the infinite volume scattering phase shifts through a weighted integral. The system is first discretized in a finite box with periodic boundary conditions, and the formalism in real time is verified by employing a contact interaction potential with exact solutions. Quantum circuits are then designed and constructed to implement the formalism on current quantum computing architectures. To overcome the fast oscillatory behavior of the integrated correlation functions in real-time simulation, different methods of postdata analysis are proposed and discussed. Test results on IBM hardware show that good agreement can be achieved with two qubits, but complete failure ensues with three qubits due to two-qubit gate operation errors and thermal relaxation errors.

Guo, Peng [Dakota State Univ., Madison, SD (United↗