Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Celeritas R&D Report: Accelerating Geant4

Celeritas is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a Graphics Processing Unit (GPU)-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.4, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling straightforward integration of Celeritas into the high energy physics (HEP) experimental frameworks CMSSW and ATLAS FullSimLight. On the Perlmutter supercomputer, Celeritas performs EM physics between 3× and 18× faster using the machine’s Nvidia GPUs compared to using only CPUs, corresponding to an electrical power efficiency up to a factor of 5. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.7–2.2× on GPU and 1.2× on CPU. In a CMS test application using tt¯ events and the prototype Run 4 configuration, compared to Geant4 CPU, Celeritas with a Nvidia A100 improves overall throughput up to a factor of 2.7× but cannot be efficiently shared with more than 8 cores.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Heterogeneous Computing

To leverage the increasing heterogeneity in modern computing resources, Geant4 incorporates advanced software tools and a task-based framework (G4Tasking) that enables efficient parallelism at event, sub-event, and track levels. Ongoing R&D efforts focus on integrating GPUs into high-energy physics (HEP) simulations, including optical photon simulation with Opticks/NVIDIA OptiX, offloading electromagnetic particle transport using G4HepEM/AdePT and Celeritas, and employing advanced surface-based geometry models such as VecGeom2.0 and ORANGE. As Geant4 continues evolving toward high-performance computing (HPC) and heterogeneous architectures, it remains a key tool for large-scale simulations in HEP and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Probabilistic simulation of quantum circuits using a deep-learning architecture

The fundamental question of how to best simulate quantum systems using conventional computational resources lies at the forefront of condensed matter and quantum computation. It impacts both our understanding of quantum materials and our ability to emulate quantum circuits. Here we present an exact formulation of quantum dynamics via factorized generalized measurements which maps quantum states to probability distributions with the advantage that local unitary dynamics and quantum channels map to local quasistochastic matrices. This representation provides a general framework for using state-of-the-art probabilistic models in machine learning for the simulation of quantum many-body dynamics. Using this framework, we have developed a practical algorithm to simulate quantum circuits using an attention network based on a powerful neural network ansatz responsible for the most recent breakthroughs in natural language processing. We demonstrate our approach by simulating circuits that build Greenberger-Horne-Zeilinger and linear graph states of up to 60 qubits, as well as a variational quantum eigensolver circuit for preparing the ground state of the transverse field Ising model on several system sizes. Our methodology constitutes a modern machine learning approach to the simulation of quantum physics with applicability both to quantum circuits as well as other quantum many-body systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Neural Architecture Search Benchmarks: Insights and Survey

Neural Architecture Search (NAS), a promising and fast-moving research field, aims to automate the architectural design of Deep Neural Networks (DNNs) to achieve better performance on the given task and dataset. NAS methods have been very successful in discovering efficient models for various Computer Vision, Natural Language Processing, etc. The major obstacles to the advancement of NAS techniques are the demand for large computation resources and fair evaluation of various search methods. The differences in training pipeline and setting make it challenging to compare the efficiency of two NAS algorithms. A large number of NAS Benchmarks to simulate the architecture evaluation in seconds have been released over the last few years to ease the computation burden of training neural networks and can aid in the unbiased assessment of different search methods. This paper provides an extensive review of several publicly available NAS Benchmarks in the literature. We provide technical details and a deeper understanding of each benchmark and point out future directions.

97 MATHEMATICS AND COMPUTING↗

Surrogate modeling of Cellular-Potts agent-based models as a segmentation task using the U-Net neural network architecture

The Cellular-Potts model is a powerful and ubiquitous framework for developing computational models for simulating complex multicellular biological systems. Cellular-Potts models (CPMs) are often computationally expensive due to the explicit modeling of interactions among large numbers of individual model agents and diffusive fields described by partial differential equations (PDEs). In this work, we develop a convolutional neural network (CNN) surrogate model using a U-Net architecture that accounts for periodic boundary conditions. We use this model to accelerate the evaluation of a mechanistic CPM previously used to investigate in vitro vasculogenesis. The surrogate model was trained to predict 100 computational steps ahead (Monte-Carlo steps, MCS), accelerating simulation evaluations by a factor of 562 times compared to single-core CPM code execution on CPU. Over short timescales of up to 3 recursive evaluations, or 300 MCS, our model captures the emergent behaviors demonstrated by the original Cellular-Potts model such as vessel sprouting, extension and anastomosis, and contraction of vascular lacunae. This approach demonstrates the potential for deep learning to serve as a step toward efficient surrogate models for CPM simulations, enabling faster evaluation of computationally expensive CPM simulations of biological processes.

97 MATHEMATICS AND COMPUTING↗

Teacher-student training improves the accuracy and efficiency of machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) are revolutionizing the field of molecular dynamics (MD) simulations. Recent MLIPs have tended towards more complex architectures trained on larger datasets. The resulting increase in computational and memory costs may prohibit the application of these MLIPs to perform large-scale MD simulations. Herein, we present a teacher-student training framework in which the latent knowledge from the teacher (atomic energies) is used to augment the students' training. We show that the light-weight student MLIPs have faster MD speeds at a fraction of the memory footprint compared to the teacher models. Remarkably, the student models can even surpass the accuracy of the teachers, even though both are trained on the same quantum chemistry dataset. Our work highlights a practical method for MLIPs to reduce the resources required for large-scale MD simulations.

36 MATERIALS SCIENCE↗

ToPolyAgent: AI agents for coarse-grained bead-spring topological polymer simulations

We introduce ToPolyAgent, a multi-agent AI framework for performing coarse-grained molecular dynamics (MD) simulations of topological polymers through natural language instructions. By integrating large language models (LLMs) with domain-specific computational tools, ToPolyAgent supports both interactive and autonomous simulation workflows across diverse polymer architectures, including linear, ring, brush, and star polymers, as well as dendrimers. The system consists of four LLM-powered agents: a Config Agent for generating initial polymer–solvent configurations, a Simulation Agent for executing LAMMPS-based MD simulations and conformational analyses, a Report Agent for compiling markdown reports, and a Workflow Agent for streamlined autonomous operations. Interactive mode incorporates user feedback loops for iterative refinements, while autonomous mode enables end-to-end task execution from detailed prompts. We demonstrate ToPolyAgent's versatility through case studies involving diverse polymer architectures under varying solvent conditions, thermostats, and simulation lengths. Furthermore, we highlight its potential as a research assistant by directing it to investigate the effect of interaction parameters on the linear polymer conformation, and the influence of grafting density on the persistence length of the brush polymer. By coupling natural language interfaces with rigorous simulation tools, ToPolyAgent lowers barriers to complex computational workflows and advances AI-driven materials discovery in polymer science. It lays the foundation for autonomous and extensible multi-agent scientific research ecosystems.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

A GPU‐Based Ocean Dynamical Core for Routine Mesoscale‐Resolving Climate Simulations

Abstract We describe an ocean hydrostatic dynamical core implemented in Oceananigans optimized for Graphical Processing Unit (GPU) architectures. On 64 A100 GPUs, equivalent to 16 computational nodes in current state‐of‐the‐art supercomputers, our dynamical core can simulate a decade of near‐global ocean dynamics per wall‐clock day at an 8‐km horizontal resolution; a resolution adequate to resolve the ocean's mesoscale eddy field. Such efficiency, achieved with relatively modest hardware resources, suggests that climate simulations on GPUs can incorporate fully eddy‐resolving ocean models. This removes a major source of systematic bias in current IPCC coupled model projections, the parameterization of ocean eddies, and represents a major advance in climate modeling. We discuss the computational strategies, focusing on GPU‐specific optimization and numerical implementation details that enable such high performance.

Silvestri, Simone [Massachusetts Institute of Tech↗

Understanding Power and Energy Utilization in Large Scale Production Physics Simulation Codes

Power is an often-cited reason for moving to advanced architectures on the path to Exascale computing. This is due to the practical concern of delivering enough power to successfully site and operate these machines, as well as concerns over energy usage while running large simulations. Since accurate power measurements can be difficult to obtain, processor thermal design power (TDP) is a possible surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advance technology systems at Lawrence Livermore National Laboratory (LLNL) and Sandia National Laboratory, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale LLNL simulation codes are significantly more efficient than a simple processor TDP model might suggest.

97 MATHEMATICS AND COMPUTING↗

Materials genome innovation for computational software (magics) center

Functional layered material (LM) architectures will dominate nanomaterials science in this century. We have developed theory, modeling, simulation, and software and data tools that enhance understanding and AI guide synthesis, enable characterization of complex structures, and improve capabilities in the predictive design and growth of LMs. Research at the Center has focused on: Computational synthesis and characterization: AI guided synthesis and experimental synthesis of stacked LMs with tailored properties via optimized chemical vapor deposition (CVD) growth and liquid-phase exfoliation; study defects, edges, grain boundaries, wrinkling of atomic layers and their effects on chemical, mechanical, electrical, and optical properties. Far-from-equilibrium processes: Joint experimental and simulation based probe of electronic processes with NAQMD and ultrafast X-ray free-electron laser (XFEL) and ultrafast electron diffraction (UED) facilities at Stanford. Experimentally validate NAQMD by ultrafast electron diffraction and X-ray spectroscopy studies of structural and excited state dynamics, shape fluctuations, and phonon dynamics. Scalable software: Simulation engines for desktop-to-exascale platforms using low-overhead, linear-scaling QMD algorithms; divide-conquer-recombine NAQMD with electronic excitations; extended-Lagrangian reactive molecular dynamics (RMD), machine learning (ML) based neural-network quantum molecular dynamics (NNQMD), and super-state accelerated molecular dynamics (AMD) and kinetic Monte Carlo codes; thermal and electrical transport software; and design 3D architectures of LMs with desired functionality using scalable software. Distribution of software and data, and training: Software and simulation-experimental data generated within the Center are distributed to the materials science community via Berkeley Materials Project (MP) framework. We have also organized three workshops for software distribution and training at USC (Nov. 2017, Mar. 2018) and Gaithersburg, MD (Nov. 2018) to train researchers, with the last one in focused on underrepresented groups, in collaboration with Howard University which is one of the largest HBCUs. The Center supported a total of 46 personnel and 6 undergraduate students. These include 14 faculty, 11 postdoctoral research associates, 20 graduate research assistants, and mentored 6 undergraduate students. This resulted in the publications of 63 research papers that include 46 publications on Reactive and Quantum Dynamics Simulations, 13 publications on Machine Learning for Quantum Materials, and 4 publications on Quantum Computing.

2D Materials↗

TChem-atm (v2.0.0): scalable performance-portable multiphase atmospheric chemistry

We present TChem-atm, a performance-portable approach that enables efficient simulation of chemically detailed and multiphase atmospheric chemistry on modern heterogeneous computing architectures. Unlike previous efforts that rely on architecture-specific code or focus exclusively on gas-phase chemistry, TChem-atm supports fully coupled gas–aerosol systems with execution across CPUs, NVIDIA GPUs, and AMD GPUs through the Kokkos programming model. It integrates the flexible multiphase capabilities of the Community Atmospheric Model Chemistry Package (CAMP) with the high-performance kinetic routines of TChem, and includes automatic Jacobian construction with support for a range of stiff ODE solvers. In a proof-of-concept integration with the particle-resolved model PartMC, TChem-atm reproduces the existing PartMC–CAMP implementation within solver tolerances and delivers substantial GPU speedups, especially for large particle populations. Performance benchmarks reveal substantial speedups on GPU platforms, particularly for large particle populations, with consistent results across hardware backends. TChem-atm enables performance-portable execution across CPUs and GPUs, though optimal efficiency may require modest architecture-specific tuning (e.g., team and vector sizes), with up to a twofold improvement on the NVIDIA H100. It directly supports sectional and particle-resolved host models, while modal aerosol schemes require minor adaptation to provide particle-scale quantities such as representative diameters. By enabling chemically detailed, multiphase simulations with performance portability and host-model flexibility, TChem-atm facilitates the incorporation of advanced chemistry into atmospheric models.

Díaz-Ibarra, Oscar Homero [Sandia National Laborat↗

An Implicit Approach to Phase Field Modeling of Solidification for Additively Manufactured Alloys [Slides]

We are leveraging modern algorithms and computational science to provide a route to predictive simulation of microstructure evolution on emerging exascale architectures. We are utilizing the fastest supercomputers in the world for modeling and simulation of microstructure evolution for generation of data under AM conditions. Solidification conditions in AM can be tailored for the reliable design of materials to specific performance requirements. Developing computational tools to further characterize alloys and correlate the processing-structure-properties-performance (PSPP) relationship.

36 MATERIALS SCIENCE↗

An Open-Source Parallel EMT Simulation Framework

As the integration level of inverter-based resources (IBRs) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

An Open-Source Parallel EMT Simulation Framework: Preprint

As the integration level of inverter-based resources (IBR) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

Scoreboard

Emerging HPC machines have given rise to enhanced compute power that far outstrips the machine's ability to save large scale results for post-processing. To combat this, in situ data analysis techniques are slowly being adopted. With in situ data management favoring workflows composed of multiple simulations and analyses connected in transit on heterogeneous machines, scientists and engineers need a tool that enables them to create data extracts, visualizations, and interactively monitor and steer their simulations. Scoreboard Phase II is a next generation analysis software that supports composite in transit workflows on heterogeneous architectures and restores interactivity to in situ data analysis through simulation monitoring and computational steering. Scoreboard provides a simulation dashboard with graphs of metrics over time, controls for setting custom simulation steering parameters, controls for managing the set of data extracts being produced in the simulation, as well as the ability to explore data extracts, all from a web browser. Realizing the vision outlined in this project required research into making a system that integrates end to end from simulations all the way to the user. In situ tools generally suffer from complexity and excessive software dependencies. Scoreboard, by contrast, is easy to build and integrate into simulation codes and it provides first class FORTRAN support. The Scoreboard library is capable of in situ and in transit data analysis that can produce data extracts commonly needed for Computational Fluid Dynamics (CFD) analysis. Simulations can transparently stage data in transit to a Scoreboard Endpoint program, which can accept their data and produce the requested data extracts. This lets simulations return to their work while the Endpoint works on the analysis. Efficiently staging the data at scale was a topic of this research. Scoreboard provides the means to let the user manage data extracts and monitor/steer many simulations from a web browser. This area of the research focused on discovery of in transit network components to expose and control their steering parameters within an interactive browser-based user interface that includes: system topology, gathered metrics, notifications, dynamically-generated steering controls, and exploration of visualization data products.

Whitlock, BradJoseph [Intelligent Light] (00000001↗

Understanding power and energy utilization in large scale production physics simulation codes

Power is an often-cited reason for the move to advanced architectures on the path to Exascale computing. Here, this is due to practical considerations related to delivering enough power to successfully site and operate these machines, as well as concerns about energy usage while running large simulations. Since obtaining accurate power measurements can be challenging, it may be tempting to use the processor thermal design power (TDP) as a surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advanced technology systems at Lawrence Livermore and Sandia National Labs, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale Lawrence Livermore simulation codes are significantly more efficient than a simple processor TDP model might suggest.

HPC↗

Roadmap on methods and software for electronic structure based simulations in chemistry and materials

This Roadmap article provides a succinct, comprehensive overview of the state of electronic structure methods and software for molecular and materials simulations. Seventeen distinct sections collect insights by 51 leading scientists in the field. Each contribution addresses the status of a particular area, as well as current challenges and anticipated future advances, with a particular eye towards software related aspects and providing key references for further reading. Foundational sections cover density functional theory and its implementation in real-world simulation frameworks, Green's function based many-body perturbation theory, wave-function based and stochastic electronic structure approaches, relativistic effects and semiempirical electronic structure theory approaches. Subsequent sections cover nuclear quantum effects, real-time propagation of the electronic structure, challenges for computational spectroscopy simulations, and exploration of complex potential energy surfaces. The final sections summarize practical aspects, including computational workflows for complex simulation tasks, the impact of current and future high-performance computing architectures, software engineering practices, education and training to maintain and broaden the community, as well as the status of and needs for electronic structure based modeling from the vantage point of industry environments. Overall, the field of electronic structure software and method development continues to unlock immense opportunities for future scientific discovery, based on the growing ability of computations to reveal complex phenomena, processes and properties that are determined by the make-up of matter at the atomic scale, with high precision.

36 MATERIALS SCIENCE↗

Dynamic load balancing with enhanced shared-memory parallelism for particle-in-cell codes

Furthering our understanding of many of today’s interesting problems in plasma physics – including plasma based acceleration and magnetic reconnection with pair production due to quantum electrodynamic effects – requires large-scale kinetic simulations using particle-in-cell (PIC) codes. However, these simulations are extremely demanding, requiring that contemporary PIC codes be designed to efficiently use a new fleet of exascale computing architectures. To this end, the key issue of parallel load balance across computational nodes must be addressed. We discuss the implementation of dynamic load balancing by dividing the simulation space into many small, self-contained regions or ‘‘tiles,’’ along with shared-memory (e.g., OpenMP) parallelism both over many tiles and within single tiles. The load balancing algorithm can be used with three different topologies, including two space-filling curves. Here, we tested this implementation in the code Osiris and show low overhead and improved scalability with OpenMP thread number on simulations with both uniform load and severe load imbalance. Compared to other load-balancing techniques, our algorithm gives order-of-magnitude improvement in parallel scalability for simulations with severe load imbalance issues.

97 MATHEMATICS AND COMPUTING↗