Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Scattering phase shift in quantum mechanics on quantum computers

Here, we investigate the feasibility of extracting infinite volume scattering phase shift on quantum computers in a simple one-dimensional quantum mechanical model, using the formalism established in the work by Guo and Gasparian [Phys. Rev. D 108, 074504 (2023)] that relates the integrated correlation functions for a trapped system to the infinite volume scattering phase shifts through a weighted integral. The system is first discretized in a finite box with periodic boundary conditions, and the formalism in real time is verified by employing a contact interaction potential with exact solutions. Quantum circuits are then designed and constructed to implement the formalism on current quantum computing architectures. To overcome the fast oscillatory behavior of the integrated correlation functions in real-time simulation, different methods of postdata analysis are proposed and discussed. Test results on IBM hardware show that good agreement can be achieved with two qubits, but complete failure ensues with three qubits due to two-qubit gate operation errors and thermal relaxation errors.

Guo, Peng [Dakota State Univ., Madison, SD (United↗

Frontier (HPE Cray EX) Exascale Supercomputer at the Oak Ridge Leadership Computing Facility

Frontier is the HPE Cray EX exascale supercomputer deployed and operated by the Oak Ridge Leadership Computing Facility (OLCF) at Oak Ridge National Laboratory (ORNL). Frontier is designed for large-scale modeling, simulation, and AI workloads and is built from HPE Cray EX system architecture with AMD CPUs and AMD Instinct GPU accelerators connected by the HPE Slingshot interconnect. System composition (representative production configuration): Frontier is composed of approximately 74 cabinets with 128 compute nodes per cabinet (~9,400 compute nodes total). Each compute node contains one 64-core AMD EPYC CPU and four AMD Instinct MI250X GPUs. Nodes are connected using HPE Slingshot (Slingshot-200 class) networking with multiple NIC ports per node providing high injection bandwidth. Frontier is connected to the Orion parallel file system (multi-tier Lustre) providing a large, center-wide high-performance storage namespace. Operational context: Frontier entered public prominence as the first system to reach No. 1 on the TOP500 list in May 2022 (HPL benchmark), establishing the first widely recognized exascale-era performance milestone. The system supports DOE Office of Science mission workloads and enables leadership-class computational science and AI for open science users.

AMD EPYC↗

Three-Dimensional Permeability of Thick-Section Glass Fabric Reinforced Polymer Composite by Vacuum-Assisted Resin Infusion Molding

Abstract Determination of permeability of thick-section glass fabric preforms with fabric layers of different architectures is critical for manufacturing large, thick composite structures with complex geometry, such as wind turbine blades. The thick-section reinforcement permeability is inherently three-dimensional and needs to be obtained for accurate composite processing modeling and analysis. Numerical simulation of the liquid stage of vacuum-assisted resin infusion molding (VARIM) is important to advance the composite manufacturing process and reduce processing-induced defects. In this research, the 3D permeability of thick-section E-glass fabric reinforcement preforms is determined, and the results are validated by a comparison between flow front progressions from experiments and from numerical simulations using ansys fluent software. The orientation of the principal permeability axes were unknown prior to experiments. The approach used in this research differs from those in literature in that the through-thickness permeability is determined as a function of flow front positions along the principal axes and the in-plane permeabilities and is not dependent on the inlet radius. The approach was tested on reinforcements with fabric architectures which vary through-the-thickness direction, such as those in a spar cap of a wind turbine blade. The computational simulations of the flow-front progression through-the-thickness were consistent with experimental observations.

Engineering↗

PICSAR-QED: a Monte Carlo module to simulate strong-field quantum electrodynamics in particle-in-cell codes for exascale architectures

Abstract Physical scenarios where the electromagnetic fields are so strong that quantum electrodynamics (QED) plays a substantial role are one of the frontiers of contemporary plasma physics research. Investigating those scenarios requires state-of-the-art particle-in-cell (PIC) codes able to run on top high-performance computing (HPC) machines and, at the same time, able to simulate strong-field QED processes. This work presents the PICSAR-QED library, an open-source, portable implementation of a Monte Carlo module designed to provide modern PIC codes with the capability to simulate such processes, and optimized for HPC. Detailed tests and benchmarks are carried out to validate the physical models in PICSAR-QED, to study how numerical parameters affect such models, and to demonstrate its capability to run on different architectures (CPUs and GPUs). Its integration with WarpX, a state-of-the-art PIC code designed to deliver scalable performance on upcoming exascale supercomputers, is also discussed and validated against results from the existing literature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Performance Evaluation of a Two-Dimensional Flood Model on Heterogeneous High-Performance Computing Architectures

This paper describes the implementation of a two-dimensional hydrodynamic flood model with two different numerical schemes on heterogeneous high-performance computing architectures. Both schemes were able to solve the nonlinear hyperbolic shallow water equations using an explicit upwind first-order approach on finite differences and finite volumes, respectively, and were conducted using MPI and CUDA. Four different test cases were simulated on the Summit supercomputer at Oak Ridge National Laboratory. Both numerical schemes scaled up to 128 nodes (768 GPUs) with a maximum 98.2x speedup of over 1 GPU. The lowest run time for the 10 day Hurricane Harvey event simulation at 5 meter resolution (272 million grid cells) was 50 minutes. GPUDirect communication proved to be more convenient than the standard communication strategy. Both strong and weak scaling are shown.

Sharif, Md Bulbul↗

COSMIC DAWN: Distributed Analysis of Wireless at Nextscale

Distributed Analysis of Wireless at Nextscale (DAWN) is a novel simulation framework for large-scale design-space exploration (DSE) of unmodified software-defined radio (SDR) applications interacting in a scalable, high-fidelity, virtual physics environment. The software-defined nature of the coupled software-physics simulation leverages hardware emulation to permit in-depth examination and modification of not only the electromagnetic environment, including each signal in flight, but also the precise state of system software and components. DAWN supports modular, customizable physics environments allowing realistic propagation effects so that computationally efficient empirical models, reduced order/surrogate models, or large-scale, high-fidelity, site-specific simulations can be used as a propagation medium based on scenario requirements. This paper introduces DAWN’s design and initial implementation, detailing key architectural components, including the Physics Realization Engine (PhyRE), Runtime Infrastructure for Simulation Environments (RISE), and the design space exploration (DSE) suite. It concludes with demonstrations using unmodified 4G/LTE software available from srsRAN on computing resources ranging from a small cluster to ORNL’s Frontier Exascale system.

Wise, Mike [ORNL] (ORCID:0000000266120641)↗

Framework for Large-scale Implementation of Wholesale-Retail Transactive Control Mechanism

Transactive energy is a control technique that uses market mechanisms to achieve desired control objectives. Several simulation studies and field demonstrations were carried out in recent years, but all focused on small scale systems and purpose-built simplified models which are not always capable of pointing out all the advantages and shortcomings of transac- tive energy methods. This work describes a co-simulation framework built on a hierarchical control architecture that allows for conducting studies of the impacts of a very large-scale deployment of transactive energy. The hierarchical transactive control architecture adopted in this work helps with alleviating computation and communication burden to facilitate a more effective large scale real-time market operation among device level resources and the system level operators. The co-simulation framework is evaluated an integrated power sys- tem model of unprecedented scale composed of the Western Electricity Coordination Council (WECC) transmission system with tens of thousands of distribution systems deployed with flexible device-level distributed energy resources (DERs) using off-the-shelf simulators.

Transactive energy, market-based controls, DER int↗

Tusas: A fully implicit parallel approach for coupled phase-field equations

In this study, we develop a fully-coupled, fully-implicit approach for phase-field modeling of solidification in metals and alloys. Predictive simulation of solidification in pure metals and metal alloys remains a significant challenge in the field of materials science, as microstructure formation during the solidification process plays a critical role in the properties and performance of the solid material. Our simulation approach consists of a finite element spatial discretization of the fully-coupled nonlinear system of partial differential equations at the microscale, which is treated implicitly in time with a preconditioned Jacobian-free Newton-Krylov method. The approach is algorithmically scalable as well as efficient due to an effective preconditioning strategy based on algebraic multigrid and block factorization. We implement this approach in the open-source Tusas framework, which is a general, flexible tool developed in C++ for solving coupled systems of nonlinear partial differential equations. The performance of our approach is analyzed in terms of algorithmic scalability and efficiency, while the computational performance of Tusas is presented in terms of parallel scalability and efficiency on emerging heterogeneous architectures. We demonstrate that modern algorithms, discretizations, and computational science, and heterogeneous hardware provide a robust route for predictive phase-field simulation of microstructure evolution during additive manufacturing.

97 MATHEMATICS AND COMPUTING↗

High Performance, High Fidelity: A GPU‐Accelerated Doubly‐Periodic Configuration of the Simple Cloud‐Resolving E3SM Atmosphere Model Version 1 (DP‐SCREAMv1)

The development of the Simplified Cloud Resolving Energy Exascale Earth System Atmosphere Model (SCREAMv1) enables global storm-resolving simulations on modern GPU-based supercomputers. However, the high computational cost of SCREAMv1 limits its routine use for process-level studies, creating a need for efficient proxy configurations. This study addresses this gap by introducing DP-SCREAMv1, a doubly periodic cloud-resolving model designed to be fully consistent with SCREAMv1 while enabling high-resolution, long-duration simulations at significantly reduced computational expense by simulating a limited doubly periodic domain rather than the entire globe. Built on a C++/Kokkos architecture, DP-SCREAMv1 achieves exceptional performance scalability on GPU systems and includes a rich library of cases for validation and scientific exploration. In this work, we demonstrate short wall-clock times at SCREAMv1's default resolution and show that DP-SCREAMv1 supports routine execution of large-domain, high-resolution experiments that were previously challenging in practice. Furthermore, we show that DP-SCREAMv1 enables routine execution of “Giga-LES” style simulations and facilitates large-domain, high-resolution simulations that were recently considered burdensome to perform. These results document an efficient, fully consistent process-level configuration for SCREAMv1 (DP-SCREAMv1) and illustrate its use for long-duration and large-domain experiments at cloud-resolving to eddy-permitting resolution.

Environmental sciences↗

Validation of time-dependent shift using the pulsed sphere benchmarks

The detailed behavior of neutrons in a rapidly changing time-dependent physical system is a challenging computational physics problem, particularly when using Monte Carlo methods on heterogeneous high-performance computing architectures. A small number of algorithms and code implementations have been shown to be performant for time-independent (fixed source and k-eigenvalue) Monte Carlo, and there are existing simulation tools that successfully solve the time-dependent Monte Carlo problem on smaller computing platforms. To bridge this gap, a time-dependent version of ORNL’s Shift code has been recently developed. Shift’s history-based algorithm on CPUs, and its event-based algorithm on GPUs, have both been observed to scale well to very large numbers of processors, which motivated the extension of this code to solve time-dependent problems. The validation of this new capability requires a comparison with time-dependent neutron experiments. Lawrence Livermore National Laboratory’s (LLNL) pulsed sphere benchmark experiments were simulated in Shift to validate both the time-independent as well as new time-dependent features recently incorporated into Shift. A suite of pulsed-sphere models was simulated using Shift and compared to the available experimental data and simulations with MCNP. Overall results indicate that Shift accurately simulates the pulsed sphere benchmarks, and that the new time-dependent modifications of Shift are working as intended. Validated exascale neutron transport codes are essential for a wide variety of future multiphysics applications.

Palmer, Camille J.↗

Towards string order melting of spin-1 particle chains in superconducting transmons using optimal control

Utilizing optimal control to simulate a model Hamiltonian is an emerging strategy that leverages the intrinsic physics of a device with digital quantum simulation methods. Here we evaluate optimal control for probing the nonequilibrium properties of symmetry-protected topological (SPT) states simulated with superconducting hardware. Assuming a tunable transmon architecture, we cast the evolution of these SPT states as a series of one- and two-site pulse optimization problems that are solved in the presence of leakage constraints. From the generated pulses, we classically simulate the time-dependent melting of the perturbed SPT string order across a six-site model with an average state infidelity of 10 -3 . The feasibility of these pulses as well as their efficient application indicate that high-fidelity simulations of string order melting are within reach of current quantum computing systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Efficient phase-factor evaluation in quantum signal processing

Quantum signal processing (QSP) is a powerful quantum algorithm to exactly implement matrix polynomials on quantum computers. Asymptotic analysis of quantum algorithms based on QSP has shown that asymptotically optimal results can in principle be obtained for a range of tasks, such as Hamiltonian simulation and the quantum linear system problem. A further benefit of QSP is that it uses a minimal number of ancilla qubits, which facilitates its implementation on near-to-intermediate term quantum architectures. However, there is so far no classically stable algorithm allowing computation of the phase factors that are needed to build QSP circuits. Existing methods require the use of variable precision arithmetic and can only be applied to polynomials of a relatively low degree. We present here an optimization-based method that can accurately compute the phase factors using standard double precision arithmetic operations. We demonstrate the performance of this approach with applications to Hamiltonian simulation, eigenvalue filtering, and quantum linear system problems. Furthermore, our numerical results show that the optimization algorithm can find phase factors to accurately approximate polynomials of a degree larger than 10000 with errors below 10 -12 .

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

PeleLMeX: an AMR Low Mach Number Reactive Flow Simulation Code without level sub-cycling

PeleLMeX simulates chemically reacting low Mach number flows with block-structured adaptive mesh refinement (AMR). The code is built upon the AMReX library, which provides the underlying data structures and tools to manage and operate on them across massively parallel computing architectures. PeleLMeX algorithmic features are inherited from its predecessor PeleLM but key improvements allow representation of more complex physical processes. Together with its compressible flow counterpart PeleC, the thermo-chemistry library PelePhysics and the multi-physics library PeleMP, it forms the Pele suite of open-source reactive flow simulation codes.

97 MATHEMATICS AND COMPUTING↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗

Jet classification using high-level features from anatomy of top jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Modeling bacterial microcompartment architectures for enhanced cyanobacterial carbon fixation

The carboxysome is a bacterial microcompartment (BMC) which plays a central role in the cyanobacterial CO 2 -concentrating mechanism. These proteinaceous structures consist of an outer protein shell that partitions Rubisco and carbonic anhydrase from the rest of the cytosol, thereby providing a favorable microenvironment that enhances carbon fixation. The modular nature of carboxysomal architectures makes them attractive for a variety of biotechnological applications such as carbon capture and utilization. In silico approaches, such as molecular dynamics (MD) simulations, can support future carboxysome redesign efforts by providing new spatio-temporal insights on their structure and function beyond in vivo experimental limitations. However, specific computational studies on carboxysomes are limited. Fortunately, all BMC (including the carboxysome) are highly structurally conserved which allows for practical inferences to be made between classes. Here, we review simulations on BMC architectures which shed light on (1) permeation events through the shell and (2) assembly pathways. These models predict the biophysical properties surrounding the central pore in BMC-H shell subunits, which in turn dictate the efficiency of substrate diffusion. Meanwhile, simulations on BMC assembly demonstrate that assembly pathway is largely dictated kinetically by cargo interactions while final morphology is dependent on shell factors. Overall, these findings are contextualized within the wider experimental BMC literature and framed within the opportunities for carboxysome redesign for biomanufacturing and enhanced carbon fixation.

59 BASIC BIOLOGICAL SCIENCES↗

NEAMS Workbench Status and Capabilities

The mission of the US Department of Energy’s Nuclear Energy Advanced Modeling and Simulation (NEAMS) program is to develop, apply, deploy, and support state-of-the-art predictive modeling and simulation tools for the design and analysis of current and future nuclear energy systems. This is accomplished by using computing architectures that range from laptops to leadership-class facilities. The NEAMS Workbench is a new initiative that will facilitate the transition from conventional tools to highfidelity tools by providing a common user interface for model creation, review, execution, output review, and visualization for integrated codes. The Workbench can use common user input, including engineering-scale specifications that are expanded into application-specific input requirements through the use of customizable templates. The templating process can enable multifidelity analysis of a system from a common set of input data. Additionally, the common user input processor can provide an enhanced alternative application input that provides additional conveniences compared with native input, especially for legacy codes. Expansion of the integrated codes and application templates available in the Workbench will broaden the NEAMS user community and will facilitate system analysis and design. Current and planned capabilities of the NEAMS Workbench are detailed herein.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Position Papers for the 2024 ASCR Workshop on Neuromorphic Computing for Science

Engineering novel neuromorphic computing systems with functionalities, capabilities, and energy efficiency similar to biological brains is one of the most exciting and challenging scientific endeavors of our time. This workshop aims to identify key research needs, challenges, and next steps necessary to develop biologically realistic neuromorphic circuits primitives that capture the functionality of neural systems found in nature. Moreover, simulating neuromorphic computing primitives integrated into networks will be key to under standing their behavior at scale, particularly for those computing architectures where full-scale commercial fabrication is not yet readily accessible. Appropriate neuroscience datasets and metrics will have to be established to vet proposed neuromorphic circuits.

97 MATHEMATICS AND COMPUTING↗