Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

GPU-friendly surface model for Monte-Carlo detector simulations

The demands for Monte-Carlo simulation are drastically increasing with the Large Hadron Collider’s high-luminosity upgrade, and are expected to exceed the currently available compute resources. At the same time, modern high-performance computing has adopted powerful hardware accelerators, particularly GPUs. The AdePT and Celeritas projects aim to address the demanding computational needs by leveraging these heterogeneous computing architectures. While both have successfully ported realistic detector simulations to GPUs using the VecGeom library, the complexity of geometry modeling emerged as a bottleneck. Thread divergence and high register usage were degrading the GPU performance. Therefore, a new, GPU-friendly surface-based model has been introduced in the VecGeom library that decomposes the divergent code of the 3D primitive solids into simpler and more balanced surface algorithms. In this work, we present the latest developments, focusing on the additions required to efficiently model complex setups like the CMS Phase-2 geometry. This includes memory reduction techniques, and adding accelerating structures for faster traversal.

Diederichs, Severin [CERN]↗

Advances in 3D transient plasma dynamics and control through MHD and hybrid fluid-kinetic simulations with JOREK

Transient phenomena and their control are of high relevance in magnetic confinement fusion plasmas to guarantee a stable and safe plasma operation. Interpretative simulations can maximize the insights gained from experiments on present machines and predictive simulations can help in the preparation of design, mitigation techniques and operational scenarios for future devices. In this article, we provide an overview of recent advances and novel scientific results obtained with the 3D non-linear hybrid fluid-kinetic code JOREK, covering physics of plasma transients from the core to the scrape-off layer (SOL) both for tokamak and stellarator devices. Substantial progress was made in the physics understanding, model validation with experiments and experiment interpretation, thus, giving confidence for predictions to devices like DTT, ITER and DEMO. The topics addressed comprise a wide range: the edge physics of new operation scenarios and edge localized mode suppression; major disruptions with a focus on runaway electrons and vertical displacement events as well as disruption mitigation by shattered pellet injection; the physics mechanisms and operational limits of the flux pumping regime for sawtooth control; MHD limits of stellarators and work towards incorporating advanced edge/SOL/exhaust dynamics; continuing improvements of the code for more efficient hybrid simulations on conventional and accelerated high performance computing architectures.

disruptions↗

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗

Efficient Optimization of Energy Recovery From Geothermal Reservoirs With Recurrent Neural Network Predictive Models

Improving the long-term energy production performance of geothermal reservoirs can be accomplished by optimizing field development and management plans. Reliable prediction models, however, are needed to evaluate and optimize the performance of the underlying reservoirs under various operation and development strategies. In traditional frameworks, physics-based simulation models are used to predict the energy production performance of geothermal reservoirs. However, detailed simulation models are not trivial to construct, require a reliable description of the reservoir conditions and properties, and entail high computational complexity. Data-driven predictive models can offer an efficient alternative for use in optimization workflows. This paper presents an optimization framework for net power generation in geothermal reservoirs using a variant of the recurrent neural network (RNN) as a data-driven predictive model. The RNN architecture is developed and trained to replace the simulation model for computationally efficient prediction of the objective function and its gradients with respect to the well control variables. The net power generation performance of the field is optimized by automatically adjusting the mass flow rate of production and injection wells over 12 years, using a gradient-based local search algorithm. Two field-scale examples are presented to investigate the performance of the developed data-driven prediction and optimization framework. Furthermore, the prediction and optimization results from the RNN model are evaluated through comparison with the results obtained by using a numerical simulation model of a real geothermal reservoir.

15 GEOTHERMAL ENERGY↗

ERAS: Enabling the Integration of Real-World Intellectual Properties (IPs) in Architectural Simulators

Sandia National Laboratories is investigating scalable architectural simulation capabilities with a focus on simulating and evaluating highly scalable supercomputers for high performance computing applications. There is a growing demand for RTL model integration to provide the capability to simulate customized node architectures and heterogeneous systems. This report describes the first steps integrating the ESSENTial Signal Simulation Enabled by Netlist Transforms (ESSENT) tool with the Structural Simulation Toolkit (SST). ESSENT can emit C++ models from models written in FIRRTL to automatically generate components. The integration workflow will automatically generate the SST component and necessary interfaces to ’plug’ the ESSENT model into the SST framework.

97 MATHEMATICS AND COMPUTING↗

Berkeley eXtensible Environment (BXE) v3

The Berkeley eXtensible Environment (BXE) provides a cloud environment for hardware designers and computer architects to design, build, and simulate their custom architectures on an on-premises FPGA cluster. Utilizing the Chipyard, MoSAIC, and FireSim frameworks, users are provided an environment where they can assemble SoC designs from an existing library of components or import their own source code. Once their designs are ready, they can utilize the FireSim framework provided by BXE to deploy and simulate their designs on the FPGA. Users aren't limited to a single FPGA; they can deploy multiple instances across multiple FPGAs, acting like a rack of servers, or partition their large design across multiple FPGAs, ganging multiple FPGAs into a single simulated system.

Fatollahi-Fard, Farzin↗

Programmable Heisenberg interactions between Floquet qubits

Abstract The trade-off between robustness and tunability is a central challenge in the pursuit of quantum simulation and fault-tolerant quantum computation. In particular, quantum architectures are often designed to achieve high coherence at the expense of tunability. Many current qubit designs have fixed energy levels and consequently limited types of controllable interactions. Here by adiabatically transforming fixed-frequency superconducting circuits into modifiable Floquet qubits, we demonstrate an XXZ Heisenberg interaction with fully adjustable anisotropy. This interaction model can act as the primitive for an expressive set of quantum operations, but is also the basis for quantum simulations of spin systems. To illustrate the robustness and versatility of our Floquet protocol, we tailor the Heisenberg Hamiltonian and implement two-qubit iSWAP, CZ and SWAP gates with good estimated fidelities. In addition, we implement a Heisenberg interaction between higher energy levels and employ it to construct a three-qubit CCZ gate, also with a competitive fidelity. Our protocol applies to multiple fixed-frequency high-coherence platforms, providing a collection of interactions for high-performance quantum information processing. It also establishes the potential of the Floquet framework as a tool for exploring quantum electrodynamics and optimal control.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

Linear and Nonlinear Solvers for Simulating Multiphase Flow within Large-Scale Engineered Subsurface Systems

Simulation of multiphase flow in the subsurface is well-known to be computationally challenging. While there have been many studies that have explored approaches to overcoming these challenges, they often utilize relatively simple case studies. In this paper, we focus on the unique numerical challenges posed by modeling large-scale engineered subsurface systems, characterized by discrete features embedded in a heterogeneous natural subsurface setting. The man-made features such as shafts, tunnels, and barriers often cause multiple challenges in modeling the domain for multiphase porous media flow. This flow scenario can have a wide range of applications such as nuclear waste repositories, enhanced recovery of a petroleum reservoir, geothermal engineering, and carbon sequestration. An example of these severe numerical challenges is the case of performance assessment (PA) for Waste Isolation Pilot Plant (WIPP), the only operating deep geological repository in the US, which simulates extreme material properties of bedded salt rock formation and extreme contrast due to open excavation next to the formation. The models have extremes not only of permeability and porosity but also of the constitutive models needed for multiphase flow; additionally, they have process models like salt creep closure reducing porosity over time, fracturing in clay and anhydrite interbeds of the bedded salt, gas generation from the waste materials, and unintentional human borehole intrusions in some scenarios. Numerical simulations require the solution of coupled systems of nonlinear PDEs; in our work, we use the open-source simulator PFLOTRAN which is based on Finite Volume discretization. The solution of the nonlinear equations requires use of the Newton-Raphson iteration at each time step, which entails the solution of the linearized Jacobian system at each iteration. The effects of all the processes (i.e., large number of unknowns, highly nonlinear constitutive relations, large contrasts in material properties in short distances) lead to an ill-conditioned Jacobian matrix that severely challenges traditional linear solver, i.e., stabilized biconjugate gradient with block Jacobi incomplete LU preconditioner (BCGS-ILU) leading to non-convergence for traditional Newton-Raphson nonlinear solver causing unacceptably long computation time for each model. This paper presents linear solvers such as constrained pressure residual (CPR) two-stage preconditioner with alternate-block-factorization (ABF) and quasi- implicit pressure and explicit saturation (QIMPES) decouplers and flexible generalized residual solver (FGMRES). The new general-purpose nonlinear solver, Newton trust-region dogleg Cauchy (NTRDC), is also introduced to resolve extreme nonlinearities in the models. We demonstrate the effectiveness of each method relative to the default BCGS-Newton solver. The two best cases had nearly 50 times speed-up and achieved completion of a simulation in 14 hours that never completed due to non-convergence with the default solver. We also investigate the strong scalability of each method and discuss some of the deficiencies found for Block Jacobi preconditioner using parallel domain decomposition, and node packing effects of modern processor architecture.

Preconditioner, Nonlinear, Porous media, Multiphas↗

ARES v1.x - Performance Portable Tool to Simulate Supernovae based on Parthenon Framework

Historically, codes for simulating supernovae (such as Arepo, FLASH or LEAFS) have been at the forefront of scientific high-performance computing to the immense computational resources required for full 3D simulations. However, given the shift towards heterogenous HPC architectures, many current-generation codes are at the risk of losing their competitiveness as they are only designed to run on homogeneous CPU-only systems. There exist several efforts to enable these codes for GPU’s, however, these efforts only consider specific architectures or vendors (e.g., implement only CUDA or HIP), limiting themselves to a small range of exascale computing systems. Frameworks such as Kokkos aim to provide a framework which is agnostic of the targeted architecture, enabling the development of performant and portable code. In the Ares code, we develop a performance portable tool to simulate supernovae based on the Parthenon Framework, which in turn uses Kokkos in the background. Here, the Parthenon Framework provides an interface to the underlying mesh-refinement routines, which form the backbone of our code. In addition, we incorporate the already existing Singularity-EOS toolkit to provide us with various equations of state, primarily the Helmholtz equation of state. We also include the JINA Reaclib as a basis for our nuclear network solver. Finally, we implement a gravity solver to complete the required physics. This setup will provide us with a minimal code base to simulate supernova in a similar style to the tried-and-tested Arepo code, but in a futureproof performance portable framework.

Lim, Hyun↗

Ensuring statistical reproducibility of ocean model simulations in the age of hybrid computing

Novel high performance computing systems that feature hybrid architectures require large scale code refactoring to unravel underlying exploitable parallelism. Such redesign can often be accompanied with machine-precision changes as the order of computation cannot always be maintained. For chaotic systems like climate models, these round-off level differences can grow rapidly. Systematic errors may also manifest initially as machine-precision differences. Isolating genuine round off level differences from such errors remains a challenge. Here, we apply two-sample equality of distribution tests to evaluate statistical reproducibility of the ocean model component of US Department of Energy's Energy Exascale Earth System Model (E3SM). A 2-year control simulation ensemble is compared to a modified ensemble as a test case - after a known non-bit-for-bit change in a model component is introduced - to evaluate the null hypothesis that the two ensembles are statistically indistinguishable. To quantify the false negative rates of these tests, we conduct a formal power analysis using a targeted suite of short simulation ensembles. The ensemble suite contains several perturbed ensembles, each with a progressively different climate than the baseline ensemble - obtained by perturbing the magnitude of a single model tuning parameter, the Gent and McWilliams κ, in a controlled manner. The null hypothesis is evaluated for each of perturbed ensembles using these tests. The power analysis informs on the detection limits of the tests for given ensemble size allowing model developers to evaluate the impact of an introduced non-bit-for-bit change to the model.

Mahajan, Salil↗

Simulating non-native cubic interactions on noisy quantum machines

As a milestone for general-purpose computing machines, we demonstrate that quantum processors can be programed to efficiently simulate dynamics that are not native to the hardware. Moreover, on noisy devices without error correction, we show that simulation results are significantly improved when the quantum program is compiled using modular gates instead of a restricted set of standard gates. We demonstrate the general methodology by solving a cubic interaction problem, which appears in nonlinear optics, gauge theories, as well as plasma and fluid dynamics. To encode the non-native Hamiltonian evolution, we decompose the Hilbert space into a direct sum of invariant subspaces in which the nonlinear problem is mapped to a finite-dimensional Hamiltonian simulation problem. Furthermore, in a three-states example, the resultant unitary evolution is realized by a product of approximately 20 standard gates, using which approximately ten simulation steps can be carried out on state-of-the-art quantum hardware before results are corrupted by decoherence. In comparison, the simulation depth is improved by more than an order of magnitude when the unitary evolution is realized as a single cubic gate, which is compiled directly using optimal control. Alternatively, parametric gates may also be compiled by interpolating control pulses. Modular gates thus obtained provide high-fidelity building blocks for quantum Hamiltonian simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantum embedding theories to simulate condensed systems on quantum computers.

Quantum computers hold promise to improve the efficiency of quantum simulations of materials and to enable the investigation of systems and properties that are more complex than tractable at present on classical architectures. Here, we discuss computational frameworks to carry out electronic structure calculations of solids on noisy intermediate-scale quantum computers using embedding theories, and we give examples for a specific class of materials, that is, solid materials hosting spin defects. These are promising systems to build future quantum technologies, such as quantum computers, quantum sensors and quantum communication devices. Although quantum simulations on quantum architectures are in their infancy, promising results for realistic systems appear to be within reach.

Vorwerk, Christian↗

Hierarchical Materials from High Information Content Macromolecular Building Blocks: Construction, Dynamic Interventions, and Prediction

Hierarchical materials that exhibit order over multiple length scales are ubiquitous in nature. Because hierarchy gives rise to unique properties and functions, many have sought inspiration from nature when designing and fabricating hierarchical matter. More and more, however, nature’s own high-information content building blocks, proteins, peptides, and peptidomimetics, are being coopted to build hierarchy because the information that determines structure, function, and interfacial interactions can be readily encoded in these versatile macromolecules. Here, we take stock of recent progress in the rational design and characterization of hierarchical materials produced from high-information content blocks with a focus on stimuli-responsive and “smart” architectures. We also review advances in the use of computational simulations and data-driven predictions to shed light on how the side chain chemistry and conformational flexibility of macromolecular blocks drive the emergence of order and the acquisition of hierarchy and also on how ionic, solvent, and surface effects influence the outcomes of assembly. Furthermore, continued progress in the above areas will ultimately usher in an era where an understanding of designed interactions, surface effects, and solution conditions can be harnessed to achieve predictive materials synthesis across scale and drive emergent phenomena in the self-assembly and reconfiguration of high-information content building blocks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Symplectic machine learning model for fast simulation of space-charge effects

Symplectic simulation of space-charge effects is crucial for the design and operation of high-intensity particle accelerators. Traditional methods for simulating these effects are often computationally expensive, resulting in significant overhead. In this work, we introduce a generative model based on a U-Net architecture within a generative adversarial network framework to efficiently simulate space-charge effects. The model is trained to predict the transverse multiparticle space-charge Hamiltonian, which can be physically computed using a gridless spectral method. The one-step symplectic transverse transfer map for the particles is then obtained by differentiating the predicted Hamiltonian. Benchmarking results demonstrate that this generative model achieves an order of magnitude higher computational efficiency compared to the spectral method, providing a highly efficient alternative for simulating space-charge effects with a large number of particles. By maintaining symplecticity, the model effectively preserves the phase-space structure and mitigates nonphysical errors in long-term simulations. This model has been integrated into jutrack, a novel autodifferentiable accelerator modeling code developed in the julia programming language.

Beam code development & simulation techniques↗

Hybrid PDES Simulation of HPC Networks Using Zombie Packets

Although high-fidelity network simulations have proven to be reliable and cost-effective tools to peer into architectural questions for high-performance computing (HPC) networks, they incur a high resource cost. The time spent in simulating a single millisecond of network traffic in the highest detail can take hours, even for static, well-behaved traffic patterns such as uniform random. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should only be used when appropriate. Thus, there is a need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. Here, we present a surrogate model for HPC networks in which: packets bypass the network, while the network state is left untouched, i.e., suspended. To bypass the network, we use historical data to estimate the arrival time at which every packet should be scheduled at; to suspend the network, all in-flight packets are scheduled to arrive at their destinations, and are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. This light-weight surrogate obtained up to 76× speedup. Keeping the zombies in the network showed an increase in the accuracy of the high-fidelity simulation on restart when compared to restarting the network from an empty state.

HPC networks↗

Celeritas R&D Report: Accelerating Geant4

Celeritas is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a Graphics Processing Unit (GPU)-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.4, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling straightforward integration of Celeritas into the high energy physics (HEP) experimental frameworks CMSSW and ATLAS FullSimLight. On the Perlmutter supercomputer, Celeritas performs EM physics between 3× and 18× faster using the machine’s Nvidia GPUs compared to using only CPUs, corresponding to an electrical power efficiency up to a factor of 5. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.7–2.2× on GPU and 1.2× on CPU. In a CMS test application using tt¯ events and the prototype Run 4 configuration, compared to Geant4 CPU, Celeritas with a Nvidia A100 improves overall throughput up to a factor of 2.7× but cannot be efficiently shared with more than 8 cores.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Heterogeneous Computing

To leverage the increasing heterogeneity in modern computing resources, Geant4 incorporates advanced software tools and a task-based framework (G4Tasking) that enables efficient parallelism at event, sub-event, and track levels. Ongoing R&D efforts focus on integrating GPUs into high-energy physics (HEP) simulations, including optical photon simulation with Opticks/NVIDIA OptiX, offloading electromagnetic particle transport using G4HepEM/AdePT and Celeritas, and employing advanced surface-based geometry models such as VecGeom2.0 and ORANGE. As Geant4 continues evolving toward high-performance computing (HPC) and heterogeneous architectures, it remains a key tool for large-scale simulations in HEP and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗