Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Tidal disruption discs formed and fed by stream–stream and stream–disc interactions in global GRHD simulations

When a star passes close to a supermassive black hole (BH), the BH’s tidal forces rip it apart into a thin stream, leading to a tidal disruption event (TDE). In this work, we study the post-disruption phase of TDEs in general relativistic hydrodynamics (GRHD) using our GPU-accelerated code h-amr. We carry out the first grid-based simulation of a deep-penetration TDE (β = 7) with realistic system parameters: a black hole-to-star mass ratio of 10 6 , a parabolic stellar trajectory, and a non-zero BH spin. We also carry out a simulation of a tilted TDE whose stellar orbit is inclined relative to the BH midplane. We show that for our aligned TDE, an accretion disc forms due to the dissipation of orbital energy with ~20 percent of the infalling material reaching the BH. The dissipation is initially dominated by violent self-intersections and later by stream–disc interactions near the pericentre. The self-intersections completely disrupt the incoming stream, resulting in five distinct self-intersection events separated by approximately 12 h and a flaring in the accretion rate. We also find that the disc is eccentric with mean eccentricity e ≈ 0.88. For our tilted TDE, we find only partial self-intersections due to nodal precession near pericentre. Although these partial intersections eject gas out of the orbital plane, an accretion disc still forms with a similar accreted fraction of the material to the aligned case. These results have important implications for disc formation in realistic tidal disruptions. For instance, the periodicity in accretion rate induced by the complete stream disruption may explain the flaring events from Swift J1644+57.

79 ASTRONOMY AND ASTROPHYSICS↗

DAmodel: hierarchical Bayesian modelling of DA white dwarfs for spectrophotometric calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards (⁠16.5 < V , 19.5⁠) alongside three CALSPEC standards, from 912 Å to 32 μm. The framework is the first of its kind to jointly infer photometric zero points and WD parameters (surface gravity log g⁠, effective temperature T eff ⁠, extinction A V ⁠, dust relation parameter R V ) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy, and ground-based optical spectroscopy to sub-per cent precision. Photometric residuals for the sample are the lowest yet yielding < 0.004 mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory’s Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

methods: statistical↗

Multitarget Rydberg gates via spatial blockade engineering

Multi-target gates offer the potential to reduce gate depth in syndrome extraction for quantum error correction. Although neutral-atom quantum computers have demonstrated native multi-qubit gates, existing approaches that avoid additional control or multiple atomic species have been limited to single-target gates. We propose single-control-multi-target CZ^n gates on a single-species neutral-atom platform that require no extra control and have gate durations comparable to standard CZ gates. Our approach leverages tailored interatomic distances to create an asymmetric blockade between the control and target atoms. Using a GPU-accelerated pulse synthesis protocol, we design smooth control pulses for CZZ and CZZZ gates, achieving fidelities of up to 99.55% and $99.24\%$, respectively, even in the presence of simulated atom placement errors and Rydberg-state decay. Our approach is most effective for N=2 (CZZ) and N=3 targets (CZZZ); for larger N, increasing spatial crowding of the targets introduces significant challenges for maintaining the required blockade asymmetry. This work presents a practical path to implementing low-overhead multi-target gates in single-species neutral-atom systems, significantly reducing the resource overhead for syndrome extraction. To motivate the impact of these gates, we apply a greedy scheduling algorithm and we demonstrate that our proposed gates can reduce the number of atom reconfiguration costs by up to 50% for color code syndrome extraction of code distances greater than 5.

Stein, Samuel A.↗

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗

Supercomputing Pipelines Search for Therapeutics Against COVID-19

The urgent search for drugs to combat SARS-CoV-2 has included the use of supercomputers. The use of general-purpose graphical processing units (GPUs), massive parallelism, and new software for high-performance computing (HPC) has allowed researchers to search the vast chemical space of potential drugs faster than ever before. We developed a new drug discovery pipeline using the Summit supercomputer at Oak Ridge National Laboratory to help pioneer this effort, with new platforms that incorporate GPU-accelerated simulation and allow for the virtual screening of billions of potential drug compounds in days compared to weeks or months for their ability to inhibit SARS-COV-2 proteins. Here, this effort will accelerate the process of developing drugs to combat the current COVID-19 pandemic and other diseases.

60 APPLIED LIFE SCIENCES↗

Transformational Regional-Scale Earthquake Simulations with the DOE EarthQuake SIMulation Exascale Framework

Earthquakes present worldwide risk to economic and human safety. The 2023 earthquakes in Turkiye provided a reminder of the potential for catastrophic consequences with 50,700 deaths and 15.7 million people affected. The ability to predict ground motions and infrastructure damage for earthquakes continues to be a challenging problem for scientists and engineers. Until now, estimates of ground motions have been performed empirically by looking at sparse data from past earthquakes. This approach can provide statistical information on intensity amplitudes but cannot inform site-specific ground motions essential to developing the most effective resilience. Interest has grown in large-scale computational models to simulate earthquakes at regional scale. The U.S. Department of Energy EarthQuake SIMulation (EQSIM) framework was developed for regional-scale earthquake simulations at unprecedented fidelity, taking advantage of emerging GPU-accelerated systems. This article describes the EQSIM workflow and demonstrates regional-scale simulations with the new computational capability available to scientists in their quest to mitigate future disasters.

58 GEOSCIENCES↗

Continuous Emulation and Multiscale Visualization of Traffic Flow Using Stationary Roadside Sensor Data

With the advent of the next-generation traffic monitoring systems, there has been a significant increase in the spatial-temporal resolution of vehicle mobility data in many cities. Effective analysis and visualization of such data can provide transportation planners with data-driven insights, which can facilitate the understanding of multiscale traffic dynamics. In this paper, we present a web-based traffic emulator for emulating and visualizing near-real-time and historical traffic flows on highways using data from road-side sensors. To construct a continuous traffic flow, the emulator adopts an analytical pipeline that can (a) integrate traffic data collected from discrete road-side radar detection sensors, (b) interpolate traffic conditions (vehicle speed and volume) on unmeasured road segments based on traffic flow theory, and (c) generate lane-specific vehicle trajectories and movements using a mathematically optimized representation of the road network. Our app also provides an integrated visual workflow that allows users to explore the interconnected traffic dynamics using an appropriate traffic flow visualization selected based on the level of detail. We devise two innovative geo-visualization techniques that utilize an animated strips-network representation and a lane usage matrix to visualize lane performances. To ensure a smooth emulation of large-scale traffic flow in an easy-to-access web environment, we implement the emulator using client-side GPU-accelerated techniques. Lastly, we close with a case study that visualizes traffic dynamics of two scenarios - an afternoon peak hour and a traffic accident - in Chattanooga, Tennessee. Our app visualizes the responses of traffic dynamics during different traffic conditions, and to the presence of the traffic accident at different spatial scales.

42 ENGINEERING↗

Performance of a geometric deep learning pipeline for HL-LHC particle tracking

The Exa.TrkX project has applied geometric learning concepts such as metric learning and graph neural networks to HEP particle tracking. Exa.TrkX’s tracking pipeline groups detector measurements to form track candidates and filters them. The pipeline, originally developed using the TrackML dataset (a simulation of an LHC-inspired tracking detector), has been demonstrated on other detectors, including DUNE Liquid Argon TPC and CMS High-Granularity Calorimeter. This paper documents new developments needed to study the physics and computing performance of the Exa.TrkX pipeline on the full TrackML dataset, a first step towards validating the pipeline using ATLAS and CMS data. The pipeline achieves tracking efficiency and purity similar to production tracking algorithms. Crucially for future HEP applications, the pipeline benefits significantly from GPU acceleration, and its computational requirements scale close to linearly with the number of particles in the event.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Optimization and Portability of a Fusion OpenACC-based FORTRAN HPC Code from NVIDIA to AMD GPUs

NVIDIA has been the main provider of GPU hardware in HPC systems for over a decade. Most applications that benefit from GPUs have thus been developed and optimized for the NVIDIA software stack. Recent exascale HPC systems are, however, introducing GPUs from other vendors, e.g. with the AMD GPU-based OLCF Frontier system just becoming available. AMD GPUs cannot be directly accessed using the NVIDIA software stack, and require a porting effort by the application developers. This paper provides an overview of our experience porting and optimizing the CGYRO code, a widely-used fusion simulation tool based on FORTRAN with OpenACC-based GPU acceleration. While the porting from the NVIDIA compilers was relatively straightforward using the CRAY compilers on the AMD systems, the performance optimization required more fine-tuning. In the optimization effort, we uncovered code sections that had performed well on NVIDIA GPUs, but were unexpectedly slow on AMD GPUs. After AMD-targeted code optimizations, performance on AMD GPUs has increased to meet our expectations. Modest speed improvements were also seen on NVIDIA GPUs, which was an unexpected benefit of this exercise.

Sfiligoi, Igor↗

Revisiting Temporal Blocking Stencil Optimizations

Iterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark.

Zhang, Lingqi↗

SuperNeuro: A Fast and Scalable Simulator for Neuromorphic Computing

In many neuromorphic workflows, simulators play a vital role for important tasks such as training spiking neural networks, running neuroscience simulations, and designing, implementing, and testing neuromorphic algorithms. Currently available simulators cater to either neuroscience workflows (e.g., NEST and Brian2) or deep learning workflows (e.g., BindsNET). Problematically, the neuroscience-based simulators are slow and not very scalable, and the deep learning-based simulators do not support certain functionalities that are typical of neuromorphic workloads (e.g., synaptic delay). In this paper, we address this gap in the literature and present SuperNeuro, which is a fast and scalable simulator for neuromorphic computing capable of both homogeneous and heterogeneous simulations as well as GPU acceleration. We also present preliminary results that compare SuperNeuro to widely used neuromorphic simulators such as NEST, Brian2, and BindsNET in terms of computation times. We demonstrate that SuperNeuro can be approximately 10×--300× faster than some of the other simulators for small sparse networks. On large sparse and large dense networks, SuperNeuro can be approximately 2.2×--3.4× faster than the other simulators, respectively.

Date, Prasanna↗

Optimizing Communication in 2D Grid-Based MPI Applications at Exascale

The new reality of exascale computing faces many challenges in achieving optimal performance on large numbers of nodes. A key challenge is the efficient utilization of the message-passing interface (MPI), a critical component for process communication. This paper explores communication optimization strategies to harness the GPU-accelerated architectures of these supercomputers. We focus on MPI applications where processors form a two-dimensional process grid, a common arrangement in applications involving dense matrix operations. This configuration offers a unique opportunity to implement innovative strategies to improve performance and maintain effective load distribution. We study two applications— Dist-FW (Apsp:all-pair-shortest-path) and HPL-MxP (LU factorization with Mixed precision)—on two accelerated systems: Summit (IBM Power and NVIDIA V100) and Frontier (AMD EPYC and MI250X). These supercomputers are operated by the Oak Ridge Leadership Computing Facility (OLCF) and are currently ranked #1 and #5 on the Top500 list. We show how to scale up both applications to exascale levels and tackle the MPI challenges related to implementation, synchronization, and performance. We also compare the performance of several communication strategies at an unprecedented scale. Accurately predicting application performance becomes crucial for cost reduction as the computational scale grows. To address this, we suggest a hyperbolic model as a better alternative to the traditional one-sided asymptotic model for predicting future application performance at such large scales.

Lu, Hao↗

DiHydrogen

DiHydrogen is the second version of the Hydrogen fork of the well-known distributed linear algebra library, Elemental. DiHydrogen is a GPU-accelerated distributed multilinear algebra interface with a particular emphasis on the needs of the scalable distributed deep learning training and inference. DiHydrogen is part of the Livermore Big Artificial Neural Network (LBANN) software stack.

Maruyama, Naoya↗

MEUMAPPS (C++ Version)

Many materials, metal alloys in particular, have features on the on micrometer or nanometer scale that have a large impact on the properties of the material. These features are known as the microstructure of the material. Understanding why and how the microstructure forms in a material is of fundamental scientific interest as well as of significant technological interest. The capability to predict microstructure evolution in a material allows the intentional design of microstructures and hence the intentional design of material properties. The phase-field method is one of the leading methods for predicting microstructure evolution. One of the most significant problems for phase-field models is their computational expense. Even limited phase-field simulations can easily require thousands of CPU core-hours to complete, which significantly limits their use. This code provides both a general framework for creating scalable, GPU-accelerated phase-field model applications as well as several applications themselves. The code is capable of using hundreds of GPUs efficiently, which greatly reduces the time required to perform simulations. The code is written with an emphasis on performance portability, that is the ability for the code to run efficiently on a number of different computing architectures without modification of the source code. The performance portability of this code is primarily enabled through the use of two libraries, Kokkos (performance portable data structures and execution patterns) and heFFTe (performance portable distributed 3D fast Fourier transforms). The code consists of a core library, applications, and tests. The core library includes shared functionality between applications. This includes interfaces with fast Fourier transform (FFT) libraries such as heFFTe, data structures based on Kokkos, file input and output capabilities, and a solver for infinitesimal strain mechanical equilibrium problems. Five applications are included in the code. The flagship application is the MEUMAPPS-SS application, which implements the Kim-Kim-Suzuki phase-field model for precipitation for an arbitrary number of phases and components in a metal alloy. Five simpler applications are also included that solve the Eshelby inclusion problem, Allen-Cahn equation, the coupled Allen-Cahn and diffusion equations, and the Cahn-Hilliard equation. The code includes two applications to solve the Cahn-Hilliard equation, one with constant-step-size first-order time integration and the second with adaptive high-order time integration.

DeWitt, Stephen [Oak Ridge National Lab. (ORNL), O↗

PeleLMeX [SWR-22-48]

PeleLMeX is a solver for high fidelity reactive flow simulations, namely direct numerical simulation (DNS) and large eddy simulation (LES). The solver combines a low Mach number approach, adaptive mesh refinement (AMR), embedded boundary (EB) geometry treatment and high performance computing (HPC) to provide a flexible tool to address research questions on platforms ranging from small workstations to the world's largest GPU-accelerated supercomputers. PeleLMeX has been used to study complex flame/turbulence interactions in RCCI engines and hydrogen combustion or the effect of sustainable aviation fuel on gas turbine combustion. PeleLMeX is part of the Pele combustion Suite (https://amrex-combustion.github.io/)

Day, Marcus↗

Swift

Swift is a fast Fourier transform based spectral solver based on the MOOSE framework. It supports GPU accelerated semi-implicit solves of partial differential equations, such as those used for phase field mesoscale microstructure evolution.

Schwen, Daniel [Idaho National Laboratory (INL), I↗

OpenFerro v0.1.0

OpenFerro is a Python package for on-lattice atomistic dynamics simulation of ferroic materials. OpenFerro is based on JAX, a high-performance linear algebra package supporting auto-differentiation and GPU acceleration. OpenFerro is designed to minimize the effort required to build on-lattice Hamiltonian models, and to perform molecular dynamics (MD) and Landau-Lifshitz-Gilbert simulations. Unlike existing codes, OpenFerro provides a unified interface to model different types of local order parameters.

Xie, Pinchen [Lawrence Berkeley National Laborator↗