Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Linear solvers for power grid optimization problems: A review of GPU-accelerated linear solvers

The linear equations that arise in interior methods for constrained optimization are sparse symmetric indefinite, and they become extremely ill-conditioned as the interior method converges. These linear systems present a challenge for existing solver frameworks based on sparse LU or LDL T decompositions. Here, we benchmark five well known direct linear solver packages on CPU- and GPU-based hardware, using matrices extracted from power grid optimization problems. The achieved solution accuracy varies greatly among the packages. None of the tested packages delivers significant GPU acceleration for our test cases. For completeness of the comparison we include results for MA57, which is one of the most efficient and reliable CPU solvers for this class of problem.

97 MATHEMATICS AND COMPUTING↗

GPU-Accelerated Machine Learning Inference as a Service for Computing in Neutrino Experiments

Machine learning algorithms are becoming increasingly prevalent and performant in the reconstruction of events in accelerator-based neutrino experiments. These sophisticated algorithms can be computationally expensive. At the same time, the data volumes of such experiments are rapidly increasing. The demand to process billions of neutrino events with many machine learning algorithm inferences creates a computing challenge. We explore a computing model in which heterogeneous computing with GPU coprocessors is made available as a web service. The coprocessors can be efficiently and elastically deployed to provide the right amount of computing for a given processing task. With our approach, Services for Optimized Network Inference on Coprocessors (SONIC), we integrate GPU acceleration specifically for the ProtoDUNE-SP reconstruction chain without disrupting the native computing workflow. With our integrated framework, we accelerate the most time-consuming task, track and particle shower hit identification, by a factor of 17. This results in a factor of 2.7 reduction in the total processing time when compared with CPU-only production. For this particular task, only 1 GPU is required for every 68 CPU threads, providing a cost-effective solution.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates

Smoothed Particle Hydrodynamics (SPH) is essential for modeling complex large-deformation problems across various applications, requiring significant computational power. A major portion of SPH computation time is dedicated to the Nearest Neighboring Particle Search (NNPS) process. While advanced NNPS algorithms have been developed to enhance SPH efficiency, the potential efficiency gains from modern computation hardware remain underexplored. Here, this study investigates the impact of GPU parallel architecture, low-precision computing on GPUs, and GPU memory management on NNPS efficiency. Our approach employs a GPU-accelerated mixed-precision SPH framework, utilizing low precision float-point 16 (FP16) for NNPS while maintaining high precision for other components. To ensure FP16 accuracy in NNPS, we introduce a Relative Coordinated-based Link List (RCLL) algorithm, storing FP16 relative coordinates of particles within background cells. Our testing results show three significant speedup rounds for CPU-based NNPS algorithms. The first comes from parallel GPU computations, with up to a 1000x efficiency gain. The second is achieved through low-precision GPU computing, where the proposed FP16-based RCLL algorithm offers a 1.5x efficiency improvement over the FP64-based approach on GPUs. By optimizing GPU memory bandwidth utilization, the efficiency of the FP16 RCLL algorithm can be further boosted by 2.7x, as demonstrated in an example with 1 million particles. Our code is released at https://github.com/pnnl/lpNNPS4SPH.

97 MATHEMATICS AND COMPUTING↗

Discovery of a Gamma-Ray Black Widow Pulsar by GPU-accelerated Einstein@Home

We report the discovery of 1.97 ms period gamma-ray pulsations from the 75 minute orbital-period binary pulsar now named PSR J1653-0158. The associated Fermi Large Area Telescope gamma-ray source 4FGL J1653.6-0158 has long been expected to harbor a binary millisecond pulsar. Despite the pulsar-like gamma-ray spectrum and candidate optical/X-ray associations—whose periodic brightness modulations suggested an orbit—no radio pulsations had been found in many searches. The pulsar was discovered by directly searching the gamma-ray data using the GPU-accelerated Einstein@Home distributed volunteer computing system. The multidimensional parameter space was bounded by positional and orbital constraints obtained from the optical counterpart. More sensitive analyses of archival and new radio data using knowledge of the pulsar timing solution yield very stringent upper limits on radio emission. Any radio emission is thus either exceptionally weak, or eclipsed for a large fraction of the time. The pulsar has one of the three lowest inferred surface magnetic-field strengths of any known pulsar with B surf ≈ 4 × 10 7 G. The resulting mass function, combined with models of the companion star's optical light curve and spectra, suggests a pulsar mass gsim 2 M ⊙ . The companion is lightweight with mass ~0.01 M ⊙ , and the orbital period is the shortest known for any rotation-powered binary pulsar. This discovery demonstrates the Fermi Large Area Telescope's potential to discover extreme pulsars that would otherwise remain undetected.

79 ASTRONOMY AND ASTROPHYSICS↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

thornado-transport: Anderson- and GPU-accelerated nonlinear solvers for neutrino-matter coupling

Algorithms for neutrino-matter coupling in core-collapse supernovae (CCSNe) are investigated in the context of a spectral two-moment model, which is discretized in space with the discontinuous Galerkin method, integrated in time with implicit-explicit (IMEX) methods, and implemented in the toolkit for high-order neutrino-radiation hydrodynamics (thornado). The model considers electron neutrinos and antineutrinos and tabulated opacities from Bruenn (1985), which includes neutrino-electron scattering and pair processes. The nonlinear system arising from implicit time discretization of the equations governing neutrino-matter coupling is iterated to convergence using Anderson-accelerated fixed-point methods, which avoid formation of Jacobians and inversion of dense linear systems. Numerical experiments show that, for a given tolerance, a nested iteration scheme which aims to reduce opacity evaluations can lower the computational cost. Our initial port to GPUs, using both OpenMP and OpenACC, shows an overall speedup of up to ~ 100× when compared to results using a single CPU core. These results indicate that the algorithms implemented in thornado are well-suited to GPU acceleration.

Laiu, Paul↗

Domain decomposition in the GPU-accelerated Shift Monte Carlo code

The GPU solver within the Shift continuous-energy Monte Carlo neutron transport code has been extended to provide domain decomposition in addition to domain replication to enable the solution of problems with memory requirements exceeding the capacity of a single GPU. The strategy follows the Multiple Set, Overlapping Domain (MSOD) approach that is used in Shift’s CPU solver and integrates into the event-based algorithm used for Shift’s GPU solver. Furthermore, the ability to assign processors to spatial domains non-uniformly has been maintained. In this work, two different approaches for communicating particle data between domains are considered, and multiple criteria for load balancing problems have been investigated. Numerical results are presented for both fresh and depleted small modular nuclear reactor (SMR) cores. A parallel efficiency of approximately 80% was achieved with up to 16 spatial domains measured relative to full domain replication. A scaling study on the Summit supercomputer demonstrates a weak scaling parallel efficiency of over 90% on over 24000 GPUs.

97 MATHEMATICS AND COMPUTING↗

GPU-accelerated multitiered iterative phasing algorithm for fluctuation X-ray scattering

The multitiered iterative phasing (MTIP) algorithm is used to determine the biological structures of macromolecules from fluctuation scattering data. It is an iterative algorithm that reconstructs the electron density of the sample by matching the computed fluctuation X-ray scattering data to the external observations, and by simultaneously enforcing constraints in real and Fourier space. This paper presents the first ever MTIP algorithm acceleration efforts on contemporary graphics processing units (GPUs). The Compute Unified Device Architecture (CUDA) programming model is used to accelerate the MTIP algorithm on NVIDIA GPUs. The computational performance of the CUDA-based MTIP algorithm implementation outperforms the CPU-based version by an order of magnitude. Furthermore, the Heterogeneous-Compute Interface for Portability (HIP) runtime APIs are used to demonstrate portability by accelerating the MTIP algorithm across NVIDIA and AMD GPUs.

97 MATHEMATICS AND COMPUTING↗

Highly-scalable GPU-accelerated compressible reacting flow solver for modeling high-speed flows

Emerging supercomputing systems utilize a combination of central processing units (CPUs) and graphics processing units (GPUs) in an effort to reach exascale capabilities while minimizing the energy footprint of operating such systems. Such heterogeneous machines introduce new challenges for fluids solvers because the hardware architecture and operation of a GPU are fundamentally different from conventional CPUs. In this work, a general approach for efficient implementation of finite-volume based reacting flow solvers on such heterogeneous systems is presented. Three main challenges, namely, data access pattern, thread divergence, and thread safety, are addressed. Since compressible reacting flows require special methods to deal with chemical reactions, hyperbolic and nonlinear convection terms, and the presence of turbulence, specific algorithms that ensure GPU-based efficiency are developed. The approach is demonstrated on the widely available OpenFOAM open source software by modifying core algorithms for GPU accessibility. The scalability of the resulting solver, is demonstrated using practical test cases, including flow through a scramjet engine and the dynamics of a rotating detonation engine. Here, the solver provides near-ideal scaleup on a large number of GPUs (>3000), and extremely efficient use of the GPUs, with throughput nearly a constant even when processing a large number of control volumes.

42 ENGINEERING↗

GPU acceleration of Swendsen–Wang dynamics

When simulating a lattice system near its critical temperature, local algorithms for modeling the system’s evolution can introduce very large autocorrelation times into sampled data. Here, this critical slowing down places restrictions on the analysis that can be completed in a timely manner of the behavior of systems around the critical point. Because it is often desirable to study such systems around this point, a new algorithm must be introduced. Therefore, we turn to cluster algorithms, such as the Swendsen–Wang algorithm and the Wolff clustering algorithm. They incorporate global updates which generate new lattice configurations with little correlation to previous states, even near the critical point. We look to accelerate the rate at which these algorithm are capable of running by implementing and benchmarking a parallel implementation of each algorithm designed to run on GPUs under NVIDIA’s CUDA framework. A 17 and 90 fold increase in the computational rate was, respectively, experienced when measured against the equivalent algorithm implemented in serial code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybrid simulation of energetic particles interacting with magnetohydrodynamics using a slow manifold algorithm and GPU acceleration

The hybrid method combining particle-in-cell and magnetohydrodynamics can be used to study the interaction between energetic particles and global plasma modes. In this paper we introduce the M3D-C1-K code, which is developed based on the M3D-C1 finite element code solving the magnetohydrodynamics equations, with a newly developed kinetic module simulating energetic particles. The particle pushing is done using a new algorithm by applying the Boris pusher to the classical Pauli particles to simulate the slow-manifold of particle orbits, with long-term accuracy and fidelity. The particle pushing can be accelerated using GPUs with a significant speedup. The moments of the particles are calculated using the δƒ method, and are coupled into the magnetohydrodynamics simulation through pressure or current coupling schemes. Several linear simulations of magnetohydrodynamics modes driven by energetic particles have been conducted using M3D-C1-K with the δƒ method, including fishbone, toroidal Alfvén eigenmodes and reversed shear Alfvén eigenmodes. Good agreement with previous results from other eigenvalue, kinetic and hybrid codes have been achieved.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NekRS, a GPU-accelerated spectral element Navier–Stokes solver

The development of NekRS, a GPU-oriented thermal-fluids simulation code based on the spectral element method (SEM) is described. For performance portability, the code is based on the open concurrent compute abstraction and leverages scalable developments in the SEM code Nek5000 and in libParanumal, which is a library of high-performance kernels for high-order discretizations and PDE-based miniapps. Critical performance sections of the Navier–Stokes time advancement are addressed. Performance results on several platforms are presented here, including scaling to 27,648 V100s on OLCF Summit, for calculations of up to 60B gridpoints.

97 MATHEMATICS AND COMPUTING↗

Performance-Portable GPU Acceleration of the EFIT Tokamak Plasma Equilibrium Reconstruction Code

This paper presents the steps followed to GPU-offload parts of the core solver of EFIT-AI, an equilibrium reconstruction code suitable for tokamak experiments and burning plasmas. For this work, we will focus on the fitting procedure that consists of a Grad–Shafranov (GS) equation inverse solver that calculates equilibrium reconstructions on a grid. We will show profiling results of the original code (CPU-baseline), as well as the directives used to GPU-offload the most time-consuming function, initially to compare OpenACC and OpenMP on NVIDIA and AMD GPUs and later on to assess OpenMP performance portability on NVIDIA, AMD and Intel GPUs. We will make a performance comparison for different spatial grid sizes and show the speedup achieved on NVIDIA A100 (Perlmutter-NERSC), AMD MI250X (Frontier-OLCF) and Intel PVC GPUs (Sunspot-ALCF). Finally, we will draw some conclusions and recommendations to achieve high-performance portability for an equilibrium reconstruction code on the new HPC architectures

GPU↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

GPU-accelerated DNS of compressible turbulent flows

Here, this paper explores strategies to transform an existing CPU-based high-performance computational fluid dynamics solver, HyPar, for compressible flow simulations on emerging exascale heterogeneous (CPU+GPU) computing platforms. The scientific motivation for developing a GPU-enhanced version of HyPar is to simulate canonical turbulent flows at the highest resolution possible on such platforms. We show that optimizing memory operations and thread blocks results in 200x speedup of computationally intensive kernels compared with a CPU core. Using multiple GPUs and CUDA-aware MPI communication, we demonstrate both strong and weak scaling of our GPU-based HyPar implementation on the NVIDIA Volta V100 GPUs. We simulate the decay of homogeneous isotropic turbulence in a triply periodic box on grids with up to 1024 3 points (5.3 billion degrees of freedom) and on up to 1,024 GPUs. We compare the wall times for CPU-only and CPU+GPU simulations. The results presented in the paper are obtained on the Summit and Lassen supercomputers at Oak Ridge and Lawrence Livermore National Laboratories, respectively.

97 MATHEMATICS AND COMPUTING↗

High Performance, High Fidelity: A GPU‐Accelerated Doubly‐Periodic Configuration of the Simple Cloud‐Resolving E3SM Atmosphere Model Version 1 (DP‐SCREAMv1)

The development of the Simplified Cloud Resolving Energy Exascale Earth System Atmosphere Model (SCREAMv1) enables global storm-resolving simulations on modern GPU-based supercomputers. However, the high computational cost of SCREAMv1 limits its routine use for process-level studies, creating a need for efficient proxy configurations. This study addresses this gap by introducing DP-SCREAMv1, a doubly periodic cloud-resolving model designed to be fully consistent with SCREAMv1 while enabling high-resolution, long-duration simulations at significantly reduced computational expense by simulating a limited doubly periodic domain rather than the entire globe. Built on a C++/Kokkos architecture, DP-SCREAMv1 achieves exceptional performance scalability on GPU systems and includes a rich library of cases for validation and scientific exploration. In this work, we demonstrate short wall-clock times at SCREAMv1's default resolution and show that DP-SCREAMv1 supports routine execution of large-domain, high-resolution experiments that were previously challenging in practice. Furthermore, we show that DP-SCREAMv1 enables routine execution of “Giga-LES” style simulations and facilitates large-domain, high-resolution simulations that were recently considered burdensome to perform. These results document an efficient, fully consistent process-level configuration for SCREAMv1 (DP-SCREAMv1) and illustrate its use for long-duration and large-domain experiments at cloud-resolving to eddy-permitting resolution.

Environmental sciences↗

Implementation of Relativistic Coupled Cluster Theory for Massively Parallel GPU-Accelerated Computing Architectures

In this paper, we report reimplementation of the core algorithms of relativistic coupled cluster theory aimed at modern heterogeneous high-performance computational infrastructures. The code is designed for parallel execution on many compute nodes with optional GPU coprocessing, accomplished via the new ExaTENSOR back end. The resulting ExaCorr module is primarily intended for calculations of molecules with one or more heavy elements, as relativistic effects on the electronic structure are included from the outset. In the current work, we thereby focus on exact two-component methods and demonstrate the accuracy and performance of the software. The module can be used as a stand-alone program requiring a set of molecular orbital coefficients as the starting point, but it is also interfaced to the DIRAC program that can be used to generate these. We therefore also briefly discuss an improvement of the parallel computing aspects of the relativistic self-consistent field algorithm of the DIRAC program.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗