Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphical processing unit”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Introduction to Graphics Processing Units [Slides]

Graphics Processing Units are designed for fast graphics processing. Graphics are a form of arithemetic and have gradually evolved a design that is also usefule for non-graphics computing. They are not standalone, but work alongside a CPU(host)-coprocessor.

97 MATHEMATICS AND COMPUTING↗

Towards Efficient Alternating Current Optimal Power Flow Analysis on Graphical Processing Units

We present a solution of sparse ACOPF analysis on GPU. In particular, we discuss the performance bottlenecks and detail our efforts to accelerate the linear solver, a core component of ACOPF that dominates the computational time. ACOPF solutions of two large-scale systems, synthetic Northeast (25,000 buses) and Eastern (70,000 buses) \cite{birchfield2017tamu-cases} on GPU show promising speed-up compared to CPU based solution using a state-of-the-art solver. To our knowledge, this is the first result demonstrating acceleration of sparse ACOPF on GPUs.

Power grid analysis, GPU↗

A Graphics Processing Unit–Based, Industrial Grade Compositional Reservoir Simulator

Summary Recently, graphics processing units (GPUs) have been demonstrated to provide a significant performance benefit for black-oil reservoir simulation, as well as flash calculations that serve an important role in compositional simulation. A comprehensive approach to compositional simulation based on GPUs has yet to emerge, and the question remains as to whether the benefits observed in black-oil simulation persist with a more complex fluid description. We present a positive answer to this question through the extension of a commercial GPU-based black-oil simulator to include a compositional description based on standard cubic equations of state (EOSs). We describe the motivations for the selected nonlinear formulation, including the choice of primary variables and iteration scheme, and support for both fully implicit methods (FIMs) and adaptive implicit methods (AIMs). We then present performance results on an example sector model and simplified synthetic case designed to allow a detailed examination of runtime and memory scaling with respect to the number of hydrocarbon components and model size, as well as the number of processors. We finally show results from two complex asset models (synthetic and real) and examine performance scaling with respect to GPU generation, demonstrating that performance correlates strongly with GPU memory bandwidth. NOTE: This paper is also published as part of the 2021 SPE Reservoir Simulation Conference Special Issue.

Engineering↗

Accelerating the density-functional tight-binding method using graphical processing units

Acceleration of the density-functional tight-binding (DFTB) method on single and multiple graphical processing units (GPUs) was accomplished using the MAGMA linear algebra library. Herein two major computational bottlenecks of DFTB ground-state calculations were addressed in our implementation: the Hamiltonian matrix diagonalization and the density matrix construction. The code was implemented and benchmarked on two different computer systems: (1) the SUMMIT IBM Power9 supercomputer at the Oak Ridge National Laboratory Leadership Computing Facility with 1–6 NVIDIA Volta V100 GPUs per computer node and (2) an in-house Intel Xeon computer with 1–2 NVIDIA Tesla P100 GPUs. The performance and parallel scalability were measured for three molecular models of 1-, 2-, and 3-dimensional chemical systems, represented by carbon nanotubes, covalent organic frameworks, and water clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit↗

Evaluation of AC optimal power flow on graphical processing units

This paper investigates the performance of alternating current optimal power flow (ACOPF) on hardware accelerators such as graphical processing units (GPUs). We describe the strategies employed and the software used to port the ACOPF application to GPU. Through reorganizing the flow of fundamental calculations, restructuring data organization for the GPUs, and using portability libraries, maximum utilization of GPU is attempted. We present details of our efforts with representative results on 200, 500, and 2000-bus networks.

Abhyankar, Shrirang G.↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

Porting Fragmentation Methods to Graphical Processing Units Using an OpenMP Application Programming Interface: Offloading the Fock Build for Low Angular Momentum Functions

Here, a framework to offload four-index two-electron repulsion integrals to graphical processing units (GPUs) using OpenMP is discussed. The method has been applied to the Fock build for low angular momentum s and p functions in both the restricted Hartree–Fock (RHF) and in the effective fragment molecular orbital (EFMO) framework. Benchmark calculations for the GPU code for the pure RHF method show an increasing speedup relative to the existing OpenMP CPU code in GAMESS from 1.04 to 52× for clusters of 70–569 water molecules. The parallel efficiency on 24 NVIDIA V100 GPU boards also increases when increasing the system size: from 75 to 94% for water clusters that contain 303–1120 molecules. In the EFMO framework, the GPU Fock build shows a high linear scalability up to 4608 V100s with a parallel efficiency of 96% for calculations on a solvated mesoporous silica nanoparticle system with ~67,000 basis functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interactive Quantum Chemistry Enabled by Machine Learning, Graphical Processing Units, and Cloud Computing

Modern quantum chemistry algorithms are increasingly able to accurately predict molecular properties that are useful for chemists in research and education. Despite this progress, performing such calculations is currently unattainable to the wider chemistry community, as they often require domain expertise, computer programming skills, and powerful computer hardware. In this review, we outline methods to eliminate these barriers using cutting-edge technologies. We discuss the ingredients needed to create accessible platforms that can compute quantum chemistry properties in real time, including graphical processing units–accelerated quantum chemistry in the cloud, artificial intelligence–driven natural molecule input methods, and extended reality visualization. We end by highlighting a series of exciting applications that assemble these components to create uniquely interactive platforms for computing and visualizing spectra, 3D structures, molecular orbitals, and many other chemical properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Performance Portable Graphics Processing Unit Acceleration of a High-Order Finite Element Multiphysics Application

The Lawrence Livermore National Laboratory (LLNL) will soon have in place the El Capitan exascale supercomputer, based on advanced micro devices (AMD) graphics processing units (GPUs). As part of a multiyear effort under the National Nuclear Security Administration (NNSA) Advanced Simulation and Computing (ASC) program, we have been developing marbl, a next generation, performance portable multiphysics application based on high-order finite elements. In previous years, we successfully ported the Arbitrary Lagrangian–Eulerian (ALE), multimaterial, compressible flow capabilities of marbl to nvidia GPUs as described in Vargas et al. Here, in this paper, we describe our ongoing effort in extending marbl's GPU capabilities with additional physics, including multigroup radiation diffusion and thermonuclear burn for high energy density physics (HEDP) and fusion modeling. We also describe how our portability abstraction approach based on the raja Portability Suite and the mfem finite element discretization library has enabled us to achieve high performance on AMD based GPUs with minimal effort in hardware-specific porting. Throughout this work, we highlight numerical and algorithmic developments that were required to achieve GPU performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analytical derivatives of the individual state energies in ensemble density functional theory. II. Implementation on graphical processing units (GPUs)

Conical intersections control excited state reactivity, and thus, elucidating and predicting their geometric and energetic characteristics are crucial for understanding photochemistry. Locating these intersections requires accurate and efficient electronic structure methods. Unfortunately, the most accurate methods (e.g., multireference perturbation theories such as XMS-CASPT2) are computationally challenging for large molecules. The state-interaction state-averaged restricted ensemble referenced Kohn–Sham (SI-SA-REKS) method is a computationally efficient alternative. The application of SI-SA-REKS to photochemistry was previously hampered by a lack of analytical nuclear gradients and nonadiabatic coupling matrix elements. We have recently derived analytical energy derivatives for the SI-SA-REKS method and implemented the method effectively on graphical processing units. We demonstrate that our implementation gives the correct conical intersection topography and energetics for several examples. Furthermore, our implementation of SI-SA-REKS is computationally efficient, with observed sub-quadratic scaling as a function of molecular size. This demonstrates the promise of SI-SA-REKS for excited state dynamics of large molecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-speed feedback control of an oscillating magnetic helicity injector using a graphics processing unit

A real-time control system has been developed to control the amplitude, phase, and offset of bulk plasma parameters inside an oscillating magnetic helicity injector. Control software running entirely on an Nvidia Tesla P40 graphical processing unit is able to receive digitizer inputs and send response patterns to a Pulse Width Modulation (PWM) controller with a minimum control loop period of 12.8μs. With an input digitization rate of 10 MS/s, a three-parameter proportional integral differential controller is shown to be sufficient to inform the PWM controller to drive the desired oscillating plasma waveform with a frequency of 16.6 kHz that is located near the resonance of a coupled RLC circuit. In particular, the temporal phase of the injector waveform is held within 10 degrees of the target value. Control is demonstrated over the toroidal modal structure of the imposed magnetic perturbations of the helicity injection system, allowing a new class of discharges to be studied.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Graphics Processing Unit (GPU) Approach to Large Eddy Simulation (LES) for Transport and Contaminant Dispersion

Recent advances in the development of large eddy simulation (LES) atmospheric models with corresponding atmospheric transport and dispersion (AT&D) modeling capabilities have made it possible to simulate short, time-averaged, single realizations of pollutant dispersion at the spatial and temporal resolution necessary for common atmospheric dispersion needs, such as designing air sampling networks, assessing pollutant sensor system performance, and characterizing the impact of airborne materials on human health. The high computational burden required to form an ensemble of single-realization dispersion solutions using an LES and coupled AT&D model has, until recently, limited its use to a few proof-of-concept studies. An example of an LES model that can meet the temporal and spatial resolution and computational requirements of these applications is the joint outdoor-indoor urban large eddy simulation (JOULES). A key enabling element within JOULES is the computationally efficient graphics processing unit (GPU)-based LES, which is on the order of 150 times faster than if the LES contaminant dispersion simulations were executed on a central processing unit (CPU) computing platform. JOULES is capable of resolving the turbulence components at a suitable scale for both open terrain and urban landscapes, e.g., owing to varying environmental conditions and a diverse building topology. In this paper, we describe the JOULES modeling system, prior efforts to validate the accuracy of its meteorological simulations, and current results from an evaluation that uses ensembles of dispersion solutions for unstable, neutral, and stable static stability conditions in an open terrain environment.

54 ENVIRONMENTAL SCIENCES↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

FULL RANGE TUNE SCAN STUDIES USING GRAPHICS PROCESSING UNITS WITH CUDA IN EIC BEAM-BEAM SIMULATIONS

The hadron beam in the Electron-Ion Collider (EIC) suffers high order betatron and synchro-betatron resonances. In this paper, we present a weak-strong full range (0.0 ~ 0.5) fractional tune scan with a step size as small as 0.001. Multiple Graphics Processing Units (GPUs) are used to speed up the simulation. A code parallelized with MPI and CUDA is implemented. The good tune region from weak-strong scan is further checked by the self-consistent strong-strong simulation. This study provides beam dynamics guidance in choosing proper working points for the future EIC.

43 PARTICLE ACCELERATORS↗

Discrete element simulation of Pebble Bed Reactors on graphics processing units

Prediction of pebble positions in a Pebble Bed Reactor (PBR) is necessary for both reactor physics and thermal hydraulics simulations as the arrangement of pebbles has a significant impact on the resulting core power, coolant flow, and fuel temperature. Knowledge of pebble movement as the fuel is cycled through the core is also critical for predicting the fuel residence time and subsequently, the fuel burnup. Simulation with the Discrete Element Method (DEM) can provide knowledge of both the fuel packing and the fuel movement during cycling. Previous works that have performed 3D full-core DEM simulation of PBRs have used simplified models that neglect reflector wall features. This work employs a graphics processing unit (GPU)-enabled DEM code, Project Chrono, to analyze the differences in pebble packing and pebble velocities between a simplified smooth PBR reflector and a more realistic reflector that includes circular wall features. Additionally, a sensitivity study is performed on the depth of the wall features to ensure that crystallization is prevented. Project Chrono is also validated for PBR cycling applications using experimental data. It is found that wall features with a depth of at least 0.5 pebble diameters significantly reduce crystallization in the near-wall region, leading to discrepancies in both packing fraction and pebble velocity in this region compared to the simplified reflector models. These discrepancies are found to lead to roughly a 5–10% difference in the prediction of the near-wall porosity and a 10% difference in the prediction of the velocity of pebbles near the wall. As a result of these discrepancies, it is suggested that future DEM simulations of PBRs include wall features to reduce modeling errors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗