Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphic processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Interactive Quantum Chemistry Enabled by Machine Learning, Graphical Processing Units, and Cloud Computing

Modern quantum chemistry algorithms are increasingly able to accurately predict molecular properties that are useful for chemists in research and education. Despite this progress, performing such calculations is currently unattainable to the wider chemistry community, as they often require domain expertise, computer programming skills, and powerful computer hardware. In this review, we outline methods to eliminate these barriers using cutting-edge technologies. We discuss the ingredients needed to create accessible platforms that can compute quantum chemistry properties in real time, including graphical processing units–accelerated quantum chemistry in the cloud, artificial intelligence–driven natural molecule input methods, and extended reality visualization. We end by highlighting a series of exciting applications that assemble these components to create uniquely interactive platforms for computing and visualizing spectra, 3D structures, molecular orbitals, and many other chemical properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Performance Portable Graphics Processing Unit Acceleration of a High-Order Finite Element Multiphysics Application

The Lawrence Livermore National Laboratory (LLNL) will soon have in place the El Capitan exascale supercomputer, based on advanced micro devices (AMD) graphics processing units (GPUs). As part of a multiyear effort under the National Nuclear Security Administration (NNSA) Advanced Simulation and Computing (ASC) program, we have been developing marbl, a next generation, performance portable multiphysics application based on high-order finite elements. In previous years, we successfully ported the Arbitrary Lagrangian–Eulerian (ALE), multimaterial, compressible flow capabilities of marbl to nvidia GPUs as described in Vargas et al. Here, in this paper, we describe our ongoing effort in extending marbl's GPU capabilities with additional physics, including multigroup radiation diffusion and thermonuclear burn for high energy density physics (HEDP) and fusion modeling. We also describe how our portability abstraction approach based on the raja Portability Suite and the mfem finite element discretization library has enabled us to achieve high performance on AMD based GPUs with minimal effort in hardware-specific porting. Throughout this work, we highlight numerical and algorithmic developments that were required to achieve GPU performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analytical derivatives of the individual state energies in ensemble density functional theory. II. Implementation on graphical processing units (GPUs)

Conical intersections control excited state reactivity, and thus, elucidating and predicting their geometric and energetic characteristics are crucial for understanding photochemistry. Locating these intersections requires accurate and efficient electronic structure methods. Unfortunately, the most accurate methods (e.g., multireference perturbation theories such as XMS-CASPT2) are computationally challenging for large molecules. The state-interaction state-averaged restricted ensemble referenced Kohn–Sham (SI-SA-REKS) method is a computationally efficient alternative. The application of SI-SA-REKS to photochemistry was previously hampered by a lack of analytical nuclear gradients and nonadiabatic coupling matrix elements. We have recently derived analytical energy derivatives for the SI-SA-REKS method and implemented the method effectively on graphical processing units. We demonstrate that our implementation gives the correct conical intersection topography and energetics for several examples. Furthermore, our implementation of SI-SA-REKS is computationally efficient, with observed sub-quadratic scaling as a function of molecular size. This demonstrates the promise of SI-SA-REKS for excited state dynamics of large molecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-speed feedback control of an oscillating magnetic helicity injector using a graphics processing unit

A real-time control system has been developed to control the amplitude, phase, and offset of bulk plasma parameters inside an oscillating magnetic helicity injector. Control software running entirely on an Nvidia Tesla P40 graphical processing unit is able to receive digitizer inputs and send response patterns to a Pulse Width Modulation (PWM) controller with a minimum control loop period of 12.8μs. With an input digitization rate of 10 MS/s, a three-parameter proportional integral differential controller is shown to be sufficient to inform the PWM controller to drive the desired oscillating plasma waveform with a frequency of 16.6 kHz that is located near the resonance of a coupled RLC circuit. In particular, the temporal phase of the injector waveform is held within 10 degrees of the target value. Control is demonstrated over the toroidal modal structure of the imposed magnetic perturbations of the helicity injection system, allowing a new class of discharges to be studied.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Graphics Processing Unit (GPU) Approach to Large Eddy Simulation (LES) for Transport and Contaminant Dispersion

Recent advances in the development of large eddy simulation (LES) atmospheric models with corresponding atmospheric transport and dispersion (AT&D) modeling capabilities have made it possible to simulate short, time-averaged, single realizations of pollutant dispersion at the spatial and temporal resolution necessary for common atmospheric dispersion needs, such as designing air sampling networks, assessing pollutant sensor system performance, and characterizing the impact of airborne materials on human health. The high computational burden required to form an ensemble of single-realization dispersion solutions using an LES and coupled AT&D model has, until recently, limited its use to a few proof-of-concept studies. An example of an LES model that can meet the temporal and spatial resolution and computational requirements of these applications is the joint outdoor-indoor urban large eddy simulation (JOULES). A key enabling element within JOULES is the computationally efficient graphics processing unit (GPU)-based LES, which is on the order of 150 times faster than if the LES contaminant dispersion simulations were executed on a central processing unit (CPU) computing platform. JOULES is capable of resolving the turbulence components at a suitable scale for both open terrain and urban landscapes, e.g., owing to varying environmental conditions and a diverse building topology. In this paper, we describe the JOULES modeling system, prior efforts to validate the accuracy of its meteorological simulations, and current results from an evaluation that uses ensembles of dispersion solutions for unstable, neutral, and stable static stability conditions in an open terrain environment.

54 ENVIRONMENTAL SCIENCES↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

FULL RANGE TUNE SCAN STUDIES USING GRAPHICS PROCESSING UNITS WITH CUDA IN EIC BEAM-BEAM SIMULATIONS

The hadron beam in the Electron-Ion Collider (EIC) suffers high order betatron and synchro-betatron resonances. In this paper, we present a weak-strong full range (0.0 ~ 0.5) fractional tune scan with a step size as small as 0.001. Multiple Graphics Processing Units (GPUs) are used to speed up the simulation. A code parallelized with MPI and CUDA is implemented. The good tune region from weak-strong scan is further checked by the self-consistent strong-strong simulation. This study provides beam dynamics guidance in choosing proper working points for the future EIC.

43 PARTICLE ACCELERATORS↗

Graphics Processing Unit (GPU) Acceleration of the Goddard Earth Observing System Atmospheric Model

The Goddard Earth Observing System 5 (GEOS-5) is the atmospheric model used by the Global Modeling and Assimilation Office (GMAO) for a variety of applications, from long-term climate prediction at relatively coarse resolution, to data assimilation and numerical weather prediction, to very high-resolution cloud-resolving simulations. GEOS-5 is being ported to a graphics processing unit (GPU) cluster at the NASA Center for Climate Simulation (NCCS). By utilizing GPU co-processor technology, we expect to increase the throughput of GEOS-5 by at least an order of magnitude, and accelerate the process of scientific exploration across all scales of global modeling, including: The large-scale, high-end application of non-hydrostatic, global, cloud-resolving modeling at 10- to I-kilometer (km) global resolutions Intermediate-resolution seasonal climate and weather prediction at 50- to 25-km on small clusters of GPUs Long-range, coarse-resolution climate modeling, enabled on a small box of GPUs for the individual researcher After being ported to the GPU cluster, the primary physics components and the dynamical core of GEOS-5 have demonstrated a potential speedup of 15-40 times over conventional processor cores. Performance improvements of this magnitude reduce the required scalability of 1-km, global, cloud-resolving models from an unfathomable 6 million cores to an attainable 200,000 GPU-enabled cores.

Putnam, Williama↗

An Optimized Multicolor Point-Implicit Solver for Unstructured Grid Applications on Graphics Processing Units

In the field of computational fluid dynamics, the Navier-Stokes equations are often solved using an unstructuredgrid approach to accommodate geometric complexity. Implicit solution methodologies for such spatial discretizations generally require frequent solution of large tightly-coupled systems of block-sparse linear equations. The multicolor point-implicit solver used in the current work typically requires a significant fraction of the overall application run time. In this work, an efficient implementation of the solver for graphics processing units is proposed. Several factors present unique challenges to achieving an efficient implementation in this environment. These include the variable amount of parallelism available in different kernel calls, indirect memory access patterns, low arithmetic intensity, and the requirement to support variable block sizes. In this work, the solver is reformulated to use standard sparse and dense Basic Linear Algebra Subprograms (BLAS) functions. However, numerical experiments show that the performance of the BLAS functions available in existing CUDA libraries is suboptimal for matrices representative of those encountered in actual simulations. Instead, optimized versions of these functions are developed. Depending on block size, the new implementations show performance gains of up to 7x over the existing CUDA library functions.

Zubair, Mohammad↗

Discrete element simulation of Pebble Bed Reactors on graphics processing units

Prediction of pebble positions in a Pebble Bed Reactor (PBR) is necessary for both reactor physics and thermal hydraulics simulations as the arrangement of pebbles has a significant impact on the resulting core power, coolant flow, and fuel temperature. Knowledge of pebble movement as the fuel is cycled through the core is also critical for predicting the fuel residence time and subsequently, the fuel burnup. Simulation with the Discrete Element Method (DEM) can provide knowledge of both the fuel packing and the fuel movement during cycling. Previous works that have performed 3D full-core DEM simulation of PBRs have used simplified models that neglect reflector wall features. This work employs a graphics processing unit (GPU)-enabled DEM code, Project Chrono, to analyze the differences in pebble packing and pebble velocities between a simplified smooth PBR reflector and a more realistic reflector that includes circular wall features. Additionally, a sensitivity study is performed on the depth of the wall features to ensure that crystallization is prevented. Project Chrono is also validated for PBR cycling applications using experimental data. It is found that wall features with a depth of at least 0.5 pebble diameters significantly reduce crystallization in the near-wall region, leading to discrepancies in both packing fraction and pebble velocity in this region compared to the simplified reflector models. These discrepancies are found to lead to roughly a 5–10% difference in the prediction of the near-wall porosity and a 10% difference in the prediction of the velocity of pebbles near the wall. As a result of these discrepancies, it is suggested that future DEM simulations of PBRs include wall features to reduce modeling errors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

High-Performance Semiempirical Excited-State Molecular Dynamics Powered by Graphics Processing Units

Here, this Letter introduces excited-state molecular dynamics in PYSEQM, a GPU-accelerated semiempirical quantum chemistry engine implemented in PyTorch. The new module enables Born–Oppenheimer molecular dynamics (BOMD) using configuration-interaction singles and random phase approximation for excited states, allowing long trajectories and large statistical ensembles to be simulated efficiently on a single GPU. We also implement an extended Lagrangian excited-state BOMD (XL-ESMD) scheme that propagates auxiliary electronic variables, enabling relaxed ground and excited-state convergence thresholds without compromising energy conservation. The excited-state BOMD implementation scales smoothly from small chromophores to a nearly 900-atom dendrimer (taking 6.5 s per MD step). PYSEQM also supports batched execution, allowing many geometries or trajectories to be evaluated in a single GPU launch, substantially increasing throughput and making ensemble-based protocols routine. As a demonstration, we compute absorption, emission, and infrared spectra from trajectories propagated on the ground and first excited states. The XL-ESMD scheme yields identical spectra at significantly lower computational cost, establishing the role of extended Lagrangian based dynamics for efficient excited-state BOMD simulations. Beyond raw performance, PYSEQM’s PyTorch foundation provides automatic differentiation for forces, efficient GPU batching, and seamless interfacing with machine learning models. These capabilities position PYSEQM as a practical platform for machine learning-augmented excited-state dynamics and lay the foundation for future data-driven nonadiabatic excited-state dynamics modeling of ultrafast spectroscopic probes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Acceleration of the Parameterization of Unified Microphysics Across Scales (PUMAS) on the Graphics Processing Unit (GPU) With Directive-Based Methods

Cloud microphysics is one of the most time-consuming components in a climate model. In this study, we port the cloud microphysics parameterization in the Community Atmosphere Model (CAM), known as Parameterization of Unified Microphysics Across Scales (PUMAS), from CPU to GPU to seek a computational speedup. The directive-based methods (OpenACC and OpenMP target offload) are determined as the best fit specifically for our development practices, which enable a single version of source code to run either on the CPU or GPU, and yield a better portability and maintainability. Their performance is first examined in a PUMAS stand-alone kernel and the directive-based methods can outperform a CPU node as long as there is enough computational burden on the GPU. A consistent behavior is observed when we run PUMAS on the GPU in a practical CAM simulation. A 3.6× speedup of the PUMAS execution time, including data movement between CPU and GPU, is achieved at a coarse horizontal resolution (8 NVIDIA V100 GPUs against 36 Intel Skylake CPU cores). This speedup further increases up to 5.4× at a high resolution (24 NVIDIA V100 GPUs against 108 Intel Skylake CPU cores), which highlights the fact that GPU favors larger problem size. This study demonstrates that using GPU in a CAM simulation can save noticeable computational costs even with a small portion of code being GPU-enabled. Therefore, we are encouraged to port more parameterizations to GPU to take advantage of its computational benefit.

54 ENVIRONMENTAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Graphics Processing Units (GPU) and the Goddard Earth Observing System atmospheric model (GEOS-5): Implementation and Potential Applications

Earth system models like the Goddard Earth Observing System model (GEOS-5) have been pushing the limits of large clusters of multi-core microprocessors, producing breath-taking fidelity in resolving cloud systems at a global scale. GPU computing presents an opportunity for improving the efficiency of these leading edge models. A GPU implementation of GEOS-5 will facilitate the use of cloud-system resolving resolutions in data assimilation and weather prediction, at resolutions near 3.5 km, improving our ability to extract detailed information from high-resolution satellite observations and ultimately produce better weather and climate predictions

Putnam, William M.↗

Hardware accelerated dynamic work creation on a graphics processing unit

A processor core is configured to execute a parent task that is described by a data structure stored in a memory. A coprocessor is configured to dispatch a child task to the at least one processor core in response to the coprocessor receiving a request from the parent task concurrently with the parent task executing on the at least one processor core. In some cases, the parent task registers the child task in a task pool and the child task is a future task that is configured to monitor a completion object and enqueue another task associated with the future task in response to detecting the completion object. The future task is configured to self-enqueue by adding a continuation future task to a continuation queue for subsequent execution in response to the future task failing to detect the completion object.

Gutierrez, Anthony↗

Hardware accelerated dynamic work creation on a graphics processing unit

A processor core is configured to execute a parent task that is described by a data structure stored in a memory. A coprocessor is configured to dispatch a child task to the at least one processor core in response to the coprocessor receiving a request from the parent task concurrently with the parent task executing on the at least one processor core. In some cases, the parent task registers the child task in a task pool and the child task is a future task that is configured to monitor a completion object and enqueue another task associated with the future task in response to detecting the completion object. The future task is configured to self-enqueue by adding a continuation future task to a continuation queue for subsequent execution in response to the future task failing to detect the completion object.

Gutierrez, Anthony T.↗