Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Multinode Multi-GPU Two-Electron Integrals: Code Generation Using the Regent Language

The computation of two-electron repulsion integrals (ERIs) is often the most expensive step of integral-direct self-consistent field methods. Formally it scales as O(N 4 ), where N is the number of Gaussian basis functions used to represent the molecular wave function. In practice, this scaling can be reduced to O(N 2 ) or less by neglecting small integrals with screening methods. The contributions of the ERIs to the Fock matrix are of Coulomb (J) and exchange (K) type and require separate algorithms to compute matrix elements efficiently. We previously implemented highly efficient GPU-accelerated J-matrix and K-matrix algorithms in the electronic structure code TeraChem. Although these implementations supported the use of multiple GPUs on a node, they did not support the use of multiple nodes. This presents a key bottleneck to cutting-edge ab initio simulations of large systems, e.g., excited state dynamics of photoactive proteins. We present our implementation of multinode multi-GPU J- and K-matrix algorithms in TeraChem using the Regent programming language. Regent directly supports distributed computation in a task-based model and can generate code for a variety of architectures, including NVIDIA GPUs. We demonstrate multinode scaling up to 45 GPUs (3 nodes) and benchmark against hand-coded TeraChem integral code. Finally, we also outline our metaprogrammed Regent implementation, which enables flexible code generation for integrals of different angular momenta.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations: Preprint

The US Department of Energy has identified Exascale-Class wind farm simulation tools as critical to wind energy scientific discovery. A primary objective of the Exawind project is to build high-performance, predictive Computational Fluid Dynamics tools that satisfy these modeling needs. GPU accelerators will serve as the computational thoroughbreds of next generation, Exascale-Class, platforms. Here, we report on our efforts for preparing the Exawind unstructured mesh solver, Nalu-Wind, for Exascale-Class machines. For computing at this scale, a simple port of the incompressible-flow algorithms to GPUs is not sufficient. One needs novel algorithms that are application aware, memory efficient, and optimized for latest generation GPU devices to get high-performance. The result of our efforts are unstructured mesh simulations of wind turbines that use 1/6 the compute resources of Summit supercomputer at Oak Ridge National Lab. In particular, we demonstrate a first-of-its-kind, simulation using Algebraic Multigrid solvers on over 4000 GPUs.

algebraic multigrid↗

Developments in Performance and Portability for MadGraph5_aMC@NLO

Event generators simulate particle interactions using Monte Carlo techniques, providing the primary connection between experiment and theory in experimental high energy physics. These software packages, which are the first step in the simulation worflow of collider experiments, represent approximately 5 to 20% of the annual WLCG usage for the ATLAS and CMS experiments. With computing architectures becoming more heterogeneous, it is important to ensure that these key software frameworks can be run on future systems, large and small. In this contribution, recent progress on porting and speeding up the Madgraph5_aMC@NLO event generator on hybrid architectures, i.e. CPU with GPU accelerators, is discussed. The main focus of this work has been in the calculation of scattering amplitudes and "matrix elements", which is the computational bottleneck of an event generation application. For physics processes limited to QCD leading order, the code generation toolkit has been expanded to produce matrix element calculations using C++ vector instructions on CPUs and using CUDA for NVidia GPUs, as well as using Alpaka, Kokkos and SYCL for multiple CPU and GPU architectures. Performance is reported in terms of matrix element calculations per time on NVidia, Intel, and AMD devices. The status and outlook for the integration of this work into a production release usable by the LHC experiments, with the same functionalities and very similar user interfaces as the current Fortran version, is also described.

Valassi, Andrea↗

Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations

The U.S. Department of Energy has identified exascale-class wind farm simulation as critical to wind energy scientific discovery. A primary objective of the ExaWind project is to build high-performance, predictive computational fluid dynamics (CFD) tools that satisfy these modeling needs. GPU accelerators will serve as the computational thoroughbreds of next-generation, exascale-class supercomputers. Here, we report on our efforts in preparing the ExaWind unstructured mesh solver, Nalu-Wind, for exascale-class machines. For computing at this scale, a simple port of the incompressible-flow algorithms to GPUs is insufficient. To achieve high performance, one needs novel algorithms that are application aware, memory efficient, and optimized for the latest-generation GPU devices. The result of our efforts are unstructured-mesh simulations of wind turbines that can effectively leverage thousands of GPUs. In particular, we demonstrate a first-of-its-kind, incompressible-flow simulation using Algebraic Multigrid solvers that strong scales to more than 4000 GPUs on the Summit supercomputer.

algebraic multigrid↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

Farpoint: A High-resolution Cosmology Simulation at the Gigaparsec Scale

Abstract In this paper we introduce the Farpoint simulation, the latest member of the Hardware/Hybrid Accelerated Cosmology Code (HACC) gravity-only simulation family. The domain covers a volume of (1000 h −1 Mpc) 3 and evolves close to two trillion particles, corresponding to a mass resolution of m p ∼ 4.6 × 10 7 h −1 M ⊙ . These specifications enable comprehensive investigations of the galaxy–halo connection, capturing halos down to small masses. Further, the large volume resolves scales typical of modern surveys with good statistical coverage of high-mass halos. The simulation was carried out on the GPU-accelerated system Summit, one of the fastest supercomputers currently available. We provide specifics about the Farpoint run and present an initial set of results. The high mass resolution facilitates precise measurements of important global statistics, such as the halo concentration–mass relation and the correlation function down to small scales. Selected subsets of the simulation data products are publicly available via the HACC Simulation Data Portal.

79 ASTRONOMY AND ASTROPHYSICS↗

PYSEQM

PYSEQM is a package for performing semi-empirical quantum mechanical (SEQM) simulations on molecular systems utilizing PyTorch. SEQM simulations determine molecular properties (energy, electron density, dipole moment, ect.) by solving an approximate Schrödinger equation for the motions of electrons in a molecule. The use of PyTorch provides three specific advantages. First, it allows the calculations to be offloaded to GPU accelerators, giving an order of magnitude increase in speed. Second, back propagation is used to get atomic forces (derivative of the total energy with respect to atomic position) at the same computational cost as the energy calculation itself. Finally, the use of PyTorch makes for a natural interface to modern machine learning methods, which can be used to adjust the semi-empirical parameters build into SEQM methods. Additionally, PYSEQM implements various other optimizations for performing quantum mechanics based molecular dynamics, including SP2 for rapid GPU based solution of the self-consistent field algorithm, and the extended Lagrangian method for rapid QM-MD.

Nebgen, Benjamin↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Data refactoring↗

A Length Adaptive Algorithm-Hardware Co-design of Transformer on FPGA Through Sparse Attention and Dynamic Pipelining

Transformers are considered one of the most important deep learning models since 2018, in part because it establishes state-of-the-art (SOTA) records and could potentially replace existing Deep Neural Networks (DNNs). Despite the remarkable triumphs, the prolonged turnaround time of Transformer models is a widely recognized roadblock. The variety of sequence lengths imposes additional computing overhead where inputs need to be zero-padded to the maximum sentence length in the batch to accommodate the parallel computing platforms. This paper targets the field-programmable gate array (FPGA) and proposes a coherent sequence length adaptive algorithm–hardware co-design for Transformer acceleration. Particularly, we develop a hardware-friendly sparse attention operator and a length-aware hardware resource scheduling algorithm. The proposed sparse attention operator brings the complexity of attention-based models down to linear complexity and alleviates the off-chip memory traffic. The proposed length-aware resource hardware scheduling algorithm dynamically allocates the hardware resources to fill up the pipeline slots and eliminates bubbles for NLP tasks. Experiments show that our design has very small accuracy loss and has 80.2 × and 2.6 × speedup compared to CPU and GPU implementation, and 4 × higher energy efficiency than state-of-the-art GPU accelerator optimized via CUBLAS GEMM.

Peng, Hongwu↗

Nek5000/RS Performance on Advanced GPU Architectures

We demonstrate NekRS performance results on various advanced GPU architectures. NekRS is a GPU-accelerated version of Nek5000 that is targeting high performance on forthcoming exascale platforms. It is being developed in DOE’s Center of Efficient Exascale Discretizations (CEED), which is one of the co-design centers under the Exascale Computing Project (ECP). In this report, we consider Frontier, Crusher, Spock, Polaris, Perlmutter, ThetaGPU and Summit. Simulations are performed using ExaSMR’s 17x17 rod-bundle geometries with different problem sizes. The report focuses on strong scaling performance and analysis. Many of the results shown in this report are the outcome from participation in the ALCF GPU Hackathon on Polaris, which was held on 7/19/22, and 7/26–7/28/22. Members of the Nek5000/RS team included Misun Min, Yu-Hsiang Lan, Paul Fischer and Thilina Rathnayake. Mentors were Kris Rowe (ALCF) and Peng Wang (NVIDIA). Also presented in this report are results on Frontier, which were obtained in collaboration with John Holmen at OLCF.

97 MATHEMATICS AND COMPUTING↗

Early experiences evaluating the HPE/Cray ecosystem for AMD GPUs

Summary The Oak Ridge Leadership Computing Facility (OLCF) has a long history of supporting and promoting GPU‐accelerated computing starting with the deployment of the Titan supercomputer in 2021 and continuing with the Summit supercomputer which has a theoretical peak performance of approximately 200 petaflops. Because the majority of Summit's computational power comes from its 27,972 GPUs, users must port their applications to one of the supported programming models in order to make efficient use of the system. To prepare the transition to Frontier, the OLCF's exascale supercomputer, users will need to adapt to an entirely new ecosystem which will include new hardware and software technologies. First, users will need to familiarize themselves with the AMD Radeon GPU architecture. Furthermore, users who have been previously relying on CUDA will need to transition to the Heterogeneous‐Computing Interface for Portability (HIP) or one of the other supported programming models (e.g., OpenMP, OpenACC). In this work, we describe our initial experiences and lessons learned in porting three applications or proxy apps currently running on Summit to the HPE/Cray ecosystem to leverage the compute power from AMD GPUs: minisweep, GenASiS, and Sparkler. Each one is representative of current production workloads utilized at the OLCF, different programming languages, and different programming models.

Melesse Vergara, Verónica G.↗

Optimizing Error-Bounded Lossy Compression for Scientific Data on GPUs

Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. With ever-emerging heterogeneous high-performance computing (HPC) architecture, GPU-accelerated error-bounded compressors (such as CUSZ and cuZFP) have been developed. However, they suffer from either low performance or low compression ratios. To this end, we propose CUSZ+ to target both high compression ratios and throughputs. We identify that data sparsity and data smoothness are key factors for high compression throughputs. Our key contributions in this work are fourfold: (1) We propose an efficient compression workflow to adaptively perform run-length encoding and/or variable-length encoding. (2) We derive Lorenzo reconstruction in decompression as multidimensional partial-sum computation and propose a fine-grained Lorenzo reconstruction algorithm for GPU architectures. (3) We carefully optimize each of CUSZ kernels by leveraging state-of-the-art CUDA parallel primitives. (4) We evaluate CUSZ+ using seven real-world HPC application datasets on V100 and A100 GPUs. Experiments show CUSZ+ improves the compression throughputs and ratios by up to 18.4x and 5.3x, respectively, over CUSZ on the tested datasets.

Tian, Jiannan↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Chen, Jieyang↗

Simulating Hydrodynamics in Cosmology with CRK-HACC

Abstract We introduce CRK-HACC, an extension of the Hardware/Hybrid Accelerated Cosmology Code (HACC), to resolve gas hydrodynamics in large-scale structure formation simulations of the universe. The new framework couples the HACC gravitational N -body solver with a modern smoothed-particle hydrodynamics (SPH) approach called conservative reproducing kernel SPH (CRKSPH). CRKSPH utilizes smoothing functions that exactly interpolate linear fields while manifestly preserving conservation laws (momentum, mass, and energy). The CRKSPH method has been incorporated to accurately model baryonic effects in cosmology simulations—an important addition targeting the generation of precise synthetic sky predictions for upcoming observational surveys. CRK-HACC inherits the codesign strategies of the HACC solver and is built to run on modern GPU-accelerated supercomputers. In this work, we summarize the primary solver components and present a number of standard validation tests to demonstrate code accuracy, including idealized hydrodynamic and cosmological setups, as well as self-similarity measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

SQMBox: Interfacing a semiempirical integral library to modular ab initio electronic structure enables new semiempirical methods

Ab initio and semiempirical electronic structure methods are usually implemented in separate software packages or use entirely different code paths. As a result, it can be time-consuming to transfer an established ab initio electronic structure scheme to a semiempirical Hamiltonian. Here we present an approach to unify ab initio and semiempirical electronic structure code paths based on a separation of the wavefunction ansatz and the needed matrix representations of operators. With this separation, the Hamiltonian can refer to either an ab initio or semiempirical treatment of the resulting integrals. We built a semiempirical integral library and interfaced it to the GPU-accelerated electronic structure code TeraChem. Equivalency between ab initio and semiempirical tight-binding Hamiltonian terms is assigned according to their dependence on the one-electron density matrix. The new library provides semiempirical equivalents of the Hamiltonian matrix and gradient intermediates, corresponding to those provided by the ab initio integral library. This enables the straightforward combination of semiempirical Hamiltonians with the full pre-existing ground and excited state functionality of the ab initio electronic structure code. We demonstrate the capability of this approach by combining the extended tight-binding method GFN1-xTB with both spin-restricted ensemble-referenced Kohn–Sham and complete active space methods. We also present a highly efficient GPU implementation of the semiempirical Mulliken-approximated Fock exchange. The additional computational cost for this term becomes negligible even on consumer-grade GPUs, enabling Mulliken-approximated exchange in tight-binding methods for essentially no additional cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bolt: A Fast Solver for Kinetic Theories Using a High-Resolution Constrained Transport Scheme [Slides]

Our understanding of collisionless and semi-collisional plasmas in the nonlinear regime is limited by the expense of computing solutions numerically. Bolt is a fast, GPU-accelerated code for rapidly computing such solutions with accurate transport and an approximate collision operator. Such calculations are relevant to both problems in astrophysics, such as heat conduction and magnetic reconnection in accretion disks around black holes, and also to programmatic interests at LANL. The bolt code paper, demonstrating accuracy via a suite of test problems calculated on kodiak, is currently in preparation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

ERF: Energy Research and Forecasting Model

High performance computing (HPC) architectures have undergone rapid development in recent years. As a result, established software suites face an ever increasing challenge to remain performant on and portable across modern systems. Many of the widely adopted atmospheric modeling codes cannot fully (or in some cases, at all) leverage the acceleration provided by General-Purpose Graphics Processing Units, leaving users of those codes constrained to increasingly limited HPC resources. Energy Research and Forecasting (ERF) is a regional atmospheric modeling code that leverages the latest HPC architectures, whether composed of only Central Processing Units (CPUs) or incorporating GPUs. ERF contains many of the standard discretizations and basic features needed to model general atmospheric dynamics. The modular design of ERF provides a flexible platform for exploring different physics parameterizations and numerical strategies. ERF is built on a state-of-the-art, well-supported, software framework (AMReX) that provides a performance portable interface and ensures ERF's long-term sustainability on next generation computing systems. This paper details the numerical methodology of ERF, presents results for a series of verification/validation cases, and documents ERF's performance on current HPC systems. The roughly 5× speed up of ERF (using GPUs) over Weather Research and Forecasting (CPUs only) for a 3D squall line test case highlights the significance of leveraging GPU acceleration.

17 WIND ENERGY↗