Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphic processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The Pele Simulation Suite for Reacting Flows at Exascale

In this work, we present the Pele suite of software tools for compressible and incompressible reacting flows. The Pele suite leverages several different libraries, notably AMReX and SUNDIALS, to achieve performance portability on heterogeneous computing architectures across the supercomputing landscape. The Pele suite is comprised of PeleC, a compressible reacting flow block-structured adaptive mesh refinement solver, PeleLMeX, a low-Mach number reacting flow block-structured adaptive mesh refinement solver, Pele-Physics, a library for transport, thermodynamics, finite rate chemistry, soot, spray and radiation physics. The objective of this paper is (i) to present the code development efforts necessary to achieve highly effective and scalable applications for exascale machines and (ii) to detail the performance results of the Combustion-Pele project applications on Oak Ridge National Laboratory's Frontier. We show good weak and strong scaling results for both PeleC and PeleLMeX up to more than 50 billion cells on more than 4096 Frontier graphics processing unit nodes. We also present a capability demonstration simulation of a dual-fuel pulse compression ignition engine (six adaptive mesh refinement levels, and 60 billion cells or 2.1 trillion degrees of freedom) on Frontier, to date one of the largest simulations performed on the first exascale-class supercomputer.

adaptive mesh refinement↗

The LISE package: solvers for static and time-dependent superfluid local density approximation equations in three dimensions

Nuclear implementation of the density functional theory (DFT) is at present the only microscopic framework applicable to the whole nuclear landscape. The extension of DFT to superfluid systems in the spirit of the Kohn-Sham approach, the superfluid local density approximation (SLDA) and its extension to time-dependent situations, time-dependent superfluid local density ap- proximation (TDSLDA), have been extensively used to describe various static and dynamical problems in nuclear physics, neutron star crust, and cold atom systems. In this paper, we present the codes that solve the static and time-dependent SLDA equations in three-dimensional coordinate space without any symmetry restriction. These codes are fully parallelized with the message passing interface (MPI) library and take advantage of graphic processing units (GPU) for accelerating execution. The dynamic codes have checkpoint/restart capabilities and for initial conditions one can use any generalized Slater determinant type of wave function. The code can describe a large number of physical problems: nuclear fission, collisions of heavy ions, the interaction of quantized vor- tices with nuclei in the nuclear star crust, excitation of superfluid fermion systems by time dependent external fields, quantum shock waves, domain wall generation and propagation, the dynamics of the Anderson-Bogoliubov-Higgs mode, dynamics of fragmented condensates, vortex rings dynamics, generation and dynamics of quantized vortices, their crossing and recombinations and the incipient phases of quantum turbulence.

Jin, Shi↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

DeePMD-kit v2: A software package for deep potential models

DeePMD-kit is a powerful open-source software package that facilitates molecular dynamics simulations using machine learning potentials known as Deep Potential (DP) models. This package, which was released in 2017, has been widely used in the fields of physics, chemistry, biology, and material science for studying atomistic systems. The current version of DeePMD-kit offers numerous advanced features, such as DeepPot-SE, attention-based and hybrid descriptors, the ability to fit tensile properties, type embedding, model deviation, DP-range correction, DP long range, graphics processing unit support for customized operators, model compression, non-von Neumann molecular dynamics, and improved usability, including documentation, compiled binary packages, graphical user interfaces, and application programming interfaces. This article presents an overview of the current major version of the DeePMD-kit package, highlighting its features and technical details. Additionally, this article presents a comprehensive procedure for conducting molecular dynamics as a representative application, benchmarks the accuracy and efficiency of different models, and discusses ongoing developments.

97 MATHEMATICS AND COMPUTING↗

Toward active disruption avoidance via real-time estimation of the safe operating region and disruption proximity in tokamaks

This paper describes a real-time capable algorithm for identifying the safe operating region around a tokamak operating point. The region is defined by a convex set of linear constraints, from which the distance of a point from a disruptive boundary can be calculated. The disruptivity of points is calculated from an empirical machine learning predictor that generates the likelihood of disruption. While the likelihood generated by such empirical models can be compared to a threshold to trigger a disruption mitigation system, the safe operating region calculation enables active optimization of the operating point to maintain a safe margin from disruptive boundaries. The proposed algorithm is tested using a random forest disruption predictor fit on data from DIII-D. The safe operating region identification algorithm is applied to historical data from DIII-D showing the evolution of disruptive boundaries and the potential impact of optimization of the operating point. Real-time relevant execution times are made possible by parallelizing many of the calculation steps and implementing the algorithm on a graphics processing unit. Lastly, a real-time capable algorithm for optimizing the target operating point within the identified constraints is also proposed and simulated.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Vidyut3d: A GPU accelerated fluid solver for non-equilibrium plasmas on adaptive grids

We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure three-electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate ~ 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

GPU acceleration of hybrid functional calculations in the SPARC electronic structure code

We present a Graphics Processing Unit (GPU)-accelerated version of the real-space SPARC electronic structure code for performing hybrid functional calculations in generalized Kohn–Sham density functional theory. In particular, we develop a batch variant of the recently formulated Kronecker product-based linear solver for the simultaneous solution of multiple linear systems. We then develop a modular, math kernel based implementation for hybrid functionals on NVIDIA architectures, where computationally intensive operations are offloaded to the GPUs, while the remaining workload is handled by the central processing units (CPUs). Considering bulk and slab examples, we demonstrate that GPUs enable up to 8× speedup in node-hours and 80× in core-hours compared to CPU-only execution, reducing the time to solution on V100 GPUs to around 300 s for a metallic system with over 6000 electrons, and significantly reducing the computational resources required for a given wall time.

Kohn-Sham density functional theory↗

Thermo4PFM: Facilitating Phase-field simulations of alloys with thermodynamic driving forces

Phase-field modeling is a popular front-tracking approach used to model solidification. Its time-evolution equations are often coupled to alloy composition and/or thermal diffusion in high-resolution multiphysics approaches. Materials thermodynamic properties tabulated in CALPHAD databases can be used for phase-field modeling to parameterize bulk energies of alloys. In addition, they can be naturally integrated into models such as the Kim-Kim-Suzuki (KKS) model where driving forces depend on the differences between chemical potentials of co-existing phases. In that case, a small system of coupled nonlinear equations needs to be solved at every point in space where the phase-field order parameter is to be updated and evolved in time. Here we present Thermo4PFM, a solver for the KKS equations for binary and ternary alloys, with two or three phases, and parameterized with CALPHAD models. Thermo4PFM is open source, written in C++, and can take advantage of Graphics Processing Units (GPU) accelerators. Using OpenMP offload capabilities for C++ classes, an excellent performance is demonstrated on GPU using the LLVM compiler. CALPHAD data is read from simple JSON files using an open source parser from the boost library.

36 MATERIALS SCIENCE↗

A survey of software implementations used by application codes in the Exascale Computing Project

The US Department of Energy Office of Science and the National Nuclear Security Administration initiated the Exascale Computing Project (ECP) in 2016 to prepare mission-relevant applications and scientific software for the delivery of the exascale computers starting in 2023. The ECP currently supports 24 efforts directed at specific applications and six supporting co-design projects. These 24 application projects contain 62 application codes that are implemented in three high-level languages—C, C++, and Fortran—and use 22 combinations of graphical processing unit programming models. The most common implementation language is C++, which is used in 53 different application codes. The most common programming models across ECP applications are CUDA and Kokkos, which are employed in 15 and 14 applications, respectively. This article provides a survey of the programming languages and models used in the ECP applications codebase that will be used to achieve performance on the future exascale hardware platforms.

97 MATHEMATICS AND COMPUTING↗

A parallel and performance portable implementation of a full-field crystal plasticity model

We have developed a parallel implementation of an Elasto-Viscoplastic Fast Fourier Transform-based (EVPFFT) micromechanical solver to enable computationally efficient crystal plasticity modeling for polycrystalline materials. Our primary focus lies in achieving performance portability, allowing a single EVPFFT implementation to run optimally on various homogeneous architectures, including multi-core Central Processing Units (CPUs), as well as on heterogeneous computer architectures comprising multi-core CPUs and Graphics Processing Units (GPUs) from different vendors. To accomplish this goal, we have leveraged MATAR, a C++ software library that simplifies the creation and utilization of multidimensional dense or sparse matrix and array data structures. These data structures are designed to be portable across diverse architectures through the use of Kokkos, a performance-portable library. Additionally, we have employed the Message Passing Interface (MPI) to efficiently distribute the computational workload among processors. The heFFTe (Highly Efficient FFT for Exascale) library is used to facilitate the performance portability of the fast Fourier transforms (FFTs) computation. The computational performance of EVPFFT is evaluated and presented in terms of parallel scalability and simulation runtime on different high-performance computing (HPC) architectures. As a result, the utility of the developed framework to efficiently simulate the micro-mechanical fields in polycrystalline microstructures in engineering applications is discussed.

36 MATERIALS SCIENCE↗

Mesoscopic Modeling and Rapid Simulation of Incremental Changes in Epidemic Scenarios on GPUs

In simulation-based studies and analyses of epidemics, a major challenge lies in resolving the conflict between fidelity of models and the speed of their simulation. Another related challenge arises in dealing with the large number of what–if scenarios that need to be explored. Here, we describe new computational methods that together provide an approach to dealing with both challenges. A mesoscopic modeling approach is described that strikes a middle ground between macroscopic models based on coupled differential equations and microscopic models built on fine-grained behaviors at the individual entity level. The mesoscopic approach offers the ability to incorporate complex compositions of multiple layers of dynamics even while retaining the potential for aggregate behaviors at varying levels. It also is an excellent match to the accelerator-based architectures of modern computing platforms in which graphical processing units (GPUs) can be exploited for fast simulation via the parallel execution mode of single instruction multiple thread (SIMT). The challenge of simulating a large number of scenarios is addressed via a method of sharing model state and computation across a tree of what–if scenarios that are localized, incremental changes to a large base simulation. A combination of the mesoscopic modeling approach and the incremental what–if scenario tree evaluation has been implemented in the software on modern GPUs. Synthetic simulation scenarios are presented to demonstrate the computational characteristics of our approach. Results from the experiments with large population data, including USA, UK, and India, illustrate the modeling methodology and computational performance on thousands of synthetically generated what–if scenarios. Execution of our implementation scaled to 8192 GPUs of supercomputing platforms demonstrates the ability to rapidly evaluate what–if scenarios several orders of magnitude faster than the conventional methods.

97 MATHEMATICS AND COMPUTING↗

Fast GPU-Based Generation of Large Graph Networks From Degree Distributions

Synthetically generated, large graph networks serve as useful proxies to real-world networks for many graph-based applications. The ability to generate such networks helps overcome several limitations of real-world networks regarding their number, availability, and access. Here, we present the design, implementation, and performance study of a novel network generator that can produce very large graph networks conforming to any desired degree distribution. The generator is designed and implemented for efficient execution on modern graphics processing units (GPUs). Given an array of desired vertex degrees and number of vertices for each desired degree, our algorithm generates the edges of a random graph that satisfies the input degree distribution. Multiple runtime variants are implemented and tested: 1) a uniform static work assignment using a fixed thread launch scheme, 2) a load-balanced static work assignment also with fixed thread launch but with cost-aware task-to-thread mapping, and 3) a dynamic scheme with multiple GPU kernels asynchronously launched from the CPU. The generation is tested on a range of popular networks such as Twitter and Facebook, representing different scales and skews in degree distributions. Results show that, using our algorithm on a single modern GPU (NVIDIA Volta V100), it is possible to generate large-scale graph networks at rates exceeding 50 billion edges per second for a 69 billion-edge network. GPU profiling confirms high utilization and low branching divergence of our implementation from small to large network sizes. For networks with scattered distributions, we provide a coarsening method that further increases the GPU-based generation speed by up to a factor of 4 on tested input networks with over 45 billion edges.

97 MATHEMATICS AND COMPUTING↗

Efficient GPU Implementation of Automatic Differentiation for Computational Fluid Dynamics

Many scientific and engineering applications require repeated calculations of derivatives of output functions with respect to input parameters. Automatic Differentiation (AD) is a method that automates derivative calculations and can significantly speed up code development. In Computational Fluid Dynamics (CFD), derivatives of flux functions with respect to state variables (Jacobian) are needed for efficient solutions of the nonlinear governing equations. AD of flux functions on graphics processing units (GPUs) is challenging as flux computations involve many intermediate variables that create high register pressure and require significant memory traffic because of the need to store the derivatives. This paper presents a forward-mode AD method based on multivariate dual numbers that addresses these challenges and simultaneously reduces the floating-point operation count. The dimension of the multivariate dual numbers is optimized for performance. The flux computations are restructured to minimize the number of temporary variables and reduce register pressure. For effective utilization of memory bandwidth, shared memory is used to store the local flux Jacobian. This AD implementation is compared with several other Jacobian implementations on an NVIDIA V100 GPU (V100). For three-dimensional perfect-gas compressible-flow equations implemented in a practical CFD code, the AD implementation of a flux Jacobian based on multivariate dual numbers of dimension 5 outperforms all other GPU AD implementations on V100. Its performance is comparable with the optimized hand-differentiated version. Finally, the implementation achieves 75% of the peak floating-point throughput and 61 % of the peak global device memory bandwidth usage.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Development of a Performance Portable Non-Equilibrium Plasma Fluid Solver on Adaptive Grids

This presentation will describe the numerical techniques, programming paradigms, verification, and performance of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures. Our plasma fluid model solves the conservation equations for self-consistent electrostatic Poisson, electron and heavy species transport, and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive mesh management library, AMReX (Zhang et al., JOSS, 4 (37) 1370, 2019), and can be built and run on widely available vendor specific GPU architectures (NVIDIA/AMD/Intel). We utilize a non-subcycled second order semi-implicit time-stepping method where all adaptive mesh refinement (AMR) levels are advanced with the same time step. The composite multi-level multigrid solver from within AMReX is used for each of the governing equations that are cast into a Helmholtz equation form. We have also developed a python based chemical mechanism parser framework that uses a similar format as CANTERA (Goodwin et al., Zenodo, 2018) yaml files as input. Our custom parser reads the yaml file and provides C++ files with transport and production rate functions that can be executed on both host (CPU) and device (GPU). We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on low-pressure capacitive and high-pressure streamer discharges. Our initial performance studies indicate 10X speed-up using 20 NVIDIA GPUs versus 200 CPUs for an atmospheric streamer discharge problem solved on a 512 x 1024 x 512 grid.

graphics processing units↗

Large-Scale Welding Process Simulation by GPU Parallelized Computing

The computational design of industrially relevant welded structures is extremely time consuming due to coupled physics and high nonlinearity. Previously, most welding distortion and residual stress simulations have been limited to small coupons and reduced order (from three-dimensional [3D] to two-dimensional [2D]), or inherent strain approximations were used for large structures. In this current study, an explicit finite element code based on a graphics processing unit was utilized to perform 3D transient thermomechanical simulation of structural components during welding. Laser brazing of aluminum alloy panels as representative of automotive manufacturing scenarios was simulated to predict out-of-plane distortion under different clamping conditions. The predicted deformation pattern and magnitude were validated by laser scanning data of physical assemblies. In addition, the code was used to investigate residual stresses developed during multipass arc welding of a nuclear industry pressurizer surge nozzle and subsequent welding repair where a 3D simulation was necessary. Taking the experimental data as reference, the 3D model predicted better residual stress distribution than a typical 2D asymmetrical model. Stress evolution in welding repair was also presented and discussed in this study. Furthermore, the efficient numerical model made it feasible to use integrated computational welding engineering to simulate welding processes for large-scale structures.

97 MATHEMATICS AND COMPUTING↗

Software for the frontiers of quantum chemistry: An overview of developments in the Q-Chem 5 package

This article summarizes technical advances contained in the fifth major release of the Q-Chem quantum chemistry program package, covering developments since 2015. A comprehensive library of exchange-correlation functionals, along with a suite of correlated many-body methods, continues to be a hallmark of the Q-Chem software. The many-body methods include novel variants of both coupled-cluster and configuration-interaction approaches along with methods based on the algebraic diagrammatic construction and variational reduced density-matrix methods. Methods highlighted in Q-Chem 5 include a suite of tools for modeling core-level spectroscopy, methods for describing metastable resonances, methods for computing vibronic spectra, the nuclear-electronic orbital method, and several different energy decomposition analysis techniques. High-performance capabilities including multithreaded parallelism and support for calculations on graphics processing units are described. Q-Chem boasts a community of well over 100 active academic developers, and the continuing evolution of the software is supported by an "open teamware" model and an increasingly modular design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Propagation Pattern for Moment Representation of the Lattice Boltzmann Method

A propagation pattern for the moment representation of the regularized lattice Boltzmann method (LBM) in three dimensions is presented. Using effectively lossless compression, the simulation state is stored as a set of moments of the lattice Boltzmann distribution function, instead of the distribution function itself. An efficient cache-aware propagation pattern for this moment representation has the effect of substantially reducing both the storage and memory bandwidth required for LBM simulations. This article extends recent work with the moment representation by expanding the performance analysis on central processing unit (CPU) architectures, considering how boundary conditions are implemented, and demonstrating the effectiveness of the moment representation on a graphics processing unit (GPU) architecture.

42 ENGINEERING↗

The ICON-A model for direct QBO simulations on GPUs (version icon-cscs:baf28a514)

Abstract. Classical numerical models for the global atmosphere, as used for numerical weather forecasting or climate research, have been developed for conventional central processing unit (CPU) architectures. This hinders the employment of such models on current top-performing supercomputers, which achieve their computing power with hybrid architectures, mostly using graphics processing units (GPUs). Thus also scientific applications of such models are restricted to the lesser computer power of CPUs. Here we present the development of a GPU-enabled version of the ICON atmosphere model (ICON-A), motivated by a research project on the quasi-biennial oscillation (QBO), a global-scale wind oscillation in the equatorial stratosphere that depends on a broad spectrum of atmospheric waves, which originates from tropical deep convection. Resolving the relevant scales, from a few kilometers to the size of the globe, is a formidable computational problem, which can only be realized now on top-performing supercomputers. This motivated porting ICON-A, in the specific configuration needed for the research project, in a first step to the GPU architecture of the Piz Daint computer at the Swiss National Supercomputing Centre and in a second step to the JUWELS Booster computer at the Forschungszentrum Jülich. On Piz Daint, the ported code achieves a single-node GPU vs. CPU speedup factor of 6.4 and allows for global experiments at a horizontal resolution of 5 km on 1024 computing nodes with 1 GPU per node with a turnover of 48 simulated days per day. On JUWELS Booster, the more modern hardware in combination with an upgraded code base allows for simulations at the same resolution on 128 computing nodes with 4 GPUs per node and a turnover of 133 simulated days per day. Additionally, the code still remains functional on CPUs, as is demonstrated by additional experiments on the Levante compute system at the German Climate Computing Center. While the application shows good weak scaling over the tested 16-fold increase in grid size and node count, making also higher resolved global simulations possible, the strong scaling on GPUs is relatively poor, which limits the options to increase turnover with more nodes. Initial experiments demonstrate that the ICON-A model can simulate downward-propagating QBO jets, which are driven by wave–mean flow interaction.

54 ENVIRONMENTAL SCIENCES↗