Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphics processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

An automated and portable method for selecting an optimal GPU frequency

Power consumption poses a significant challenge in current and emerging graphics processing unit (GPU) enabled high-performance computing systems. In modern GPUs, dynamic voltage frequency scaling (DVFS) appears to be a reliable control to regulate power consumption and performance. However, the DVFS design space is large - hence, brute-force approaches are infeasible to select the optimal frequency. Furthermore, no single frequency can be universally optimal for applications with varying computational intensities. Thus, the application's complexity and the availability of a wide range of frequency settings are a challenge in selecting the optimal frequency configuration for a given GPU workload. To that end, this paper proposes a systematic approach that consists of three steps. The feature characterization study identifies the fine-grain GPU utilization metrics that influence the power consumption and execution time of a given workload. To understand the performance, power, and energy consumption behaviors of a workload across GPU's DVFS design space, we derived analytical power and performance models using the identified fine-grain features. Here, it is shown that the same set of GPU utilization metrics can estimate both the power consumption and execution time while being agnostic of changes to frequency and input sizes. Applying a power control with the single objective of reducing power may cause performance degradation, leading to more energy consumption. A multi-objective approach is proposed to select the optimal GPU DVFS configuration for a workload that reduces power consumption with negligible degradation in performance. The evaluation was conducted using SPEC ACCEL benchmarks and three real applications - NAMD LAMMPS, and LSTM on NVIDIA GV100, GA100, and AMD MI210 GPUs. On average, real applications showed 29.6% energy savings with a performance loss of 5.2% on GA100 and 22.6% energy savings with a performance loss of 4.7% on GV100. Moreover, the proposed models are portable to real applications, GPU architectures, and vendors, and require metric collection at only the default frequency rather than all supported DVFS configurations. Additionally, we conducted a comparison between our models and the GPU assembly instructions (PTX)-based static models. The results revealed a significant reduction in the average error rates, with a decrease from 19.7% to 3.1% for power models and from 29.4% to 5.2% for performance models.

97 MATHEMATICS AND COMPUTING↗

GPU-resident sparse direct linear solvers for alternating current optimal power flow analysis

Integrating renewable resources within the transmission grid at a wide scale poses significant challenges for economic dispatch as it requires analysis with more optimization parameters, constraints, and sources of uncertainty. This motivates the investigation of more efficient computational methods, especially those for solving the underlying linear systems, which typically take more than half of the overall computation time. In this paper, we present our work on sparse linear solvers that take advantage of hardware accelerators, such as graphical processing units (GPUs), and improve the overall performance when used within economic dispatch computations. We treat the problems as sparse, which allows for faster execution but also makes the implementation of numerical methods more challenging. We present the first GPU-native sparse direct solver that can execute on both AMD and NVIDIA GPUs. We demonstrate significant performance improvements when using high-performance linear solvers within alternating current optimal power flow (ACOPF) analysis. Furthermore, we demonstrate the feasibility of getting significant performance improvements by executing the entire computation on GPU-based hardware. Finally, we identify outstanding research issues and opportunities for even better utilization of heterogeneous systems, including those equipped with GPUs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Stabilized bases for high-order, interpolation semi-Lagrangian, element-based tracer transport

In a computational fluid model of the atmosphere, the advective transport of trace species, or tracers, can be computationally expensive. For efficiency, models often use semi-Lagrangian advection methods. High-order interpolation semi-Lagrangian (ISL) methods, in particular, can be extremely efficient, if the problem of property preservation specific to them can be addressed. Atmosphere models often use geometrically and logically nonuniform grids for efficiency and, as a result, element-based discretizations. Such grids and discretizations make stability a particular problem for ISL methods. Generally, high-order, element-based ISL methods that use the natural polynomial interpolant associated with a nodal finite-element discretization are unstable. Here, we derive new bases having order of accuracy up to nine, with positive nodal weights, that stabilize the element-based ISL method. We use these bases to construct the linear advection operator in the property-preserving Interpolation Semi-Lagrangian Element-based Transport (Islet) method. Then we discuss key software implementation details. Finally, we show performance results for the Energy Exascale Earth System Model's atmosphere dynamical core, comparing the original and new transport methods. These simulations used up to 27,600 Graphical Processing Units (GPU) on the Oak Ridge Leadership Computing Facility's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

A new efficient grain growth model using a random Gaussian-sampled mode filter

This paper presents the use of a Gaussian neighborhood mode filter for predicting grain growth in a manner similar to the solutions obtained by a Monte Carlo Potts model. This flexible grain growth model can quickly utilize modern, computationally optimized data science strategies on graphics processing units to simulate grain growth up to 100 times faster than the state-of-the-art, publicly available Monte Carlo Potts model. We show that, given the correct neighborhood, the mode filter can replicate normal grain growth in two or three dimensions. In addition, the paper briefly demonstrates the ability to model limited anisotropic in grain boundary energy and mobility. Anisotropic grain boundary energy is modeled by defining a weighted mode filter operation. Anisotropic grain boundary mobility is modeled by scaling and orienting the Gaussian neighborhood in a particular direction.

Anisotropy↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

SPH modeling of biomass granular flow: Theoretical implementation and experimental validation

The commercialization of biomass-derived energy is impeded by flowability challenges arising from the feeding and handling of granular biomass materials in full-scale biorefineries. To overcome these obstacles, a robust and accurate model to simulate the flow of granular biomass is indispensable. However, conventional mesh-based numerical codes are limited by inherent mesh distortion in simulating large deformation that commonly occurs in granular biomass handling. Here, in this study, we propose a graphics processing unit (GPU)-accelerated meshless Smoothed Particle Hydrodynamics (SPH) code to model the flow of granular biomass materials. A modified void ratio-based mass conversation, a hybrid particle-to-particle/surface frictional boundary treatment, and a hypoplastic constitutive model are implemented. Four numerical examples, an elastic block sliding on inclined planes, sand column collapse, Angle of Repose, and axial compression tests for pine chips, were simulated using the developed SPH code. The results demonstrate good agreement between numerical predictions and analytical and experimental data for all four examples, validating the SPH code and increasing confidence that it can be applied to simulate more complex granular biomass handling processes, such as hopper feeding or auger conveyance.

09 BIOMASS FUELS↗

A High-Performance Discrete-Element Framework for Simulating Flow and Jamming of Moisture Bearing Biomass Feedstocks

We developed and verified a high-performance open-source discrete element method (DEM) solver with simultaneously-supported feedstock-specific interaction models, including bonded-sphere, liquid bridge, cohesion, and non-linear contact models. Our solver uses parallel data structures on hybrid central and graphics processing unit (CPU/GPU) architectures, with favorable strong scaling performance observed for large problem sizes comprised of (100 M particles), and 4X single-node GPU speedup. The particles for corn stover feedstock were conceptualized and calibrated based on experimental measurements and results. Sensitivity analyses demonstrate that the mass flow rate from a wedge hopper is governed primarily by moisture content, friction coefficient, and cohesion energy density. The model is used to reproduce experimentally observed hopper jamming results, highlighting that the experimental no-flow trends can only be achieved by using non-spherical particles, liquid bridge and cohesion models, highlighting the importance of using concurrent feedstock specialized models for the effective representation of biomass material handling problems.

bioenergy↗

Uncertainty quantification for Multiphase-CFD simulations of bubbly flows: a machine learning-based Bayesian approach supported by high-resolution experiments

In this paper, we developed a machine learning-based Bayesian approach to inversely quantify and reduce the uncertainties of multiphase computational fluid dynamics (MCFD) simulations for bubbly flows. The proposed approach is supported by high-resolution two-phase flow measurements, including those by double-sensor conductivity probes, high-speed imaging, and particle image velocimetry. Local distributions of key physical quantities of interest (QoIs), including the void fraction and phasic velocities, are obtained to support the Bayesian inference. In the process, the epistemic uncertainties of the closure relations are inversely quantified while the aleatory uncertainties from stochastic fluctuations of the system are evaluated based on experimental uncertainty analysis. The combined uncertainties are then propagated through the MCFD solver to obtain uncertainties of the QoIs, based on which probability-boxes are constructed for validation. The proposed approach relies on three machine learning methods: feedforward neural networks and principal component analysis for surrogate modeling, and Gaussian processes for model form uncertainty modeling. The whole process is implemented within the framework of an open-source deep learning library PyTorch with graphics processing unit (GPU) acceleration, thus ensuring the efficiency of the computation. The results demonstrate that with the support of high-resolution data, the uncertainties of MCFD simulations can be significantly reduced. The proposed approach has the potential for other applications that involve numerical models with empirical parameters.

42 ENGINEERING↗

Interpretable artificial intelligence and exascale molecular dynamics simulations to reveal kinetics: Applications to Alzheimer's disease

The rapid increase in computing power, especially with the integration of graphics processing units, has dramatically increased the capabilities of molecular dynamics simulations. To date, these capabilities extend from running very long simulations (tens to hundreds of microseconds) to thousands of short simulations. However, the expansive data generated in these simulations must be made interpretable not only by the investigator who performs them but also by others as well. Here, we demonstrate how integrating learning techniques, such as artificial intelligence, machine learning, and neural networks, into analysis pipelines can reveal the kinetics of Alzheimer's disease (AD) protein aggregation. Finally, we review select AD targets, describe current simulation methods, and introduce learning concepts and their application in AD, highlighting limitations and potential solutions.

59 BASIC BIOLOGICAL SCIENCES↗

From NWChem to NWChemEx: Evolving with the Computational Chemistry Landscape

Since the advent of the first computers, chemists have been at the forefront of using computers to understand and solve complex chemical problems. As the hardware and software have evolved, so have the theoretical and computational chemistry methods and algorithms. Parallel computers clearly changed the common computing paradigm in the late 1970s and 80s, and the field has again seen a paradigm shift with the advent of graphical processing units. This review explores the challenges and some of the solutions in transforming software from the terascale to the petascale and now to the upcoming exascale computers. While discussing the field in general, NWChem and its redesign, NWChemEx, will be highlighted as one of the early codesign projects to take advantage of massively parallel computers and emerging software standards to enable large scientific challenges to be tackled.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerators for Classical Molecular Dynamics Simulations of Biomolecules

Atomistic Molecular Dynamics (MD) simulations provide researchers the ability to model biomolecular structures such as proteins and their interactions with drug-like small molecules with greater spatiotemporal resolution than is otherwise possible using experimental methods. MD simulations are notoriously expensive computational endeavors that have traditionally required massive investment in specialized hardware to access biologically relevant spatiotemporal scales. Our goal is to summarize the fundamental algorithms that are employed in the literature to then highlight the challenges that have affected accelerator implementations in practice. We consider three broad categories of accelerators: Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Application Specific Integrated Circuits (ASICs). These categories are comparatively studied to facilitate discussion of their relative trade-offs and to gain context for the current state of the art. We conclude by providing insights into the potential of emerging hardware platforms and algorithms for MD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

I ntera C hem : Exploring Excited States in Virtual Reality with Ab Initio Interactive Molecular Dynamics

InteraChem is an ab initio interactive molecular dynamics (AI-IMD) visualizer that leverages recent advances in virtual reality hardware and software, as well as the graphical processing unit (GPU)-accelerated TeraChem electronic structure package, in order to render quantum chemistry in real time. We introduce the exploration of electronically excited states via AI-IMD using the floating occupation molecular orbital-complete active space configuration interaction method. The optimization tools in InteraChem enable identification of excited state minima as well as minimum energy conical intersections for further characterization of excited state chemistry in small- to medium-sized systems. We demonstrate that finite-temperature Hartree–Fock theory is an efficient method to perform ground state AI-IMD. InteraChem allows users to track electronic properties such as molecular orbitals and bond order in real time, resulting in an interactive visualization tool that aids in the interpretation of excited state chemistry data and makes quantum chemistry more accessible for both research and educational purposes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The General Atomic and Molecular Electronic Structure System (GAMESS): Novel Methods on Novel Architectures

The primary focus of GAMESS over the last 5 years has been the development of new high-performance codes that are able to take effective and efficient advantage of the most advanced computer architectures, both CPU and accelerators. These efforts include employing density fitting and fragmentation methods to reduce the high scaling of well-correlated (e.g., coupled-cluster) methods as well as developing novel codes that can take optimal advantage of graphical processing units and other modern accelerators. Because accurate wave functions can be very complex, an important new functionality in GAMESS is the quasi-atomic orbital analysis, an unbiased approach to the understanding of covalent bonds embedded in the wave function. Finally, best practices for the maintenance and distribution of GAMESS are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Speeding Up Hartree–Fock in JuliaChem with Density Fitting

In this work, the density fitting (DF) approximation is added to the restricted Hartree–Fock (RHF) implementation in the JuliaChem computational chemistry code. Utilizing a DF algorithm that uses symmetry and integral screening, a significant reduction in time to compute the Fock matrix is achieved. The symmetry and screening DF-RHF techniques were adapted to be performed on graphics processing units (GPUs), which are well suited to perform the matrix multiplications that comprise the bulk of the Fock build time in DF-RHF. The JuliaChem DF-RHF GPU algorithm employs a novel approach that automatically switches between two DF-RHF algorithms depending on the number of basis functions in the calculation. The JuliaChem GPU DF-RHF implementation demonstrates up to 2× speedup for Fock build times compared to the existing best-in-class GPU DF-RHF implementation by operating directly on screened intermediate matrices. Due to the high portability of the Julia language code, the JuliaChem CPU and GPU DF-RHF implementations could be benchmarked on a variety of CPU and GPU architectures from multiple hardware vendors.

Hayes, John J. [Ames Laboratory, and Iowa State Un↗

Simulating the Excited-State Dynamics of Polaritons with Ab Initio Multiple Spawning

Over the past decade, there has been a growth of interest in polaritonic chemistry, where the formation of hybrid light–matter states (polaritons) can alter the course of photochemical reactions. These hybrid states are created by strong coupling between molecules and photons in resonant optical cavities and can even occur in the absence of light when the molecule is strongly coupled with the electromagnetic fluctuations of the vacuum field. Here, we present a first-principles model to simulate nonadiabatic dynamics of such polaritonic states inside optical cavities by leveraging graphical processing units (GPUs). Our first implementation of this model is specialized for a single molecule coupled to a single-photon mode confined inside the optical cavity but with any number of excited states computed using complete active space configuration interaction (CASCI) and a Jaynes–Cummings-type Hamiltonian. Using this model, we have simulated the excited-state dynamics of a single salicylideneaniline (SA) molecule strongly coupled to a cavity photon with the ab initio multiple spawning (AIMS) method. We demonstrate how the branching ratios of the photodeactivation pathways for this molecule can be manipulated by coupling to the cavity. We also show how one can stop the photoreaction from happening inside of an optical cavity. Finally, we also investigate cavity-based control of the ordering of two excited states (one optically bright and the other optically dark) inside a cavity for a set of molecules, where the dark and bright states are close in energy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗