Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphical processing unit”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Electron microscopy holdings of the Protein Data Bank: the impact of the resolution revolution, new validation tools, and implications for the future

Abstract As a discipline, structural biology has been transformed by the three-dimensional electron microscopy (3DEM) “Resolution Revolution” made possible by convergence of robust cryo-preservation of vitrified biological materials, sample handling systems, and measurement stages operating a liquid nitrogen temperature, improvements in electron optics that preserve phase information at the atomic level, direct electron detectors (DEDs), high-speed computing with graphics processing units, and rapid advances in data acquisition and processing software. 3DEM structure information (atomic coordinates and related metadata) are archived in the open-access Protein Data Bank (PDB), which currently holds more than 11,000 3DEM structures of proteins and nucleic acids, and their complexes with one another and small-molecule ligands (~ 6% of the archive). Underlying experimental data (3DEM density maps and related metadata) are stored in the Electron Microscopy Data Bank (EMDB), which currently holds more than 21,000 3DEM density maps. After describing the history of the PDB and the Worldwide Protein Data Bank (wwPDB) partnership, which jointly manages both the PDB and EMDB archives, this review examines the origins of the resolution revolution and analyzes its impact on structural biology viewed through the lens of PDB holdings. Six areas of focus exemplifying the impact of 3DEM across the biosciences are discussed in detail (icosahedral viruses, ribosomes, integral membrane proteins, SARS-CoV-2 spike proteins, cryogenic electron tomography, and integrative structure determination combining 3DEM with complementary biophysical measurement techniques), followed by a review of 3DEM structure validation by the wwPDB that underscores the importance of community engagement.

Burley, Stephen K. (ORCID:0000000224879713)↗

A Comparative Study of Layer Heating and Continuous Heating Methods on Prediction Accuracy of Residual Stresses in Selective Laser Melted Tube Samples

Thermal distortion and residual stresses are two important factors that affect the quality and reliability of steel parts manufactured by laser powder bed fusion (LPBF) processes. A cost-effective model for evaluation of those heat effects is needed to refine the manufacturing process and provides insights into the product design and heat treatment. In this study, the layer heating method and sophisticated track-layer scanning method were applied to simulate the thermo-mechanical response of IN625 tube parts built by LPBF. Based on the similarity of temperature field in each layer deposit, a swept mesh was constructed to perform the thermal analysis for top layer, with the rest of layers referring to the temperature by node number offsetting. A novel explicit finite element analysis code accelerated by graphics processing unit was used for the massive-element numerical analysis. The computational accuracy and efficiency of the layer heating and track-layer scanning methods were compared in detail. It is shown that layer heating method can efficiently capture the pattern of stress distribution with reasonable accuracy in stress magnitude. The grouped track-layer scanning method can predict the residual stress and strain more accurately at a higher cost (5 ~ 10×). The elastic strain distribution was compared with the measurement by X-ray diffraction, confirming the accuracy of residual stress prediction.

36 MATERIALS SCIENCE↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗

Direct numerical simulations of turbulent reacting flows with shock waves and stiff chemistry using many-core/GPU acceleration

Compressible reacting flows may display sharp spatial variation related to shocks, contact discontinuities or reactive zones embedded within relatively smooth regions. The presence of such phenomena emphasizes the relevance of shock-capturing schemes such as the weighted essentially non-oscillatory (WENO) scheme as an essential ingredient of the numerical solver. However, these schemes are complex and have more computational cost than the simple high-order compact or non-compact schemes. In this paper, we present the implementation of a seventh-order, minimally-dissipative mapped WENO (WENO7M) scheme in a newly developed direct numerical simulation (DNS) code called KAUST Adaptive Reactive Flows Solver (KARFS). In order to make efficient use of the computer resources and reduce the solution time, without compromising the resolution requirement, the WENO routines are accelerated via graphics processing unit (GPU) computation. The performance characteristics and scalability of the code are studied using different grid sizes and block decomposition. Furthermore, the performance portability of KARFS is demonstrated on a variety of architectures including NVIDIA Tesla P100 GPUs and NVIDIA Kepler K20X GPUs. In addition, the capability and potential of the newly implemented WENO7M scheme in KARFS to perform DNS of compressible flows is also demonstrated with model problems involving shocks, isotropic turbulence, detonations and flame propagation into a stratified mixture with complex chemical kinetics.

97 MATHEMATICS AND COMPUTING↗

Recent advances in describing and driving crystal nucleation using machine learning and artificial intelligence

With the advent of faster computer processors and especially graphics processing units (GPUs) over the last few decades, the use of data-intensive machine learning (ML) and artificial intelligence (AI) has increased greatly, and the study of crystal nucleation has been one of the beneficiaries. In this study, we outline how ML and AI have been applied to address four outstanding difficulties of crystal nucleation: how to discover better reaction coordinates (RCs) for describing accurately non-classical nucleation situations; the development of more accurate force fields for describing the nucleation of multiple polymorphs or phases for a single system; more robust identification methods for determining crystal phases and structures; and as a method to yield improved course-grained models for studying nucleation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerated impurity solver for DMFT and its diagrammatic extensions

Here, we present ComCTQMC, a GPU accelerated quantum impurity solver. It uses the continuous-time quantum Monte Carlo (CTQMC) algorithm wherein the partition function is expanded in terms of the hybridisation function (CT-HYB). ComCTQMC supports both partition and worm-space measurements, and it uses improved estimators and the reduced density matrix to improve observable measurements whenever possible. ComCTQMC efficiently measures all one and two-particle Green's functions, all static observables which commute with the local Hamiltonian, and the occupation of each impurity orbital. ComCTQMC can solve complex-valued impurities with crystal fields that are hybridized to both fermionic and bosonic baths. Most importantly, ComCTQMC utilizes graphical processing units (GPUs), if available, to dramatically accelerate the CTQMC algorithm when the Hilbert space is sufficiently large. We demonstrate acceleration by a factor of over 600 (100) in a simulation of δ-Pu at 600 K with (without) crystal fields. In easier problems, the GPU offers less impressive acceleration or even decelerates the CTQMC. Here we describe the theory, algorithms, and structure used by ComCTQMC in order to achieve this set of features and level of acceleration.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

TensorBNN: Bayesian inference for neural networks using TensorFlow

We report that TensorBNN is a new package based on TensorFlow that implements Bayesian inference for modern neural network models. The posterior density of neural network model parameters is represented as a point cloud sampled using Hamiltonian Monte Carlo. The TensorBNN package leverages TensorFlow's architecture and its ability to use modern graphics processing units in both the training and prediction stages.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

TTDFT: A GPU accelerated Tucker tensor DFT code for large-scale Kohn-Sham DFT calculations

We present the Tucker tensor DFT (TTDFT) code which uses a tensor-structured algorithm with graphic processing unit (GPU) acceleration for conducting ground-state DFT calculations on large-scale systems. The Tucker tensor DFT algorithm uses a localized Tucker tensor basis computed from an additive separable approximation to the Kohn-Sham Hamiltonian. The discrete Kohn-Sham problem is solved using Chebyshev filtered subspace iteration method that relies on matrix-matrix multiplications of a sparse symmetric Hamiltonian matrix and a dense wavefunction matrix, expressed in the localized Tucker tensor basis. These matrix-matrix multiplication operations, which constitute the most computationally intensive step of the solution procedure, are GPU accelerated providing ~8-fold GPU-CPU speedup for these operations on the largest systems studied. In conclusion, the computational performance of the TTDFT code is presented using benchmark studies on aluminum nano-particles and silicon quantum dots with system sizes ranging up to ~7,000 atoms.

97 MATHEMATICS AND COMPUTING↗

hPIC2: A hardware-accelerated, hybrid particle-in-cell code for dynamic plasma-material interactions

The exascale era of high performance computing promises to bring the field of computational plasma physics ever closer to the goal of accurate multiscale modeling. Such computers will rely on hardware acceleration to offload work to dedicated components, notably general-purpose graphics processing units (GPUs). However, devices from different manufacturers require software to be written with different parallel programming models, greatly increasing the code maintenance burden of applications designed to perform on more than one such device. hPIC2 is a hybrid plasma simulation code developed with the Kokkos performance portability framework to target the architectures that will drive exascale computing for the foreseeable future. As a hybrid simulation code, hPIC2 investigates the simultaneous use of various plasma models on the same domain, at the same time. hPIC2 also optionally couples to RustBCA, which accurately models ion-material interactions using the binary collision approximation (BCA) method. In conclusion, hPIC2 therefore achieves scalable performance on a variety of computing architectures when simulating complex and diverse plasmas, particularly near plasma-material interfaces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

GPU-acceleration of tensor renormalization with PyTorch using CUDA

We show that numerical computations based on tensor renormalization group (TRG) methods can be significantly accelerated with PyTorch on graphics processing units (GPUs) by leveraging NVIDIA's Compute Unified Device Architecture (CUDA). Here we find improvement in the runtime and its scaling with bond dimension for two-dimensional systems. Our results establish that the utilization of GPU resources is essential for future precision computations with TRG.

97 MATHEMATICS AND COMPUTING↗

Harnessing distributed GPU computing for generalizable graph convolutional networks in power grid reliability assessments

Although machine learning (ML) has emerged as a powerful tool for rapidly assessing grid contingencies, prior studies have largely considered a static grid topology in their analyses. This limits their application, since they need to be re-trained for every new topology. Here, this paper explores the development of generalizable graph convolutional network (GCN) models by pre-training them across a range of grid topologies and contingency types. We found that a GCN model with auto-regressive moving average (ARMA) layers with a line graph representation of the grid offered the best predictive performance in predicting voltage magnitudes (VM) and voltage angles (VA). We introduced the concept of phantom nodes to consider disparate grid topologies with a varying number of nodes and lines. For pre-training the GCN ARMA model across a variety of topologies, distributed graphics processing unit (GPU) computing afforded us significant training scalability. The predictive performance of this model on grid topologies that were part of the training data is substantially better than the direct current (DC) approximation. Although direct application of the pre-trained model to topologies that are not part of the grid is not particularly satisfactory, fine-tuning with small amounts of data from a specific topology of interest significantly improves predictive performance. In general, this paper highlights the feasibility of training large-scale GNN models to assess the reliability of power grids by considering a wide variety of grid topologies and contingency types. With the advent of foundational models in ML and the exponential increase in GPU computing clusters, generalizable ML models will significantly enhance how utilities manage power systems and make decisions in real-time or near-real-time.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An automated and portable method for selecting an optimal GPU frequency

Power consumption poses a significant challenge in current and emerging graphics processing unit (GPU) enabled high-performance computing systems. In modern GPUs, dynamic voltage frequency scaling (DVFS) appears to be a reliable control to regulate power consumption and performance. However, the DVFS design space is large - hence, brute-force approaches are infeasible to select the optimal frequency. Furthermore, no single frequency can be universally optimal for applications with varying computational intensities. Thus, the application's complexity and the availability of a wide range of frequency settings are a challenge in selecting the optimal frequency configuration for a given GPU workload. To that end, this paper proposes a systematic approach that consists of three steps. The feature characterization study identifies the fine-grain GPU utilization metrics that influence the power consumption and execution time of a given workload. To understand the performance, power, and energy consumption behaviors of a workload across GPU's DVFS design space, we derived analytical power and performance models using the identified fine-grain features. Here, it is shown that the same set of GPU utilization metrics can estimate both the power consumption and execution time while being agnostic of changes to frequency and input sizes. Applying a power control with the single objective of reducing power may cause performance degradation, leading to more energy consumption. A multi-objective approach is proposed to select the optimal GPU DVFS configuration for a workload that reduces power consumption with negligible degradation in performance. The evaluation was conducted using SPEC ACCEL benchmarks and three real applications - NAMD LAMMPS, and LSTM on NVIDIA GV100, GA100, and AMD MI210 GPUs. On average, real applications showed 29.6% energy savings with a performance loss of 5.2% on GA100 and 22.6% energy savings with a performance loss of 4.7% on GV100. Moreover, the proposed models are portable to real applications, GPU architectures, and vendors, and require metric collection at only the default frequency rather than all supported DVFS configurations. Additionally, we conducted a comparison between our models and the GPU assembly instructions (PTX)-based static models. The results revealed a significant reduction in the average error rates, with a decrease from 19.7% to 3.1% for power models and from 29.4% to 5.2% for performance models.

97 MATHEMATICS AND COMPUTING↗

GPU-resident sparse direct linear solvers for alternating current optimal power flow analysis

Integrating renewable resources within the transmission grid at a wide scale poses significant challenges for economic dispatch as it requires analysis with more optimization parameters, constraints, and sources of uncertainty. This motivates the investigation of more efficient computational methods, especially those for solving the underlying linear systems, which typically take more than half of the overall computation time. In this paper, we present our work on sparse linear solvers that take advantage of hardware accelerators, such as graphical processing units (GPUs), and improve the overall performance when used within economic dispatch computations. We treat the problems as sparse, which allows for faster execution but also makes the implementation of numerical methods more challenging. We present the first GPU-native sparse direct solver that can execute on both AMD and NVIDIA GPUs. We demonstrate significant performance improvements when using high-performance linear solvers within alternating current optimal power flow (ACOPF) analysis. Furthermore, we demonstrate the feasibility of getting significant performance improvements by executing the entire computation on GPU-based hardware. Finally, we identify outstanding research issues and opportunities for even better utilization of heterogeneous systems, including those equipped with GPUs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Stabilized bases for high-order, interpolation semi-Lagrangian, element-based tracer transport

In a computational fluid model of the atmosphere, the advective transport of trace species, or tracers, can be computationally expensive. For efficiency, models often use semi-Lagrangian advection methods. High-order interpolation semi-Lagrangian (ISL) methods, in particular, can be extremely efficient, if the problem of property preservation specific to them can be addressed. Atmosphere models often use geometrically and logically nonuniform grids for efficiency and, as a result, element-based discretizations. Such grids and discretizations make stability a particular problem for ISL methods. Generally, high-order, element-based ISL methods that use the natural polynomial interpolant associated with a nodal finite-element discretization are unstable. Here, we derive new bases having order of accuracy up to nine, with positive nodal weights, that stabilize the element-based ISL method. We use these bases to construct the linear advection operator in the property-preserving Interpolation Semi-Lagrangian Element-based Transport (Islet) method. Then we discuss key software implementation details. Finally, we show performance results for the Energy Exascale Earth System Model's atmosphere dynamical core, comparing the original and new transport methods. These simulations used up to 27,600 Graphical Processing Units (GPU) on the Oak Ridge Leadership Computing Facility's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗