Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gpu”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

PREEMPT: Scalable Epidemic Interventions Using Submodular Optimization on Multi-GPU Systems

Preventing and slowing the spread of epidemics is achieved through techniques such as vaccination and social distancing. Given practical limitations on the number of vaccines and cost of administration, optimization becomes a necessity. Previous approaches using mathematical programming methods have shown to be effective but are limited by computational costs. In this work, we make several contributions: First, we present a new approach for intervention via maximizing the influence of vaccinated nodes on the network. We call this method \preempt. Next, we prove submodular properties associated with the objective function of our method so that it aids in construction of an efficient greedy approximation strategy. Consequently, we present a new parallel algorithm based on greedy hill climbing for \preempt, and present an efficient parallel implementation for distributed CPU-GPU heterogeneous platforms. Our results demonstrate that \preempt{} is able to achieve a significant reduction (up to 6.75$\times$) in the percentage of people infected on a city-scale network. We also show strong scaling results of \preempt{} on 128 nodes of the Summit supercomputer. Our parallel implementation is able to significantly reduce time to solution, from hours to minutes on large networks. This work represents a first-of-its-kind effort in parallelizing greedy hill climbing and applying it toward devising effective interventions for epidemics.

Minutoli, Marco↗

thornado-transport: Anderson- and GPU-accelerated nonlinear solvers for neutrino-matter coupling

Algorithms for neutrino-matter coupling in core-collapse supernovae (CCSNe) are investigated in the context of a spectral two-moment model, which is discretized in space with the discontinuous Galerkin method, integrated in time with implicit-explicit (IMEX) methods, and implemented in the toolkit for high-order neutrino-radiation hydrodynamics (thornado). The model considers electron neutrinos and antineutrinos and tabulated opacities from Bruenn (1985), which includes neutrino-electron scattering and pair processes. The nonlinear system arising from implicit time discretization of the equations governing neutrino-matter coupling is iterated to convergence using Anderson-accelerated fixed-point methods, which avoid formation of Jacobians and inversion of dense linear systems. Numerical experiments show that, for a given tolerance, a nested iteration scheme which aims to reduce opacity evaluations can lower the computational cost. Our initial port to GPUs, using both OpenMP and OpenACC, shows an overall speedup of up to ~ 100× when compared to results using a single CPU core. These results indicate that the algorithms implemented in thornado are well-suited to GPU acceleration.

Laiu, Paul↗

GPU-accelerated multitiered iterative phasing algorithm for fluctuation X-ray scattering

The multitiered iterative phasing (MTIP) algorithm is used to determine the biological structures of macromolecules from fluctuation scattering data. It is an iterative algorithm that reconstructs the electron density of the sample by matching the computed fluctuation X-ray scattering data to the external observations, and by simultaneously enforcing constraints in real and Fourier space. This paper presents the first ever MTIP algorithm acceleration efforts on contemporary graphics processing units (GPUs). The Compute Unified Device Architecture (CUDA) programming model is used to accelerate the MTIP algorithm on NVIDIA GPUs. The computational performance of the CUDA-based MTIP algorithm implementation outperforms the CPU-based version by an order of magnitude. Furthermore, the Heterogeneous-Compute Interface for Portability (HIP) runtime APIs are used to demonstrate portability by accelerating the MTIP algorithm across NVIDIA and AMD GPUs.

97 MATHEMATICS AND COMPUTING↗

Hybrid simulation of energetic particles interacting with magnetohydrodynamics using a slow manifold algorithm and GPU acceleration

The hybrid method combining particle-in-cell and magnetohydrodynamics can be used to study the interaction between energetic particles and global plasma modes. In this paper we introduce the M3D-C1-K code, which is developed based on the M3D-C1 finite element code solving the magnetohydrodynamics equations, with a newly developed kinetic module simulating energetic particles. The particle pushing is done using a new algorithm by applying the Boris pusher to the classical Pauli particles to simulate the slow-manifold of particle orbits, with long-term accuracy and fidelity. The particle pushing can be accelerated using GPUs with a significant speedup. The moments of the particles are calculated using the δƒ method, and are coupled into the magnetohydrodynamics simulation through pressure or current coupling schemes. Several linear simulations of magnetohydrodynamics modes driven by energetic particles have been conducted using M3D-C1-K with the δƒ method, including fishbone, toroidal Alfvén eigenmodes and reversed shear Alfvén eigenmodes. Good agreement with previous results from other eigenvalue, kinetic and hybrid codes have been achieved.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

Asynchronous GPU-based DEM solver embedded in commercial CFD software with polyhedral mesh support

A novel graphical processing unit-based discrete element method solver is introduced to improve stability, performance, and provide seamless integration into commercial or open-source computational fluid dynamics software. A key innovation is eliminating a need for network communication between solvers, which was previously required for cross-platform coupling. This is accomplished by a direct coupling method that employs dynamic-linked libraries. Furthermore, the solver optimizes memory usage by streamlining the particle-cell search algorithm by eliminating the cells' searching grid. This ensures the solver is compatible with a wide range of mesh types, providing high geometric flexibility. The approach simplifies the simulation process by directly incorporating computational fluid dynamics mesh information into the discrete element method solver. The performance analysis indicates about sixteen times boost in computational speed compared to benchmark central processing unit-based solvers. Finally, the solver's compatibility with polyhedral meshes, a vital advantage for complex geometries, is tested against a referenced study regarding the simulation of an immersed-tube fluidized bed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

GPU acceleration of Swendsen–Wang dynamics

When simulating a lattice system near its critical temperature, local algorithms for modeling the system’s evolution can introduce very large autocorrelation times into sampled data. Here, this critical slowing down places restrictions on the analysis that can be completed in a timely manner of the behavior of systems around the critical point. Because it is often desirable to study such systems around this point, a new algorithm must be introduced. Therefore, we turn to cluster algorithms, such as the Swendsen–Wang algorithm and the Wolff clustering algorithm. They incorporate global updates which generate new lattice configurations with little correlation to previous states, even near the critical point. We look to accelerate the rate at which these algorithm are capable of running by implementing and benchmarking a parallel implementation of each algorithm designed to run on GPUs under NVIDIA’s CUDA framework. A 17 and 90 fold increase in the computational rate was, respectively, experienced when measured against the equivalent algorithm implemented in serial code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

bolt: A GPU-Accelerated Solver for Marginally Collisional Kinetic Physics Using a High- Resolution Constrained Transport Scheme

Collisionless and marginally collisional plasmas are very expensive to simulate due to high dimensionality and large scale separations, especially in global problems with additional length scales. bolt addresses this challenge by achieving high performance on GPUs using the arrayfire performance portability library. We have demonstrated accuracy on a range of test problems, and the bolt code paper is currently in preparation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large-Scale Welding Process Simulation by GPU Parallelized Computing

The computational design of industrially relevant welded structures is extremely time consuming due to coupled physics and high nonlinearity. Previously, most welding distortion and residual stress simulations have been limited to small coupons and reduced order (from three-dimensional [3D] to two-dimensional [2D]), or inherent strain approximations were used for large structures. In this current study, an explicit finite element code based on a graphics processing unit was utilized to perform 3D transient thermomechanical simulation of structural components during welding. Laser brazing of aluminum alloy panels as representative of automotive manufacturing scenarios was simulated to predict out-of-plane distortion under different clamping conditions. The predicted deformation pattern and magnitude were validated by laser scanning data of physical assemblies. In addition, the code was used to investigate residual stresses developed during multipass arc welding of a nuclear industry pressurizer surge nozzle and subsequent welding repair where a 3D simulation was necessary. Taking the experimental data as reference, the 3D model predicted better residual stress distribution than a typical 2D asymmetrical model. Stress evolution in welding repair was also presented and discussed in this study. Furthermore, the efficient numerical model made it feasible to use integrated computational welding engineering to simulate welding processes for large-scale structures.

97 MATHEMATICS AND COMPUTING↗

OpenABLext: An automatic code generation framework for agent-based simulations on CPU-GPU-FPGA heterogeneous platforms

The execution of agent-based simulations (ABSs) on hardware accelerator devices such as graphics processing units (GPUs) has been shown to offer great performance potentials. However, in heterogeneous hardware environments, it can become increasingly difficult to find viable partitions of the simulation and provide implementations for different hardware devices. To automate this process, we present OpenABLext, an extension to OpenABL, a model specification language for ABSs. By providing a device-aware OpenCL backend, OpenABLext enables the co-execution of ABS on heterogeneous hardware platforms consisting of central processing units, GPUs, and field programmable gate arrays (FPGAs).We present a novel online dispatching method that efficiently profiles partitions of the simulation during run-time to optimize the hardware assignment while using the profiling results to advance the simulation itself. In addition, OpenABLext features automated conflict resolution based on user-specified rules, supports graph-based simulation spaces, and utilizes an efficient neighbor search algorithm. We show the improved performance of OpenABLext and demonstrate the potential of FPGAs in the context of ABS. We illustrate how co-execution can be used to further lower execution times. OpenABLext can be seen as an enabler to tap the computing power of heterogeneous hardware platforms for ABS.

97 MATHEMATICS AND COMPUTING↗

A new coupling of a GPU-resident large-eddy simulation code with a multiphysics wind turbine simulation tool

The development of new wind farm control strategies can benefit from combined analysis of flow dynamics in the farm and the behavior of individual turbines within one simulation environment. In this work, we present such an environment by developing a new coupling between the large-eddy simulation (LES) code GRASP and the multiphysics wind turbine simulation tool OpenFAST via an actuator line model (ALM). In addition, the implementation of the recently proposed filtered actuator line model (FALM) within the coupling is described. The new ALM implementation is cross-verified with results from four other commonly used research LES codes. The results for the blade loads and the near wake obtained with the new coupling are consistent with the other codes. Deviations are observed in the far wake. The results further indicate that the FALM is able to reduce the lift and power overprediction from which the traditional ALM suffers on coarse LES grids. This new simulation environment paves the way for future wind farm simulations under realistic weather conditions by leveraging GRASP's ability to impose data from large-scale meteorological models as boundary conditions.

17 WIND ENERGY↗

Developing an ELM Ecosystem Dynamics Model on GPU with OpenACC

Porting a complex scientific code, such as the E3SM land model (ELM), onto a new computing architecture is challenging. The paper presents design strategies and technical approaches to develop an ELM ecosystem dynamics model with compiler directives (OpenACC) on NVIDIA GPUs. The code has been refactored with advanced OpenACC features (such as deepcopy and routine directives) to reduce memory consumption and to increase the levels of parallelism through parallel loop reconstruction and new data structures. As a result, the optimized parallel implementation achieved more than a 140-time speedup (50 ms vs 7600 ms), compared to a naive implementation that uses OpenACC routine directive and parallelizes the code across existing loops on a single NVIDIA V100. On a fully loaded computing node with 44 CPUs and 6 GPUs, the code achieved over a 3.0-times speedup, compared to the original code on the CPU. Furthermore, the memory footprint of the optimized parallel implementation is 300 MB, which is around 15% of the 2.15 GB of memory consumed by a naive implementation. This study is the first effort to develop the ELM component on GPUs efficiently to support ultra-high-resolution land simulations at continental scales.

Schwartz, Peter↗