Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphics processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Simulations of the shock-driven Kelvin–Helmholtz instability in inclined gas curtains

In this paper, we present simulation results for the two-dimensional, shock-driven Kelvin–Helmholtz instability. Simulations are performed with a Mach 2.0 shock propagating through a finite-thickness curtain of gas inclined at an angle α0=30° with respect to the shock plane. After the passage of the shock, the gas curtain is accelerated along its axis. A perturbation develops due to shock reflection near the lower wall, and a Kelvin–Helmholtz instability forms near the vertical center of the curtain. This is the first known numerical reproduction of these phenomena that have previously been observed in experiments with an inclined cylindrical gas column. The effects of varying Mach number and column width were explored in detail to complement experimental data. Additionally, the dependence of the Kelvin–Helmholtz wavelength on Mach number closely matches the relationship observed in experiments. This supports the notion that the observed instability is effectively two-dimensional and inviscid (like classical Kelvin–Helmholtz). The growth rate of the perturbations in the gas curtain was also found to be similar for different Mach numbers. The perturbation at the curtain foot, previously unreported in experiments, was found to have a similar relationship to Mach number as the Kelvin–Helmholtz instability. Both perturbation wavelengths are found to be proportional to layer width. Simulations were performed with the fast interfaces and transport in the atmosphere, an exascale ready, graphics processing unit-accelerated compressible flow solver developed at the University of New Mexico.

42 ENGINEERING↗

Cross-correlation image analysis for real-time single particle tracking

Accurately measuring the translations of objects between images is essential in many fields, including biology, medicine, chemistry, and physics. One important application is tracking one or more particles by measuring their apparent displacements in a series of images. Popular methods, such as the center of mass, often require idealized scenarios to reach the shot noise limit of particle tracking and, therefore, are not generally applicable to multiple image types. More general methods, such as maximum likelihood estimation, reliably approach the shot noise limit, but are too computationally intense for use in real-time applications. These limitations are significant, as real-time, shot-noise-limited particle tracking is of paramount importance for feedback control systems. To fill this gap, we introduce a new cross-correlation-based algorithm that approaches shot-noise-limited displacement detection and a graphics processing unit-based implementation for real-time image analysis of a single particle.

Instruments & Instrumentation↗

Accuracy, transferability, and computational efficiency of interatomic potentials for simulations of carbon under extreme conditions

Large-scale atomistic molecular dynamics (MD) simulations provide an exceptional opportunity to advance the fundamental understanding of carbon under extreme conditions of high pressures and temperatures. However, the fidelity of these simulations depends heavily on the accuracy of classical interatomic potentials governing the dynamics of many-atom systems. Here, this study critically assesses several popular empirical potentials for carbon, as well as machine learning interatomic potentials (MLIPs), in their ability to simulate a range of physical properties at high pressures and temperatures, including the diamond equation of state, its melting line, shock Hugoniot, uniaxial compressions, and the structure of liquid carbon. Empirical potentials fail to accurately predict the behavior of carbon under high pressure–temperature conditions. In contrast, MLIPs demonstrate quantum accuracy, with Spectral Neighbor Analysis Potential (SNAP) and atomic cluster expansion (ACE) being the most accurate in reproducing the density functional theory results. ACE displays remarkable transferability despite not being specifically trained for extreme conditions. Furthermore, ACE and SNAP exhibit superior computational performance on graphics processing unit-based systems in billion atom MD simulations, with SNAP emerging as the fastest. In addition to offering practical guidance in selecting an interatomic potential with a fine balance of accuracy, transferability, and computational efficiency, this work also highlights transformative opportunities for groundbreaking scientific discoveries facilitated by quantum-accurate MD simulations with MLIPs on emerging exascale supercomputers.

36 MATERIALS SCIENCE↗

Calibrating a finite-strain phase-field model of fracture for bonded granular materials with uncertainty quantification

To study the mechanical behavior of mock high explosives, an experimental and simulation program was developed to calibrate, with quantified uncertainty, a material model of the bonded granular material Idoxuridine and nitroplasticized Estane-5703. This paper reports on the efficacy of such a framework as a generalizable methodology for calibrating material models against experimental data with uncertainty quantification. Additionally, this paper studies the effect of two manufacturing temperatures and three initial granular configurations on the unconfined compressive behavior of the resulting bonded granular materials. In each of these cases, the same calibration framework was used; in that, hundreds of high-fidelity direct numerical simulations using a new, graphics processing unit-enabled, high-performance finite element method software, Ratel, were run to calibrate a finite-strain phase-field fracture model against experimental data. It was found that manufacturing temperature influenced the elastic response of the mock high explosives, with higher temperatures yielding a stiffer response. By contrast, it was found that the initial configuration of the grains had a negligible impact on the overall behavior of the mock high explosives though it remains possible that local damage accumulation within the specimens could be altered by the initial configurations. Overall, the calibration framework was successful at creating well-calibrated models, showing its usefulness as an engineering and scientific tool.

36 MATERIALS SCIENCE↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

DeepMC: a deep learning method for efficient Monte Carlo beamlet dose calculation by predictive denoising in magnetic resonance-guided radiotherapy

Abstract Emerging magnetic resonance (MR) guided radiotherapy affords significantly improved anatomy visualization and, subsequently, more effective personalized treatment. The new therapy paradigm imposes significant demands on radiation dose calculation quality and speed, creating an unmet need for the acceleration of Monte Carlo (MC) dose calculation. Existing deep learning approaches to denoise the final plan MC dose fail to achieve the accuracy and speed requirements of large-scale beamlet dose calculation in the presence of a strong magnetic field for online adaptive radiotherapy planning. Our deep learning dose calculation method, DeepMC, addresses these needs by predicting low-noise dose from extremely noisy (but fast) MC-simulated dose and anatomical inputs, thus enabling significant acceleration. DeepMC simultaneously reduces MC sampling noise and predicts corrupted dose buildup at tissue-air material interfaces resulting from MR-field induced electron return effects. Here we demonstrate our model’s ability to accelerate dose calculation for daily treatment planning by a factor of 38 over traditional low-noise MC simulation with clinically meaningful accuracy in deliverable dose and treatment delivery parameters. As a post-processing approach, DeepMC provides compounded acceleration of large-scale dose calculation when used alongside established MC acceleration techniques in variance reduction and graphics processing unit-based MC simulation.

Engineering↗

New constraints on warm dark matter from the Lyman- α forest power spectrum

The forest of Lyman-α absorption lines detected in the spectra of distant quasars encodes information on the nature and properties of dark matter and the thermodynamics of diffuse baryonic material. Its main observable—the 1D flux power spectrum (FPS)—should exhibit a suppression on small scales and an enhancement on large scales in warm dark matter (WDM) cosmologies compared to standard Λ⁢CDM. Here, we present an unprecedented suite of 1080 high-resolution cosmological hydrodynamical simulations run with the graphics processing unit-accelerated code cholla to study the evolution of the Lyman-α forest under a wide range of physically motivated gas thermal histories along with different free-streaming lengths of WDM thermal relics in the early Universe. A statistical comparison of synthetic data with the forest FPS measured down to the smallest velocity scales ever probed at redshifts 4.0≲z≲5.2 [E. Boera et al., Revealing reionization with the thermal history of the intergalactic medium: New constraints from the Ly⁢α flux power spectrum, Astrophys. J. 872, 101 (2019)] yields a lower-limit m WDM >3.1 keV (95% C.L.) for the WDM particle mass and constrains the amplitude and spectrum of the photoheating and photoionizing background produced by star-forming galaxies and active galactic nuclei at these redshifts. Interestingly, our Bayesian inference analysis appears to weakly favor WDM models with a peak likelihood value at the thermal relic mass of m WDM =4.5 keV. In conclusion, we find that the suppression of the FPS from free-streaming saturates at k≳0.1 s km -1 because of peculiar velocity smearing, and this saturated suppression combined with a slightly lower gas temperature provides a moderately better fit to the observed small-scale FPS for WDM cosmologies.

79 ASTRONOMY AND ASTROPHYSICS↗

A High-Throughput Solver for Marginalized Graph Kernels on GPU

Here, we present the design and optimization of a solver for efficient and high-throughput computation of the marginalized graph kernel on General Purpose GPUs. The graph kernel is computed using the conjugate gradient method to solve a generalized Laplacian of the tensor product between a pair of graphs. To cope with the large gap between the instruction throughput and the memory bandwidth of the GPUs, our solver forms the graph tensor product on-the-fly without storing it in memory. This is achieved by using threads in a warp cooperatively to stream the adjacency and edge label matrices of individual graphs by small square matrix blocks called tiles, which are then staged in registers and the shared memory for later reuse. Warps across a thread block can further share tiles via the shared memory to increase data reuse. We exploit the sparsity of the graphs hierarchically by storing only non-empty tiles using a coordinate format and nonzero elements within each tile using bitmaps. We propose a new partition-based reordering algorithm for aggregating nonzero elements of the graphs into fewer but denser tiles to further exploit sparsity. We carry out extensive theoretical analyses on the graph tensor product primitives for tiles of various density and evaluate their performance on synthetic and real-world datasets. Our solver delivers three to four orders of magnitude speedup over existing CPU-based solvers such as GraKeL and GraphKernels. The capability of the solver enables kernel-based learning tasks at unprecedented scales.

97 MATHEMATICS AND COMPUTING↗

Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture

Here, our exploratory work finds that the SambaNova Reconfigurable Dataflow Architecture (RDA) along with the SambaFlow software stack provides for an attractive system and solution to accelerate AI for science workloads. We have observed the efficacy of using the system with a diverse set of science applications and reasoned their suitability for performance gains over traditional hardware. As the Data-Scale system provides for a very large memory capacity, the system can be used to train models that typically do not fit in a GPU. The architecture also provides for deeper integration with upcoming supercomputers at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science user facility, to help advance science insights.

97 MATHEMATICS AND COMPUTING↗

Performance Portability in the Exascale Computing Project: Exploration Through a Panel Series

Performance portability is a critical issue for the Exascale Computing Project (ECP) because of nontrivial architectural differences between machines available today and those expected at exascale. Many ECP project teams are working toward performance portability, and would expect to benefit from sharing lessons learned, identifying gaps, and discovering opportunities for partnerships. To facilitate this communication, the IDEAS-ECP project partnered with the three focus areas of ECP (application development, software technology, and hardware and integration), and Department of Energy computing facilities, to lead a series of panel discussions on performance portability. The panels were organized around broadly common themes of algorithmic and data locality challenges. In this article, we describe the panel series, its objectives, and perspectives from the various areas of the project. Finally, we also discuss use cases that are distinctive, as well as conclusions drawn from the collective experience of the participants.

97 MATHEMATICS AND COMPUTING↗

Co-design for Particle Applications at Exascale

Co-design across the Exascale Computing Project (ECP) has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing (HPC) architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and PROGRESS/BML libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from atomistic, mesoscale, continuum, through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate separation of concerns and directly support science runs.

97 MATHEMATICS AND COMPUTING↗

Real-Time Interactive 4D-STEM Phase-Contrast Imaging From Electron Event Representation Data: Less computation with the right representation

The arrival of direct electron detectors (DED) with high frame-rates in the field of scanning transmission electron microscopy has enabled many experimental techniques that require collection of a full diffraction pattern at each scan position, a field which is subsumed under the name four dimensional-scanning transmission electron microscopy (4D-STEM). DED frame rates approaching 100 kHz require data transmission rates and data storage capabilities that exceed commonly available computing infrastructure. Current commercial DEDs allow the user to make compromises in pixel bit depth, detector binning or windowing to reduce the per-frame file size and allow higher frame rates. This change in detector specifications requires decisions to be made before data acquisition that may reduce or lose information that could have been advantageous during data analysis. The 4D Camera, a DED with 87 kHz frame-rate developed at Lawrence Berkeley National Laboratory, reduces the raw data to a linear-index encoded electron event representation (EER). Here we show with experimental data from the 4D Camera that linear-index encoded EER and its direct use in 4D-STEM phase contrast imaging methods enables real-time, interactive phase-contrast from large-area 4D-STEM datasets. Furthermore, we detail the computational complexity advantages of the EER and the necessary computational steps to achieve real-time interactive ptychography and center-of-mass differential phase contrast using commonly available hardware accelerators.

4D-STEM↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Simulation of Hurricane Harvey flood event through coupled hydrologic-hydraulic models: Challenges and next steps

Using the 2017 Hurricane Harvey flood event as a test case, this study set up a series of sensitivity analyses to highlight three challenges associated with large-scale flood inundation modeling, including (a) model parameterization, (b) errors in digital elevation models, and (c) effects of reservoir retention. Driven by radar-based hourly rainfall data, a series of hydrologic-hydraulic models including the VIC hydrologic model, RAPID routing model, and Flood2D-GPU hydrodynamic model are set up over Harris County, Texas, to simulate flood inundation and hazards. The results demonstrate the importance of hydrologic parameters in improving flood modeling. For a large flood event such as Hurricane Harvey, the effect of the initial water depths is insignificant. The Manning's n values may increase the peak water depth by ~1%, the flood extents by 65km 2 , and the high danger zone by ~6%. On the contrary, the bathymetry correction factors may reduce the flood extent by ~1.4% and the high-danger zone by ~4%. Reducing the reservoir storage capacity to 1% may increase the flood extent by ~4% and the high-danger zone by ~17%. This study may provide supporting information to guide and prioritize the development of future high-performance computing hydrodynamic large-scale flood simulations.

54 ENVIRONMENTAL SCIENCES↗

Multiphysics coupling in the Exascale computing project

Multiphysics coupling presents a significant challenge in terms of both computational accuracy and performance. Achieving high performance on coupled simulations can be particularly challenging in a high-performance computing context. The US Department of Energy Exascale Computing Project has the mission to prepare mission-relevant applications for the delivery of the exascale computers starting in 2023. Many of these applications require multiphysics coupling, and the implementations must be performant on exascale hardware. In this special issue we feature six articles performing advanced multiphysics coupling that span the computational science domains in the Exascale Computing Project.

97 MATHEMATICS AND COMPUTING↗

PeleC: An adaptive mesh refinement solver for compressible reacting flows

Reacting flow simulations for combustion applications require extensive computing capabilities. Leveraging the AMReX library, the Pele suite of combustion simulation tools targets the largest supercomputers available and future exascale machines. We introduce PeleC, the compressible solver in the Pele suite, and detail its capabilities, including complex geometry representation, chemistry integration, and discretization. We present a comparison of development efforts using both OpenACC and AMReX’s C++ performance portability framework for execution on multiple GPU architectures. We discuss relevant details that have allowed PeleC to achieve high performance and scalability. PeleC’s performance characteristics are measured through relevant simulations on multiple supercomputers. The success of PeleC’s design for exascale is exhibited through demonstration of a 160 billion cell simulation and weak scaling onto 100% of Summit, an NVIDIA-based GPU supercomputer at Oak Ridge National Laboratory. Our results provide confidence that PeleC will enable future combustion science simulations with unprecedented fidelity.

97 MATHEMATICS AND COMPUTING↗

Matrix Product (GEMM) Performance Data from GPUs

Timing data for mixed precision GEMM matrix product operations on several GPU models, including NVIDIA V100 and A100, AMD MI100 and Intel P580. Also data from machine learning model training on this data using Scikit-learn.

97 MATHEMATICS AND COMPUTING↗

TxDOT Road Elevation Model Dataset

This dataset provides three formats of Road Elevation Model (REM) data: 3D road line/polygon GeoPackage (GPKG), road lidar LAZ and COPC LAZ, and road digital surface model (DSM) GeoTIFF. Data are produced from the ~50TB TxGIO (formerly TNRIS) state lidar collections. This dataset is currently organized by maintenance section in each TxDOT district. Computation is done on GPU computing resources at Oak Ridge National Laboratory (ORNL), through a Strategic Partnership Project with UT Austin and an NSF ACCESS computing allocation award that enables fast massive data movement between TACC Corral and ORNL CADES/OLCF using Globus. In addition to this release from ORNL, a copy of this dataset can also be downloaded at https://web.corral.tacc.utexas.edu/nfiedata/road3d/.

13 HYDRO ENERGY↗