Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graphical processing unit”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Calibrating a finite-strain phase-field model of fracture for bonded granular materials with uncertainty quantification

To study the mechanical behavior of mock high explosives, an experimental and simulation program was developed to calibrate, with quantified uncertainty, a material model of the bonded granular material Idoxuridine and nitroplasticized Estane-5703. This paper reports on the efficacy of such a framework as a generalizable methodology for calibrating material models against experimental data with uncertainty quantification. Additionally, this paper studies the effect of two manufacturing temperatures and three initial granular configurations on the unconfined compressive behavior of the resulting bonded granular materials. In each of these cases, the same calibration framework was used; in that, hundreds of high-fidelity direct numerical simulations using a new, graphics processing unit-enabled, high-performance finite element method software, Ratel, were run to calibrate a finite-strain phase-field fracture model against experimental data. It was found that manufacturing temperature influenced the elastic response of the mock high explosives, with higher temperatures yielding a stiffer response. By contrast, it was found that the initial configuration of the grains had a negligible impact on the overall behavior of the mock high explosives though it remains possible that local damage accumulation within the specimens could be altered by the initial configurations. Overall, the calibration framework was successful at creating well-calibrated models, showing its usefulness as an engineering and scientific tool.

36 MATERIALS SCIENCE↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

DeepMC: a deep learning method for efficient Monte Carlo beamlet dose calculation by predictive denoising in magnetic resonance-guided radiotherapy

Abstract Emerging magnetic resonance (MR) guided radiotherapy affords significantly improved anatomy visualization and, subsequently, more effective personalized treatment. The new therapy paradigm imposes significant demands on radiation dose calculation quality and speed, creating an unmet need for the acceleration of Monte Carlo (MC) dose calculation. Existing deep learning approaches to denoise the final plan MC dose fail to achieve the accuracy and speed requirements of large-scale beamlet dose calculation in the presence of a strong magnetic field for online adaptive radiotherapy planning. Our deep learning dose calculation method, DeepMC, addresses these needs by predicting low-noise dose from extremely noisy (but fast) MC-simulated dose and anatomical inputs, thus enabling significant acceleration. DeepMC simultaneously reduces MC sampling noise and predicts corrupted dose buildup at tissue-air material interfaces resulting from MR-field induced electron return effects. Here we demonstrate our model’s ability to accelerate dose calculation for daily treatment planning by a factor of 38 over traditional low-noise MC simulation with clinically meaningful accuracy in deliverable dose and treatment delivery parameters. As a post-processing approach, DeepMC provides compounded acceleration of large-scale dose calculation when used alongside established MC acceleration techniques in variance reduction and graphics processing unit-based MC simulation.

Engineering↗

New constraints on warm dark matter from the Lyman- α forest power spectrum

The forest of Lyman-α absorption lines detected in the spectra of distant quasars encodes information on the nature and properties of dark matter and the thermodynamics of diffuse baryonic material. Its main observable—the 1D flux power spectrum (FPS)—should exhibit a suppression on small scales and an enhancement on large scales in warm dark matter (WDM) cosmologies compared to standard Λ⁢CDM. Here, we present an unprecedented suite of 1080 high-resolution cosmological hydrodynamical simulations run with the graphics processing unit-accelerated code cholla to study the evolution of the Lyman-α forest under a wide range of physically motivated gas thermal histories along with different free-streaming lengths of WDM thermal relics in the early Universe. A statistical comparison of synthetic data with the forest FPS measured down to the smallest velocity scales ever probed at redshifts 4.0≲z≲5.2 [E. Boera et al., Revealing reionization with the thermal history of the intergalactic medium: New constraints from the Ly⁢α flux power spectrum, Astrophys. J. 872, 101 (2019)] yields a lower-limit m WDM >3.1 keV (95% C.L.) for the WDM particle mass and constrains the amplitude and spectrum of the photoheating and photoionizing background produced by star-forming galaxies and active galactic nuclei at these redshifts. Interestingly, our Bayesian inference analysis appears to weakly favor WDM models with a peak likelihood value at the thermal relic mass of m WDM =4.5 keV. In conclusion, we find that the suppression of the FPS from free-streaming saturates at k≳0.1 s km -1 because of peculiar velocity smearing, and this saturated suppression combined with a slightly lower gas temperature provides a moderately better fit to the observed small-scale FPS for WDM cosmologies.

79 ASTRONOMY AND ASTROPHYSICS↗

A High-Throughput Solver for Marginalized Graph Kernels on GPU

Here, we present the design and optimization of a solver for efficient and high-throughput computation of the marginalized graph kernel on General Purpose GPUs. The graph kernel is computed using the conjugate gradient method to solve a generalized Laplacian of the tensor product between a pair of graphs. To cope with the large gap between the instruction throughput and the memory bandwidth of the GPUs, our solver forms the graph tensor product on-the-fly without storing it in memory. This is achieved by using threads in a warp cooperatively to stream the adjacency and edge label matrices of individual graphs by small square matrix blocks called tiles, which are then staged in registers and the shared memory for later reuse. Warps across a thread block can further share tiles via the shared memory to increase data reuse. We exploit the sparsity of the graphs hierarchically by storing only non-empty tiles using a coordinate format and nonzero elements within each tile using bitmaps. We propose a new partition-based reordering algorithm for aggregating nonzero elements of the graphs into fewer but denser tiles to further exploit sparsity. We carry out extensive theoretical analyses on the graph tensor product primitives for tiles of various density and evaluate their performance on synthetic and real-world datasets. Our solver delivers three to four orders of magnitude speedup over existing CPU-based solvers such as GraKeL and GraphKernels. The capability of the solver enables kernel-based learning tasks at unprecedented scales.

97 MATHEMATICS AND COMPUTING↗

Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture

Here, our exploratory work finds that the SambaNova Reconfigurable Dataflow Architecture (RDA) along with the SambaFlow software stack provides for an attractive system and solution to accelerate AI for science workloads. We have observed the efficacy of using the system with a diverse set of science applications and reasoned their suitability for performance gains over traditional hardware. As the Data-Scale system provides for a very large memory capacity, the system can be used to train models that typically do not fit in a GPU. The architecture also provides for deeper integration with upcoming supercomputers at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science user facility, to help advance science insights.

97 MATHEMATICS AND COMPUTING↗

Performance Portability in the Exascale Computing Project: Exploration Through a Panel Series

Performance portability is a critical issue for the Exascale Computing Project (ECP) because of nontrivial architectural differences between machines available today and those expected at exascale. Many ECP project teams are working toward performance portability, and would expect to benefit from sharing lessons learned, identifying gaps, and discovering opportunities for partnerships. To facilitate this communication, the IDEAS-ECP project partnered with the three focus areas of ECP (application development, software technology, and hardware and integration), and Department of Energy computing facilities, to lead a series of panel discussions on performance portability. The panels were organized around broadly common themes of algorithmic and data locality challenges. In this article, we describe the panel series, its objectives, and perspectives from the various areas of the project. Finally, we also discuss use cases that are distinctive, as well as conclusions drawn from the collective experience of the participants.

97 MATHEMATICS AND COMPUTING↗

Co-design for Particle Applications at Exascale

Co-design across the Exascale Computing Project (ECP) has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing (HPC) architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and PROGRESS/BML libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from atomistic, mesoscale, continuum, through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate separation of concerns and directly support science runs.

97 MATHEMATICS AND COMPUTING↗

Real-Time Interactive 4D-STEM Phase-Contrast Imaging From Electron Event Representation Data: Less computation with the right representation

The arrival of direct electron detectors (DED) with high frame-rates in the field of scanning transmission electron microscopy has enabled many experimental techniques that require collection of a full diffraction pattern at each scan position, a field which is subsumed under the name four dimensional-scanning transmission electron microscopy (4D-STEM). DED frame rates approaching 100 kHz require data transmission rates and data storage capabilities that exceed commonly available computing infrastructure. Current commercial DEDs allow the user to make compromises in pixel bit depth, detector binning or windowing to reduce the per-frame file size and allow higher frame rates. This change in detector specifications requires decisions to be made before data acquisition that may reduce or lose information that could have been advantageous during data analysis. The 4D Camera, a DED with 87 kHz frame-rate developed at Lawrence Berkeley National Laboratory, reduces the raw data to a linear-index encoded electron event representation (EER). Here we show with experimental data from the 4D Camera that linear-index encoded EER and its direct use in 4D-STEM phase contrast imaging methods enables real-time, interactive phase-contrast from large-area 4D-STEM datasets. Furthermore, we detail the computational complexity advantages of the EER and the necessary computational steps to achieve real-time interactive ptychography and center-of-mass differential phase contrast using commonly available hardware accelerators.

4D-STEM↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Simulation of Hurricane Harvey flood event through coupled hydrologic-hydraulic models: Challenges and next steps

Using the 2017 Hurricane Harvey flood event as a test case, this study set up a series of sensitivity analyses to highlight three challenges associated with large-scale flood inundation modeling, including (a) model parameterization, (b) errors in digital elevation models, and (c) effects of reservoir retention. Driven by radar-based hourly rainfall data, a series of hydrologic-hydraulic models including the VIC hydrologic model, RAPID routing model, and Flood2D-GPU hydrodynamic model are set up over Harris County, Texas, to simulate flood inundation and hazards. The results demonstrate the importance of hydrologic parameters in improving flood modeling. For a large flood event such as Hurricane Harvey, the effect of the initial water depths is insignificant. The Manning's n values may increase the peak water depth by ~1%, the flood extents by 65km 2 , and the high danger zone by ~6%. On the contrary, the bathymetry correction factors may reduce the flood extent by ~1.4% and the high-danger zone by ~4%. Reducing the reservoir storage capacity to 1% may increase the flood extent by ~4% and the high-danger zone by ~17%. This study may provide supporting information to guide and prioritize the development of future high-performance computing hydrodynamic large-scale flood simulations.

54 ENVIRONMENTAL SCIENCES↗

Multiphysics coupling in the Exascale computing project

Multiphysics coupling presents a significant challenge in terms of both computational accuracy and performance. Achieving high performance on coupled simulations can be particularly challenging in a high-performance computing context. The US Department of Energy Exascale Computing Project has the mission to prepare mission-relevant applications for the delivery of the exascale computers starting in 2023. Many of these applications require multiphysics coupling, and the implementations must be performant on exascale hardware. In this special issue we feature six articles performing advanced multiphysics coupling that span the computational science domains in the Exascale Computing Project.

97 MATHEMATICS AND COMPUTING↗

PeleC: An adaptive mesh refinement solver for compressible reacting flows

Reacting flow simulations for combustion applications require extensive computing capabilities. Leveraging the AMReX library, the Pele suite of combustion simulation tools targets the largest supercomputers available and future exascale machines. We introduce PeleC, the compressible solver in the Pele suite, and detail its capabilities, including complex geometry representation, chemistry integration, and discretization. We present a comparison of development efforts using both OpenACC and AMReX’s C++ performance portability framework for execution on multiple GPU architectures. We discuss relevant details that have allowed PeleC to achieve high performance and scalability. PeleC’s performance characteristics are measured through relevant simulations on multiple supercomputers. The success of PeleC’s design for exascale is exhibited through demonstration of a 160 billion cell simulation and weak scaling onto 100% of Summit, an NVIDIA-based GPU supercomputer at Oak Ridge National Laboratory. Our results provide confidence that PeleC will enable future combustion science simulations with unprecedented fidelity.

97 MATHEMATICS AND COMPUTING↗

Matrix Product (GEMM) Performance Data from GPUs

Timing data for mixed precision GEMM matrix product operations on several GPU models, including NVIDIA V100 and A100, AMD MI100 and Intel P580. Also data from machine learning model training on this data using Scikit-learn.

97 MATHEMATICS AND COMPUTING↗

TxDOT Road Elevation Model Dataset

This dataset provides three formats of Road Elevation Model (REM) data: 3D road line/polygon GeoPackage (GPKG), road lidar LAZ and COPC LAZ, and road digital surface model (DSM) GeoTIFF. Data are produced from the ~50TB TxGIO (formerly TNRIS) state lidar collections. This dataset is currently organized by maintenance section in each TxDOT district. Computation is done on GPU computing resources at Oak Ridge National Laboratory (ORNL), through a Strategic Partnership Project with UT Austin and an NSF ACCESS computing allocation award that enables fast massive data movement between TACC Corral and ORNL CADES/OLCF using Globus. In addition to this release from ORNL, a copy of this dataset can also be downloaded at https://web.corral.tacc.utexas.edu/nfiedata/road3d/.

13 HYDRO ENERGY↗

A Multireference Approach to Electron and Electron–Nuclear Dynamics in Nanomaterials (Final Report)

Many important chemical and physical phenomena involve dynamics on large number of electronic states. Thus, there is a critical need to develop methods to simulate dynamics in dense manifolds of states. Towards this end, we have: a) developed the multiple cloning in dense manifolds of states (MCDMS) method, which is capable of accurately modeling the quantum mechanical coherence between populations on a large number of electronic states, b) implemented MCDMS into the free, open-source PySpawn software package, c) developed graphics processing unit-accelerated algorithms modeling electron dynamics in light fields via Floquet time-dependent configuration interaction (F-TDCI), and d) critically compared different orbital bases in order to achieve an accurate and efficient F-TDCI expansion. This grant ended in August 2020, when our group moved to from Michigan State University to Stony Brook University, where this project continues under grant number DE-SC0021643.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large Terrain Modeling and Visualization for Planets

Physics-based simulations are actively used in the design, testing, and operations phases of surface and near-surface planetary space missions. One of the challenges in realtime simulations is the ability to handle large multi-resolution terrain data sets within models as well as for visualization. In this paper, we describe special techniques that we have developed for visualization, paging, and data storage for dealing with these large data sets. The visualization technique uses a real-time GPU-based continuous level-of-detail technique that delivers multiple frames a second performance even for planetary scale terrain model sizes.

digital elevation map↗

Considerations for GPU SEE Testing

This presentation will discuss the considerations an engineer should take to perform Single Event Effects (SEE) testing on GPU devices. Notable topics will include setup complexity, architecture insight which permits cross platform normalization, acquiring a reasonable detail of information from the test suite, and a few lessons learned from preliminary testing.

NASA Electronic Parts and Packaging (NEPP) Program↗