Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

GASNet-EX Memory Kinds: Support for Device Memory in PGAS Programming Models

There is an emerging need for adaptive, lightweight communication in irregular HPC applications at exascale, where GPU accelerators provide the majority of available compute cycles. To address this need, Lawrence Berkeley National Lab is developing a programming system to support distributed-memory HPC application development using the Partitioned Global Address Space (PGAS) model. This work includes two major components: UPC++ and GASNet-EX. UPC++ is a C++ template library providing Remote Memory Access (RMA) and Remote Procedure Call (RPC) communication interfaces. GASNet-EX is a portable, high-performance communication middleware library, used by the implementations of UPC++ and many other PGAS programming models. We describe recent advances in GASNet-EX to efficiently implement zero-copy Remote Memory Access (RMA) communication to and from memory on accelerator devices such as GPUs. We demonstrate performance improvements via benchmark results from UPC++ (on Summit) and the Legion programming system (on DGX-1), both using GASNet-EX for communication.

Hargrove, Paul H↗

Sensitivity of Simulations of Double-detonation Type Ia Supernovae to Integration Methodology

Abstract We study the coupling of hydrodynamics and reactions in simulations of the double-detonation model for Type Ia supernovae. When assessing the convergence of simulations, the focus is usually on spatial resolution; however, the method of coupling the physics together as well as the tolerances used in integrating a reaction network also play an important role. In this paper, we explore how the choices made in both coupling and integrating the reaction portion of a simulation (operator/Strang splitting versus the simplified spectral deferred corrections method we introduced previously) influences the accuracy, efficiency, and nucleosynthesis of simulations of double detonations. We find no need to limit reaction rates or reduce the simulation time step to the reaction timescale. The entire simulation methodology used here is GPU-accelerated and made freely available as part of the Castro simulation code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

PCMS: Parallel Coupler For Multimodel Simulations

This paper presents the Parallel Coupler for Multimodel Simulations (PCMS), a new GPU accelerated generalized coupling framework for coupling simulation codes on leadership class supercomputers. PCMS includes distributed control and field mapping methods for up to five dimensions. For field mapping PCMS can utilize discretization and field information to accommodate physics constraints. PCMS is demonstrated with a coupling of the gyrokinetic microturbulence code XGC with a Monte Carlo neutral transport code DEGAS2 and with a 5D distribution function coupling of an energetic particle transport code (GNET) to a gyrokinetic microturbulence code (GTC). Weak scaling is also demonstrated on up to 2,080 GPUs of Frontier with a weak scaling efficiency of 85%.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards multiscale modeling of ocean surface turbulent mixing using coupled MPAS-Ocean v6.3 and PALM v5.0

A multiscale modeling approach for studying the ocean surface turbulent mixing is explored by coupling an ocean general circulation model (GCM) MPAS-Ocean with the Parallelized Large Eddy Simulation Model (PALM). The coupling approach is similar to the superparameterization approach that has been used to represent the effects of deep convection in atmospheric GCMs. However, the focus of this multiscale modeling approach is on the small-scale turbulent mixing and their interactions with the larger-scale processes in the ocean, so that a more flexible coupling strategy is used. To reduce the computational cost, a customized version of PALM is ported on the general-purpose graphics processing unit (GPU) with OpenACC, achieving 10–16 times overall speedup as compared to running on a single CPU. Even with the GPU-acceleration technique, a superparameterization-like approach to represent the ocean surface turbulent mixing in GCMs using embedded high fidelity and three-dimensional large eddy simulations (LESs) over the global ocean is still computationally intensive and infeasible for long simulations. However, running PALM regionally on selected MPAS-Ocean grid cells is shown to be a promising approach moving forward. The flexible coupling between MPAS-Ocean and PALM allows further exploration of the interactions between the ocean surface turbulent mixing and larger-scale processes, as well as future development and improvement of ocean surface turbulent mixing parameterizations for GCMs.

54 ENVIRONMENTAL SCIENCES↗

Importance of ice nucleation and precipitation on climate with the Parameterization of Unified Microphysics Across Scales version 1 (PUMASv1)

Cloud microphysics is critical for weather and climate prediction. In this work, we document updates and corrections to the cloud microphysical scheme used in the Community Earth System Model (CESM) and other models. These updates include a new nomenclature for the scheme, now called Parameterization of Unified Microphysics Across Scales (PUMAS), and the ability to run the scheme on graphics processing units (GPUs). The main science changes include refactoring an ice number limiter and associated changes to ice nucleation, adding vapor deposition onto snow, and introducing an implicit numerical treatment for sedimentation. We also detail the improvements in computational performance that can be achieved with GPU acceleration. We then show the impact of these scheme changes on the (a) mean state climate, (b) cloud feedback response to warming, and (c) aerosol forcing. We find that corrections are needed to the immersion freezing parameterization and that ice nucleation has important impacts on climate. We also find that the revised scheme produces less cloud liquid and ice but that this can be adjusted by changing the loss process for cloud liquid (autoconversion). Furthermore, there are few discernible effects of the PUMAS changes on cloud feedbacks but some reductions in the magnitude of aerosol–cloud interactions (ACIs). Small cloud feedback changes appear to be related to the implicit sedimentation scheme, with a number of factors affecting ACIs.

54 ENVIRONMENTAL SCIENCES↗

Comparing the Performance of Julia on CPUs versus GPUs and Julia-MPI versus Fortran-MPI: a case study with MPAS-Ocean (Version 7.1)

Abstract. Some programming languages are easy to develop at the cost of slow execution, while others are fast at runtime but much more difficult to write. Julia is a programming language that aims to be the best of both worlds – a development and production language at the same time. To test Julia's utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, the Model for Prediction Across Scales–Ocean (MPAS-Ocean), as well as a Python shallow water code. Three versions of the Julia shallow water code were created: for single-core CPU, graphics processing unit (GPU), and Message Passing Interface (MPI) CPU clusters. Comparing identical simulations revealed that our first version of the Julia model was 13 times faster than Python using NumPy, where both used an unthreaded single-core CPU. Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20× speed-up of the single-core CPU Julia model. The GPU-accelerated Julia code was almost identical in terms of performance to the MPI parallelized code on 64 processes, an unexpected result for such different architectures. Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts and ranges from 2× faster to 2× slower for higher processor counts. Our experience is that Julia development is fast and convenient for prototyping but that Julia requires further investment and expertise to be competitive with compiled codes. We provide advice on Julia code optimization for HPC systems.

54 ENVIRONMENTAL SCIENCES↗

Zooming in: SCREAM at 100 m using regional refinement over the San Francisco Bay Area

Pushing global climate models to large-eddy simulation (LES) scales over complex terrain has remained a major challenge. This study presents the first known implementation of a global model – SCREAM (Simple Cloud-Resolving E3SM Atmosphere Model) – at 100 m horizontal resolution using a regionally refined mesh (RRM) over the San Francisco Bay Area. Two hindcast simulations were conducted to test performance under both strong synoptic forcing and weak, boundary-layer-driven conditions. We demonstrate that SCREAM can stably run at LES scales while realistically capturing topography, surface heterogeneity, and coastal processes. The 100 m SCREAM-RRM substantially improves near-surface wind speed, temperature, humidity, and pressure biases compared to the baseline 3.25 km simulation, and better reproduces fine-scale wind oscillations and boundary-layer structures. These advances leverage SCREAM's scale-aware SHOC turbulence parameterization, which transitions smoothly across scales without tuning. Performance tests show that while CPU-only simulations remain costly, GPU acceleration with SCREAMv1 on NERSC's Perlmutter system enables two-day hindcasts to complete in under two wall-clock days. Our results open the door to LES-scale studies of orographic flows, boundary-layer turbulence, and coastal clouds within a fully comprehensive global modeling framework.

Geosciences↗

Assessing climate-change-induced flood risk in the Conasauga River watershed: an application of ensemble hydrodynamic inundation modeling

Abstract. This study evaluates the impact of potential future climate change on flood regimes, floodplain protection, and electricity infrastructures across the Conasauga River watershed in the southeastern United States through ensemble hydrodynamic inundation modeling. The ensemble streamflow scenarios were simulated by the Distributed Hydrology Soil Vegetation Model (DHSVM) driven by (1) 1981–2012 Daymet meteorological observations and (2) 11 sets of downscaled global climate models (GCMs) during the 1966–2005 historical and 2011–2050 future periods. Surface inundation was simulated using a GPU-accelerated Two-dimensional Runoff Inundation Toolkit for Operational Needs (TRITON) hydrodynamic model. A total of 9 out of the 11 GCMs exhibit an increase in the mean ensemble flood inundation areas. Moreover, at the 1 % annual exceedance probability level, the flood inundation frequency curves indicate a ∼ 16 km2 increase in floodplain area. The assessment also shows that even after flood-proofing, four of the substations could still be affected in the projected future period. The increase in floodplain area and substation vulnerability highlights the need to account for climate change in floodplain management. Overall, this study provides a proof-of-concept demonstration of how the computationally intensive hydrodynamic inundation modeling can be used to enhance flood frequency maps and vulnerability assessment under the changing climatic conditions.

54 ENVIRONMENTAL SCIENCES↗

Optimization of Selected Remote Sensing Algorithms for Embedded NVIDIA Kepler GPU Architecture

This paper evaluates the potential of embedded Graphic Processing Units in the Nvidias Tegra K1 for onboard processing. The performance is compared to a general purpose multi-core CPU and full fledge GPU accelerator. This study uses two algorithms: Wavelet Spectral Dimension Reduction of Hyperspectral Imagery and Automated Cloud-Cover Assessment (ACCA) Algorithm. Tegra K1 achieved 51 for ACCA algorithm and 20 for the dimension reduction algorithm, as compared to the performance of the high-end 8-core server Intel Xeon CPU with 13.5 times higher power consumption.

Riha, Lubomir↗

QSPIN: A High Level Java API for Quantum Computing Experimentation

QSPIN is a high level Java language API for experimentation in QC models used in the calculation of Ising spin glass ground states and related quadratic unconstrained binary optimization (QUBO) problems. The Java API is intended to facilitate research in advanced QC algorithms such as hybrid quantum-classical solvers, automatic selection of constraint and optimization parameters, and techniques for the correction and mitigation of model and solution errors. QSPIN includes high level solver objects tailored to the D-Wave quantum annealing architecture that implement hybrid quantum-classical algorithms [Booth et al.] for solving large problems on small quantum devices, elimination of variables via roof duality, and classical computing optimization methods such as GPU accelerated simulated annealing and tabu search for comparison. A test suite of documented NP-complete applications ranging from graph coloring, covering, and partitioning to integer programming and scheduling are provided to demonstrate current capabilities.

Quantu↗

AIRNOISE: A Tool for Preliminary Noise-Abatement Terminal Approach Route Design

Noise from aircraft in the airport vicinity is one of the leading aviation-induced environmental issues. The FAA developed the Integrated Noise Model (INM) and its replacement Aviation Environmental Design Tool (AEDT) software to assess noise impact resulting from all aviation activities. However, a software tool is needed that is simple to use for terminal route modification, quick and reasonably accurate for preliminary noise impact evaluation and flexible to be used for iterative design of optimal noise-abatement terminal routes. In this paper, we extend our previous work on developing a noise-abatement terminal approach route design tool, named AIRNOISE, to satisfy this criterion. First, software efficiency has been significantly increased by over tenfold using the C programming language instead of MATLAB. Moreover, a state-of-the-art high performance GPU-accelerated computing module is implemented that was tested to be hundreds time faster than the C implementation. Secondly, a Graphical User Interface (GUI) was developed allowing users to import current terminal approach routes and modify the routes interactively to design new terminal approach routes. The corresponding noise impacts are then calculated and displayed in the GUI in seconds. Finally, AIRNOISE was applied to Baltimore-Washington International Airport terminal approach route to demonstrate its usage.

aircraft noise↗

Jumping the Queue: From NASA to the Commercial Cloud

NASA's High-End Computing Capability (HECC) Project has made it possible for its users to run on commercial cloud resources in a seamless way. In the first of three phases, we implemented a pilot project for a few users, enabling them to “jump the queue” and burst jobs from the HECC environment to Amazon Web Services (AWS). By using GPU-accelerated nodes at AWS, the users were able to make significant advances in their research. The second phase of the project made AWS access available to all HECC users and added accounting to make users responsible for cloud charges. We are also enabling export-controlled work through the use of AWS GovCloud. In the third phase, we will add web-based mechanisms to permit non-HECC users to access cloud resources for their HPC projects.

Hood, Robert↗

Effects of Spatial Resolution on Retropropulsion Aerodynamics in an Atmospheric Environment

Development of a powered descent capability for atmospheric environments is heavily reliant on computational simulation. The prohibitive computational cost of such simulations motivates an improvement in the understanding of the minimum computational fidelity re-quired to accurately characterize aerodynamic-propulsive interference for such applications. This work examines the applicability of detached eddy simulation methods for retropropulsion in atmospheric environments through utilization of a GPU-accelerated computational framework, yielding data that are largely unachievable with conventional high-performance computing resources. This effort was specifically designed to quantitatively assess the effects of spatial resolution on vehicle aerodynamics for nominal operation of a low lift-to-drag ratio, human-scale Mars lander concept. The test matrix and scaling approach span relevant nozzle expansion conditions as well as mid-supersonic to high-subsonic operating conditions. Solutions were generated using computational grids ranging from 143 million to 1.14 billion grid points (degrees of freedom). This paper will provide an overview of the computational campaign, approach, and discussion of preliminary results focused on a range of operating conditions for a conceptual low lift-to-drag, human-scale Mars lander.

Ashley M Korzun↗

Computational Investigation of Retropropulsion Operating Environments with a GPU-Enabled Detached Eddy Simulation Approach

Human exploration of the surface of Mars will require an extended powered descent phase of flight, during which aerodynamic-propulsive interference effects can be significant. Characterization of these environments to enable implementation of this technology into a flight vehicle will rely heavily on computational simulation. This work advances the understanding of retropropulsion aerodynamics through application of a massively parallel detached eddy simulation approach on a GPU-accelerated computational framework, yielding data that are largely unachievable with conventional high-performance computing resources. This work includes time-dependent and time-averaged forces and moments on a conceptual, full-scale vehicle in environments and operating conditions relevant to human Mars exploration. Conditions are examined where the engine exhaust flow transitions between over-expanded and under-expanded flow structures, and flight operation will require the ability to maintain control of the vehicle during such a transition. These transitions occur as the vehicle decelerates, and as such, this investigation includes supersonic, transonic, and subsonic flight conditions. Options for vehicle control during powered flight include differential throttling of the engines. This paper provides an overview of the computational campaign, approach, and discussion of results in characterizing the resulting aerodynamics for differential throttling with retropropulsion in atmospheric environments.

EDL↗

FUN3D Manual: 13.7

This manual describes the installation and execution of FUN3D version 13.7, including optional dependent packages. FUN3D is a suite of computational fluid dynamics simulation and design tools that uses mixed-element unstructured grids in a large number of formats, including structured multiblock and overset grid systems. A discretely-exact adjoint solver may be used for formal design optimization, error estimation, and mesh adaptation. FUN3D also offers a reacting, real-gas capability and provides GPU acceleration of many common simulation options.

FUN3D↗

All-sky search in early O3 LIGO data for continuous gravitational-wave signals from unknown neutron stars in binary systems

Rapidly spinning neutron stars are promising sources of continuous gravitational waves. Detecting such a signal would allow probing of the physical properties of matter under extreme conditions. A significant fraction of the known pulsar population belongs to binary systems. Searching for unknown neutron stars in binary systems requires specialized algorithms to address unknown orbital frequency modulations. We present a search for continuous gravitational waves emitted by neutron stars in binary systems in early data from the third observing run of the Advanced LIGO and Advanced Virgo detectors using the semicoherent, GPU-accelerated, binaryskyhough pipeline. The search analyzes the most sensitive frequency band of the LIGO detectors, 50–300 Hz. Binary orbital parameters are split into four regions, comprising orbital periods of three to 45 days and projected semimajor axes of two to 40 light seconds. No detections are reported. We estimate the sensitivity of the search using simulated continuous wave signals, achieving the most sensitive results to date across the analyzed parameter space.

Gravitation↗

Adding GPU Support to the Markov Chain Monte Carlo Code Catmip

In geophysics, we are confronted with many under-determined inverse problems. For example, all of our observations of earthquakes are made at the Earth’s surface. So, when we try to infer how slip during an earthquake evolves in space and time, we find that there are many potential slip histories that are consistent with our limited observations and our understanding of earthquake physics. One way to approach these problems is with Bayesian analysis which allows us to infer the ensemble of all potential slip models that satisfy the observations and our prior knowledge of earthquake physics. In Bayesian analysis, our prior knowledge is known as the prior probability density function or prior PDF, the fit to the data is known as the data likelihood, and the target PDF that satisfies both the prior PDF and data likelihood is known as the posterior PDF. However, simulating the posterior PDF typically requires using Markov Chain Monte Carlo (MCMC) to draw tens of billions of random realizations of earthquake slip models, which may not be computationally feasible. To make this and similar geophysical inversions computationally tractable, we developed the Cascading Adaptive Transitional Metropolis In Parallel (CATMIP) algorithm. CATMIP is an efficient parallel Markov Chain Monte Carlo (MCMC) sampler that is used for model fitting and uncertainty quantification in geophysics. Example use cases are earthquake rupture modeling, determining mineral composition on Mars, reconstructing the history of ocean salinity, and historical earthquake relocation. CATMIP employs many parallel instances of the Metropolis algorithm for sampling in a transitioning framework. Transitioning is a process in which a set of random samples at equilibrium with a known probability density function (PDF) are used as seeds for the Markov chains to sample successive target PDFs that incrementally move the distribution from the starting seeds to the final desired PDF that describes the relative plausibility of potential values for the model parameters. The algorithm is implemented as a Master-Worker model employing MPI for communication. The worker processes are loosely coupled with global parameters periodically optimized by the master process. This provides a very high amount of parallelism with little communication between updates. During the presentation we will discuss the history of the algorithm and elaborate the earthquake rupture modeling use case for the CATMIP package. Our first step toward GPU optimization was to optimize the code for the CPU. CPU profiling revealed that most of the compute time is spent in calls to level 2 BLAS routines and calls to GSL random number generators. We revised the algorithm to employ level 3 BLAS routines instead. In our presentation we will describe how this was accomplished. Adding GPU support to CATMIP consisted mostly of replacing the calls to GSL with calls to GPU vendor-provided library routines. A small number of loops were directly implemented in CUDA. In the presentation will provide implementation details. Finally, we will discuss methods for profiling and opportunities for further optimizing GPU execution. By creating a code with the flexibility to run on either a CPU or GPU architecture, CATMIP can be used on systems ranging from large CPU-based HPC environments to single servers with GPU acceleration and everything in between.

HECC↗

The Additive Manufacturing Moment Measure (AM3) Approach to Predictions of Solid Cooling Rate and Time Above Melt

Qualification of a laser powder bed fusion additive manufacturing (LPBF-AM) process requires knowledge of the multi-scale material physics during the process, per part. As the LPBF-AM build occurs, each moment is influenced by the process history. Knowledge of the build sequence can be used to generate a discretized time-space-condition point field that when coupled with a nearest neighbors’ calculation results in a generalized and fully parallel process model computation. This GPU accelerated approach was developed for part-scale analysis of build files along with in-situ process monitoring sensor data and is termed the “Additive Manufacturing Moment Measure” (AM3). The AM3 approach will be presented and then used to evaluate an AM Bench relevant geometry with synchronized in-situ process data, ex-situ nondestructive evaluation, and optical microscopy observations. These comparisons permit a better understanding of how the process actions can affect the LPBF-AM build quality and the signals generated during in-situ process monitoring.

Additive Manufacturing↗