Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “CUDA-aware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

GPU-accelerated DNS of compressible turbulent flows

Here, this paper explores strategies to transform an existing CPU-based high-performance computational fluid dynamics solver, HyPar, for compressible flow simulations on emerging exascale heterogeneous (CPU+GPU) computing platforms. The scientific motivation for developing a GPU-enhanced version of HyPar is to simulate canonical turbulent flows at the highest resolution possible on such platforms. We show that optimizing memory operations and thread blocks results in 200x speedup of computationally intensive kernels compared with a CPU core. Using multiple GPUs and CUDA-aware MPI communication, we demonstrate both strong and weak scaling of our GPU-based HyPar implementation on the NVIDIA Volta V100 GPUs. We simulate the decay of homogeneous isotropic turbulence in a triply periodic box on grids with up to 1024 3 points (5.3 billion degrees of freedom) and on up to 1,024 GPUs. We compare the wall times for CPU-only and CPU+GPU simulations. The results presented in the paper are obtained on the Summit and Lassen supercomputers at Oak Ridge and Lawrence Livermore National Laboratories, respectively.

97 MATHEMATICS AND COMPUTING↗

Accelerating the Lagrangian Particle Tracking in Hydrologic Modeling to Continental-Scale

Unprecedented climate change and anthropogenic activities have induced increasing ecohydrological problems, which have motivated the development of large-scale modeling for solutions. Water age/quality is as important as water quantity for understanding the water cycle. However, current scientific progress in tracking water parcels at large-scale with high spatiotemporal resolutions is far behind that in water balance/quantity owing to the lack of powerful tools. EcoSLIM is a particle tracking model that works with the hydrologic model ParFlow-CLM, which couples surface-subsurface hydrology with land surface processes. Here, we demonstrate a parallel framework to accelerate EcoSLIM to continental-scale on a distributed, multi-GPU platform with CUDA-Aware MPI. In tests from catchment-, to regional-, and then to continental-scale using 25-million to 1.6-billion particles, EcoSLIM shows significant speedup and excellent parallel performance. The parallel framework is portable to atmospheric and oceanic particle tracking models, where parallelization is inadequate and a standard parallel framework is absent. Parallelized EcoSLIM is a promising tool to accelerate our understanding of the terrestrial water cycle and the upscaling of subsurface hydrology to Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Venado acceptance: results and tips [Slides]

Nvidia compiler support is not available through cray-mpich/compiler wrapper interface. Adjust CMAKE files to use the correct COMPILER_ID in conditionals and explicit variables to package flags. Use pinned host memory in cray-libsci_acc and cublasXt calls. Set a large blockDim for cublasXt calls. Try MPS and/or explicit numactl binding if performance is lackluster. Use CRAY_MALLOPT_OFF=1 if unexpected OOM errors appear using cce. Use MPICH_SMP_SINGLE_COPY_MODE=CMA for xpmem issues. Use MPICH_OPT_THREAD_SYNC=0 for MPI_THREAD issues. Poor CUDA-aware MPI performance remains an issue.

97 MATHEMATICS AND COMPUTING↗

Porting OVERFLOW CFD Code to GPUs: To Hackathons and Beyond!

OVERFLOW is an overset, structured computational fluid dynamics (CFD) code written in Fortran which is widely used in the government, industry, and academia. Over the last several years the OVERFLOW developers have been working to port miniapps based on computationally expensive parts of OVERFLOW to run on GPUs, primarily using OpenACC. This effort started at our first hackathon in 2019 and since then the OVERFLOW team has attended two additional hackathons (virtually). These hackathon environments have provided a great place to collaborate with others and learn from experts. These learning experiences enabled porting two miniapps to run effectively on NVIDIA GPUs using OpenACC. The first miniapp focused on motifs found in the solver itself and the final ported version runs three times fast ona single V100 compared to a 40 core, dual-socket Intel Skylake node. The speed up in this solverminiapp required multiple design changes including increasing the amount of parallelism available and the amount of work performed in each kernel. The second miniapp focused on overset MPI communication, also saw significant speedups over the CPU implementation using a CUDA-aware MPI implementation through OpenACC. This presentation will discuss our experience at the hackathons, our process of porting the miniapps to run on the GPUs, and several lessons learned throughout.

OpenACC↗

Characterizing the performance of node-aware strategies for irregular point-to-point communication on heterogeneous architectures

Supercomputer architectures are trending toward higher computational throughput due to the inclusion of heterogeneous compute nodes. These multi-GPU nodes increase on-node computational efficiency, while also increasing the amount of data to be communicated and the number of potential data flow paths. In this work, we characterize the performance of irregular point-to-point communication with MPI on heterogeneous compute environments through performance modeling, demonstrating the limitations of standard communication strategies for both device-aware and staging-through-host communication techniques. Presented models suggest staging communicated data through host processes then using node-aware communication strategies for high inter-node message counts. Notably, the models also predict that node-aware communication utilizing all available CPU cores to communicate inter-node data leads to the most performant strategy when communicating with a high number of nodes. Furthermore, model validation is provided via a case study of irregular point-to-point communication patterns in distributed sparse matrix–vector products. Importantly, we include a discussion on the implications model predictions have on communication strategy design for emerging supercomputer architectures.

97 MATHEMATICS AND COMPUTING↗

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗