Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Summit supercomputer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

High-throughput virtual laboratory for drug discovery using massive datasets

Time-to-solution for structure-based screening of massive chemical databases for COVID-19 drug discovery has been decreased by an order of magnitude, and a virtual laboratory has been deployed at scale on up to 27,612 GPUs on the Summit supercomputer, allowing an average molecular docking of 19,028 compounds per second. Over one billion compounds were docked to two SARS-CoV-2 protein structures with full optimization of ligand position and 20 poses per docking, each in under 24 hours. GPU acceleration and high-throughput optimizations of the docking program produced 350× mean speedup over the CPU version (50× speedup per node). GPU acceleration of both feature calculation for machine-learning based scoring and distributed database queries reduced processing of the 2.4 TB output by orders of magnitude. The resulting 50× speedup for the full pipeline reduces an initial 43 day runtime to 21 hours per protein for providing high-scoring compounds to experimental collaborators for validation assays.

97 MATHEMATICS AND COMPUTING↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Towards exascale for wind energy simulations

We examine large-eddy-simulation modeling approaches and computational performance of two open-source computational fluid dynamics codes for the simulation of atmospheric boundary layer flows that are of direct relevance to wind energy production. The first code, NekRS, is a high-order, unstructured-grid, spectral element code. The second code, AMR-Wind, is a second-order, block-structured, finite-volume code with adaptive mesh refinement capabilities. The objective of this study is to co-develop these codes in order to improve model fidelity and performance for each. These features will be critical for running ABL-based applications such as wind farm analysis on advanced computing architectures. To this end, we investigate the performance of NekRS and AMR-Wind on the Oak Ridge Leadership Facility supercomputers Summit, using 4 to 800 nodes (24 to 4,800 NVIDIA V100 GPUs), and Crusher, the testbed for the Frontier exascale system, using 18 to 384 Graphics Compute Dies on AMD MI250X GPUs. We compare strong- and weak-scaling capabilities, linear solver performance, and time to solution. We also identify leading inhibitors to parallel scaling.

17 WIND ENERGY↗

FeSi binary alloy electronic structure low-Si dataset (1024 atoms)

This dataset contaims the calculated atomic charge density, atomic magnetic moment, and total energy for 1600 configurations of iron-silicon (Fe-Si) binary alloys body-centered cubic (BCC) structures at 3, 6, and 9% Si. These large scale (1024 atom) ab initio calculations were produced with the LSMS code on the OLCF Summit supercomputer. LSMS GitHub repository: https://github.com/mstsuite/lsms

36 MATERIALS SCIENCE↗

Compare linear-system solver and preconditioner stacks with emphasis on GPU performance and propose phase-2 NGP solver development pathway

The goal of the ExaWind project is to enable predictive simulations of wind farms comprised of many megawatt-scale turbines situated in complex terrain. Predictive simulations will require computational fluid dynamics (CFD) simulations for which the mesh resolves the geometry of the turbines and captures the rotation and large deflections of blades. Whereas such simulations for a single turbine are arguably petascale class, multi-turbine wind farm simulations will require exascale-class resources. The primary physics codes in the ExaWind project are Nalu-Wind, which is an unstructured-grid solver for the acoustically incompressible Navier-Stokes equations, and OpenFAST, which is a whole-turbine simulation code. The Nalu-Wind model consists of the mass-continuity Poisson-type equation for pressure and a momentum equation for the velocity. For such modeling approaches, simulation times are dominated by linear-system setup and solution for the continuity and momentum systems. For the ExaWind challenge problem, the moving meshes greatly affect overall solver costs as reinitialization of matrices and recomputation of preconditioners is required at every time step. In this report we evaluated GPU-performance baselines for the linear solvers in the Trilinos and hypre solver stacks using two representative Nalu-Wind simulations: an atmospheric boundary layer precursor simulation on a structured mesh, and a fixed-wing simulation using unstructured overset meshes. Both strong-scaling and weak-scaling experiments were conducted on the OLCF supercomputer Summit and similar proxy clusters. We focused on the performance of multi-threaded Gauss-Seidel and two-stage Gauss-Seidel that are extensions of classical Gauss-Seidel; of one-reduce GMRES, a communication-reducing variant of the Krylov GMRES; and algebraic multigrid methods that incorporate the afore-mentioned methods. The team has established that AMG methods are capable of solving linear systems arising from the fixed-wing overset meshes on CPU, a critical intermediate result for ExaWind FY20 Q3 and Q4 milestones. For the fixed-wing strong-scaling study (model with 3M grid-points), the team identified that Nalu-Wind simulations with the new Trilinos and hypre solvers scale to modest GPU counts, maintaining above 70% efficiency up to 6 GPUs. However, there still remain significant bottlenecks to performance: matrix assembly (hypre), AMG setup (hypre and Trilinos) In the weak-scaling experiments (going from 0.4M to 211M gridpoints), it's shown that the solver apply phases are faster on GPUs, but that Nalu-Wind simulation times grow, primarily due to the multigrid-setup process. Finally, based on the report outcomes, we propose a linear solver path-forward for the remainder of the ExaWind project. Near term, the NREL team will continue their work on GPU-based linear-system assembly. They will also investigate how the use of alternatives to the NVIDIA UVM (unified virtual memory) paradigm affects performance. Longer term, the NREL team will evaluate algorithmic performance on other types of accelerators and merge their improvements back to the main hypre repository branch. Near term, the Trilinos team will address performance bottlenecks identified in this milestone, such as implementing a GPU-based segregated momentum solve and reusing matrix graphs across linear-system assembly phases. Longer term, the Trilinos team will do detailed analysis and optimization of multigrid setup.

17 WIND ENERGY↗

Initial full core SMR simulations with NekRS

This document describes the completion of a recent milestone of the ExaSMR program concerning "Full core simulations". The ExaSMR project is developing tools for coupling Monte Carlo (MC) radiation transport solvers to a computational fluid dynamics (CFD) solver. Work is based on the Shift MC, OpenMC MC, and Nek5000/NekRS CFD codes. As part of this milestone, a novel set of pin-resolved CFD full-core simulations have been performed for the first time. These simulations represent a significant increase in capability in what is now possible with CFD on pre-Exascale systems. The simulations have been performed with the spectral element solvers NekRS on the supercomputer Summit. Multiple simulations campaigns have been conducted: (1) LES simulations in bare bundle, (2) RANS simulations in bare bundles, (3) RANS simulations with momentum sources and (4) Conjugate heat transfer calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Distributed Training for High Resolution Images: A Domain and Spatial Decomposition Approach

In this work we developed two Pytorch libraries using the PyTorch RPC interface for distributed deep learning approaches on high resolution images. The spatial decomposition library allows for distributedtraining on very large images, which otherwise won’t be possible on a single GPU. The domain parallelism library allows for distributed training across multiple domain unlabeled data, by leveraging the domain separation architecture. Both of those libraries were tested on the Summit supercomputer at a moderate scale, and we are releasing the code for both of them.

97 MATHEMATICS AND COMPUTING↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

Inference-Optimized AI and High Performance Computing for Gravitational Wave Detection at Scale

We introduce an ensemble of artificial intelligence models for gravitational wave detection that we trained in the Summit supercomputer using 32 nodes, equivalent to 192 NVIDIA V100 GPUs, within 2 h. Once fully trained, we optimized these models for accelerated inference using NVIDIA TensorRT. We deployed our inference-optimized AI ensemble in the ThetaGPU supercomputer at Argonne Leadership Computer Facility to conduct distributed inference. Using the entire ThetaGPU supercomputer, consisting of 20 nodes each of which has 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, our NVIDIA TensorRT-optimized AI ensemble processed an entire month of advanced LIGO data (including Hanford and Livingston data streams) within 50 s. Our inference-optimized AI ensemble retains the same sensitivity of traditional AI models, namely, it identifies all known binary black hole mergers previously identified in this advanced LIGO dataset and reports no misclassifications, while also providing a 3X inference speedup compared to traditional artificial intelligence models. We used time slides to quantify the performance of our AI ensemble to process up to 5 years worth of advanced LIGO data. In this synthetically enhanced dataset, our AI ensemble reports an average of one misclassification for every month of searched advanced LIGO data. We also present the receiver operating characteristic curve of our AI ensemble using this 5 year long advanced LIGO dataset. This approach provides the required tools to conduct accelerated, AI-driven gravitational wave detection at scale.

97 MATHEMATICS AND COMPUTING↗

H-AMR: A New GPU-accelerated GRMHD Code for Exascale Computing with 3D Adaptive Mesh Refinement and Local Adaptive Time Stepping

General relativistic magnetohydrodynamic (GRMHD) simulations have revolutionized our understanding of black hole accretion. Here, we present a GPU-accelerated GRMHD code H-AMR with multifaceted optimizations that, collectively, accelerate computation by 2–5 orders of magnitude for a wide range of applications. First, it introduces a spherical grid with 3D adaptive mesh refinement that operates in each of the three dimensions independently. This allows us to circumvent the Courant condition near the polar singularity, which otherwise cripples high-resolution computational performance. Second, we demonstrate that local adaptive time stepping on a logarithmic spherical-polar grid accelerates computation by a factor of ≲10 compared to traditional hierarchical time-stepping approaches. Jointly, these unique features lead to an effective speed of ~10 9 zone cycles per second per node on 5400 NVIDIA V100 GPUs (i.e., 900 nodes of the OLCF Summit supercomputer). We illustrate H-AMR's computational performance by presenting the first GRMHD simulation of a tilted thin accretion disk threaded by a toroidal magnetic field around a rapidly spinning black hole. With an effective resolution of 13,440 × 4608 × 8092 cells and a total of ≲22 billion cells and ~0.65 × 10 8 time steps, it is among the largest astrophysical simulations ever performed. We find that frame dragging by the black hole tears up the disk into two independently precessing subdisks. The innermost subdisk rotation axis intermittently aligns with the black hole spin, demonstrating for the first time that such long-sought alignment is possible in the absence of large-scale poloidal magnetic fields.

79 ASTRONOMY AND ASTROPHYSICS↗

AI ensemble for signal detection of higher order gravitational wave modes of quasi-circular, spinning, non-precessing binary black hole mergers

We introduce spatiotemporal-graph models that concurrently process data from the twin advanced LIGO detectors and the advanced Virgo detector. We trained these AI classifiers with 2.4 million IMRPhenomXPHM waveforms that describe quasi-circular, spinning, non-precessing binary black hole mergers with component masses m{1,2}∈[3M⊙,50M⊙], and individual spins sz{1,2}∈[−0.9,0.9]; and which include the (ℓ,|m|)={(2,2),(2,1),(3,3),(3,2),(4,4)} modes, and mode mixing effects in the ℓ=3,|m|=2 harmonics. We trained these AI classifiers within 22 hours using distributed training over 96 NVIDIA V100 GPUs in the Summit supercomputer. We then used transfer learning to create AI predictors that estimate the total mass of potential binary black holes identified by all AI classifiers in the ensemble. We used this ensemble, 3 classifiers for signal detection and 2 total mass predictors, to process a year-long test set in which we injected 300,000 signals. This year-long test set was processed within 5.19 minutes using 1024 NVIDIA A100 GPUs in the Polaris supercomputer (for AI inference) and 128 CPU nodes in the ThetaKNL supercomputer (for post-processing of noise triggers), housed at the Argonne Leadership Computing Facility. These studies indicate that our AI ensemble provides state-of-the-art signal detection accuracy, and reports 2 misclassifications for every year of searched data. This is the first AI ensemble designed to search for and find higher order gravitational wave mode signals.

Tian, Minyang↗

VisMetHack2022: Visualizing winds and surface variables from the ECMWF IFS 1-km nature run

This data collection was contributed to the Visualisation Hackathon 2022 (#VisMetHack2022), in conjunction with the Using ECMWF's Forecasts (UEF2022) workshop. The European Center for Medium-Range Weather Forecasts (ECMWF) and the Oak Ridge National Laboratory (ORNL) are pleased to announce access to the data collection from global 1-km nature run (NR) simulations using the Integrated Forecast System (IFS) with explicit convection. We invite you to join us in exploring this precursor to a digital twin of the earth! The NR simulations reveal unprecedented detail of the earth’s atmosphere, and the then outgoing Editor-in-Chief of AGU JAMES commended the project as one of “stunning ambitions,” enabled by computational capacity at scale. The project also won the 2020 HPCwire Readers Choice Award for Best Use of HPC in Physical Sciences. A set of two NR seasonal simulations have been completed, one corresponding to the northern hemispheric winter months (NDJF) and the other for the North Atlantic tropical cyclone season (ASO). The project used the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). The simulations were facilitated with an INCITE award from the US Department of Energy Office of Science. For the first seasonal run of four months (NDJF), the hydrostatic IFS model was initialized at 00Z on 1 November 2018. The NR for the TC season (AS) was initialized at 00Z on 1 August 2019. The NR simulations were constrained only by sea surface temperatures (SST) at the lower boundary. The IFS output was saved every 3 hours. After feedback and interest from the scientific community, the simulations were rerun for four specific extreme events, with output every 15 minutes. The special cases include a tropical cycle and three severe storm events over the continental USA.

54 ENVIRONMENTAL SCIENCES↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

Workflow Submit Nodes as a Service on Leadership Class Systems

DOE scientists, today, have access to high performance computing (HPC) facilities with very powerful systems that enable them to execute their computations faster, more efficiently, and at greater scales than ever before. To further their knowledge and produce new discoveries, scientists rely on workflows - sometimes very complex - that provide them with an easy way to automate, reproduce and verify their computations. However, historically, creating workflow submission environments in large HPC facilities has been cumbersome, requires expertise and many man-hours of effort due to the peculiarities, policies, and the restrictions that these systems present. In this paper we discuss the approach a large DOE facility (OLCF) is taking in order to provide containers as a service to its users. This capability is used to create Pegasus workflow management system submit nodes as a service (WSaaS) at the Oak Ridge Leadership Computing Facilities (OLCF), targeting the Summit supercomputer. This deployment builds upon the Kubernetes/Openshift cluster (Slate) that exists within OLCF’s DMZ and its automation triggers. Additionally, we evaluate our approach’s overhead and effort to deploy the solution as compared to previous solutions, such as setting up a Pegasus submission environment on OLCF’s login nodes or submitting jobs remotely via the rvGAHP.

Papadimitriou, George↗

PREEMPT: Scalable Epidemic Interventions Using Submodular Optimization on Multi-GPU Systems

Preventing and slowing the spread of epidemics is achieved through techniques such as vaccination and social distancing. Given practical limitations on the number of vaccines and cost of administration, optimization becomes a necessity. Previous approaches using mathematical programming methods have shown to be effective but are limited by computational costs. In this work, we make several contributions: First, we present a new approach for intervention via maximizing the influence of vaccinated nodes on the network. We call this method \preempt. Next, we prove submodular properties associated with the objective function of our method so that it aids in construction of an efficient greedy approximation strategy. Consequently, we present a new parallel algorithm based on greedy hill climbing for \preempt, and present an efficient parallel implementation for distributed CPU-GPU heterogeneous platforms. Our results demonstrate that \preempt{} is able to achieve a significant reduction (up to 6.75$\times$) in the percentage of people infected on a city-scale network. We also show strong scaling results of \preempt{} on 128 nodes of the Summit supercomputer. Our parallel implementation is able to significantly reduce time to solution, from hours to minutes on large networks. This work represents a first-of-its-kind effort in parallelizing greedy hill climbing and applying it toward devising effective interventions for epidemics.

Minutoli, Marco↗

Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters

We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.

Sao, Piyush↗

Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUs

Rapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput—83% of theoretical peak—on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software.

Chen, Jieyang↗

Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations: Preprint

The US Department of Energy has identified Exascale-Class wind farm simulation tools as critical to wind energy scientific discovery. A primary objective of the Exawind project is to build high-performance, predictive Computational Fluid Dynamics tools that satisfy these modeling needs. GPU accelerators will serve as the computational thoroughbreds of next generation, Exascale-Class, platforms. Here, we report on our efforts for preparing the Exawind unstructured mesh solver, Nalu-Wind, for Exascale-Class machines. For computing at this scale, a simple port of the incompressible-flow algorithms to GPUs is not sufficient. One needs novel algorithms that are application aware, memory efficient, and optimized for latest generation GPU devices to get high-performance. The result of our efforts are unstructured mesh simulations of wind turbines that use 1/6 the compute resources of Summit supercomputer at Oak Ridge National Lab. In particular, we demonstrate a first-of-its-kind, simulation using Algebraic Multigrid solvers on over 4000 GPUs.

algebraic multigrid↗