Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Exascale applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Transforming Science Through Software: Improving While Delivering 100×

The U.S. Department of Energy (DOE) Exascale Computing Project (ECP) funded the development of new (and the transformation of important existing) applications, libraries, and tools that realized improvement in performance and capabilities of often 100 times or more on emerging exascale computers. This exceptional gain inspired the title of this special issue: Transforming Science through Software: Improving while delivering 100X. The term 100X refers to advancing capabilities in modeling, simulation, and analysis by a factor of 100 or more using some combination of new algorithms, optimization techniques, software libraries, and programming models, coupled with the next generation of hardware for high-performance computing (HPC). The papers in this issue share experiences with the practice and science of scientific software development, with an emphasis on developing a coherent, portable, and sustainable HPC software ecosystem for next-generation computational science. Finally, we hope to foster expanded community efforts related to the fundamental role of sustainable scientific software ecosystems in advancing the computing sciences.

97 MATHEMATICS AND COMPUTING↗

Performance Portable Graphics Processing Unit Acceleration of a High-Order Finite Element Multiphysics Application

The Lawrence Livermore National Laboratory (LLNL) will soon have in place the El Capitan exascale supercomputer, based on advanced micro devices (AMD) graphics processing units (GPUs). As part of a multiyear effort under the National Nuclear Security Administration (NNSA) Advanced Simulation and Computing (ASC) program, we have been developing marbl, a next generation, performance portable multiphysics application based on high-order finite elements. In previous years, we successfully ported the Arbitrary Lagrangian–Eulerian (ALE), multimaterial, compressible flow capabilities of marbl to nvidia GPUs as described in Vargas et al. Here, in this paper, we describe our ongoing effort in extending marbl's GPU capabilities with additional physics, including multigroup radiation diffusion and thermonuclear burn for high energy density physics (HEDP) and fusion modeling. We also describe how our portability abstraction approach based on the raja Portability Suite and the mfem finite element discretization library has enabled us to achieve high performance on AMD based GPUs with minimal effort in hardware-specific porting. Throughout this work, we highlight numerical and algorithmic developments that were required to achieve GPU performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Energy Exascale Computational Fluid Dynamics Simulations With the Spectral Element Method

Development and application of the open-source GPU-based fluid-thermal simulation code, NekRS, are described. Time advancement is based on an efficient kth-order accurate timesplit formulation coupled with scalable iterative solvers. Spatial discretization is based on the high-order spectral element method (SEM), which affords the use of fast, low-memory, matrix-free operator evaluation. Further, recent developments include support for nonconforming meshes using overset grids and for GPU-based Lagrangian particle tracking. Results of large-eddy simulations of atmospheric boundary layers for wind-energy applications as well as extensive nuclear energy applications are presented.

42 ENGINEERING↗

Leveraging Compiler-Based Translation to Evaluate a Diversity of Exascale Platforms

Accelerator-based heterogeneous computing is the de facto standard in current and upcoming exascale machines. These heterogeneous resources empower computational scientists to select a machine or platform well-suited to their domain or applications. However, this diversity of machines also poses challenges related to programming model selection: inconsistent availability of programming models across different exascale systems, lack of performance portability for those programming models that do span several systems, and inconsistent performance between different models on a single platform. We explore these challenges on exascale-similar hardware, including AMD MI100 and NVIDIA A100 GPUs. By extending the sourceto-source compiler OpenARC, we demonstrate the power of automated translation of applications written in a single frontend programming model (OpenACC) into a variety of backend models (OpenMP, OpenCL, CUDA, HIP) that span the upcoming exascale environments. This translation enables us to compare performance within and across devices and to analyze programming model behavior with profiling tools.

Lambert, Jacob↗

Position Papers for the ASCR Workshop on Reimagining Codesign

On behalf of the Advanced Scientific Computing Research (ASCR) program in the US Department of Energy (DOE) Office of Science, we are organizing a Workshop on Reimagining Codesign (ReCoDe). Codesign is the process of jointly designing interoperating components of a computing system—in particular: applications, algorithms, system software, programming models, and the hardware on which they run. The goal is to maximize the overall performance, efficiency, and other desirable qualities of the system as a whole. Codesign is a standard methodology in the embedded-systems community, where space, power, and cost constraints are commonly pitted against execution speed for a tightly constrained feature set. Over the last decade, the DOE has invested in codesign efforts to foster the development of exascale computing systems for broad classes of scientific and engineering applications. The ReCoDe workshop hopes to explore how scientific applications of interest to the DOE can be accelerated through close interactions with hardware designers and software-stack developers, in which all components adapt to each other’s requirements and constraints. We want to answer the question of what are the key tools and methodologies for accomplishing codesign in today’s computing landscape, and what will be the highest impact targets for meeting DOE’s emerging mission requirements. This workshop aims to bring together DOE, industry, and academia to identify opportunities to build on past codesign successes and identify new areas that are either emerging or that may need reimagining for the future. We want to continue to find opportunities that can be pursued as a joint effort and continue to break down the traditional customer/vendor dichotomy with true partnerships. From this work, DOE will benefit from increased application performance relative to what stock hardware or existing general-purpose roadmaps can provide, and vendors will benefit from expanding their hardware’s capabilities to address needs they might have not otherwise anticipated and thereby create more widespread interest in their products. The workshop will be structured around a set of breakout sessions, with every attendee expected to participate actively in the discussions. Afterward, workshop attendees—from DOE, industry, and academia—will produce a report for ASCR that summarizes the findings made during the workshop.

97 MATHEMATICS AND COMPUTING↗

Science & Technology Review: The Road to Exascale Computing

At Lawrence Livermore National Laboratory, we focus on science and technology research to ensure our nation’s security. We also apply that expertise to solve other important national problems in energy, bioscience, and the environment. Science & Technology Review is published eight times a year to communicate, to a broad audience, the Laboratory’s scientific and technological accomplishments in fulfilling its primary missions. The publication’s goal is to help readers understand these accomplishments and appreciate their value to the individual citizen, the nation, and the world. The Department of Energy’s Exascale Computing Project (ECP) and Lawrence Livermore’s RADIUSS (Rapid Application Development via an Institutional Universal Software Stack) initiative benefit from strategically developed software tools. The front cover shows a simulation of advection under twisting rotation that uses high-order finite elements from Livermore’s Modular Finite Element Methods (MFEM) software library and GLVis visualization tool. On the back cover, the logo (also created with GLVis) for the MFEM project illustrates the curved mesh and sub-element resolution used in high-order simulations. MFEM and GLVis are key components of the ECP’s co-design Center for Efficient Exascale Discretizations (CEED) and RADIUSS. MFEM is also part of ECP’s Extreme-Scale Scientific Software Development Kit (xSDK).

97 MATHEMATICS AND COMPUTING↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗

Scaling the memory wall using mixed-precision - HPG-MxP on an exascale-class machine

Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.

Kashi, Aditya [ORNL] (ORCID:0000000325893792)↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputer

As we approach the exascale computing era, the focused understanding of power consumption and its overall constraint on HPC architectures and applications are becoming increasingly paramount. Summit, located at the Oak Ridge Leadership Computing Facility (OLCF), is one of the fastest and largest pre-exascale platforms in operation today. This paper provides a first-order examination and analysis of power consumption at the component-level, node-level, and system-level, from all 4,626 Summit compute nodes, each with over 100 metrics at 1Hz frequency over the entire year of 2020. We also investigate the power characteristics and energy efficiency of over 840k Summit jobs and 250k GPU failure logs for further operational insights. To the best of our knowledge, this is the first systematic analysis of power data of HPC system at this scale.

Shin, Woong↗

Porting a fast-running wildfire plume model for use in national security and climate applications

In this project, the code for the fast-running Freitas wildfire plume model was successfully ported into the NARAC system, creating an operations-ready capability for predicting the transport of hazardous smoke from large wildfires. “Porting” refers to copying a section of code from one software package to another while accounting for unique dependencies in each software package. In this case, the code was ported from the Weather Research and Forecasting (WRF) Model, which is used extensively by the project team. To verify correct implementation, model results from the updated NARAC system were compared to those from WRF and the High-Resolution Rapid Refresh (HRRR) forecast model, which also employs the Freitas scheme. The NARAC User Interface (UI) and code documentation were also updated to include the new wildfire plume option, and a wildfire case was added to the NARAC test suite. The code was used in a real-time, semi-operational setup to predict smoke transport during the Cerro Pelado fire near Los Alamos National Laboratory in May 2022. The Freitas plume model was also tested for LLNL climate applications. Although the code was not ported into the Energy Exascale Earth System Model (E3SM), it was tested offline using E3SM atmospheric inputs, and results were compared to existing high-resolution plume predictions from large-eddy simulations in WRF. The code implementation and testing completed during this project will enable future work predicting smoke transport from large wildfires in LLNL national security and climate mission spaces.

54 ENVIRONMENTAL SCIENCES↗

ExaWorks: Workflows for Exascale

Exascale computers will offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. These software combinations and integrations, however, are difficult to achieve due to challenges of coordination and deployment of heterogeneous software components on diverse and massive platforms. We present the ExaWorks project, which can address many of these challenges: ExaWorks is leading a co-design process to create a workflow Software Development Toolkit (SDK) consisting of a wide range of workflow management tools that can be composed and interoperate through common interfaces. We describe the initial set of tools and interfaces supported by the SDK, efforts to make them easier to apply to complex science challenges, and examples of their application to exemplar cases. Furthermore, we discuss how our project is working with the workflows community, large computing facilities as well as HPC platform vendors to sustainably address the requirements of workflows at the exascale.

97 MATHEMATICS AND COMPUTING↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

PeleMP: The Multiphysics Solver for the Combustion Pele Adaptive Mesh Refinement Code Suite

Combustion encompasses multiscale, multiphase reacting flow physics spanning a wide range of scales from the molecular scales, where chemical reactions occur, to the device scales, where the turbulent flow is affected by the geometry of the combustor. This scale disparity and the limited measurement capabilities from experiments make modeling combustion a significant challenge. Recent advancements in high-performance computing (HPC), particularly with the Department of Energy's Exascale Computing Project (ECP), have enabled high-fidelity simulations of practical applications to be performed. The major physics submodels, including chemical reactions, turbulence, sprays, soot, and thermal radiation, exhibit distinctive computational characteristics that need to be examined separately to ensure efficient utilization of computational resources. This paper presents the multiphysics solver for the Pele code suite, called PeleMP, which consists of models for spray, soot, and thermal radiation. Here, the mathematical and algorithmic aspects of the model implementations are described in detail as well as the verification process. The computational performance of these models is benchmarked on multiple supercomputers, including Frontier, an exascale machine. Results are presented from production simulations of a turbulent sooting ethylene flame and a bluff-body swirl stabilized spray flame with sustainable aviation fuels to demonstrate the capability of the Pele codes for modeling practical combustion problems with multiphysics. This work is an important step toward the exascale computing era for high-fidelity combustion simulations providing physical insights and data for predictive modeling of real-world devices.

42 ENGINEERING↗

Julia as a unifying end-to-end workflow language on the Frontier exascale system

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.

Godoy, William↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of PIConGPU: Particle-in-Cell on GPUs using NVIDIA’s NSight Systems and NSight Compute

PIConGPU, Particle In Cell on GPUs, is an open source simulations framework for plasma and laser-plasma physics used to develop advanced particle accelerators for radiation therapy of cancer, high energy physics and photon science. While PIConGPU has been optimized for at least 5 years to run well on NVIDIA GPU-based clusters, there has been limited exploration by the development team of potential scalability bottlenecks using recently updated and new tools including NVIDIA’s NVProf tool and the brand-new NVIDIA NSight Suite (Systems and Compute) tools. PIConGPU is a highly optimized application that runs production jobs at scale on a system Oak Ridge Leadership Facility’s (OLCF) Summit supercomputer (using the full machine at 4600 nodes; at 98% of GPU utilization on all ~28000 NVIDIA Volta GPUs). PIConGPU has been selected as one of the the eight applications for OLCF’s coveted Center for Accelerated Application Readiness (CAAR) program aimed at the facility’s Frontier supercomputer (OLCF’s first exascale system to launch in 2021), to partner with our vendors (primary vendors: AMD and Cray/HPE) ensuring that Frontier will be able to perform large-scale science when it opens to users in 2022. To this effect, performance engineers on the PIConGPU team wanted to dive deep into the application to understand at the finest granularity, which portions of the code could be further optimized to exploit the hardware on Summit at it’s maximum potential and also to elucidate which key kernels should be tracked and optimized for the CAAR effort to port this code to Frontier. Any bottlenecks that are observed via performance profiling on Summit are likely to also impact scalability on the Frontier-dev system and the Frontier Early Access (EA) system. Additionally, the engineers wanted to take a closer look at the newest NVIDIA profiling tools which allows us to identify the most useful features on these tools and will provide an opportunity to compare it to new AMD and Cray’s performance analysis tool releases and provide feedback to our vendor partners on what features are most important and mission critical for CAAR efforts. The primary goal of this report is to focus on the evaluation of PIConGPU’s most time-intensive kernels using NVProf and NSight Suite. Three kernels, Current Deposition (also known as Compute Current), Particle Push (Move and Mark), and Shift Particles are known to be some of the most time-consuming kernels in PIConGPU. The Current xi Deposition kernel and Particle Push kernel both set up the particle attributes for running any physics simulation with PIConGPU, so it is crucial to improve the performance of these two kernels. In this report, we measure single GPU metrics for the three kernels, offer high level takeaways from the conducted analysis, and compare the profiling data from NSight Compute to that of NVProf. This analysis was performed using a grid size of 240 x 272 x 224, and 10 time steps with the Mid-November Figure of Merit (FOM) run setup. The Traveling Wave Electron Acceleration (TWEAC) science case used in this run is a representative science case for PIConGPU. This execution can also be used for baseline analysis on AMD MI50/ MI60 systems. As of the time of writing, the PIConGPU application has limited use for features of NSight Systems, so this report will mainly focus on insights garnered from NSight Compute. For this analysis, we run the “full” metric set available in NSight Compute version 2020.1.2 and use NSight Systems version 2020.3.1 to generate the application timeline.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Strategies for Working Remotely: Responding to Pandemic-Driven Change.

In response to the COVID-19 pandemic, the Exascale Computing Project’s (ECP) Interoperable Design of Extreme-scale Application Software (IDEAS) productivity team launched the panel series Strategies for Working Remotely to facilitate informal, cross-organizational dialog in the absence of face-to-face meetings. In a time of pandemic, organizations increasingly need to reach across perceived boundaries to learn from each other, so that we can move beyond stand-alone silos to more connected multidisciplinary and multiorganizational configurations. The present paper argues that the unplanned transition to remote work, overuse of electronic communication, and need to unlearn habits associated with an overreliance on face-to-face, created unique opportunities to learn from the situation and accelerate cross-institutional cooperation and collaboration through online community dialog facilitated by informal panel discussions. Recommendations for facilitating online panel discussions to foster cross-organizational dialog are provided by applying the Simulation Experience Design Method.

Raybourn, Elaine M.↗