Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “exascale computing project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Co-design for Particle Applications at Exascale

Co-design across the Exascale Computing Project (ECP) has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing (HPC) architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and PROGRESS/BML libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from atomistic, mesoscale, continuum, through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate separation of concerns and directly support science runs.

97 MATHEMATICS AND COMPUTING↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

Application Results on Early Exascale Hardware

This Exascale Computing Project (ECP) milestone report summarizes the status of 27 of the 31 ECP Applications Development (AD) subprojects at the end of FY21. In November and December of 2021, a comprehensive assessment of AD projects was conducted by the ECP leadership along with external subject matter experts (SMEs). (NNSA application projects are reviewed separately using the ASC milestone process.) The AD review committee—consisting of the AD lead, AD deputy, Level 3 (L3), and at least one external project SME—was tasked with evaluating each project’s progress relative to ECP project goals specified in the FY21 timeline. Key areas of focus were code maturity and performance on pre-exascale systems, an in-depth analysis of final key performance parameter (KPP) verification contracts, and future R&D priorities in the final year of ECP and beyond. As such, this report contains not only an accurate snapshot of each subproject’s current status but also represents a broad account of successes and challenges in porting large scientific applications to DOE’s next-generation high-performance computing architectures – the Frontier and Aurora systems.

97 MATHEMATICS AND COMPUTING↗

Science & Technology Review: The Road to Exascale Computing

At Lawrence Livermore National Laboratory, we focus on science and technology research to ensure our nation’s security. We also apply that expertise to solve other important national problems in energy, bioscience, and the environment. Science & Technology Review is published eight times a year to communicate, to a broad audience, the Laboratory’s scientific and technological accomplishments in fulfilling its primary missions. The publication’s goal is to help readers understand these accomplishments and appreciate their value to the individual citizen, the nation, and the world. The Department of Energy’s Exascale Computing Project (ECP) and Lawrence Livermore’s RADIUSS (Rapid Application Development via an Institutional Universal Software Stack) initiative benefit from strategically developed software tools. The front cover shows a simulation of advection under twisting rotation that uses high-order finite elements from Livermore’s Modular Finite Element Methods (MFEM) software library and GLVis visualization tool. On the back cover, the logo (also created with GLVis) for the MFEM project illustrates the curved mesh and sub-element resolution used in high-order simulations. MFEM and GLVis are key components of the ECP’s co-design Center for Efficient Exascale Discretizations (CEED) and RADIUSS. MFEM is also part of ECP’s Extreme-Scale Scientific Software Development Kit (xSDK).

97 MATHEMATICS AND COMPUTING↗

Intro to HPC Bootcamp: Engaging New Communities Through Energy Justice Projects

The U.S. Department of Energy (DOE) is a long-standing leader in research and development of high-performance computing (HPC) in the pursuit of science. However, we face daunting challenges in fostering a robust and diverse HPC workforce. Basic HPC is not typically taught at early stages of students' academic careers, and the capacity and knowledge of HPC at many institutions are limited. Even so, such topics are prerequisites for advanced training programs, internships, graduate school, and ultimately for careers in HPC. To help address this challenge, as part of the DOE Exascale Computing Project's Broadening Participation Initiative, we recently launched the Introduction to HPC Training and Workforce Pipeline Program to provide accessible introductory material on HPC, scalable AI, and analytics. We describe the Intro to HPC Bootcamp, an immersive program designed to engage students from underrepresented groups as they learn foundational HPC skills. Here, the program takes a novel approach to HPC training by turning the traditional curriculum upside down. Instead of focusing on technology and its applications, the bootcamp focuses on energy justice to motivate the training of HPC skills through project-based pedagogy and real-life science stories. Additionally, the bootcamp prepares students for internships and future careers at DOE labs. The first bootcamp, hosted by the advanced computing facilities at Argonne, Lawrence Berkeley, and Oak Ridge National Labs and organized by Sustainable Horizons Institute, took place in August 2023.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Cabana: A Performance Portable Library for Particle-Based Simulations

Particle-based simulations are ubiquitous throughout many fields of computational science and engineering, spanning the atomistic level with molecular dynamics (MD), to mesoscale particle-in-cell (PIC) simulations for solid mechanics, device-scale modeling with PIC methods for plasma physics, and massive N-body cosmology simulations of galaxy structures, with many other methods in between (Hockney & Eastwood, 1989). While these methods use particles to represent significantly different entities with completely different physical models, many low-level details are shared including performant algorithms for short- and/or long-range particle interactions, multi-node particle communication patterns, and other data management tasks such as particle sorting and neighbor list construction. Cabana is a performance portable library for particle-based simulations, developed as part of the Co-Design Center for Particle Applications (CoPA) within the Exascale Computing Project (ECP) (Alexander et al., 2020). The CoPA project and its full development scope, including ECP partner applications, algorithm development, and similar software libraries for quantum MD, is described in (Mniszewski et al., 2021). Cabana uses the Kokkos library for on-node parallelism (Edwards et al., 2014; Trott et al., 2022), enabling simulation on multi-core CPU and GPU architectures, and MPI for GPU-aware, multi-node communication. Cabana provides particle simulation capabilities on almost all current Kokkos backends, including serial execution, OpenMP (including OpenMP-Target for GPUs), CUDA (NVIDIA GPUs), HIP (AMD GPUs), and SYCL (Intel GPUs), providing a clear path for the coming generation of accelerator-based exascale hardware. Cabana builds on Kokkos by providing new particle data structures and particle algorithms resulting in a similar execution policy-based, node-level programming model that is intended to be used in addition to the core Kokkos library within an application. Cabana is designed as an application and physics agnostic, but particle-specific toolkit which can either be used to generate a new application, or to be used as needed in existing applications at various levels of invasiveness including through interfaces that wrap user memory in existing data structures.

97 MATHEMATICS AND COMPUTING↗

PeleMP: The Multiphysics Solver for the Combustion Pele Adaptive Mesh Refinement Code Suite

Combustion encompasses multiscale, multiphase reacting flow physics spanning a wide range of scales from the molecular scales, where chemical reactions occur, to the device scales, where the turbulent flow is affected by the geometry of the combustor. This scale disparity and the limited measurement capabilities from experiments make modeling combustion a significant challenge. Recent advancements in high-performance computing (HPC), particularly with the Department of Energy's Exascale Computing Project (ECP), have enabled high-fidelity simulations of practical applications to be performed. The major physics submodels, including chemical reactions, turbulence, sprays, soot, and thermal radiation, exhibit distinctive computational characteristics that need to be examined separately to ensure efficient utilization of computational resources. This paper presents the multiphysics solver for the Pele code suite, called PeleMP, which consists of models for spray, soot, and thermal radiation. Here, the mathematical and algorithmic aspects of the model implementations are described in detail as well as the verification process. The computational performance of these models is benchmarked on multiple supercomputers, including Frontier, an exascale machine. Results are presented from production simulations of a turbulent sooting ethylene flame and a bluff-body swirl stabilized spray flame with sustainable aviation fuels to demonstrate the capability of the Pele codes for modeling practical combustion problems with multiphysics. This work is an important step toward the exascale computing era for high-fidelity combustion simulations providing physical insights and data for predictive modeling of real-world devices.

42 ENGINEERING↗

GPU algorithms for Efficient Exascale Discretizations

In this paper we describe the research and development activities in the Center for Efficient Exascale Discretization within the US Exascale Computing Project, targeting state-of-the-art high-order finite-element algorithms for high-order applications on GPU-accelerated platforms. Furthermore, we discuss the GPU developments in several components of the CEED software stack, including the libCEED, MAGMA, MFEM, libParanumal, and Nek projects. We report performance and capability improvements in several CEED-enabled applications on both NVIDIA and AMD GPU systems.

97 MATHEMATICS AND COMPUTING↗

Building a Diverse and Inclusive HPC Community for Mission-Driven Team Science

The U.S. Department of Energy (DOE) has been a long-standing leader in driving advances in science and technology through advanced computing. However, DOE laboratories are currently facing urgent workforce challenges, particularly in terms of underrepresentation from key communities, including people of color, women, persons with disabilities, and first-generation scholars. This paper introduces the work carried out as part of the Exascale Computing Project (ECP) Broadening Participation Initiative, which aims to address workforce challenges through a lens that considers the distinct needs and culture of high-performance computing (HPC). The work focuses on three main efforts: hosting Intro to HPC Bootcamps, expanding the Sustainable Research Pathways (SRP) internship and workforce development program, and establishing an HPC Workforce Development and Retention Action Group. Finally, the paper also highlights various workforce efforts throughout the computational science community and explores opportunities for future work aimed at broadening participation in HPC.

97 MATHEMATICS AND COMPUTING↗

Nek5000 developments in support of industry and the NRC

This year, the Nuclear Energy Advanced Modeling Simulation program (NEAMS) thermal-hydraulics verification and validation (V&V) work has focused in three areas of Nek5000 V&V-driven development. First, in a close collaborative effort with the U. S. Nuclear Regulatory Commission (NRC) staff, we have continued V&V efforts for the HYMERES-2 project using the OECD/NEA sponsored testing in the PSI PANDA facility. This year’s focus of ANL-NRC collaboration involves Nek5000 setups and validation for a range of problems relevant to and including the HYMERES-2 benchmark from PSI. The primary outcome of this year efforts is a more efficient geometry and inlet modeling simplification after a careful sensitivity study of the inlet profiles and pipe geometries. The resulting modeling choice of a short recycling/fully-developed turbulent inlet is within the experimental uncertainty estimate. This finding simplifies the next step of the cross-V&V HYMERES-2 project. In addition, the ANL team continue to provide assistance to the NRC staff in the form of Nek5000 application support in general and on the use of the HPC platforms of ALCF and INL in particular. This supports the NRC’s assessment of Nek5000 for use with the NRC Blue CRAB code suite. Second, we have implemented and tested more robust model of URANS, namely the k – τ model, a variant of the k-ω model, along with other improvements to RANS Nek5000 modeling in general. Because of its demonstrated robustness and stability, the k – τ model is the only RANS model that has been implemented in the new GPU version of the Nek5000 code, nekRS. Lastly, we report the initial implementation of Jacobian-free Newton Krylov approach to the direct Newton method for steady fluid solvers aimed at acceleration of RANS modeling and at IC improvement for LES campaigns. Also leveraging the Exascale Computing Project (ECP) ANL/CEED & SMR team’s software development effort to support NEAMS problems at large scale of the advanced computing architectures, NekRS, a GPU variant of Nek5000, built on top of kernels from libParanumal using OCCA for portability, has been successfully run on the full system of Summit (4608 nodes, 27648 GPUs).

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

ExaWind at NREL: Upping the Ante

The objective of the ExaWind component of the Exascale Computing Project is to deliver many-turbine blade-resolved simulations in complex terrain. These simulations bring new challenges to both compute and analysis of the resulting data. In this paper/video, we visually explore the impact of ExaWind on wind simulations through two studies of a small wind farm under two atmospheric conditions. We then turn to analysis and review tools that visualization researchers at NREL use to answer the challenges that ExaWind brings.

collaborative visualization↗

Providing a Flexible and Comprehensive Software Stack Via Spack, an Extreme-Scale Scientific Software Stack, and Software Development Kits

To manage the complex demands of modern high-performance computing (HPC), software applications increasingly depend on software developed by other teams, often at other institutions. An HPC software ecosystem approach is required to support dependencies on third-party scientific software. An ecosystem approach provides layers of activity above the individual software product level that promote interoperability, quality improvement, porting, testing, and deployment. The U.S. Exascale Computing Project (ECP) developed its HPC software ecosystem using a three-pronged approach. First, the ECP adopted and invested in Spack, a package manager designed to handle complex HPC package dependencies. Second, the ECP created the Extreme Scale Scientific Software Stack, an effort that supports developing, deploying, and running scientific applications on HPC platforms. Third, the ECP supported software product communities, or software development kits, to develop and promote best practices, improve software interoperability, and other collaborative efforts. This article describes ECP contributions to HPC software ecosystem challenges.

97 MATHEMATICS AND COMPUTING↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Feasibility of full-core pin resolved CFD simulations of small modular reactor with momentum sources

Complex flow structure interactions and heat transfer processes take place in nuclear reactor cores. Given the extreme pressure/temperature and radioactive conditions inside the core, numerical simulations offer an attractive and sometimes more feasible approach to study the related flow and heat transfer phenomena in addition to the experiments. Under the Exascale Computing Project, the full-core simulation of a small modular reactor (SMR) has been pursued coupling Computational Fluid Dynamics (CFD) and neutronics. A key aspect of the modeling of SMR fuel assemblies is the presence of spacer grids and the mixing promoted by mixing vanes or the equivalent. A reduced order methodology is adopted based on momentum sources to mimic the mixing of the vanes. The momentum sources have been carefully calibrated with detailed Large Eddy Simulations (LES) of spacer grids performed with Nek5000. Modeling the spacer grid and mixing vanes (SGMV) effect without body-fitted computational grid avoids the excessive costs in resolving the local geometric details, and thus supports the simulation to be scaled up to the full core. Besides the progress on momentum source modeling, this paper also features the first full-core pin resolved CFD simulation ever performed to the authors' knowledge. This represents a significant advancement in capability for the CFD of nuclear reactors, which will hopefully serve as an inspiration for further integrating high-fidelity numerical simulations in actual engineering designs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Measurement and analysis of GPU-accelerated applications with HPCToolkit

To address the challenge of performance analysis on the US DOE’s forthcoming exascale supercomputers, Rice University has been extending its HPCToolkit performance tools to support measurement and analysis of GPU-accelerated applications. To help developers understand the performance of accelerated applications as a whole, HPCToolkit’s measurement and analysis tools attribute metrics to calling contexts that span both CPUs and GPUs. To measure GPU-accelerated applications efficiently, HPCToolkit employs a novel wait-free data structure to coordinate monitoring and attribution of GPU performance. To help developers understand the performance of complex GPU code generated from high-level programming models, HPCToolkit constructs sophisticated approximations of call path profiles for GPU computations. To support fine-grained analysis and tuning, HPCToolkit uses PC sampling and instrumentation to measure and attribute GPU performance metrics to source lines, loops, and inlined code. To supplement fine-grained measurements, HPCToolkit can measure GPU kernel executions using hardware performance counters. To provide a view of how an execution evolves over time, HPCToolkit can collect, analyze, and visualize call path traces within and across nodes. Finally, on NVIDIA GPUs, HPCToolkit can derive and attribute a collection of useful performance metrics based on measurements using GPU PC samples. Here, we illustrate HPCToolkit’s new capabilities for analyzing GPU-accelerated applications with several codes developed as part of the Exascale Computing Project.

97 MATHEMATICS AND COMPUTING↗

Accelerated kinetic model for global macro stability studies of high-beta fusion reactors

The field reversed configuration (FRC), such as studied in the C-2W experiment at TAE Technologies, is an attractive candidate for realizing a nuclear fusion reactor. In an FRC, kinetic ion effects play the majority role in macroscopic stability, which allows global stability studies to make use of fluid-kinetic hybrid (also referred to as Ohm's law) models wherein ions are treated kinetically while electrons are treated as a fluid. The development and validation of such a hybrid particle-in-cell algorithm in the Exascale Computing Project code WarpX are reported here. Implementation of this model in the WarpX framework benefits from the numerical efficiency of WarpX as well as its scalability on large HPC systems and portability to different architectures. Performance benchmarks of the new algorithm for large, 3-dimensional, full device simulations from the Perlmutter supercomputer are presented. Results of a series of FRC simulations are discussed in which the impact of two-fluid effects on the tilt-mode growth rate was studied. It was observed that, in agreement with previous Hall-MHD studies, two-fluid effects have a stabilizing impact on the tilt mode.

Physics↗

Creating Continuous Integration Infrastructure for Software Development on U.S. Department of Energy High-Performance Computing Systems

The Exascale Computing Project (ECP) software deployment effort developed and advanced DevOps capabilities. One goal was to enable robust continuous integration (CI) workflows that span the protected high performance computing (HPC) environments found within many of the Department of Energy’s (DOE) national laboratories. This article highlights several challenges encountered with enabling automation, such as charging models for CI jobs, and meeting individualized security requirements that revolve around strongly associating running code with a human identity. Here, it also describes how the Jacamar CI tool evolved to meet latter requirements and became a key aspect of the solutions currently offered. Derived from this experience, we offer a conceptual framework for understanding current and future CI challenges at DOE facilities and offer suggestions for long-term solutions.

97 MATHEMATICS AND COMPUTING↗

Frontier: Exploring Exascale

As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.

Atchley, Scott {Leadership Computing}↗