Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “exascale computing project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Feasibility of full-core pin resolved CFD simulations of small modular reactor with momentum sources

Complex flow structure interactions and heat transfer processes take place in nuclear reactor cores. Given the extreme pressure/temperature and radioactive conditions inside the core, numerical simulations offer an attractive and sometimes more feasible approach to study the related flow and heat transfer phenomena in addition to the experiments. Under the Exascale Computing Project, the full-core simulation of a small modular reactor (SMR) has been pursued coupling Computational Fluid Dynamics (CFD) and neutronics. A key aspect of the modeling of SMR fuel assemblies is the presence of spacer grids and the mixing promoted by mixing vanes or the equivalent. A reduced order methodology is adopted based on momentum sources to mimic the mixing of the vanes. The momentum sources have been carefully calibrated with detailed Large Eddy Simulations (LES) of spacer grids performed with Nek5000. Modeling the spacer grid and mixing vanes (SGMV) effect without body-fitted computational grid avoids the excessive costs in resolving the local geometric details, and thus supports the simulation to be scaled up to the full core. Besides the progress on momentum source modeling, this paper also features the first full-core pin resolved CFD simulation ever performed to the authors' knowledge. This represents a significant advancement in capability for the CFD of nuclear reactors, which will hopefully serve as an inspiration for further integrating high-fidelity numerical simulations in actual engineering designs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Measurement and analysis of GPU-accelerated applications with HPCToolkit

To address the challenge of performance analysis on the US DOE’s forthcoming exascale supercomputers, Rice University has been extending its HPCToolkit performance tools to support measurement and analysis of GPU-accelerated applications. To help developers understand the performance of accelerated applications as a whole, HPCToolkit’s measurement and analysis tools attribute metrics to calling contexts that span both CPUs and GPUs. To measure GPU-accelerated applications efficiently, HPCToolkit employs a novel wait-free data structure to coordinate monitoring and attribution of GPU performance. To help developers understand the performance of complex GPU code generated from high-level programming models, HPCToolkit constructs sophisticated approximations of call path profiles for GPU computations. To support fine-grained analysis and tuning, HPCToolkit uses PC sampling and instrumentation to measure and attribute GPU performance metrics to source lines, loops, and inlined code. To supplement fine-grained measurements, HPCToolkit can measure GPU kernel executions using hardware performance counters. To provide a view of how an execution evolves over time, HPCToolkit can collect, analyze, and visualize call path traces within and across nodes. Finally, on NVIDIA GPUs, HPCToolkit can derive and attribute a collection of useful performance metrics based on measurements using GPU PC samples. Here, we illustrate HPCToolkit’s new capabilities for analyzing GPU-accelerated applications with several codes developed as part of the Exascale Computing Project.

97 MATHEMATICS AND COMPUTING↗

Accelerated kinetic model for global macro stability studies of high-beta fusion reactors

The field reversed configuration (FRC), such as studied in the C-2W experiment at TAE Technologies, is an attractive candidate for realizing a nuclear fusion reactor. In an FRC, kinetic ion effects play the majority role in macroscopic stability, which allows global stability studies to make use of fluid-kinetic hybrid (also referred to as Ohm's law) models wherein ions are treated kinetically while electrons are treated as a fluid. The development and validation of such a hybrid particle-in-cell algorithm in the Exascale Computing Project code WarpX are reported here. Implementation of this model in the WarpX framework benefits from the numerical efficiency of WarpX as well as its scalability on large HPC systems and portability to different architectures. Performance benchmarks of the new algorithm for large, 3-dimensional, full device simulations from the Perlmutter supercomputer are presented. Results of a series of FRC simulations are discussed in which the impact of two-fluid effects on the tilt-mode growth rate was studied. It was observed that, in agreement with previous Hall-MHD studies, two-fluid effects have a stabilizing impact on the tilt mode.

Physics↗

Creating Continuous Integration Infrastructure for Software Development on U.S. Department of Energy High-Performance Computing Systems

The Exascale Computing Project (ECP) software deployment effort developed and advanced DevOps capabilities. One goal was to enable robust continuous integration (CI) workflows that span the protected high performance computing (HPC) environments found within many of the Department of Energy’s (DOE) national laboratories. This article highlights several challenges encountered with enabling automation, such as charging models for CI jobs, and meeting individualized security requirements that revolve around strongly associating running code with a human identity. Here, it also describes how the Jacamar CI tool evolved to meet latter requirements and became a key aspect of the solutions currently offered. Derived from this experience, we offer a conceptual framework for understanding current and future CI challenges at DOE facilities and offer suggestions for long-term solutions.

97 MATHEMATICS AND COMPUTING↗

Frontier: Exploring Exascale

As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.

Atchley, Scott {Leadership Computing}↗

miniGAN: a proxy application for generative adversarial networks

miniGAN is a python-based machine learning proxy application for generative adversarial networks, developed through the Exascale Computing Project's (ECP) ExaLearn project. It will be included in the main ECP proxy application and the machine learning proxy application suite. It is a proxy for ECP cosmological(CosmoFlow, ExaGAN) and wind energy(ExaWind) applications. miniGAN will be distributed to ECP hardware vendors as part of hardware codesign. miniGAN uses the Numpy/PyTorch/TensorFlow/Keras/Horovod frameworks and libraries. It also relies on the Kokkos and Kokkos-Kernels packages developed here at Sandia Labs. SAND2020-2038 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Ellis, John↗

AMR-Wind [SWR-20-85]

AMR-Wind is a massively parallel, block-structured adaptive-mesh, incompressible flow solver for wind turbine and wind farm simulations. The solver is built on top of the AMReX library. AMReX is developed at LBNL , NREL , and ANL as part of the Block-Structured AMR Co-Design Center in DOE's Exascale Computing Project. AMReX library provides the mesh data structures, mesh adaptivity, as well as the linear solvers used for solving the governing equations. The primary applications for AMR-Wind are: performing large-eddy simulations (LES) of atmospheric boundary layer (ABL) flows, simulating wind farm turbine-wake interactions using actuator disk or actuator line models for turbines, and as a background solver when coupled with a near-body solver with overset methodology to perform blade-resolved simulations of multiple wind turbines within a wind farm.

Ananthan, Shreyas↗

Reposcanner

SAND2023-05455O Reposcanner provides a highly modular, extensible framework for defining routines for mining data from software repositories and performing analyses on that data to yield valuable insights on team behaviors. Reposcanner features seamless support for different version control platforms like GitHub, Gitlab, and Bitbucket; smart parsing of URLs; intelligent credential management capabilities; and a comprehensive test suite. Reposcanner is connected to the Exascale Computing Project and is intended for research purposes.

Mundt, Miranda↗

The Exascale Framework for High Fidelity coupled Simulations (EFFIS): Enabling whole device modeling in fusion science

We present the Exascale Framework for High Fidelity coupled Simulations (EFFIS), a workflow and code coupling framework developed as part of the Whole Device Modeling Application (WDMApp) in the Exascale Computing Project. EFFIS consists of a library, command line utilities, and a collection of run-time daemons. Together, these software products enable users to easily compose and execute workflows that include: strong or weak coupling, in situ (or offline) analysis/visualization/monitoring, command-and-control actions, remote dashboard integration, and more. We describe WDMApp physics coupling cases and computer science requirements that motivate the design of the EFFIS framework. Furthermore, we explain the essential enabling technology that EFFIS leverages: ADIOS for performant data movement, PerfStubs/TAU for performance monitoring, and an advanced COUPLER for transforming coupling data from its native format to the representation needed by another application. Finally, we demonstrate EFFIS using coupled multi-simulation WDMApp workflows and exemplify how the framework supports the project’s needs. We show that EFFIS and its associated services for data movement, visualization, and performance collection does not introduce appreciable overhead to the WDMApp workflow and that the resource-dominant application’s idle time while waiting for data is minimal.

97 MATHEMATICS AND COMPUTING↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers part I: Computational models and workflow

Computational simulations have become central to the seismic analysis and design of major infrastructure over the past several decades. Most major structures are now “proof tested” virtually through representative simulations of earthquake-induced response. More recently, with the advancement of high-performance computing (HPC) platforms and the associated massively parallel computational ecosystems, simulation is beginning to play a role in increased understanding and prediction of ground motions for earthquake hazard assessments. However, the computational requirements for regional-scale geophysics-based ground motion simulations are extreme, which has restricted the frequency resolution of direct simulations and limited the ability to perform the large number of simulations required to numerically explore the problem parametric space. In this article, recent developments toward an integrated, multidisciplinary earth science-engineering computational framework for the regional-scale simulation of both ground motions and resulting structural response are described with a particular emphasis on advancing simulations to frequencies relevant to engineered systems. This multidisciplinary computational development is being carried out as part of the US Department of Energy (DOE) Exascale Computing Project with the goal of achieving a computational framework poised to exploit emerging DOE exaflop computer platforms scheduled for the 2022–2023 timeframe.

58 GEOSCIENCES↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

Advancing Scientific Productivity through Better Scientific Software: Developer Productivity and Software Sustainability Report

The Exascale Computing Project (ECP) provides a unique opportunity to advance computational science and engineering (CSE) through an accelerated growth phase in extreme-scale computing. Central to the project is the development of next-generation applications and software technologies that can exploit emerging architectures for optimal performance and provide high-fidelity, multiphysics, multiscale capabilities. However, disruptive changes in computer architectures and the complexities of tackling new frontiers in extreme-scale modeling, simulation, and analysis present daunting challenges to the productivity of software developers and the sustainability of software artifacts. Members of the CSE community - especially at extreme scales but more broadly at all scales of computing - face an urgent need to improve developer productivity, positively impacting product quality, development time, and staffing resources, and software sustainability, reducing the cost of maintaining, sustaining, and evolving software capabilities.

97 MATHEMATICS AND COMPUTING↗