Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “exascale computing project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Exascale Was Not Inevitable; Neither Is What Comes Next

The long, steady advance of supercomputing capabilities historically makes it tempting to assume that similar future improvements are inevitable. In fact, the successful path to exascale systems was far from certain, requiring a shift to accelerator-based designs, an associated end-to-end rethinking of traditional computational approaches, and a fundamental change in how scientists collaborate and interact. Here in this article, we share lessons from the U.S. Department of Energy Exascale Computing Project along with perspectives on how to design and implement scientific applications in the exascale and post-exascale eras.

97 MATHEMATICS AND COMPUTING↗

Then and Now: Improving Software Portability, Productivity, and 100× Performance

The US Exascale Computing Project (ECP) has succeeded in preparing applications to run efficiently on the first reported Exascale supercomputers in the world. To achieve this, it modernized the whole leadership software stack, from libraries to simulation codes. In this article, we contrast selected leadership software before and after ECP. We discuss how sustainable research software development for leadership computing can embrace the conversation with the hardware vendors, the leadership computing facilities, the software community, and the domain scientists who are the application developers and integrators of software products. We elaborate on how software needs to take portability as a central design principle and to benefit from interdependent teams; we also demonstrate how moving to programming languages with high momentum, like modern C++, can help improve the sustainability, interoperability, and performance of research software. Finally, we showcase how cross-institutional efforts can enable algorithm advances that are beyond incremental performance optimization.

97 MATHEMATICS AND COMPUTING↗

ECP libraries and tools: An overview

The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.

97 MATHEMATICS AND COMPUTING↗

ECP Software Technology Capability Assessment Report V3.0

The Exascale Computing Project (ECP) Software Technology (ST) focus area is responsible for (1) developing critical software capabilities that will enable the successful execution of ECP applications and (2) providing key components of a productive and sustainable exascale computing ecosystem that will position the US Department of Energy (DOE) and the broader high-performance computing (HPC) community with a firm foundation for future extreme-scale computing capabilities. This ECP ST Capability Assessment Report (CAR) provides an overview and assessment of current ECP ST capabilities and activities, giving stakeholders and the broader HPC community information that can be used to assess ECP ST progress and plan their own efforts accordingly. ECP ST leaders commit to updating this document on regular basis (every 6–12 months). Highlights from this version of the report are presented here. This version of the CAR contains the following updates relative to the previous revision: (1) This report highlights the progress with the Extreme-scale Scientific Software Stack (E4S) efforts. In particular, this report discusses how E4S continues to gain traction as a first-class entity in the HPC ecosystem, enabling new conversations with users, facilities, vendors, other US agencies, and international partners. (2) The several-page summaries of each ECP Level 4 project were updated to reflect recent progress and next steps (Section 4). Of particular note are the experiences of our teams on early-access systems for Frontier. (3) The E4S is described further. E4S is now updated via quarterly releases. E4S is the primary integration and delivery vehicle for ECP ST capabilities (Section 2.1.1). (4) The ECP ST software development kit (SDK) effort further refined its groupings (Section 2.1.2). The ECP ST focus area represents the key bridge between exascale systems and the scientists developing applications that will run on those platforms. ECP ST efforts contribute to approximately 70 software products (Section 2.1.3) in six technical areas (Table 1). Since publishing the previous revision of the CAR, the team has continued to evolve the product dictionary of official product names, which enables more rigorous mapping of ECP ST deliverables to stakeholders (Section 2.1.4).

97 MATHEMATICS AND COMPUTING↗

2.3.3.01- xSDK-Batched Final Report for Subcontract Partner KIT

The xSDK batched effort as focus effort of the ECP project xSDK focused on the development, deployment, and dissemination of batched functionality in the US Exascale Computing Project. Over a duration of almost three years, KIT as subcontract to LLNL provided technology development, functionality deployment, integration support, and consulting on batched iterative solvers and batched preconditioners. This final report accumulates the quarterly progress reports and most significant contributions.

97 MATHEMATICS AND COMPUTING↗

MAGMA: Enabling exascale performance with accelerated BLAS and LAPACK for diverse GPU architectures

MAGMA (Matrix Algebra for GPU and Multicore Architectures) is a pivotal open-source library in the landscape of GPU-enabled dense and sparse linear algebra computations. With a repertoire of approximately 750 numerical routines across four precisions, MAGMA is deeply ingrained in the DOE software stack, playing a crucial role in high-performance computing. Notable projects such as ExaConstit, HiOP, MARBL, and STRUMPACK, among others, directly harness the capabilities of MAGMA. In addition, the MAGMA development team has been acknowledged multiple times for contributing to the vendors’ numerical software stacks. Looking back over the time of the Exascale Computing Project (ECP), we highlight how MAGMA has adapted to recent changes in modern HPC systems, especially the growing gap between CPU and GPU compute capabilities, as well as the introduction of low precision arithmetic in modern GPUs. We also describe MAGMA’s direct impact on several ECP projects. Maintaining portable performance across NVIDIA and AMD GPUs, and with current efforts toward supporting Intel GPUs, MAGMA ensures its adaptability and relevance in the ever-evolving landscape of GPU architectures.

97 MATHEMATICS AND COMPUTING↗

Transforming Science Through Software: Improving While Delivering 100×

The U.S. Department of Energy (DOE) Exascale Computing Project (ECP) funded the development of new (and the transformation of important existing) applications, libraries, and tools that realized improvement in performance and capabilities of often 100 times or more on emerging exascale computers. This exceptional gain inspired the title of this special issue: Transforming Science through Software: Improving while delivering 100X. The term 100X refers to advancing capabilities in modeling, simulation, and analysis by a factor of 100 or more using some combination of new algorithms, optimization techniques, software libraries, and programming models, coupled with the next generation of hardware for high-performance computing (HPC). The papers in this issue share experiences with the practice and science of scientific software development, with an emphasis on developing a coherent, portable, and sustainable HPC software ecosystem for next-generation computational science. Finally, we hope to foster expanded community efforts related to the fundamental role of sustainable scientific software ecosystems in advancing the computing sciences.

97 MATHEMATICS AND COMPUTING↗

ECP Software Technology Capability Assessment Report

The Exascale Computing Project (ECP) Software Technology (ST) Focus Area is responsible for developing critical software capabilities that will enable successful execution of ECP applications, and for providing key components of a productive and sustainable Exascale computing ecosystem that will position the US Department of Energy (DOE) and the broader high performance (HPC) community with a firm foundation for future extreme-scale computing capabilities. This ECP ST Capability Assessment Report (CAR) provides an overview and assessment of current ECP ST capabilities and activities, giving stakeholders and the broader HPC community information that can be used to assess ECP ST progress and plan their own efforts accordingly. ECP ST leaders commit to updating this document on regular basis (every six to 12 months).

97 MATHEMATICS AND COMPUTING↗

2.3.3.01- xSDK-Multiprecision Final Report for Subcontract Partner KIT

The xSDK Multiprecision effort as focus effort of the ECP project xSDK focused on the development, deployment, and dissemination of mixed precision functionality in the US Exascale Computing Project. Over a duration of four years, KIT as subcontract to LLNL provided technology development, functionality deployment, integration support, and consulting on mixed precision functionality. The subcontract was modified several times over the project duration. All deliverables of the subcontract have been completed. This final report accumulates the quarterly progress reports and most significant contributions.

97 MATHEMATICS AND COMPUTING↗

xSDK Multiprecision (Final Report)

The xSDK Multiprecision effort as focus effort of the ECP project xSDK focused on the development, deployment, and dissemination of mixed precision functionality in the US Exascale Computing Project. Over a duration of four years, KIT as subcontract to LLNL provided technology development, functionality deployment, integration support, and consulting on mixed precision functionality. The subcontract was modified several times over the project duration. All deliverables of the subcontract have been completed. This final report accumulates the quarterly progress reports and most significant contributions.

97 MATHEMATICS AND COMPUTING↗

The Multiphysics on Advanced Platforms Project

In 2015, the Lawrence Livermore National Laboratory started development of next-generation multiphysics simulation capabilities for the National Nuclear Security Administration under the Advanced Technologies Development and Mitigation (ATDM) element of the Advanced Simulation and Computing program in collaboration with the Exascale Computing Project (ECP). A key driver for this effort across the NNSA tri-lab was the emergence of advanced high performance computing (HPC) architectures based on heterogeneous compute capabilities, including GPU based systems, as part of the national drive toward exascale computing platforms at multiple Department of Energy (DOE) facilities. Developing a multiphysics code capable of meeting the various simulation needs of the NNSA as defined by the current generation of integrated codes (or ICs), initially developed as part of the Accelerated Strategic Computing Initiative (ASCI) program beginning in 1996, and able to scale to the current 100 petaflop class pre-exascale systems, as well the forthcoming exaflop class computers, is a daunting challenge. To accomplish this ambitious goal, LLNL has embraced two key themes: use of high-order numerical methods and a modular approach to code development. The LLNL next generation effort is organized under the Multi-Physics on Advanced Platforms Project (MAPP). A foundational component of MAPP is the Axom computer science (CS) toolkit which provides infrastructure for the development of modular, performance portable, multi-physics application codes. MARBL is a next-generation application code built on the Axom base to address the modeling needs of the high energy density physics (HEDP) community for simulating high-explosive, magnetic or laser driven experiments such as inertial confinement fusion (ICF), pulsed-power magneto-hydrodynamics (MHD), equation of state (EOS) and material strength studies as part of the NNSA’s stockpile stewardship program (SSP).

97 MATHEMATICS AND COMPUTING↗

Modeling of a chain of three plasma accelerator stages with the WarpX electromagnetic PIC code on GPUs

The fully electromagnetic particle-in-cell code WarpX is being developed by a team of the U.S. DOE Exascale Computing Project (with additional non-U.S. collaborators on part of the code) to enable the modeling of chains of tens to hundreds of plasma accelerator stages on exascale supercomputers, for future collider designs. The code is combining the latest algorithmic advances (e.g., Lorentz boosted frame and pseudo-spectral Maxwell solvers) with mesh refinement and runs on the latest computer processing unit and graphical processing unit (GPU) architectures. In this paper, we summarize the strategy that was adopted to port WarpX to GPUs, report on the weak parallel scaling of the pseudo-spectral electromagnetic solver, and then present solutions for decreasing the time spent in data exchanges from guard regions between subdomains. In Sec. IV, we demonstrate the simulations of a chain of three consecutive multi-GeV laser-driven plasma accelerator stages.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advances in ArborX to support exascale applications

ArborX is a performance portable geometric search library developed as part of the Exascale Computing Project (ECP). In this paper, we explore a collaboration between ArborX and a cosmological simulation code HACC. Large cosmological simulations on exascale platforms encounter a bottleneck due to the in-situ analysis requirements of halo finding, a problem of identifying dense clusters of dark matter (halos). This problem is solved by using a density-based DBSCAN clustering algorithm. With each MPI rank handling hundreds of millions of particles, it is imperative for the DBSCAN implementation to be efficient. In addition, the requirement to support exascale supercomputers from different vendors necessitates performance portability of the algorithm. We describe how this challenge problem guided ArborX development, and enhanced the performance and the scope of the library. We explore the improvements in the basic algorithms for the underlying search index to improve the performance, and describe several implementations of DBSCAN in ArborX. Further, we report the history of the changes in ArborX and their effect on the time to solve a representative benchmark problem, as well as demonstrate the real world impact on production end-to-end cosmology simulations.

97 MATHEMATICS AND COMPUTING↗

E3SM: Improved Climate Prediction with Exascale Capability

The Energy Exascale Earth System Model (E3SM) project is an ongoing, state-of-the-science earth system modeling, simulation, and prediction effort that optimizes Department of Energy (DOE) computing resources to meet the science needs of the nation and the agency’s mission objectives. Climate simulation has become a proven tool for identifying and quantifying the impacts of climate change, but even greater accuracy is required at all levels to improve forecast precision. Understanding the impact of climate change on global and regional water cycles is one of the highest priorities and most difficult challenges in climate change prediction. As part of a subproject of DOE’s Exascale Computing Project, a multidisciplinary team including geophysical and computational scientists developed a multiscale modeling framework (MMF) to refine cloud representation in E3SM climate simulation on GPU accelerated supercomputers, making higher resolution, more computationally efficient predictions possible.

54 ENVIRONMENTAL SCIENCES↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

A massively parallel time-domain coupled electrodynamics–micromagnetics solver

We present a high-performance coupled electrodynamics–micromagnetics solver for full physical modeling of signals in microelectronic circuitry. The overall strategy couples a finite-difference time-domain approach for Maxwell’s equations to a magnetization model described by the Landau–Lifshitz–Gilbert equation. The algorithm is implemented in the Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. Furthermore, the code leverages ongoing developments of the Exascale Application Code, WarpX, which is primarily being developed for plasma wakefield accelerator modeling. Our temporal coupling scheme provides second-order accuracy in space and time by combining the integration steps for the magnetic field and magnetization into an iterative sub-step that includes a trapezoidal temporal discretization for the magnetization. The performance of the algorithm is demonstrated by the excellent scaling results on NERSC multicore and GPU systems, with a significant (59×) speedup on the GPU using a node-by-node comparison. We demonstrate the utility of our code by performing simulations of an electromagnetic waveguide and a magnetically tunable filter.

97 MATHEMATICS AND COMPUTING↗

AMReX: Block-structured adaptive mesh refinement for multiphysics applications

Block-structured adaptive mesh refinement (AMR) provides the basis for the temporal and spatial discretization strategy for a number of Exascale Computing Project applications in the areas of accelerator design, additive manufacturing, astrophysics, combustion, cosmology, multiphase flow, and wind plant modeling. AMReX is a software framework that provides a unified infrastructure with the functionality needed for these and other AMR applications to be able to effectively and efficiently utilize machines from laptops to exascale architectures. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving accurate descriptions of different physical processes in complex multiphysics algorithms. AMReX supports algorithms that solve systems of partial differential equations in simple or complex geometries and those that use particles and/or particle–mesh operations to represent component physical processes. In this article, we will discuss the core elements of the AMReX framework such as data containers and iterators as well as several specialized operations to meet the needs of the application projects. In addition, we will highlight the strategy that the AMReX team is pursuing to achieve highly performant code across a range of accelerator-based architectures for a variety of different applications.

Zhang, Weiqun↗

Differentiable Multiphysics Codes: A Breakthrough Technology for Simulation and Computing

This document summarizes the findings of a strategic planning exercise commissioned by the Weapons Simulation and Computing, Computational Physics (WSC/CP) program at the Lawrence Livermore National Laboratory (LLNL) in FY24. During the year, the committee met with multiple stakeholder communities to gather input, opinions, suggestions and concerns which have been incorporated throughout this document. The key findings from this exercise are summarized: • The development of multiphysics modelling and simulation (mod/sim) codes and software technologies, their deployment on exascale compute platforms, and their broad adoption across the NNSA is a major success of the Advanced Simulation and Computing (ASC) program and the Exascale Computing Project (ECP). Sustained investment in these core technologies is essential. • Today’s state of the art involves running ensembles of O(100K) simulations to perform uncertainty quantification (UQ) and design studies using multiple statistical methods such as Bayesian optimization to understand sensitivities of our models and explore parameterized design spaces. Even with exascale computing, we are practically limited to O(10) parameters in these studies since the number of simulations required to sample the space scales exponentially with the number of design parameters. • The data from these simulation ensembles is increasingly being used to train machine learned (ML) surrogates (or reduced order models, ROMs) which can then be used for optimization or real time design exploration. However, the trained surrogates are still limited in the number of parameters they can represent due to the sampling limitations previously noted. • Augmenting our suite of integrated multiphysics simulation codes, both current and emerging, with the ability to compute gradients (solution derivatives) of arbitrary simulation outputs with respect to (some or all) simulation inputs would be a breakthrough technology, opening the door to a new era of efficient and automated inverse design based on verified and validated mod/sim capabilities. • This capability, which we refer to as differentiable multiphysics codes (DMCs), would revolutionize both UQ and optimization studies by breaking the curse of dimensionality that presently limits our “gradient-free” ensemble based computing approach. A similar breakthrough occurred in the AI/ML community once the ability to compute gradients of arbitrary loss functions using back-propagation became commonplace. Gradient information from the multiphysics codes can also be used to dramatically improve the efficiency and scale of training of ML/ROM surrogates for rapid assessments. • Achieving this in our suite of codes will be a grand challenge, similar to the amount of effort that was required to transition from CPU to GPU computing. It will require buy-in from the entire WSC/CP program and beyond, including all integrated codes, physics and engineering models, third-party library dependencies and performance portability abstractions. It will also require investment in research and development of numerical methods for computing adjoints of coupled physics across multiple adaptively refined moving meshes and of stochastic (Monte Carlo) and mesh free (SPH) methods. • New software and numerical techniques, largely pioneered by the AI/ML community, make this feasible. Chief among these is automatic differentiation (AD), the ability to employ AD at point-wise locations in a physics calculation (instead of traditional black-box approaches) and the ability to perform “back-propagation in time” (or reverse mode AD) for non-linear partial differential equations (PDEs). Fundamentally, the conclusion of this strategic planning exercise is that the time is right to undertake a large scale effort in WSC, centered on the existing integrated codes, to continue the natural evolution of mod/sim in the age of AI/ML. Instead of attempting to replace mod/sim with purely data driven AI/ML models, we believe the key to success is to integrate AI/ML by building on top of the decades of hard-won knowledge and the verified/validated multiphysics modelling capability that is the hallmark of the ASC program.

97 MATHEMATICS AND COMPUTING↗