Navigating the Road to Successfully Manage a Large-Scale Research and Development Project: The Exascale Computing Project (ECP) Experience
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The recent arrival of the Frontier Supercomputer at Oak Ridge National Laboratory officially marked the dawn of the exascale computing era. Its successful deployment coincided with the culmination of the U.S. Department of Energy Exascale Computing Project (ECP), an ambitious, complex, and risky research and development effort that integrated contributions from a broad and diverse subset of the high-performance computing community. The success of ECP will ultimately be judged by the scientific and engineering advances that it enabled. In conclusion, this Special Issue is focused on showcasing early successes in the use of exascale resources to enable breakthroughs in key areas of science in engineering.
Multiphysics coupling presents a significant challenge in terms of both computational accuracy and performance. Achieving high performance on coupled simulations can be particularly challenging in a high-performance computing context. The US Department of Energy Exascale Computing Project has the mission to prepare mission-relevant applications for the delivery of the exascale computers starting in 2023. Many of these applications require multiphysics coupling, and the implementations must be performant on exascale hardware. In this special issue we feature six articles performing advanced multiphysics coupling that span the computational science domains in the Exascale Computing Project.
Performance portability is a critical issue for the Exascale Computing Project (ECP) because of nontrivial architectural differences between machines available today and those expected at exascale. Many ECP project teams are working toward performance portability, and would expect to benefit from sharing lessons learned, identifying gaps, and discovering opportunities for partnerships. To facilitate this communication, the IDEAS-ECP project partnered with the three focus areas of ECP (application development, software technology, and hardware and integration), and Department of Energy computing facilities, to lead a series of panel discussions on performance portability. The panels were organized around broadly common themes of algorithmic and data locality challenges. In this article, we describe the panel series, its objectives, and perspectives from the various areas of the project. Finally, we also discuss use cases that are distinctive, as well as conclusions drawn from the collective experience of the participants.
The US Department of Energy Office of Science and the National Nuclear Security Administration initiated the Exascale Computing Project (ECP) in 2016 to prepare mission-relevant applications and scientific software for the delivery of the exascale computers starting in 2023. The ECP currently supports 24 efforts directed at specific applications and six supporting co-design projects. These 24 application projects contain 62 application codes that are implemented in three high-level languages—C, C++, and Fortran—and use 22 combinations of graphical processing unit programming models. The most common implementation language is C++, which is used in 53 different application codes. The most common programming models across ECP applications are CUDA and Kokkos, which are employed in 15 and 14 applications, respectively. This article provides a survey of the programming languages and models used in the ECP applications codebase that will be used to achieve performance on the future exascale hardware platforms.
Collaboration and team science are emerging areas of interest in software production. Historically, multi-institutional research collaborations are difficult to initiate and maintain, negatively impacting communication, negotiation, and dialogue between industry, government, and academic researchers. The Exascale Computing Project (ECP), a massive, multi-team, high-stakes initiative, facilitated broader research collaboration under a shared funding structure and extended timeline to support scientific discovery. Here, we conducted interviews with ECP teams, representing a variety of domain specialties, research institutions, and programming backgrounds. Using thematic analysis, we assessed how ECP’s structure created an environment of increased trust among projects and how software shared between teams facilitated sustained collaboration. We found that the expectation of future collaboration, i.e., the shadow of the future, greatly enhanced trust among teams and the quality of scientific software produced. Based on our findings within ECP projects, we connect to the existing literature on trust in software engineering and share recommendations for sustainable multi-institutional collaboration and shared best software practices.
AMReX is a software framework for the development of block-structured mesh applications with adaptive mesh refinement (AMR). AMReX was initially developed and supported by the AMReX Co-Design Center as part of the U.S. DOE Exascale Computing Project (ECP), and is continuing to grow post-ECP. In addition to adding new functionality and performance improvements to the core AMReX framework, we have also developed a Python binding, pyAMReX, that provides a bridge between AMReX-based application codes and the data science ecosystem. pyAMReX provides zero-copy application GPU data access for AI/ML, in situ analysis and application coupling, and enables rapid, massively parallel prototyping. In this paper we review the overall functionality of AMReX and pyAMReX, focusing on new developments, new functionality, and optimizations of key operations. We also summarize capabilities of ECP projects that used AMReX and provide an overview of new, non-ECP applications.
Large-scale simulations require efficient computation across the entire computing hierarchy. A challenge of the Exascale Computing Project (ECP) was to reconcile highly heterogeneous hardware with the myriad of applications that were required to run on these supercomputers. Mathematical software forms the backbone of almost all scientific applications, providing efficient abstractions and operations that are crucial to harness the performance of computing systems. Ginkgo is one such mathematical software library, nurtured by ECP, providing high-performance, user-friendly, and performance portable interfaces for applications in ECP and beyond. In this paper, we elaborate on Ginkgo’s philosophy of high-performance software that is sustainable, reproducible, and easy to use. We showcase the wide feature set of solvers and preconditioners available in Ginkgo and the central concepts involved in their design. We elaborate on four different ECP software integrations: MFEM, PeleLM + SUNDIALS, XGC, and ExaSGD that use Ginkgo to accelerate their science runs. Performance studies of different problems from these applications highlight the effectiveness of Ginkgo and the benefits incurred by these ECP applications.
Content is mostly from an external Exascale Project worked on prior to joining Sandia.
The use of multiple types of precision in mathematical software has the potential to increase its performance on new heterogeneous architectures. The xSDK project focuses both on the investigation and development of multiprecision algorithms as well as their inclusion into xSDK member libraries. This report summarizes current efforts on including and/or using mixed precision capabilities in the math libraries Ginkgo, heFFTe, hypre, MAGMA, PETSc/TAO, SLATE, SuperLU, and Trilinos, including KokkosKernels. It contains both numerical results from libraries that already provide mixed precision capabilities, as well as descriptions of the strategies to incorporate multiprecision into established libraries.
This Exascale Computing Project (ECP) milestone report summarizes the status of all 30 ECP Applications Development (AD) subprojects at the end of FY20. In October and November of 2020, a comprehensive assessment of AD projects was conducted by the ECP leadership. Reviews occurred virtually between October 27, 2020 and November 12, 2020. The review committee—consisting of the AD lead, deputy, and L3—was tasked with evaluating each subproject’s progress in porting their codes to early exascale architectures considered precursors to the planned exascale machines. This includes characterizing which modules have been ported to multi-accelerator nodes, initial performance analyses, the status of software integration, and a current vision of successes, obstacles, and next steps. As such, this report contains not only an accurate snapshot of each subproject’s current status but also represents an unprecedentedly broad account of experiences in porting large scientific applications to next-generation high-performance computing architectures.
We present high-fidelity large-eddy-simulation (LES) modeling approaches for the turbulent atmospheric boundary layer (ABL) flows. Wind energy is a prime example of an application driven by ABL. Generation of electrical energy from farms of wind turbines at night in the stable ABL is a particularly interesting situation. In this report, we consider the well-known GEWEX (Global Energy and Water Cycle Experiment) Atmospheric Boundary Layer Study (GABLS) stably stratified benchmark LES case. We use a high-order spectral element code Nek5000/RS, which is supported under the DOE's Exascale Computing Project (ECP) Center for Efficient Exascale Discretizations (CEED) project, targeting application simulations on various acceleration-device based exascale computing platforms. In our earlier ANL report, we demonstrated our newly developed subgrid-scale (SGS) models based on high-pass filter (HPF), mean-field eddy viscosity (MFEV), and Smagorinsky (SMG) with no-slip and traction boundary conditions, provided with low-order statistics, convergence and turbulent structure analysis. In this report, we extend the range of our SGS modeling approaches in the context of the mean-field eddy viscosity (MFEV), to include the solution of an SGS turbulent kinetic energy equation (TKE). We demonstrate the model fidelity of Nek5000/RS in comparison to that of AMR-Wind, a block-structured second-order finite-volume code with adaptive-mesh-refinement capabilities, with which we studied scaling performance for both codes in comparison on DOE's leadership computing platforms.
Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. Department of Energy’s Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities—building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software. Members of the Interoperable Design of Extreme-scale Application Software project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This article discusses how these synergistic activities are advancing scientific discovery—mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.
Simulations of core-collapse supernovae, and other astrophysical phenomena, are quintessential extreme-scale computing challenges. For core-collapse supernova simulations to be carried out by the ExaStar project under the Exascale Computing Project umbrella, a robust, efficient, and state-of-the-art magnetohydrodynamics solver is a critical requirement. In Flash-X, the primary software instrument for ExaStar, a new magnetohydrodynamics solver has been designed and implemented from the ground up to achieve accuracy and efficiency for simulations of complex astrophysical flows. This new solver, dubbed Spark, uses high-order spatial reconstruction, Runge-Kutta time integration, and an efficient cell-centered approach to satisfying the divergence-free condition for the magnetic fields. Spark was written to be optimized for data locality in cache hierarchy of CPUs. Since data locality optimizations for cache hierarchy are not directly compatible with those of accelerators, we have taken the approach of using program synthesis to avoid massive amounts of code replication that would be necessary if we were to maintain two different versions of the solver. Our program synthesis relies on a simple key-dictionary approach, implemented in python, that enables us to assemble the version of the solver suitable for the target hardware from code fragments identified by specific keys. In this work, we describe the data locality optimizations of the solver for CPUs and accelerators and the program synthesis tools that enable this portability. We also detail the parallel performance of Spark for both CPUs and accelerators.
This poster presents the status of the Pele Combustion project, an applications project in the Exascale Computing Project. It summarizes the goal of the projects, developments added to the Pele simulation capabilities over the last year, some example performance figures that target ultimate application on the Frontier supercomputer, and show a number of published and in-progress simulation results, by us and our collaborators.
The Open MPI for Exascale (OMPI-X) project was one of two in the Exascale Computing Project (ECP) focused on advancing the MPI ecosystem. The OMPI-X team worked with other MPI Forum members to champion several important features for inclusion in the MPI 4.0, 4.1, and upcoming 5.0 MPI standard versions, in support of the needs of exascale applications and systems. The team also worked with the larger Open MPI community to bring implementations of these new features and other enhancements into Open MPI, one of the leading open-source implementations of the MPI interface. Here, this paper describes the motivation for the work of the OMPI-X project in the context of exascale computing needs, the nature of the resulting new capabilities in the MPI standard, and how they were implemented in the Open MPI library. Features include improved support for “MPI + X” programming models through partitioned communications and support for user-level threading, sessions, fault tolerance through the user-level fault mitigation (ULFM) and Reinit models, and other features. We also discuss enhancements to Open MPI providing improved performance and scalability for existing features, such as collective operations, one-sided operations, support for the Slingshot-11 interconnect of the initial exascale systems, and how the OMPI-X team worked to improve quality assurance for the Open MPI library, particularly on platforms of interest to the Department of Energy community.
PeleC is an Exascale Computing Project application for simulating compressible combustion in complex geometries. It has been built on top of the popular AMReX library. In the beginning of the Exascale Computing Project, PeleC was focused on KNL. It uses a mixture of C++, C, and kernels written in Fortran to obtain performance by focusing on vectorization. Recently we have taken two approaches in deciding PeleC's future for obtaining performance on exascale GPU machines. In the first programming model, we decorated the Fortran kernels with OpenACC directives. This expedited our ability to run at large scales on Summit's GPUs, where we achieved a significant speedup over the CPUs on Summit. The second programming model involved rewriting the Fortran kernels in C++ and using AMReX's Kokkos-like lambda abstractions for running on the GPU. This resulted in similar speedups on Summit's GPUs over merely utilizing the CPUs. Both approaches involved AMReX's management of memory transfers between the device and host. In this work, we compare and contrast the benefits and pitfalls to both programming approaches regarding performance, performance portability, and productivity. We also discuss advantages we have found in taking the time to modernize our code and why have chosen a specific pathway to prepare our code for the future DOE exascale machines.
The applications being developed within the U.S. Exascale Computing Project (ECP) to run on imminent Exascale computers will generate scientific results with unprecedented fidelity and record turn-around time. Many of these codes are based on particle-mesh methods and use advanced algorithms, especially dynamic load-balancing and mesh-refinement, to achieve high performance on Exascale machines. Yet, as such algorithms improve parallel application efficiency, they raise new challenges for I/O logic due to their irregular and dynamic data distributions. Thus, while the enormous data rates of Exascale simulations already challenge existing file system write strategies, the need for efficient read and processing of generated data introduces additional constraints on the data layout strategies that can be used when writing data to secondary storage. We review these I/O challenges and introduce two online data layout reorganization approaches for achieving good tradeoffs between read and write performance. We demonstrate the benefits of using these two approaches for the ECP particle-in-cell simulation WarpX, which serves as a motif for a large class of important Exascale applications. Here, we show that by understanding application I/O patterns and carefully designing data layouts we can increase read performance by more than 80 percent.