Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “exascale computing project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Frontier: Exploring Exascale

As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.

Atchley, Scott {Leadership Computing}↗

AMR-Wind [SWR-20-85]

AMR-Wind is a massively parallel, block-structured adaptive-mesh, incompressible flow solver for wind turbine and wind farm simulations. The solver is built on top of the AMReX library. AMReX is developed at LBNL , NREL , and ANL as part of the Block-Structured AMR Co-Design Center in DOE's Exascale Computing Project. AMReX library provides the mesh data structures, mesh adaptivity, as well as the linear solvers used for solving the governing equations. The primary applications for AMR-Wind are: performing large-eddy simulations (LES) of atmospheric boundary layer (ABL) flows, simulating wind farm turbine-wake interactions using actuator disk or actuator line models for turbines, and as a background solver when coupled with a near-body solver with overset methodology to perform blade-resolved simulations of multiple wind turbines within a wind farm.

Ananthan, Shreyas↗

Reposcanner

SAND2023-05455O Reposcanner provides a highly modular, extensible framework for defining routines for mining data from software repositories and performing analyses on that data to yield valuable insights on team behaviors. Reposcanner features seamless support for different version control platforms like GitHub, Gitlab, and Bitbucket; smart parsing of URLs; intelligent credential management capabilities; and a comprehensive test suite. Reposcanner is connected to the Exascale Computing Project and is intended for research purposes.

Mundt, Miranda↗

The Exascale Framework for High Fidelity coupled Simulations (EFFIS): Enabling whole device modeling in fusion science

We present the Exascale Framework for High Fidelity coupled Simulations (EFFIS), a workflow and code coupling framework developed as part of the Whole Device Modeling Application (WDMApp) in the Exascale Computing Project. EFFIS consists of a library, command line utilities, and a collection of run-time daemons. Together, these software products enable users to easily compose and execute workflows that include: strong or weak coupling, in situ (or offline) analysis/visualization/monitoring, command-and-control actions, remote dashboard integration, and more. We describe WDMApp physics coupling cases and computer science requirements that motivate the design of the EFFIS framework. Furthermore, we explain the essential enabling technology that EFFIS leverages: ADIOS for performant data movement, PerfStubs/TAU for performance monitoring, and an advanced COUPLER for transforming coupling data from its native format to the representation needed by another application. Finally, we demonstrate EFFIS using coupled multi-simulation WDMApp workflows and exemplify how the framework supports the project’s needs. We show that EFFIS and its associated services for data movement, visualization, and performance collection does not introduce appreciable overhead to the WDMApp workflow and that the resource-dominant application’s idle time while waiting for data is minimal.

97 MATHEMATICS AND COMPUTING↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers part I: Computational models and workflow

Computational simulations have become central to the seismic analysis and design of major infrastructure over the past several decades. Most major structures are now “proof tested” virtually through representative simulations of earthquake-induced response. More recently, with the advancement of high-performance computing (HPC) platforms and the associated massively parallel computational ecosystems, simulation is beginning to play a role in increased understanding and prediction of ground motions for earthquake hazard assessments. However, the computational requirements for regional-scale geophysics-based ground motion simulations are extreme, which has restricted the frequency resolution of direct simulations and limited the ability to perform the large number of simulations required to numerically explore the problem parametric space. In this article, recent developments toward an integrated, multidisciplinary earth science-engineering computational framework for the regional-scale simulation of both ground motions and resulting structural response are described with a particular emphasis on advancing simulations to frequencies relevant to engineered systems. This multidisciplinary computational development is being carried out as part of the US Department of Energy (DOE) Exascale Computing Project with the goal of achieving a computational framework poised to exploit emerging DOE exaflop computer platforms scheduled for the 2022–2023 timeframe.

58 GEOSCIENCES↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

LLNL Response to the DOE ASCR RFI, "Stewardship of Software for Scientific and High-Performance Computing"

For decades, Lawrence Livermore National Laboratory (LLNL) has been engaged in significant research, development, and support for software to enable scientific computing and, particularly, the use of high performance computing (HPC) in the NNSA mission space. In particular, the move in the mid-1990’s to simulation as a leading component of stockpile stewardship through the ASCI and the successor ASC programs, as well as the need for reliable data acquisition and control software for the National Ignition Facility, have been important drivers in building expertise in production-quality software development at LLNL. LLNL has also been a leader in the DOE SciDAC FASTMath Institute and the DOE Exascale Computing Project (ECP), both of which have striven to make scientific computing software – in particular, the enabling technologies underpinning simulation capabilities – more widely adopted and sustainable. As such, we believe that our experience can inform the broader goal of software stewardship for scientific and high-performance computing. LLNL strongly supports the formation of a new DOE ASCR program element in software stewardship and sustainment. Historically, DOE ASCR has funded applied mathematics and computer science research that has led to the development of important new capabilities and algorithms that are expressed as artifacts in research software. Such frameworks, libraries, and tools have seldom been directly funded to address the important issues of code maintenance, documentation, robustness, and community building. Software engineering and support have typically been done on the side in support of the ASCR-driven research products. DOE funding priorities have been slow to recognize that good software engineering, the kind that ensures research investments have more adoption and longevity, requires significant resources. Based upon our experiences, we have prepared this response to highlight the concerns and issues we believe to be important as DOE ASCR considers its role in scientific software stewardship. We believe that role is important and will require a significant investment of new funding to legitimately support the technologies past and future DOE ASCR investments have and will produce to facilitate their uptake and adoption in the broader scientific computing community. Following a summary of our involvement in scientific software development, the remainder our response is organized around the nine topics specifically identified in the RFI.

97 MATHEMATICS AND COMPUTING↗

Nek5000/RS Performance on Advanced GPU Architectures

We demonstrate NekRS performance results on various advanced GPU architectures. NekRS is a GPU-accelerated version of Nek5000 that is targeting high performance on forthcoming exascale platforms. It is being developed in DOE’s Center of Efficient Exascale Discretizations (CEED), which is one of the co-design centers under the Exascale Computing Project (ECP). In this report, we consider Frontier, Crusher, Spock, Polaris, Perlmutter, ThetaGPU and Summit. Simulations are performed using ExaSMR’s 17x17 rod-bundle geometries with different problem sizes. The report focuses on strong scaling performance and analysis. Many of the results shown in this report are the outcome from participation in the ALCF GPU Hackathon on Polaris, which was held on 7/19/22, and 7/26–7/28/22. Members of the Nek5000/RS team included Misun Min, Yu-Hsiang Lan, Paul Fischer and Thilina Rathnayake. Mentors were Kris Rowe (ALCF) and Peng Wang (NVIDIA). Also presented in this report are results on Frontier, which were obtained in collaboration with John Holmen at OLCF.

97 MATHEMATICS AND COMPUTING↗

Five years of ForTrilinos ECP

The ForTrilinos subproject of the Exascale Computing Project (ECP) was initiated to bring the capabilities and scalability of the Trilinos numerical solver collection to Fortran scientific application codes. A novel Fortran extension to the Simplified Wrapper and Interface Generator (SWIG) tool, which automatically generates Fortran bindings from existing C/C++ library code, has been applied to key Trilinos solver libraries to generate the new ForTrilinos libraries. SWIG-Fortran has additionally been used to generate new Fortran compatibility layers for additional scientific libraries and applications. This report summarizes the products and impact of the ForTrilinos subproject.

97 MATHEMATICS AND COMPUTING↗

Advanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 Workshops

Over the past decade, fundamental changes in artificial intelligence (AI)—from foundational to applied—have delivered dramatic insights across a wide breadth of U.S. Department of Energy (DOE) mission space. AI is helping to augment and improve scientific and engineering workflows (e.g., for control, design, and dramatic performance gains through surrogate models) in national security, the Office of Science, and DOE’s applied energy programs. The progress and potential for AI in DOE science was captured in the 2020 “AI for Science” report from the DOE laboratory community in collaboration with academia and industry. Specific scientific areas ready to further leverage the power of AI ranged from the scale and performance of computational models to data analysis to creating new classes of observations using computer vision. Since that report, the scale and scope of scientific AI have accelerated, revealing new, emergent properties that yield insights that go beyond enabling opportunities to being potentially transformative in the way that scientific problems are posed and solved. Thus, under the guidance of both the Office of Science (SC) and the National Nuclear Security Administration (NNSA), the DOE national laboratories organized a series of workshops in 2022 to gather input on new and rapidly emerging opportunities and challenges of scientific AI. This 2023 report is a synthesis of those workshops. The scientific community believes AI can have a foundational impact on a broad range of DOE missions, including science, energy, and national security. Further, DOE has unique capabilities that enable the community to drive progress in scientific use of AI, building on long-standing DOE strengths and investments in computation, data, and communications infrastructure, spanning the Energy Sciences Network (ESnet), the Exascale Computing Project (ECP), and integrative programs such as the NNSA Office of Defense Programs Advanced Simulation and Computing (ASC) and the SC Scientific Discovery through Advanced Computing (SciDAC) programs.

97 MATHEMATICS AND COMPUTING↗

Enterprise Risks for Scientific Software in the Post Exascale Era

The Department of Energy’s Office of Science’s Advanced Scientific Computing Research (ASCR) Office hosts three high performance computing Facilities and one high performance networking Facility. ASCR’s Facilities Division and these four Facilities established a joint working group to address the most severe threats to the software capabilities that can prevent the Facilities from supporting a diverse array of mission critical scientific research carried out by thousands of researchers from national labs, academia, and industry. This report outlines the risk categories, threat matrix calculations, and number of risks in each of the categories along with the threat levels. Two categories—programming environment/tools and system software—account for 13 out of the 20 risks encountered in the risk register. A significant risk to ASCR’s software ecosystem also arises from the possible loss of personnel and expertise past the end of the exascale computing project. The working group recommends that the software enterprise risks are reevaluated regularly, and are considered in future decision making regarding system acquisition, knowledge sharing, partnership development, and strategic planning

97 MATHEMATICS AND COMPUTING↗