Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “exascale applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications

As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences exponential growth. Adaptive Mesh Refinement (AMR) has emerged as an effective solution to address these two challenges. Concurrently, error-bounded lossy compression is recognized as one of the most efficient approaches to tackle the latter issue. Despite their respective advantages, few attempts have been made to investigate how AMR and error-bounded lossy compression can function together. To this end, this study presents a novel in-situ lossy compression framework that employs the HDF5 filter to improve both I/O costs and boost compression quality for AMR applications. We implement our solution into the AMReX framework and evaluate on two real-world AMR applications, Nyx and WarpX, on the Summit supercomputer. Experiments with 512 cores demonstrate that AMRIC improves the compression ratio by 81X and the I/O performance by 39X over AMReX's original compression solution.

Wang, Daoce↗

Advanced Turbulence Models for Large-Scale Atmospheric Boundary Layer Flows

We present high-fidelity large-eddy-simulation (LES) modeling approaches for the turbulent atmospheric boundary layer (ABL) flows. Wind energy is a prime example of an application driven by ABL. Generation of electrical energy from farms of wind turbines at night in the stable ABL is a particularly interesting situation. In this report, we consider the well-known GEWEX (Global Energy and Water Cycle Experiment) Atmospheric Boundary Layer Study (GABLS) stably stratified benchmark LES case. We use a high-order spectral element code Nek5000/RS, which is supported under the DOE's Exascale Computing Project (ECP) Center for Efficient Exascale Discretizations (CEED) project, targeting application simulations on various acceleration-device based exascale computing platforms. In our earlier ANL report, we demonstrated our newly developed subgrid-scale (SGS) models based on high-pass filter (HPF), mean-field eddy viscosity (MFEV), and Smagorinsky (SMG) with no-slip and traction boundary conditions, provided with low-order statistics, convergence and turbulent structure analysis. In this report, we extend the range of our SGS modeling approaches in the context of the mean-field eddy viscosity (MFEV), to include the solution of an SGS turbulent kinetic energy equation (TKE). We demonstrate the model fidelity of Nek5000/RS in comparison to that of AMR-Wind, a block-structured second-order finite-volume code with adaptive-mesh-refinement capabilities, with which we studied scaling performance for both codes in comparison on DOE's leadership computing platforms.

17 WIND ENERGY↗

A Co-design Framework for Online Data Analysis and Reduction

Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation-analysis-reduction pipelines coupled in-memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application- and system-specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. Furthermore, this toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next-generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers.

97 MATHEMATICS AND COMPUTING↗

Toward Energy-Efficient HPC: Insights from Power Profiling a Cloud-Resolving Earth System Model

Power is a fundamental constraint as supercomputing advances to exascale. Efficient operation within strict power budgets requires application-aware power management based on a detailed understanding of application-level power behavior. This work analyzes the Energy Exascale Earth System Model (E3SM) atmosphere component, SCREAM, on Perlmutter (NERSC) and Frontier (OLCF). We characterize power variation across inputs, concurrency levels, and power caps, evaluate the energy impact of code optimizations, and attribute energy within the code using a newly developed GPU energy model. Results show that SCREAM’s peak power remains stable during its core execution phase and decreases gradually as concurrency increases. Power capping experiments reveal a performance–energy "sweet spot". On Perlmutter, limiting GPU power to 50% of thermal design power (TDP) achieves up to 15% energy savings with a 7% performance penalty. On Frontier, a 40% TDP cap yields up to 10% energy savings with less than 10% performance loss. Code optimizations reduce SCREAM energy by shortening run time without increasing power. Modeling reveals a critical insight: data movement accounts for approximately 70% of SCREAM’s GPU energy. This fundamentally shifts the optimization focus from FLOPS to data transfer reduction for this class of applications, offering the most impactful strategy for improving energy efficiency. This work establishes a foundation for practical, application-aware power management at exascale.

Zhao, Zhengji [Lawrence Berkeley National Laborato↗

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

HPC I/O innovations in the exascale era

As high performance computing architecture evolves to deliver ever-increasing performance, the middleware tools also need to adapt in order for applications to better use these higher-performance features. Here, the Adaptable Input Output System (ADIOS), which provides scalable IO performance for exascale HPC applications is one such middleware. During the Exascale Computing Project (ECP), key portions of the ADIOS environment were adapted to respond to ongoing developments in exascale computing and the stresses and opportunities inherent in those changes. This paper examines those changes and where appropriate compares them to pre-exascale implementations.

ADIOS↗

Evolving HPC and Application Design Toward a Coupled Data Assimilation System at NASA Suitable for Emerging Exascale Platforms

The prediction capabilities of global models have continuously evolved from the traditional medium-range global weather prediction application to span scales in support of hourly prediction of convective scale storms to seasonal Earth system prediction. This evolution has increased the demands on the system infrastructure design and workflow to achieve the required performance on modern high-performance computing (HPC) platforms. The planned evolution of the Goddard Earth Observing System (GEOS) modeling and assimilation system will stress the capabilities of conventional HPC overwhelming the available compute cycles at the NASA Center for Climate Simulation (NCCS) at the NASA Goddard Space Flight Center in the coming 5-10 years. This has led to the re-design of key elements of the assimilation and modeling systems to achieve significant gains in performance on anticipated Exacale platforms. The transition of the assimilation system to the Joint Effort for Data assimilation Integration (JEDI) framework has positioned GEOS to exploit new efficient algorithms for data assimilation (DA) in a fully-coupled Earth system context. The suitability of the GEOS model to leverage a domain specific language (DSL) approach and artificial intelligence (AI) is being explored to accelerate computational performance and data exchange efficiency of the coupled Earth system model. The storage and processing of large data volumes produced by these advance systems is being redesigned with a data-centric cloud-based approach. We will highlight the recent efforts in these areas and emphasize the demand for further development and re-design to achieve the science objectives in support of NASA's Earth system modeling and assimilation missions.

Putman, Bill↗

Co-design Center for Exascale Machine Learning Technologies (ExaLearn)

We report rapid growth in data, computational methods, and computing power is driving a remarkable revolution in what variously is termed machine learning (ML), statistical learning, computational learning, and artificial intelligence. In addition to highly visible successes in machine-based natural language translation, playing the game Go, and self-driving cars, these new technologies also have profound implications for computational and experimental science and engineering, as well as for the exascale computing systems that the Department of Energy (DOE) is developing to support those disciplines. Not only do these learning technologies open up exciting opportunities for scientific discovery on exascale systems, they also appear poised to have important implications for the design and use of exascale computers themselves, including high-performance computing (HPC) for ML and ML for HPC. The overarching goal of the ExaLearn co-design project is to provide exascale ML software for use by Exascale Computing Project (ECP) applications, other ECP co-design centers, and DOE experimental facilities and leadership class computing facilities.

97 MATHEMATICS AND COMPUTING↗

Then and Now: Improving Software Portability, Productivity, and 100× Performance

The US Exascale Computing Project (ECP) has succeeded in preparing applications to run efficiently on the first reported Exascale supercomputers in the world. To achieve this, it modernized the whole leadership software stack, from libraries to simulation codes. In this article, we contrast selected leadership software before and after ECP. We discuss how sustainable research software development for leadership computing can embrace the conversation with the hardware vendors, the leadership computing facilities, the software community, and the domain scientists who are the application developers and integrators of software products. We elaborate on how software needs to take portability as a central design principle and to benefit from interdependent teams; we also demonstrate how moving to programming languages with high momentum, like modern C++, can help improve the sustainability, interoperability, and performance of research software. Finally, we showcase how cross-institutional efforts can enable algorithm advances that are beyond incremental performance optimization.

97 MATHEMATICS AND COMPUTING↗

An Adaptive-Mesh-Refinement Based Computational Tool for Simulating Catalysis at Mesoscale

In this work, we present a computational tool for mesoscale applications using open-source exascale- computing compatible adaptive-mesh-refinement (AMR) library, AMReX [2]. AMReX is software library that enables development of application solvers with block-structured Cartesian AMR. Our tool has capabilities to include realistic geometry representation, chemical species transport, reactions and thermodynamics that are critical for capturing mesoscale physics. A significant achievement is the ability of our solver to automatically import electron microscopy data in the form of a stereolithography (STL) or pixelated file format (mrc, tiff) without undergoing the tedious task of unstructured mesh generation. This feature allows for rapid simulation of catalyst particles with complex morphologies using an immersed-boundary formulation. The use of AMR allows for higher resolutions at catalyst surface interfaces, which in turn provides an accurate description of surface reactions and transport. Our solver uses a hybrid distributed and shared memory parallelism (OpenMP/GPU-based) with which strong scaling up to 10,000 processors for realistic catalyst particle simulations have been demonstrated.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

TChem-atm v1.0

SAND2024-11300O TChem-atm is a software library that was developed to solve complex kinetic models for atmospheric chemistry applications. TChem-atm interface employs a hierarchical parallelism design to exploit the massive parallelism available from modern computing platforms. It also supports gas atmospheric chemistry applications, e.g., the energy exascale earth system model. TChem can be used as a box model or coupled with a climate model to compute the time evolution of gas tracer species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

CSRI Summer Proceedings 2021

The Computer Science Research Institute (CSRI) brings university faculty and students to Sandia National Laboratories for focused collaborative research on Department of Energy (DOE) computer and computational science problems. The institute provides an opportunity for university researches to learn about problems in computer and computational science at DOE laboratories, and help transfer results of their research to programs at the labs. Some specific CSRI research interest areas are: scalable solvers, optimization, algebraic preconditioners, graph-based, discrete, and combinatorial algorithms, uncertainty estimation, validation and verification methods, mesh generation, dynamic load-balancing, virus and other malicious-code defense, visualization, scalable cluster computers, beyond Moore’s Law computing, exascale computing tools and application design, reduced order and multiscale modeling, parallel input/output, and theoretical computer science. The CSRI Summer Program is organized by CSRI and includes a weekly seminar series and the publication of a summer proceedings.

97 MATHEMATICS AND COMPUTING↗

CSRI Summer Proceedings 2021

The Computer Science Research Institute (CSRI) brings university faculty and students to Sandia National Laboratories for focused collaborative research on Department of Energy (DOE) computer and computational science problems. The institute provides an opportunity for university researches to learn about problems in computer and computational science at DOE laboratories, and help transfer results of their research to programs at the labs. Some specific CSRI research interest areas are: scalable solvers, optimization, algebraic preconditioners, graph-based, discrete, and combinatorial algorithms, uncertainty estimation, validation and verification methods, mesh generation, dynamic load-balancing, virus and other malicious-code defense, visualization, scalable cluster computers, beyond Moore’s Law computing, exascale computing tools and application design, reduced order and multiscale modeling, parallel input/output, and theoretical computer science. The CSRI Summer Program is organized by CSRI and includes a weekly seminar series and the publication of a summer proceedings.

97 MATHEMATICS AND COMPUTING↗

Enabling Combustion Science Simulations for Future Exascale Machines

Reacting flow simulations for combustion applications require extensive computing capabilities. Leveraging the AMReX library, the Pele suite of combustion simulation tools targets the largest supercomputers available and future exascale machines. We introduce PeleC, the compressible solver in the Pele suite, and detail its capabilities, including complex geometry representation, chemistry integration, and discretization. We present a comparison of development efforts using both OpenACC and AMReX's C++ performance portability framework for execution on multiple GPU architectures. We discuss relevant details that have allowed PeleC to achieve high performance and scalability. PeleC's performance characteristics are measured through relevant simulations on multiple supercomputers. The success of PeleC's design for exascale is exhibited through demonstration of a 160 billion cell simulation and weak scaling onto 100\% of Summit, an NVIDIA-based GPU supercomputer at Oak Ridge National Laboratory. Our results provide confidence that PeleC will enable future combustion science simulations with unprecedented fidelity.

combustion↗

2020 Exascale Computing Project Annual Meeting (Executive Summary Report)

The Exascale Computing Project (ECP) delivers specific applications, software products, and outcomes on DOE computing facilities. Integration across these elements for specific hardware technologies for exascale system instantiations is fundamental to ECP success. The outcome of the ECP is the delivery of a capable exascale computing ecosystem to provide breakthrough solutions addressing our most critical challenges in scientific discovery, energy assurance, economic competitiveness, and national security. This outcome is not a matter of ensuring more powerful computing systems. The ECP is designed to create more valuable and rapid insights from a wide variety of applications (“capable”), which requires a much higher level of inherent efficacy in all methods, software tools, and ECP-enabled computing technologies to be acquired by DOE laboratories (“ecosystem”). The ECP annual meeting provides a unique opportunity for the core technical expertise in the United States focused on achieving this next plateau of computational science and computing performance to engage in direct discussions on project execution. Face-to-face gatherings in technical communities like this are common and needed for the exchange of scientific ideas and technical performance. The ECP annual meeting stands apart from other technical conferences and meetings in the computing community as it is uniquely and solely focused on the execution of the ECP and the integration of technical activities leading to the creation of the exascale computing ecosystem for the future. The direct interaction of key critical technical staff, who are leaders in their respective fields, and the resulting give-and-take between software, applications, and hardware and the technical co-design therein, is unique and essential to the effective execution of the ECP. The first annual meeting was held in Knoxville, Tennessee, January 31 – February 2, 2017 and brought together, for the first time, a diverse collection of researchers from 16 DOE national laboratories as well as university computer and computational science researchers to discuss shared problems and joint solutions for the development of a capable exascale computing ecosystem. These interactions resulted in focused technical plans and an energized community centered on advances for ECP. The second annual meeting was held in Knoxville, Tennessee, February 5–9, 2018. It included 643 individual thought leaders and performers in application development, software research and deployment, and hardware research and integrators, all of whom are part of the multifaceted, billion dollar HPC community. This meeting provided a platform to discuss and disseminate numerous examples where researchers with common goals and synergistic solutions came together for the first time to deliver tangible results. Additionally, at the 2018 meeting, ECP researchers had the opportunity to digest all US HPC vendor R&D product roadmaps pointing to exascale – not only to learn how their research can play a role, but, more importantly, to influence those roadmaps to ensure successful delivery on DOE applications that will contribute to (if not solve) problems of national interest in national security, science, energy, and health, as well as growing security threats. The third annual meeting was held in Houston, Texas, January 14–17, 2019. With a 19% increase in the number of registrations (768 people), and the change in location, the third annual meeting was considered the most impactful of the three at the time. The new website provided a better platform for the dissemination of the content, the new venue as a meeting hotel instead of a conference center facilitated interactions and discussions after event hours, and the addition of an award-winning mobile event conference app (Whova) transformed dramatically the attendee experience at the event. This fourth annual meeting was held in Houston, Texas, February 3-7, 2020. This meeting had an increase in the number of attendees for a total of 824 people registered (782 attendees) and included numerous enhancements based on feedback and lessons learned from previous meetings, some of which are listed here: Improved quality of the sessions, their material and the whole program.; Had more industry participation and addition of external collaborators from overseas.; Published the full agenda earlier to better accommodate attendance and travel plans based on schedule.; Centralized all sessions in one venue.; Provided additional hotels and room blocks for the attendees.; Improved communication with the audience (links, material, directions, notifications, etc.) to go paperless.; Enhanced side meeting scheduling, management and user experience.; Made available additional space and tables for impromptu meetings and side discussions.; Improved IT and A/V solutions for speakers. In addition, our final survey captured the following points as opportunities for improvement in future meetings: consider a different meeting location that is more pedestrian friendly, reduce talks during working meals to allow more collaboration and informal time, adapt the agenda to acknowledge attendees from different timezones, consider recording some of the tutorials and/or sessions to share broadly with the HPC community, have a larger poster room, provide additional power strips, and improve the WiFi.

97 MATHEMATICS AND COMPUTING↗

CALORIE: A Constraint Language and Optimizing Runtime for Exascale Power Management (Final Report)

This final technical report summarizes the key accomplishments on the CALORIE project, a DOE Early Career award received by PI Henry Hoffmann at the University of Chicago. CALORIE’s main goal was to create principled methodologies, tools, and practices to help scientists and high-performance computing (HPC) operators maximize the performance and insights obtained from scientific computing applications in the face of exascale power constraints. The project accomplished these goals by completing the following objectives: • Designing a language for describing application goals (including constraints and objectives) and system capabilities. The application goals will include things like which simulations or simulations plus in situ analysis will be run together, what the power constraints are, and what requirements there are for in situ analysis (for example, a desired frame rate for visualization). The system capabilities include all components that can be adjusted to tradeoff power and performance. • Designing a runtime system that takes specified goals and capabilities and dynamically determines what capabilities to use to meet the goals. This runtime adapts to changes in goals, application behavior, or available capabilities to automatically maintain the goals despite unexpected disturbances. • Developing a foundational understanding of how power constraints affect the problem of scheduling applications in large-scale systems. This report provides an overview of the accomplishments related to each of these key objectives.

97 MATHEMATICS AND COMPUTING↗

Exascale Multiphysics Nuclear Reactor Simulations for Advanced Designs

ENRICO is a coupled application developed under the U.S. Department of Energy's Exascale Computing Project (ECP) targeting the modeling of advanced nuclear reactors. It couples radiation transport with heat and fluid simulation, including the high-fidelity, highresolution Monte-Carlo code Shift and the Computational fluid dynamics code NekRS. NekRS is a highly-performant open-source code for simulation of incompressible and low-Mach fluid flow, heat transfer, and combustion with a particular focus on turbulent flows in complex domains. It is based on rapidly convergent high-order spectral element discretizations that feature minimal numerical dissipation and dispersion. State-of-the-art multilevel preconditioners, efficient high-order time-splitting methods, and runtime-adaptive communication strategies are built on a fast OCCA-based kernel library, libParanumal, to provide scalability and portability across the spectrum of current and future high-performance computing platforms. On Frontier, Nek5000/RS has recently achieved an unprecedented milestone in breaching over 1 billion spectral elements and 350 billion degrees of freedom. Shift has demonstrated the capability to transport upwards of 1 billion particles per second in full core nuclear reactor simulations featuring complete temperature-dependent, continuous-energy physics on Frontier. Shift achieved a weak-scaling efficiency of 97.8% on 8192 nodes of Frontier and calculated 6 reactions in 214,896 fuel pin regions below 1% statistical error yielding first-of-a-kind resolution for a Monte Carlo transport application.

Hamilton, Steven P.↗