Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Aurora”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Crystal structure of the kinase domain of a receptor tyrosine kinase from a choanoflagellate, Monosiga brevicollis

Genomic analysis of the unicellular choanoflagellate, Monosiga brevicollis (MB), revealed the remarkable presence of cell signaling and adhesion protein domains that are characteristically associated with metazoans. Strikingly, receptor tyrosine kinases, one of the most critical elements of signal transduction and communication in metazoans, are present in choanoflagellates. We determined the crystal structure at 1.95 Å resolution of the kinase domain of the M . brevicollis receptor tyrosine kinase C8 (RTKC8, a member of the choanoflagellate receptor tyrosine kinase C family) bound to the kinase inhibitor staurospaurine. The chonanoflagellate kinase domain is closely related in sequence to mammalian tyrosine kinases (~ 40% sequence identity to the human Ephrin kinase domain EphA3) and, as expected, has the canonical protein kinase fold. The kinase is structurally most similar to human Ephrin (EphA5), even though the extracellular sensor domain is completely different from that of Ephrin. The RTKC8 kinase domain is in an active conformation, with two staurosporine molecules bound to the kinase, one at the active site and another at the peptide-substrate binding site. To our knowledge this is the first example of staurospaurine binding in the Aurora A activation segment (AAS). We also show that the RTKC8 kinase domain can phosphorylate tyrosine residues in peptides from its C-terminal tail segment which is presumably the mechanism by which it transmits the extracellular stimuli to alter cellular function.

59 BASIC BIOLOGICAL SCIENCES↗

Operation Fishbowl

Fishbowl’s primary objective was to generate data on electromagnetic pulses (EMP), the creation and behavior of auroras, and the impact of a high altitude nuclear burst on radio communications. Previous high altitude tests (Teak, Orange, and Yucca) were not instrumented to provide such data, which was deemed critical to countering any Soviet high altitude bursts. The first test of Fishbowl, Starfish Prime, exploded on July 9th with a yield of 1.4 megatons. Standing on Christmas Island, Austin McGuire witnessed the effects of the event.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Flexible Siting Criteria and Staff Minimization for Micro-Reactors

The economic potential of micro-reactors is vast and underestimated. Commonly-emphasized applications include niche markets such as remote communities, mines and military bases. However, micro-reactors could be used as flexible energy generators also for larger markets, such as mobile and containerized agriculture and manufacturing facilities, district heating, micro-grids for data centers, sea ports, airports and hospitals. The implication is that micro-reactors may have to be deployed also in non-remote locations. Successful implementation of micro-reactors needs a navigable and predictable licensing process, technology-appropriate siting restrictions, risk-informed emergency and safety requirements, and practical operating and maintenance requirements. The primary goal of this project was to develop siting criteria that are tailored to micro-reactors deployable in densely-populated areas, e.g., urban environments. To achieve that goal, we compared the characteristics of the MIT research reactor (MITR) with those of leading micro-reactor concepts (e.g., eVinci, USNC, Aurora), and evaluated whether and how the MITR design basis (e.g., inherent safety features, engineered safety systems, source term, emergency planning and emergency operating procedures) and associated regulations may be applicable to these new micro-reactors as well. What makes MITR a unique analogue in this context is its small power rating (6 MWt) and physical size, mode of operations (24/7 with a somewhat more commercial flavor than typical university reactors), and especially its urban location. Of course significant differences exist, such as mission (power production vs. research) and the reactor design itself. Leveraging the MITR experience, this project was able to generate criteria that will allow micro-reactors to realize their full economic potential as flexible heat and electricity generators for a diverse portfolio of applications in non-remote locations. As such, the outcome of this project might encourage investment in and use of micro-reactors. A second goal of the project was to conceptualize a model of operations for micro-reactors that would minimize the staffing requirements, and thus reduce the cost of electricity and heat generated by these systems. Here too our approach was to systematically review the MITR experience and requirements, as well as survey the innovations in autonomous control technologies and monitoring (e.g., advanced sensors, drones, robotics, AI) that would permit a dramatic reduction in staffing at future micro-reactor installations. The scope of work was expanded after the start date to include also an evaluation of micro-reactor security, using the so-called consequence-based analysis, and the development of a methodology to perform dynamic risk assessment for micro-reactors, using system theory and modeling and simulation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

2022 Operational Assessment Report - Argonne Leadership Computing Facility

This Operational Assessment Report describes how the Argonne Leadership Computing Facility (ALCF) met or exceeded every one of its goals for the calendar year (CY) 2022. In CY 2022, ALCF operated Theta, an Intel-based Cray XC40 system (11.7 petaflops) augmented with 24 NVIDIA DGX A100-based nodes (3.9 petaflops), that supports diverse workloads, integrating data analytics with artificial intelligence (AI) training and learning in a single platform; and Polaris, a 44-petaflop AMD and NVIDIA-based HPE Apollo 6500 Gen10+ system that provides a powerful new platform to prepare applications and workloads for Aurora, Argonne National Laboratory’s (Argonne’s) upcoming Intel-Hewlett Packard Enterprise (HPE) exascale supercomputer.

97 MATHEMATICS AND COMPUTING↗

Proactive Regulatory Approaches to Electrification and Load Growth: Workshop Report

On July 10 and 11, 2024, Pacific Northwest National Laboratory and RMI led a workshop in Aurora, Colorado, to explore novel and proactive approaches to electrification and load growth while minimizing risks and costs to customers. Over the next decade, a unique opportunity exists to invest strategically in the electricity system to enable electrification across the transportation, industrial, and building sectors and respond to data and technology-based load growth. However, current utility and regulatory planning practices are insufficient to identify and enable the right investments, and work must be done to reduce the risk and decisional uncertainty faced by utility regulatory commissions and utilities. Ensuring timely electrification investments may require new approaches to address risk, uncertainty, prudence, and cost recovery. Understanding the decision-making process and information needs of utilities and regulators is critical. New policies (or application of policies), financial tools, systems analysis, regulatory mechanisms, and enhanced process transparency may be required. The workshop's goal was to identify proactive regulatory approaches for electrification and load growth that minimize costs and risks to customers. Our intention was that the conversations and the resulting solutions and takeaways would be specific and tactical rather than general and theoretical and that together we would create actionable next steps for key actors in the system, including utilities, regulators, thought leaders, researchers, and the U.S. Department of Energy (DOE). This report is intended to provide workshop attendees with a record and summary of the discussion and proposals raised at the workshop and to provide interested entities who did not attend, such as other regulators, policymakers, utilities, and U.S. DOE offices, with an understanding of what was discussed and with ideas to explore in their organizations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Accelerating transients with NekRS: GPU overlapping domain implementation and multi-rate timestepping

The simulation of nuclear transients using Computational Fluid Dynamics (CFD) presents significant computational challenges due to the inherent complexity and the wide separation in temporal scales between various flow physical phenomena. These disparities lead to high computational costs, often making the simulation of transients impractical without advanced techniques. Consequently, multiple research initiatives are being pursued by the NEAMS thermal-hydraulic area, some driven by academic institutions and some by national laboratories. Overall, they are exploring novel methods to make transient simulations more feasible and efficient. This report delves into recent advancements within the CFD code NekRS, specifically those achieved in Fiscal Year 2024 under the CONNECT effort, aimed at improving the performance and feasibility of transient simulations. The first major advancement involves the porting of NekRS to Aurora, one of the Department of Energy’s (DOE) most powerful supercomputers. Additionally, the report discusses the implementation of an overlapping domain capability within NekRS. This novel GPU-accelerated capability allows different spatial regions of the domain to be solved independently, enhancing the code’s efficiency, particularly when running large-scale simulations in complex domains. The scalability of this approach is demonstrated, highlighting its potential to transform how transients are approached in CFD simulations. Lastly, the report focuses on how this overlapping domain capability specifically accelerates transient simulations through multi-rate timestepping. By decoupling different regions and facilitating faster computations, this method offers a promising pathway to making nuclear transient simulations more computationally feasible, addressing one of the critical bottlenecks in the field. Together, these advancements represent a significant leap forward in transient simulation technology, bringing closer the possibility of handling highly complex nuclear scenarios with greater efficiency and accuracy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗

2025 Advances in NekRS: Supporting improved performance for nuclear applications

This report presents several 2025 advancements in NekRS, a high-fidelity spectral element CFD code developed at Argonne National Laboratory to support the NEAMS thermal-hydraulics program. The forthcoming v25 release consolidates several of these advances, adding new features for portability across heterogeneous GPU architectures, real-time in situ visualization, improved turbulence modeling, and conjugate heat transfer coupling. Over the past year, NekRS has demonstrated strong scalability and performance on DOE’s leading exascale platforms, including Aurora and Frontier, confirming its readiness for some of the largest and most complex simulations attempted to date. These achievements provide a powerful new platform for high-fidelity data generation, which in turn supports the development and validation of advanced closure models critical for reactor safety and design. Significant algorithmic innovations have also been introduced. A new global runtime h-refinement capability simplifies workflows by reducing mesh preparation burdens and enabling coarse-to-fine restarts. Building on this, a novel multigrid strategy was implemented to accelerate pressure and transport solves at scale, addressing long-standing bottlenecks in exascale CFD. Together, these developments improve both the efficiency and accessibility of high-fidelity simulations for reactor-relevant problems. Collectively, these enhancements represent a major step forward in simulation technology, positioning NekRS as a cornerstone of NEAMS efforts to enable accurate, efficient, and scalable high-fidelity analysis of advanced nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Optimizing Metadata Exchange: Leveraging DAOS for ADIOS Metadata I/O

In HPC I/O middleware like the Adaptable I/O System (ADIOS) often mediates data transfers between applications. The metadata I/O generated by such systems often presents significant scaling and performance limitations. This work seeks improvement opportunities for metadata I/O by leveraging the DAOS storage systems, a recent storage system solution deployed on high-end systems such as the Aurora supercomputer. We investigate the tradeoffs and the design space for integrating I/O engines for the ADIOS middleware based on the different storage mechanisms supported by DAOS. We present a new DAOS-Array-ChunkSize-aligned engine which provides up to 2.3× improved performance than when using the existing DAOS-POSIX interface, without requiring any application modifications.

Venkatesh, Ranjan Sarpangala↗

Measurement report: Understanding the seasonal cycle of Southern Ocean aerosols

Abstract. The remoteness and extreme conditions of the Southern Ocean and Antarctic region have meant that observations in this region are rare, and typically restricted to summertime during research or resupply voyages. Observations of aerosols outside of the summer season are typically limited to long-term stations, such as Kennaook / Cape Grim (KCG; 40.7∘ S, 144.7∘ E), which is situated in the northern latitudes of the Southern Ocean, and Antarctic research stations, such as the Japanese operated Syowa (SYO; 69.0∘ S, 39.6∘ E). Measurements in the midlatitudes of the Southern Ocean are important, particularly in light of recent observations that highlighted the latitudinal gradient that exists across the region in summertime. Here we present 2 years (March 2016–March 2018) of observations from Macquarie Island (MQI; 54.5∘ S, 159.0∘ E) of aerosol (condensation nuclei larger than 10 nm, CN10) and cloud condensation nuclei (CCN at various supersaturations) concentrations. This important multi-year data set is characterised, and its features are compared with the long-term data sets from KCG and SYO together with those from recent, regionally relevant voyages. CN10 concentrations were the highest at KCG by a factor of ∼50 % across all non-winter seasons compared to the other two stations, which were similar (summer medians of 530, 426 and 468 cm−3 at KCG, MQI and SYO, respectively). In wintertime, seasonal minima at KCG and MQI were similar (142 and 152 cm−3, respectively), with SYO being distinctly lower (87 cm−3), likely the result of the reduction in sea spray aerosol generation due to the sea ice ocean cover around the site. CN10 seasonal maxima were observed at the stations at different times of year, with KCG and MQI exhibiting January maxima and SYO having a distinct February high. Comparison of CCN0.5 data between KCG and MQI showed similar overall trends with summertime maxima and wintertime minima; however, KCG exhibited slightly (∼10 %) higher concentrations in summer (medians of 158 and 145 cm−3, respectively), whereas KCG showed ∼40 % lower concentrations than MQI in winter (medians of 57 and 92 cm−3, respectively). Spatial and temporal trends in the data were analysed further by contrasting data to coincident observations that occurred aboard several voyages of the RSV Aurora Australis and the RV Investigator. Results from this study are important for validating and improving our models and highlight the heterogeneity of this pristine region and the need for further long-term observations that capture the seasonal cycles.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cloud phase and macrophysical properties over the Southern Ocean during the MARCUS field campaign

To investigate the cloud phase and macrophysical properties over the Southern Ocean (SO), the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Mobile Facility (AMF2) was installed on the Australian icebreaker research vessel (R/V) Aurora Australis during the Measurements of Aerosols, Radiation, and Clouds over the Southern Ocean (MARCUS) field campaign (41 to 69° S, 60 to 160° E) from October 2017 to March 2018. To examine cloud properties over the midlatitude and polar regions, the study domain is separated into the northern (NSO) and southern (SSO) parts of the SO, with a demarcation line of 60°S. The total cloud fractions (CFs) were 77.9 %, 67.6 %, and 90.3 % for the entire domain, NSO and SSO, respectively, indicating that higher CFs were observed in the polar region. Low-level clouds and deep convective clouds are the two most common cloud types over the SO.

54 ENVIRONMENTAL SCIENCES↗

The ocean model for E3SM global applications: Omega version 0.1.0 – a new high-performance computing code for exascale architectures

This paper introduces Omega, the Ocean Model for E3SM Global Applications. Omega is a new ocean model designed to run efficiently on high performance computing (HPC) platforms, including exascale heterogeneous architectures with accelerators, such as Graphics Processing Units (GPUs). Omega is written in C and uses the Kokkos performance portability library. These were chosen because they are well-supported and will help future-proof Omega for upcoming HPC architectures. Omega will eventually replace the Model for Prediction Across Scales-Ocean (MPAS-Ocean) in the US Department of Energy's (DOE's) Energy Exascale Earth System Model (E3SM). Omega runs on unstructured horizontal meshes with variable-resolution capability and implements the same horizontal discretization as MPAS-Ocean. This work documents the design and performance of Omega Version 0.1.0 (Omega-V0), which solves the shallow water equations with passive tracers and is the first step towards the full primitive equation ocean model. On Central Processing Units (CPUs), Omega-V0 is 1.4 times faster than MPAS-Ocean with the same configuration. Omega-V0 is more efficient on GPUs than CPUs on a per-watt basis – by a factor of 5.3 on Frontier and 3.6 on Aurora, two of the world's fastest exascale computers.

54 ENVIRONMENTAL SCIENCES↗

Pele: An Exascale-Ready Suite of Combustion Codes

High fidelity simulations of realistic combustion devices are extremely demanding computationally because of the requirements to capture complex fuel chemical decomposition, its intricate interactions with turbulent, often multiphase, flows, and the wide separation of space and time scales between the thin flame and the device boundaries. Software required to carry out such computations tends to be extremely complex, particularly when designed to exploit hardware accelerators, and can be difficult to port and maintain. We present Pele, a performance portable suite of tools for the simulation of combustion systems, including codes to evolve reactive multiphase configurations in the low Mach number and compressible flow regimes, along with a set of inter-compatible post processing and in situ analysis tools. The Pele suite of tools is built on top of the AMReX framework for block-structured adaptive mesh refinement, which provides efficient data structures and algorithms that enable the development of a wide variety of efficient mesh and particle based PDE integration schemes. A hierarchical MPI+X parallelism scheme supports CPU-only and accelerated architectures, where X can be OpenMP, CUDA, and HIP based approaches for intra-node computational work distribution. The algorithms and data structures underlying the Pele simulation and analysis tools are highly scalable and performant across a wide variety of high-performance computing platforms, including DOEs newest exascale-class machines, Frontier and Aurora. The simulation and analysis tools are fully documented and freely distributed as open source via GitHub. We present key algorithmic and software challenges, solution strategies, performance and resulting set of capabilities.

AMReX↗

Developing Multiphysics, Integrated, High-Fidelity, Massively Parallel Computational Capabilities for Fusion Applications Using MOOSE

As the need for fusion as a clean, sustainable, and abundant energy source grows internationally, so does the need for multiphysics, computational tools to model, study, and predict the complex interactions between plasma, materials, and engineering processes. These tools have a crucial role to play in solving scientific and engineering challenges and accelerating fusion energy deployment. To address these needs, modeling capabilities should enable massively parallel, multiphysics, fully integrated high-fidelity simulations of fusion systems. Additional attributes, such as being open source and modular while maintaining high software quality assurance standards will maximize impact by ensuring accessibility for all and wide acceptance, rapid expansion and development, as well as reliability, efficiency, and robustness. In this paper, we describe how the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, which has a track record of success in the fission space thanks to the attributes listed above, can be leveraged in the fusion energy field. We highlight key successes of the MOOSE application in the fission space and describe how MOOSE has been and is being applied to fusion applications in the United States---e.g., Tritium Migration Analysis Program, version 8 (TMAP8), MOOSE Fusion Module, Fusion ENergy Integrated multiphys-X (FENIX)---and the United Kingdom---e.g., AURORA, Achlys, Apollo. These efforts aim to establish a suite of tools that can be further extended to accelerate fusion energy deployment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗