Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high-performance simulations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (5-Lab Demo) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, PV, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactor (SMR), control centers, and gas turbines, across five DOE national laboratories-NLR, INL, NETL, LBNL, and SNL. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance U.S. Department of Energy's (DOE) network, and controlled via a centralized energy controller hosted at NLR's ARIES facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility. SuperLab 2.0 (5-Lab Demo) showcased a major advancement in federated national laboratory collaboration, enabling real-time, cross-laboratory experimentation to coordinate geographically dispersed distributed energy resources (DERs) using various communication protocols and networks. SuperLab 2.0 (5-Lab Demo) built on previous demonstrations conducted between NLR-PNNL and NLR-INL connecting diverse assets including distant protection devices, a SMR simulator, and a high temperature electrolyzer (HTE). Previous demos were based on a single connection between two labs with minimal coordination challenges. The 5-Lab demo with a centralized controller, distributed testbeds across different geographical locations, and use of protocols-based communication represents a scenario closer to real-world grid operations that coordinate resources across a region to meet system needs. This experiment studied how local DER controllers interact with a centralized energy controller during normal and abnormal events to maintain reliability. The SuperLab team across the five labs implemented a notional power system model equivalent of transmission and distribution lines, represented by the data networks interconnecting the labs. Each lab continuously exchanged local parameters (such as P and Q) from its Hardware-In-Loop (CHIL) and Power Hardware-In-Loop (PHIL) assets through centralized energy controller at NLR, enabling real-time interaction and coordination across sites. By leveraging ESnet as the communication backbone, the team successfully operated the distributed assets as a unified power system, with each bus represented by a different laboratory. This setup mirrors how assets interact in real-world power systems across dispersed locations with various protocols and latencies. At each lab site, assets were operated using their own local controllers which were coordinated through an overarching operation and control layer of centralized energy controller, equivalent to how an energy management system (EMS) orchestrates assets across a regional or national grid. SuperLab's federated connectivity utilized a Digital Real-Time Simulators (DRTS)-type gateway to connect Controller Hardware-In-Loop (CHIL) and PHIL assets between labs. To enable this federated connection through ESnet, a deterministic network was established where latency variations were consistent. This consistency allowed the development of digital filters for the power system assets across CHIL and PHIL interfaces to avoid unstable and unreliable grid conditions. This report provides an overview of the cross-laboratory configuration and offers insights into interconnecting geographically distributed research assets to test them as if they were co-located. This experiment represents a step toward linking nine DOE national laboratories, enabling nation-wide simulations that can address utility-driven challenges with grid resilience, flexibility, and modernization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Large Polaron Screening Delays Mobility Roll-Off in Halide Perovskites under High Injection

The interplay between carrier density regimes and the various scattering processes is fundamental for understanding charge transport efficiency in high-performance semiconductors. Combining ultrafast time-resolved terahertz spectroscopy (TRTS) and transient microwave photoconductivity (TRMC), we measure carrier mobilities over 7 orders of magnitude (10 13 –10 20 carriers cm –3 ), reaching a regime where carrier–carrier scattering dominates and mobility begins to roll off. We find the Caughey–Thomas model can be effectively employed for transport modeling in 3D metal halide perovskites (MHPs), and the mobility roll-off near ∼10 19 cm –3 can be tuned by MHP composition. Carrier lifetime in MHPs starts decreasing at much lower carrier densities due to interband multi-carrier recombination, while intraband carrier mobility remains unchanged as carrier–carrier scattering is screened by the presence of polarons. The validated Caughey–Thomas model provides a practical input for device simulations, prevents overestimation of diffusion length under high injection, and suggests that MHPs are excellent candidates for unipolar device applications under high carrier densities.

14 SOLAR ENERGY↗

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

Development of a Test-Bed for Testing and Refining EarthEn’s Supercritical CO 2 Based Energy Storage System

EarthEn’s energy storage concept leverages supercritical carbon dioxide (sCO 2 ) as a working fluid and relies on compact, high-performance components operating at elevated pressures and temperatures. To accelerate component development and reduce technical risk prior to larger-scale demonstrations, Oak Ridge National Laboratory (ORNL) developed a 100 kW-scale sCO 2 test-bed under a Cooperative Research and Development Agreement with EarthEn (CRADA NO. NFE-24-10050). The objective of the work was to design and construct a flexible experimental facility capable of reproducing key thermodynamic state points and heat-transfer conditions relevant to EarthEn’s thermal energy storage (TES) cycle, with particular emphasis on enabling development and evaluation of next-generation heat exchangers and TES concepts. The test-bed consists of a closed-loop sCO 2 circulation system housed within an open-topped enclosure. In its as-installed configuration, dense-phase sCO 2 is recirculated through a printed circuit recuperator, an electrically heated section, a throttling device used to simulate turbine expansion, and a water-cooled printed circuit heat exchanger that rejects heat to the building chilled-water system before returning to the pump. The pump is driven by a variable frequency drive, enabling controlled adjustment of flow and operating point. A comprehensive instrumentation suite was integrated to support both safe operation and high-quality data collection. Installed sensors include Coriolis flow meters for sCO 2 flow rate and density, resistance temperature detectors and thermocouples distributed throughout the loop (including the heated section and key heat exchanger ports), and pressure transducers for absolute and differential pressure measurements. The facility was designed to support high-pressure (19 MPa nominal) and high-temperature (575°C nominal) operation with credited overpressure protection provided by a rupture disk. Nominal operating conditions were selected to support 100 kW-class testing while maintaining flexibility for non-heated and heated shakedown, control development, and future integration of advanced TES test sections. In parallel with facility development, a system-level thermal-hydraulic model was created using Modelica-based tools to support component sizing, anticipate performance over targeted test conditions, and establish a framework for future model calibration against experimental data. At the conclusion of the project performance period, the facility was in final assembly, and the pressure boundary was nearly completed. However, several practical challenges associated with high-pressure/high-temperature systems and specialized component procurement impacted schedule and prevented initial pump-driven operation and full commissioning within the available resources. This report documents the as-built design, operating capabilities, and instrumentation, and it summarizes key lessons learned related to heater fabrication and testing, first-of-a-kind assembly factors, specialty flange supply constraints, and fill pump corrective actions. Finally, it outlines a phased plan for future commissioning and experimental campaigns, including control and instrumentation shakedown, heater characterization, model calibration, and testing at state points representative of EarthEn’s TES cycle.

25 ENERGY STORAGE↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

Metal additively manufactured wavy fin cold-plate architecture for improved thermal-hydraulic performance

Rapid growth in artificial intelligence and data center workloads demands high-performance liquid cooling to manage increasing chip power. This study presents two metal-additive-manufactured cold plates with sinusoidal fins, constant-amplitude wavy fins and linearly variable-amplitude wavy fins and compares them against metal-additive-manufactured straight fins using experiments conducted at 1 kW heat dissipation as well as high-fidelity 3D conjugate computational fluid dynamic simulations. The cold plates were printed in AlSi10Mg material and underwent design using a Python-automated workflow prior to manufacture and testing. The experiments show that wavy fins reduce the normalized thermal resistance by 35 to 45 % at water flow rates from 1 to 4 LPM. At a fixed 20 kPa pressure drop, the variable-waviness design lowered peak surface temperature by 9 °C and thermal resistance by 51 %, while edge-channel maldistribution in the constant wavy fin design limited gains. A thermal resistance breakdown revealed that 55–63 % of the total thermal resistance in wavy designs comes from base heat conduction, 27–33 % from fin heat conduction, and 9–13 % from fin heat convection, indicating the need to address conduction bottlenecks. Parametric sweeps identify a 3 mm fin pitch as optimal, and that horizontal inlet/outlet manifolds further reduce pressure drop by 30–60 % and thermal resistance by 9–16 % relative to vertical inlet-outlet manifolds. The results yield comprehensive guidelines for fin geometry, manifold alignment, material selection and additive-manufacturing constraints to realize high-performance liquid-cooled cold plates for power-dense electronics.

3d printing↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (a five-lab demonstration) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, photovoltaics, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactors, control centers, and gas turbines, across five U.S. Department of Energy (DOE) national laboratories: National Laboratory of the Rockies (NLR), Idaho National Laboratory, National Energy Technology Laboratory, Lawrence Berkeley National Laboratory, and Sandia National Laboratories. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance DOE network, and are controlled via a centralized energy controller hosted at NLR's Advanced Research on Integrated Energy Systems facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Novel Zwitterionic Polyurethane-in-Salt Electrolytes with High Ion Conductivity, Elasticity, and Adhesion for High-Performance Solid-State Lithium Metal Batteries

This study presents a novel polymer-in-salt (PIS) zwitterionic polyurethane-based solid polymer electrolyte (zPU-SPE) that offers high ionic conductivity, strong interaction with electrodes, and excellent mechanical and electrochemical stabilities, making it promising for high-performance all solid-state lithium batteries (ASSLBs). The zPU-SPE exhibits remarkable lithium-ion (Li+) conductivity (3.7 × 10⁻⁴ S cm−1 at 25 °C), enabled by exceptionally high salt loading of up to 90 wt.% (12.6 molar ratio of Li salt to polymer unit) without phase separation. It addresses the limitations of conventional SPEs by combining high ionic conductivity with a Li+ transference number of 0.44, achieved through the incorporation of zwitterionic groups that enhance ion dissociation and transport. The high surface energy (338.4 J m−2) and elasticity ensure excellent adhesion to Li anodes, reducing interfacial resistance and ensuring uniform Li+ flux. When tested in Li||zPU||LiFePO₄ and Li||zPU||S/C cells, the zPU-SPE demonstrated remarkable cycling stability, retaining 76% capacity after 2000 cycles with the LiFePO4 cathode, and achieving 84% capacity retention after 300 cycles with the S/C cathode. Molecular simulations and a range of experimental characterizations confirm the superior structural organization of the zPU matrix, contributing to its outstanding electrochemical performance. The findings strongly suggest that zPU-SPE is a promising candidate for next-generation ASSLBs.

Wang, Kun↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Achieving high rate performance in hybrid pristine-recycled cathodes using model-informed electrode designs

Direct recycling lithium-ion battery cathodes, a process that retains the engineered oxide structures from end-of-life materials, presents a cost-effective and energy-efficient alternative to other battery recycling methods. However, while direct-recycled cathodes have demonstrated performance comparable to that of pristine materials at low cycling rates, their high-rate performance remains uncertain. Morphology changes in cathode particles, a main mode of degradation, directly impact rate performance by limiting surface kinetics and solid-phase diffusion. If direct recycling processes do not sufficiently restore pristine-like morphologies, the recycled materials may retain structural defects that hinder high-rate performance. The present work uses a physics-based pseudo-2D model to simulate hybrid electrodes with pristine and artificially “aged/recycled” NMC materials to investigate potential impacts of incorporating performance-limited aged cathode materials into cells. The study highlights how differences in transport and kinetic properties can influence rate capabilities in mixed electrodes — particularly in high-loading cells in high-demand applications. However, model results also reveal a possible mitigation strategy via dual-layer electrode architectures with lower-performing materials positioned near the current collector. Simulations of 4.0 mAh cm −2 cells cycled at 4C using a dual-layer architecture provided approximately 5%–30% more capacity in constant-current protocols compared to homogeneously blended electrode architectures with the same loadings and mixed-material compositions. These findings highlight the importance of strategic electrode design in minimizing potential performance losses and facilitating the integration of recycled materials into high-performance batteries, advancing sustainable and cost-effective battery manufacturing.

25 ENERGY STORAGE↗

Holistic energy analysis method for thermal management architectures of data centers

Modern high-performance computing (HPC) data centers (DCs), particularly those supporting energy-intensive artificial intelligence (AI) workloads, face escalating thermal management challenges that degrade performance through thermal throttling and drive up cooling power consumption and operational costs. To address this challenge, many have developed a wide variety of thermal management solutions (single-phase, two-phase, direct, indirect, hybrid, and more) which attempt to cool HPC DCs effectively while attempting to minimize overall system power consumption. However, the analysis of these solutions and methods to effectively compare one with another is lacking. Overall power usage effectiveness (PUE) and total-power usage effectiveness (TUE) provide a metric to quantify power consumption but fail to identify components in the system which require further optimization. To address this, we propose a holistic analytical framework – the waterfall diagram (WFD) – which leverages a waterfall chart methodology, offering a comprehensive visualization of both the thermal management system loop and heat flow pathways from individual server components to the outdoor ambient. Use of the WFD enables graphical estimations of power efficiency and cooling performance across each component of a DC cooling system and complements Sankey-style energy flow visualizations by additionally resolving stage-wise temperature changes and incremental TUE contributions. The framework is used in conjunction with simulation-based approaches, to conduct a detailed pressure drop and flow distribution analysis aimed at identifying the optimal coolant distribution architecture for a single-phase direct-to-chip water-cooled DC, which serves as the baseline for subsequent WFD analysis. Among the evaluated architectures, the 3 U modular coolant distribution architecture is found to demonstrate the best performance, considering minimal pressure drop and uniform flow distribution. In addition, TUE is calculated for each cooling loop component based on its associated pressure drop and corresponding pumping power, which are integrated into the WFD. This correlation between TUE and local temperature offers immediate insight into the power efficiency and thermal performance contributions of individual components, facilitating further development and optimization. Examples of WFD applications are presented under varying thermal loads and ambient conditions, demonstrating reasonable cooling strategies. Notably, the 3 U modular architecture maintains a consistent chip case temperature of 85°C, achieving a TUE of 1.016 at ambient temperature of 47°C, and a TUE of 1.026 at ambient temperature of 52°C. The WFD methodology provides an efficient, holistic, and streamlined framework for DC thermal management architecture assessment and enables design optimization which is important for addressing the thermal-fluidic energy challenges of current and next-generation DCs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

State of the art, gaps, and prospects in fusion materials theory and modelling

Advancing the theory and simulation of materials for fusion applications remains a key component of global roadmaps aimed at delivering much-needed fusion power. Especially as the drive for commercial application increases, prototypes must be designed against radiation damage before the relevant experimental data can be collected and cost reductions that are possible by testing materials in silico become even more important. Here, we summarise the state of the art as it emerged during the 7 th Fusion Materials Theory & Modelling Workshop that took place in 2024, with the aim to highlight present gaps and future directions for the fusion materials modelling community. Of particular interest were the effects of transmutations, chemical complexity with the development of novel alloys and interatomic potentials, advancements in modelling high-dose microstructures, comparison with experimental data and multiscale models for structural assessment relying on high-performance computing and virtual reality.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Harnessing Quantum Computing for Energy Materials: Opportunities and Challenges

Developing high-performance materials is critical for diverse energy applications to increase efficiency, improve sustainability and reduce costs. Classical computational methods have enabled important breakthroughs in energy materials development, but they face scaling and time-complexity limitations, particularly for high-dimensional or strongly correlated material systems. Quantum computing (QC) promises to offer a paradigm shift by exploiting quantum bits with their superposition and entanglement to address challenging problems intractable for classical approaches. This Perspective discusses the opportunities in leveraging QC to advance energy materials research and the challenges QC faces in solving complex and high-dimensional problems. We present cases on how QC, when combined with classical computing methods, can be used for the design and simulation of practical energy materials. We also outline the outlook for error-corrected, fault-tolerant QC capable of achieving predictive accuracy and quantum advantage for complex material systems.

Algorithms↗

Seed-Mediated Colloidal Synthesis of Multimetallic and High-Entropy Alloy Nanocrystal Libraries with Enhanced Catalytic Performance

Engineering colloidally stable multimetallic nanocrystals offers many benefits in a wide range of applications and allows manipulation of physical, chemical, and electronic properties of materials at the nanoscale. Synthesis routes are challenged by the chemical complexity required to temporally and spatially coordinate the reduction and alloying of multiple metal species, which has hampered the development of tunable libraries of colloidal materials to date. Here, in this work, we demonstrate a seed-mediated synthesis method to incorporate five or more metal elements into uniform, colloidally stable nanocrystals. By integrating machine learning-accelerated simulations, the synthesis of shortlisted high-entropy alloy nanocrystals was demonstrated. Multiple seed materials can be used, leading to a library of multimetallic nanocrystals with tunable electronic, physical, and alloy structures. The advantage of this synthetic protocol is highlighted in the preparation of catalytic materials that showed 2 orders of magnitude higher reaction rates than monometallic catalysts and outstanding thermal stability, thus highlighting the promise of this approach for high-performance materials in many areas.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Computational Thermal Hydraulics of a High-Performance Low-Enriched-Uranium Annular Target for HFIR Irradiation

Molybdenum-99 has historically been generated via isolation from fissioned highly enriched uranium (HEU) targets. Here, this isotope is in high demand due to its daily use across the world in radiopharmaceutical medical procedures. The primary objective of this work was to design and analyze an experimental target assembly containing one low-enriched-uranium (LEU) annular target for irradiation at the High Flux Isotope Reactor (HFIR). Efforts included incorporating spatially dependent energy sources from neutron and gamma interactions, quantifying thermal contact conductance at material interfaces, performing grid-independent studies, comparing turbulence models, and simulating various steady-state and transient scenarios relevant for irradiation qualification and eventual insertion. These models provide velocity, pressure, and temperature distributions in both space and time. Such results enable the selection of an appropriate irradiation location, fission rate density, and flow-limiting orifice size and demonstrate compliance with HFIR safety requirements such that insertion into the reactor can be approved. This analysis shows that across all scenarios, wetted surface temperatures remain below the coolant saturation temperature with no net vapor formation in the coolant. In every scenario, all components stay below 30% of the aluminum 6061 melting temperature. Computational fluid dynamics and system-level models predict peak target temperatures that agree within 4%, though the predicted axial location of the peak differs by about 10% of the heated length due to differences in flow development length. These results de-risk the irradiation of LEU (annular targets) and strengthen a domestic, HEU-independent 99 Mo supply by providing important fuel performance data to form the foundation for a robust licensing basis.

Molybdenum-99↗

Turbulence-Driven Edge-Localized-Mode-Free High-Confinement Mode with Divertor Detachment in a Metal-Wall Tokamak

We report the first demonstration of a minute-scale, edge-localized-mode-free high-confinement plasma regime compatible with divertor partial detachment and enhanced pedestal performance in a metal-wall tokamak, the Experimental Advanced Superconducting Tokamak. This regime is enabled by a newly identified mechanism: during divertor partial detachment, reduced ionization and enhanced pumping in a closed divertor lead to less cooling of the pedestal by recycling neutrals and seeding impurities, resulting in an increased pedestal temperature gradient, which excites high-frequency broadband turbulence. Gyrokinetic simulations identify the high-frequency broadband turbulence as a temperature-gradient-driven trapped electron mode (𝜂 𝑒 -TEM), which drives outward transport of particles and heat, thereby maintaining the edge-localized-mode-free state. This pedestal regime is particularly promising for the International Thermonuclear Experimental Reactor, where the anticipated lower density gradient, reduced E×B shear, and lower collisionality in the pedestal will further facilitate 𝜂 𝑒 -TEM excitation. The achieved integrated scenario with a detached divertor and turbulence-dominated pedestal thus offers a compelling solution for managing heat loads and metal impurity sources for long-pulse high-performance operation in future fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗