Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high-performance simulations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Metal additively manufactured wavy fin cold-plate architecture for improved thermal-hydraulic performance

Rapid growth in artificial intelligence and data center workloads demands high-performance liquid cooling to manage increasing chip power. This study presents two metal-additive-manufactured cold plates with sinusoidal fins, constant-amplitude wavy fins and linearly variable-amplitude wavy fins and compares them against metal-additive-manufactured straight fins using experiments conducted at 1 kW heat dissipation as well as high-fidelity 3D conjugate computational fluid dynamic simulations. The cold plates were printed in AlSi10Mg material and underwent design using a Python-automated workflow prior to manufacture and testing. The experiments show that wavy fins reduce the normalized thermal resistance by 35 to 45 % at water flow rates from 1 to 4 LPM. At a fixed 20 kPa pressure drop, the variable-waviness design lowered peak surface temperature by 9 °C and thermal resistance by 51 %, while edge-channel maldistribution in the constant wavy fin design limited gains. A thermal resistance breakdown revealed that 55–63 % of the total thermal resistance in wavy designs comes from base heat conduction, 27–33 % from fin heat conduction, and 9–13 % from fin heat convection, indicating the need to address conduction bottlenecks. Parametric sweeps identify a 3 mm fin pitch as optimal, and that horizontal inlet/outlet manifolds further reduce pressure drop by 30–60 % and thermal resistance by 9–16 % relative to vertical inlet-outlet manifolds. The results yield comprehensive guidelines for fin geometry, manifold alignment, material selection and additive-manufacturing constraints to realize high-performance liquid-cooled cold plates for power-dense electronics.

3d printing↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (a five-lab demonstration) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, photovoltaics, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactors, control centers, and gas turbines, across five U.S. Department of Energy (DOE) national laboratories: National Laboratory of the Rockies (NLR), Idaho National Laboratory, National Energy Technology Laboratory, Lawrence Berkeley National Laboratory, and Sandia National Laboratories. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance DOE network, and are controlled via a centralized energy controller hosted at NLR's Advanced Research on Integrated Energy Systems facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Revolutionizing Energy Storage: AI, Automation, and Advanced Modeling as Catalysts for Next-Generation Breakthroughs

The Presidential Symposium (PRES) at the 2025 Fall Meeting, hosted by the President’s Office and Energy and Fuels Division, American Chemical Society (ACS) in Washington, DC, brought together a diverse group of chemists, engineers, and materials scientists working in battery materials & systems, automation and artificial intelligence from academia, industry, and national laboratories. The accelerating demand for high-performance, scalable, and sustainable energy storage has catalyzed a paradigm shift in how materials are dis-covered, devices are engineered, and systems are optimized. This Presidential Symposium, entitled “Revolutionizing Energy Storage: AI, Automation, and Advanced Modeling Driving Next-Gen Breakthroughs”, brings together global leaders to unveil transformative strategies anchored in the AAA framework: Artificial Intelligence, Automation, and Advanced Modeling. Artificial Intelligence is redefining the frontiers of energy storage by enabling predictive design, real-time optimization, and intelligent control across diverse chemistries and architectures. Automation is streamlining the synthesis, characterization, and testing of battery materials, dramatically accelerating innovation cycles and unlocking scalable solutions for grid and mobility applications. Advanced Modeling, spanning atomic to system-level scales, provides unprecedented insight into electrochemical dynamics, degradation pathways, and thermal behavior, particularly when coupled with physics-informed machine learning and digital twin technologies. Digital twins, in turn, leverage the AAA framework by integrating real-time data, physics-based models, and AI predictions into dynamic virtual replicas, enabling proactive diagnostics, optimization, and system resilience. Together, these synergistic pillars are not only re-shaping the scientific landscape but also forging a new era of reproducible, data-driven, and resilient energy storage innovation. In conclusion, this symposium marks a pivotal moment in the convergence of computational intelligence and experimental rigor, charting the course for next-generation breakthroughs in lithium-ion, solid-state, and flow battery technologies.

Artificial Intelligence (AI)↗

Universal Electronic‐Structure Relationship Governing Intrinsic Magnetic Properties in Permanent Magnets

An electronic-structure-centered perspective is presented on permanent-magnet (PM) design, highlighting two key levers, that is, saturation magnetization (M s ), governed by 3d-band filling and exchange physics, and magnetocrystalline anisotropy energy (MAE), arising from spin-orbit coupling (SOC) on anisotropic orbital populations. Reviewing current practices, including DFT-based MAE/J ij extraction, atomistic-spin and micromagnetic modeling, and high-throughput machine learning (ML) pipelines, three bottlenecks limiting predictive discovery is identified that is i) electronic-structure accuracy for small MAE (sensitive to functional choice, Hubbard U, and many-body effects), ii) finite-temperature and kinetic realism (phonon/magnon renormalization, ordering kinetics), and iii) descriptor and multiscale decoupling (lack of SOC-weighted and orbital-resolved fingerprints). Deep dives into the electronic-structure of Nd─Fe─B and Fe─N show how these fingerprints govern magnetic performance, motivating DFT- and quantum-mechanics-based descriptors for discovery. Unbiased, structure-driven exploration, coupled with high-throughput simulations, ML, generative AI, and reasoning models, accelerates candidate identification and propagates insights across scales. Addressing supply-chain risks, on future needs of designing “critical-element-free” magnets with tailored microstructure and high energy products is emphasized. By integrating electronic fingerprints, AI reasoning, and multiscale modeling, a practical roadmap is provided for rare-earth-lean or rare-earth free, high-performance, sustainable PMs.

Singh, Prashant [Ames Laboratory, and Iowa State U↗

Novel Zwitterionic Polyurethane-in-Salt Electrolytes with High Ion Conductivity, Elasticity, and Adhesion for High-Performance Solid-State Lithium Metal Batteries

This study presents a novel polymer-in-salt (PIS) zwitterionic polyurethane-based solid polymer electrolyte (zPU-SPE) that offers high ionic conductivity, strong interaction with electrodes, and excellent mechanical and electrochemical stabilities, making it promising for high-performance all solid-state lithium batteries (ASSLBs). The zPU-SPE exhibits remarkable lithium-ion (Li+) conductivity (3.7 × 10⁻⁴ S cm−1 at 25 °C), enabled by exceptionally high salt loading of up to 90 wt.% (12.6 molar ratio of Li salt to polymer unit) without phase separation. It addresses the limitations of conventional SPEs by combining high ionic conductivity with a Li+ transference number of 0.44, achieved through the incorporation of zwitterionic groups that enhance ion dissociation and transport. The high surface energy (338.4 J m−2) and elasticity ensure excellent adhesion to Li anodes, reducing interfacial resistance and ensuring uniform Li+ flux. When tested in Li||zPU||LiFePO₄ and Li||zPU||S/C cells, the zPU-SPE demonstrated remarkable cycling stability, retaining 76% capacity after 2000 cycles with the LiFePO4 cathode, and achieving 84% capacity retention after 300 cycles with the S/C cathode. Molecular simulations and a range of experimental characterizations confirm the superior structural organization of the zPU matrix, contributing to its outstanding electrochemical performance. The findings strongly suggest that zPU-SPE is a promising candidate for next-generation ASSLBs.

Wang, Kun↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Achieving high rate performance in hybrid pristine-recycled cathodes using model-informed electrode designs

Direct recycling lithium-ion battery cathodes, a process that retains the engineered oxide structures from end-of-life materials, presents a cost-effective and energy-efficient alternative to other battery recycling methods. However, while direct-recycled cathodes have demonstrated performance comparable to that of pristine materials at low cycling rates, their high-rate performance remains uncertain. Morphology changes in cathode particles, a main mode of degradation, directly impact rate performance by limiting surface kinetics and solid-phase diffusion. If direct recycling processes do not sufficiently restore pristine-like morphologies, the recycled materials may retain structural defects that hinder high-rate performance. The present work uses a physics-based pseudo-2D model to simulate hybrid electrodes with pristine and artificially “aged/recycled” NMC materials to investigate potential impacts of incorporating performance-limited aged cathode materials into cells. The study highlights how differences in transport and kinetic properties can influence rate capabilities in mixed electrodes — particularly in high-loading cells in high-demand applications. However, model results also reveal a possible mitigation strategy via dual-layer electrode architectures with lower-performing materials positioned near the current collector. Simulations of 4.0 mAh cm −2 cells cycled at 4C using a dual-layer architecture provided approximately 5%–30% more capacity in constant-current protocols compared to homogeneously blended electrode architectures with the same loadings and mixed-material compositions. These findings highlight the importance of strategic electrode design in minimizing potential performance losses and facilitating the integration of recycled materials into high-performance batteries, advancing sustainable and cost-effective battery manufacturing.

25 ENERGY STORAGE↗

Holistic energy analysis method for thermal management architectures of data centers

Modern high-performance computing (HPC) data centers (DCs), particularly those supporting energy-intensive artificial intelligence (AI) workloads, face escalating thermal management challenges that degrade performance through thermal throttling and drive up cooling power consumption and operational costs. To address this challenge, many have developed a wide variety of thermal management solutions (single-phase, two-phase, direct, indirect, hybrid, and more) which attempt to cool HPC DCs effectively while attempting to minimize overall system power consumption. However, the analysis of these solutions and methods to effectively compare one with another is lacking. Overall power usage effectiveness (PUE) and total-power usage effectiveness (TUE) provide a metric to quantify power consumption but fail to identify components in the system which require further optimization. To address this, we propose a holistic analytical framework – the waterfall diagram (WFD) – which leverages a waterfall chart methodology, offering a comprehensive visualization of both the thermal management system loop and heat flow pathways from individual server components to the outdoor ambient. Use of the WFD enables graphical estimations of power efficiency and cooling performance across each component of a DC cooling system and complements Sankey-style energy flow visualizations by additionally resolving stage-wise temperature changes and incremental TUE contributions. The framework is used in conjunction with simulation-based approaches, to conduct a detailed pressure drop and flow distribution analysis aimed at identifying the optimal coolant distribution architecture for a single-phase direct-to-chip water-cooled DC, which serves as the baseline for subsequent WFD analysis. Among the evaluated architectures, the 3 U modular coolant distribution architecture is found to demonstrate the best performance, considering minimal pressure drop and uniform flow distribution. In addition, TUE is calculated for each cooling loop component based on its associated pressure drop and corresponding pumping power, which are integrated into the WFD. This correlation between TUE and local temperature offers immediate insight into the power efficiency and thermal performance contributions of individual components, facilitating further development and optimization. Examples of WFD applications are presented under varying thermal loads and ambient conditions, demonstrating reasonable cooling strategies. Notably, the 3 U modular architecture maintains a consistent chip case temperature of 85°C, achieving a TUE of 1.016 at ambient temperature of 47°C, and a TUE of 1.026 at ambient temperature of 52°C. The WFD methodology provides an efficient, holistic, and streamlined framework for DC thermal management architecture assessment and enables design optimization which is important for addressing the thermal-fluidic energy challenges of current and next-generation DCs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

State of the art, gaps, and prospects in fusion materials theory and modelling

Advancing the theory and simulation of materials for fusion applications remains a key component of global roadmaps aimed at delivering much-needed fusion power. Especially as the drive for commercial application increases, prototypes must be designed against radiation damage before the relevant experimental data can be collected and cost reductions that are possible by testing materials in silico become even more important. Here, we summarise the state of the art as it emerged during the 7 th Fusion Materials Theory & Modelling Workshop that took place in 2024, with the aim to highlight present gaps and future directions for the fusion materials modelling community. Of particular interest were the effects of transmutations, chemical complexity with the development of novel alloys and interatomic potentials, advancements in modelling high-dose microstructures, comparison with experimental data and multiscale models for structural assessment relying on high-performance computing and virtual reality.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Harnessing Quantum Computing for Energy Materials: Opportunities and Challenges

Developing high-performance materials is critical for diverse energy applications to increase efficiency, improve sustainability and reduce costs. Classical computational methods have enabled important breakthroughs in energy materials development, but they face scaling and time-complexity limitations, particularly for high-dimensional or strongly correlated material systems. Quantum computing (QC) promises to offer a paradigm shift by exploiting quantum bits with their superposition and entanglement to address challenging problems intractable for classical approaches. This Perspective discusses the opportunities in leveraging QC to advance energy materials research and the challenges QC faces in solving complex and high-dimensional problems. We present cases on how QC, when combined with classical computing methods, can be used for the design and simulation of practical energy materials. We also outline the outlook for error-corrected, fault-tolerant QC capable of achieving predictive accuracy and quantum advantage for complex material systems.

Algorithms↗

Hydrophobic Metal–Organic Frameworks Enable Superior High-Pressure Ammonia Storage through Geometric Design

Hydrophobic metal–organic frameworks (MOFs) are typically overlooked for ammonia storage due to weak host–guest interactions. Here, we demonstrate that four structurally analogous aluminum-based MOFs exhibit a counterintuitive behavior whereby framework geometry, rather than ligand hydrophilicity, determines high-pressure NH 3 adsorption performance. The hydrophobic CAU-23 achieved an exceptional capacity matching hydrophilic analogs despite its poor low-pressure uptake. This pressure-dependent enhancement stems from the unique 4-cis-4-trans geometry of CAU-23 compared to the purely cis arrangement of MIL-160 and KMF-1 and the alternating cis-trans configuration of MOF-303. Critically, CAU-23 retained 95% capacity over three high-pressure cycles, whereas hydrophilic MOFs suffered 39–46% irreversible losses due to strong NH 3 -framework interactions that compromise structural integrity. Grand canonical Monte Carlo simulations reveal that high pressure enables NH 3 clustering through intermolecular hydrogen bonding, bypassing the need for strong host–guest interactions. High-pressure powder X-ray diffraction measurements confirm the exceptional mechanical resilience of CAU-23, showing complete structural recovery upon decompression despite exhibiting the highest pressure sensitivity among the studied MOFs. An extended analog, HE-CAU-23, validates this design principle with further enhanced capacity. Furthermore, these findings reveal a paradigm shift toward hydrophobic MOFs with optimized geometry for high-performance and regenerable gas storage applications.

MOFs↗

Seed-Mediated Colloidal Synthesis of Multimetallic and High-Entropy Alloy Nanocrystal Libraries with Enhanced Catalytic Performance

Engineering colloidally stable multimetallic nanocrystals offers many benefits in a wide range of applications and allows manipulation of physical, chemical, and electronic properties of materials at the nanoscale. Synthesis routes are challenged by the chemical complexity required to temporally and spatially coordinate the reduction and alloying of multiple metal species, which has hampered the development of tunable libraries of colloidal materials to date. Here, in this work, we demonstrate a seed-mediated synthesis method to incorporate five or more metal elements into uniform, colloidally stable nanocrystals. By integrating machine learning-accelerated simulations, the synthesis of shortlisted high-entropy alloy nanocrystals was demonstrated. Multiple seed materials can be used, leading to a library of multimetallic nanocrystals with tunable electronic, physical, and alloy structures. The advantage of this synthetic protocol is highlighted in the preparation of catalytic materials that showed 2 orders of magnitude higher reaction rates than monometallic catalysts and outstanding thermal stability, thus highlighting the promise of this approach for high-performance materials in many areas.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Modeling strain and quantum confinement in GaAs/Ga x In 1−x P superlattices for spin-polarized electron sources

In this study, we systematically design and simulate a series of GaAs-based superlattice configurations aimed at enhancing heavy-hole–light-hole band splitting while simultaneously optimizing band alignment to reduce the conduction band barrier, thereby facilitating efficient electron transport. These combined effects are crucial for achieving high electron spin polarization and high quantum efficiency, the two key performance metrics of next-generation spin-polarized electron sources. We investigated three types of superlattice architectures: (1) compressively strained GaAs wells on GaInP barriers, yielding a maximum band splitting of 140 meV, (2) lattice-matched GaAs/GaInP structures, resulting in the maximum band splitting of 75 meV, and (3) tensile strained GaAs wells on GaInP barriers, with a maximum band splitting of 40 meV. The results demonstrate the tunability of heavy-hole–light-hole band splitting and establish a design framework for high-performance spin-polarized photocathodes based on a combination of strain engineering, quantum confinement, and optimized heterostructure design.

Electron sources↗

Computational Thermal Hydraulics of a High-Performance Low-Enriched-Uranium Annular Target for HFIR Irradiation

Molybdenum-99 has historically been generated via isolation from fissioned highly enriched uranium (HEU) targets. Here, this isotope is in high demand due to its daily use across the world in radiopharmaceutical medical procedures. The primary objective of this work was to design and analyze an experimental target assembly containing one low-enriched-uranium (LEU) annular target for irradiation at the High Flux Isotope Reactor (HFIR). Efforts included incorporating spatially dependent energy sources from neutron and gamma interactions, quantifying thermal contact conductance at material interfaces, performing grid-independent studies, comparing turbulence models, and simulating various steady-state and transient scenarios relevant for irradiation qualification and eventual insertion. These models provide velocity, pressure, and temperature distributions in both space and time. Such results enable the selection of an appropriate irradiation location, fission rate density, and flow-limiting orifice size and demonstrate compliance with HFIR safety requirements such that insertion into the reactor can be approved. This analysis shows that across all scenarios, wetted surface temperatures remain below the coolant saturation temperature with no net vapor formation in the coolant. In every scenario, all components stay below 30% of the aluminum 6061 melting temperature. Computational fluid dynamics and system-level models predict peak target temperatures that agree within 4%, though the predicted axial location of the peak differs by about 10% of the heated length due to differences in flow development length. These results de-risk the irradiation of LEU (annular targets) and strengthen a domestic, HEU-independent 99 Mo supply by providing important fuel performance data to form the foundation for a robust licensing basis.

Molybdenum-99↗

On FIRE mode in KSTAR

We report on the status of Fast Ion Regulated Enhancement (FIRE) mode experiments in the Korea Superconducting Tokamak Advanced Research. This regime is being developed for high-performance, steady-state operation which features a stationary ion internal transport barrier, enabling a central ion temperature approaching 10 keV to be sustained for up to 50 s, without the need for delicate profile control and with no significant impurity accumulation. As its key novelty lies in the significant contribution of fast ions that stabilize core turbulence, the regime has been named FIRE mode. To achieve this regime, neutral beam injection is applied at moderate power levels near the L–H power threshold, while maintaining low plasma density to avoid the L–H transition. The scenario is typically established in diverted magnetic configurations. The core features of FIRE mode were investigated through power balance analysis and fluctuation measurements, revealing a clear transport bifurcation in the ion channel. At the plasma edge, FIRE mode occasionally exhibits I-mode characteristics, particularly in unfavorable magnetic null configurations with q 95 ∼ 4, including the presence of weakly coherent modes. In terms of MHD activity, sawtooth oscillations are observed but appear to be stabilized during the high-performance phase. Fast ion-driven Alfvénic eigenmodes (AEs), indicated by strong frequency chirping near 200 kHz, are also observed. Additionally, lower-frequency MHD activities, distinct from the fast ion-driven AEs, are present and have some influence on plasma performance. The enhancement of core confinement is primarily attributed to fast ion effects, which were evaluated with respect to dilution, alpha stabilization, and resonant interactions with turbulence using gyrokinetic analyses. Among these, the dilution effect was found to be the most dominant. The characteristics of FIRE mode were compared with those of other hot ion plasma scenarios, such as supershot and hot ion mode. While they share many similarities, FIRE mode is distinguished by the accessibility to conditions with T i ≈ T e , presence of I-mode edge features, and its long-duration sustainment. Predictive simulations of FIRE mode were performed using integrated transport modeling with TRIASSIC, employing the TGLF anomalous transport model. These simulations confirmed the critical role of fast ions in achieving this regime. The future prospect of FIRE mode for application in fusion reactors is discussed, with an emphasis on possibly extending the regime to higher density operation.

FIRE mode↗

Turbulence-Driven Edge-Localized-Mode-Free High-Confinement Mode with Divertor Detachment in a Metal-Wall Tokamak

We report the first demonstration of a minute-scale, edge-localized-mode-free high-confinement plasma regime compatible with divertor partial detachment and enhanced pedestal performance in a metal-wall tokamak, the Experimental Advanced Superconducting Tokamak. This regime is enabled by a newly identified mechanism: during divertor partial detachment, reduced ionization and enhanced pumping in a closed divertor lead to less cooling of the pedestal by recycling neutrals and seeding impurities, resulting in an increased pedestal temperature gradient, which excites high-frequency broadband turbulence. Gyrokinetic simulations identify the high-frequency broadband turbulence as a temperature-gradient-driven trapped electron mode (𝜂 𝑒 -TEM), which drives outward transport of particles and heat, thereby maintaining the edge-localized-mode-free state. This pedestal regime is particularly promising for the International Thermonuclear Experimental Reactor, where the anticipated lower density gradient, reduced E×B shear, and lower collisionality in the pedestal will further facilitate 𝜂 𝑒 -TEM excitation. The achieved integrated scenario with a detached divertor and turbulence-dominated pedestal thus offers a compelling solution for managing heat loads and metal impurity sources for long-pulse high-performance operation in future fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗