Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high-performance simulations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Overview of the MAST Upgrade physics programme: testing novel concepts at low aspect ratio to inform future devices

The research programme performed on the Mega Amp Spherical Tokamak (MAST) Upgrade experiment has made significant advances in developing the physics understanding of low aspect ratio tokamaks in support of the operation of ITER and design of fusion powerplants. High performance plasma scenarios have been developed to facilitate a broad programme of experiments, in which confinement is constrained by the presence of m/n = 2/1 modes that cause substantial losses of fast ions. The onset of these modes coincides with the q = 2 surface residing in a local minimum in the toroidal current density profile. The maximum electron temperature at the pedestal top, T e,ped is limited with gas fuelling to ∼350 eV to maintain regular ELMs; higher T e,ped results in a transition to a non-stationary ELM-free regime. The operational space of spherical tokamaks has been expanded into small and ELM-free regimes. Strong shaping of the last closed flux surface can induce a transition from large to small ELMs, and ELM suppression with resonant magnetic perturbations has been observed for the first time in a low aspect ratio tokamak. Negative triangularity shaping has induced a transition from ELMy H-mode to a high-performance L-mode regime for the first time in a low aspect ratio tokamak. In studies of fast ion confinement, losses of fast particles due to Global Alfvén Eigenmodes have been identified. Interactions between fast ions generated by off-axis neutral beam injection and thermal neutrals can result in significant losses of fast ions. Experiments with on- and off-axis neutral beam injection exhibit a flux pumping mechanism, where the central safety factor is held to ∼1 in the absence of sawteeth. In studies of pedestal physics, it has been found that elevated main chamber neutral pressures result in an increase in the electron density and reduction in the temperature at the pedestal top. Advances in understanding plasma exhaust include the integration of a high-performance plasma core with detached outer divertors in the X-point target configuration. A newly commissioned lower divertor cryopump reduces the lower divertor neutral pressure by up to 50%, with minimal effect on the main chamber or upper divertor. New measurements and SOLPS-ITER simulations emphasise the importance of plasma–neutral interactions on divertor detachment in the conditions accessible in experiments. Real-time control of the ionisation front location in both divertor chambers independently has been demonstrated in double null experiments, enabled by the tightly baffled divertor chambers.

MAST Upgrade↗

Extended Rice–Thomson analysis and atomistic simulations revealing grain boundary effects on fracture in refractory high-entropy alloys

Significance This work serves to extend the fundamental ductile vs. brittle fracture theory, specifically the Rice–Thomson criterion, by introducing a grain boundary ahead of an initiating crack which propagates at an oblique angle to impinge the boundary. Atomistic fracture simulations on two refractory complex concentrated alloys, the brittle NbMoTaW and the ductile Nb 45 Ta 25 Ti 15 Hf 15 , demonstrate qualitative correspondence with the extended Rice–Thomson criterion and experimental observations. Abstract Understanding how grain boundaries mediate fracture remains a critical challenge in designing ductile, high-performance refractory alloys. Here, we extend the Rice–Thomson criterion to account for the angle between cracks and the impinging grain boundaries (GBs), capturing the competition between intergranular fracture and dislocation-mediated plasticity. Using machine learning interatomic potentials, we performed molecular statics simulations to probe fracture mechanisms in nanocrystalline NbMoTaW and Nb 45 Ta 25 Ti 15 Hf 15 , each with two different grain sizes, revealing trends consistent with experimental observations and the extended Rice model. Comparison with averaged R-curves for bulk samples demonstrates that GBs enhance ductility in Nb 45 Ta 25 Ti 15 Hf 15 in both grain sizes investigated. In contrast, GBs only locally improve fracture resistance in NbMoTaW when cracks are temporarily pinned at GBs inclined at high angles from the crack, but generally promote brittle intergranular fracture. These contrasting behaviors are attributed to differences in GB cohesion, reflecting clear alloying trends that align with ab-initio calculations and trends observed experimentally. Our results bridge classical fracture theory, atomistic simulations, and experimental observations, providing a comprehensive understanding of the fracture mechanisms in nanocrystalline refractory complex concentrated alloys.

36 MATERIALS SCIENCE↗

Unitary Qubit Lattice Algorithms for Plasma Physics

This final technical report summarizes research conducted under DOE Award DE-SC0021653 to develop unitary Quantum Lattice Algorithms for modeling electromagnetic wave propagation and scattering in complex media, including plasmas. The project developed and validated quantum-inspired formulations of Maxwell's equations that preserve unitary evolution and can be evaluated on classical high-performance computing systems while providing a foundation for future quantum-computing implementations. Major accomplishments include the development of two- and three-dimensional algorithms for electromagnetic scattering; scalable, distributed-memory implementations demonstrated on the Perlmutter supercomputer; formulations for nonlinear lossless fluid dynamics and cold, lossless, inhomogeneous magnetized plasmas; and an explicit quantum algorithm for a time-discretized Lorenz model. Simulations reproduced a range of characteristic wave phenomena, including transient effects that are not readily apparent in conventional frequency-domain studies, demonstrating the effectiveness of the proposed approach for modeling complex electromagnetic and plasma systems. The work establishes a unified theoretical and computational framework for quantum and quantum-inspired simulation and provides a foundation for future implementation on fault-tolerant quantum systems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A 3D-Printed Millimeter-Wave Inline Waveguide-to-Coplanar-Waveguide Transition to Enable Dense Spectrometer Arrays for Intensity Mapping Surveys

We present a 3D-printed millimeter-wave, octave-bandwidth, in-line waveguide-to-coplanar-waveguide transition designed to enable focal planes with dense arrays of on-chip spectrometers. These arrays will enable compelling surveys of the large-scale structure of the universe through millimeter-wave intensity mapping. The transition consists of a four-step ridge-waveguide transformer that couples light from a rectangular waveguide onto a coplanar waveguide via an electrical connection made with indium bump bonds. We develop a tolerance-aware optimization approach to identify high-performance transition geometries that are robust to manufacturing variations; the same formulation can be applied to other tolerance-sensitive design problems. We also describe the implementation of a custom apparatus and procedure for bump-bonding a silicon chip to a metallized 3D-printed component. We detail the fabrication of the coplanar waveguide chip and three-dimensional waveguide structure, simulations and metrology of a test device, and room temperature reflectance measurements of this device. The room temperature metrology and reflection measurements are consistent with a model that predicts a coupling efficiency of $\mathord{\sim} 95\%$ at cryogenic temperatures in the 85-170 GHz frequency range.

Stover, Austin [Chicago U.; Chicago U., KICP] (ORC↗

Bayesian optimization of laser wakefield acceleration via spectral pulse shaping

In this paper, we investigate the effect of spectral pulse shaping of the laser driver on the performance of channel-guided, laser–plasma accelerators. The study was carried out with the assistance of Bayesian optimization using particle-in-cell simulations. We used a realistic plasma profile based on a novel optical-field-ionized channel technique with ionization injection and low, on-axis plasma densities to maximize the energy gain of the electron bunch trailing the laser. Spectral shaping allows us to modify the temporal profile of the laser driver while keeping the laser energy constant, affecting the acceleration and injection processes. In addition, we consider how modifying the plasma channel parameters may affect the target outputs. Given the complexity and breadth of the parameter space in question, we used numerical optimization to identify high-performing configurations. In particular, we found laser profiles with additional spectral content that, when used with optimal plasma channel parameters, result in charge content an order of magnitude higher than the baseline Gaussian case while also increasing the mean energy of the electron bunch.

Physics - Plasma physics↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (5-Lab Demo) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, PV, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactor (SMR), control centers, and gas turbines, across five DOE national laboratories-NLR, INL, NETL, LBNL, and SNL. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance U.S. Department of Energy's (DOE) network, and controlled via a centralized energy controller hosted at NLR's ARIES facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility. SuperLab 2.0 (5-Lab Demo) showcased a major advancement in federated national laboratory collaboration, enabling real-time, cross-laboratory experimentation to coordinate geographically dispersed distributed energy resources (DERs) using various communication protocols and networks. SuperLab 2.0 (5-Lab Demo) built on previous demonstrations conducted between NLR-PNNL and NLR-INL connecting diverse assets including distant protection devices, a SMR simulator, and a high temperature electrolyzer (HTE). Previous demos were based on a single connection between two labs with minimal coordination challenges. The 5-Lab demo with a centralized controller, distributed testbeds across different geographical locations, and use of protocols-based communication represents a scenario closer to real-world grid operations that coordinate resources across a region to meet system needs. This experiment studied how local DER controllers interact with a centralized energy controller during normal and abnormal events to maintain reliability. The SuperLab team across the five labs implemented a notional power system model equivalent of transmission and distribution lines, represented by the data networks interconnecting the labs. Each lab continuously exchanged local parameters (such as P and Q) from its Hardware-In-Loop (CHIL) and Power Hardware-In-Loop (PHIL) assets through centralized energy controller at NLR, enabling real-time interaction and coordination across sites. By leveraging ESnet as the communication backbone, the team successfully operated the distributed assets as a unified power system, with each bus represented by a different laboratory. This setup mirrors how assets interact in real-world power systems across dispersed locations with various protocols and latencies. At each lab site, assets were operated using their own local controllers which were coordinated through an overarching operation and control layer of centralized energy controller, equivalent to how an energy management system (EMS) orchestrates assets across a regional or national grid. SuperLab's federated connectivity utilized a Digital Real-Time Simulators (DRTS)-type gateway to connect Controller Hardware-In-Loop (CHIL) and PHIL assets between labs. To enable this federated connection through ESnet, a deterministic network was established where latency variations were consistent. This consistency allowed the development of digital filters for the power system assets across CHIL and PHIL interfaces to avoid unstable and unreliable grid conditions. This report provides an overview of the cross-laboratory configuration and offers insights into interconnecting geographically distributed research assets to test them as if they were co-located. This experiment represents a step toward linking nine DOE national laboratories, enabling nation-wide simulations that can address utility-driven challenges with grid resilience, flexibility, and modernization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Large Polaron Screening Delays Mobility Roll-Off in Halide Perovskites under High Injection

The interplay between carrier density regimes and the various scattering processes is fundamental for understanding charge transport efficiency in high-performance semiconductors. Combining ultrafast time-resolved terahertz spectroscopy (TRTS) and transient microwave photoconductivity (TRMC), we measure carrier mobilities over 7 orders of magnitude (10 13 –10 20 carriers cm –3 ), reaching a regime where carrier–carrier scattering dominates and mobility begins to roll off. We find the Caughey–Thomas model can be effectively employed for transport modeling in 3D metal halide perovskites (MHPs), and the mobility roll-off near ∼10 19 cm –3 can be tuned by MHP composition. Carrier lifetime in MHPs starts decreasing at much lower carrier densities due to interband multi-carrier recombination, while intraband carrier mobility remains unchanged as carrier–carrier scattering is screened by the presence of polarons. The validated Caughey–Thomas model provides a practical input for device simulations, prevents overestimation of diffusion length under high injection, and suggests that MHPs are excellent candidates for unipolar device applications under high carrier densities.

14 SOLAR ENERGY↗

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (a five-lab demonstration) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, photovoltaics, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactors, control centers, and gas turbines, across five U.S. Department of Energy (DOE) national laboratories: National Laboratory of the Rockies (NLR), Idaho National Laboratory, National Energy Technology Laboratory, Lawrence Berkeley National Laboratory, and Sandia National Laboratories. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance DOE network, and are controlled via a centralized energy controller hosted at NLR's Advanced Research on Integrated Energy Systems facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Novel Zwitterionic Polyurethane-in-Salt Electrolytes with High Ion Conductivity, Elasticity, and Adhesion for High-Performance Solid-State Lithium Metal Batteries

This study presents a novel polymer-in-salt (PIS) zwitterionic polyurethane-based solid polymer electrolyte (zPU-SPE) that offers high ionic conductivity, strong interaction with electrodes, and excellent mechanical and electrochemical stabilities, making it promising for high-performance all solid-state lithium batteries (ASSLBs). The zPU-SPE exhibits remarkable lithium-ion (Li+) conductivity (3.7 × 10⁻⁴ S cm−1 at 25 °C), enabled by exceptionally high salt loading of up to 90 wt.% (12.6 molar ratio of Li salt to polymer unit) without phase separation. It addresses the limitations of conventional SPEs by combining high ionic conductivity with a Li+ transference number of 0.44, achieved through the incorporation of zwitterionic groups that enhance ion dissociation and transport. The high surface energy (338.4 J m−2) and elasticity ensure excellent adhesion to Li anodes, reducing interfacial resistance and ensuring uniform Li+ flux. When tested in Li||zPU||LiFePO₄ and Li||zPU||S/C cells, the zPU-SPE demonstrated remarkable cycling stability, retaining 76% capacity after 2000 cycles with the LiFePO4 cathode, and achieving 84% capacity retention after 300 cycles with the S/C cathode. Molecular simulations and a range of experimental characterizations confirm the superior structural organization of the zPU matrix, contributing to its outstanding electrochemical performance. The findings strongly suggest that zPU-SPE is a promising candidate for next-generation ASSLBs.

Wang, Kun↗

Holistic energy analysis method for thermal management architectures of data centers

Modern high-performance computing (HPC) data centers (DCs), particularly those supporting energy-intensive artificial intelligence (AI) workloads, face escalating thermal management challenges that degrade performance through thermal throttling and drive up cooling power consumption and operational costs. To address this challenge, many have developed a wide variety of thermal management solutions (single-phase, two-phase, direct, indirect, hybrid, and more) which attempt to cool HPC DCs effectively while attempting to minimize overall system power consumption. However, the analysis of these solutions and methods to effectively compare one with another is lacking. Overall power usage effectiveness (PUE) and total-power usage effectiveness (TUE) provide a metric to quantify power consumption but fail to identify components in the system which require further optimization. To address this, we propose a holistic analytical framework – the waterfall diagram (WFD) – which leverages a waterfall chart methodology, offering a comprehensive visualization of both the thermal management system loop and heat flow pathways from individual server components to the outdoor ambient. Use of the WFD enables graphical estimations of power efficiency and cooling performance across each component of a DC cooling system and complements Sankey-style energy flow visualizations by additionally resolving stage-wise temperature changes and incremental TUE contributions. The framework is used in conjunction with simulation-based approaches, to conduct a detailed pressure drop and flow distribution analysis aimed at identifying the optimal coolant distribution architecture for a single-phase direct-to-chip water-cooled DC, which serves as the baseline for subsequent WFD analysis. Among the evaluated architectures, the 3 U modular coolant distribution architecture is found to demonstrate the best performance, considering minimal pressure drop and uniform flow distribution. In addition, TUE is calculated for each cooling loop component based on its associated pressure drop and corresponding pumping power, which are integrated into the WFD. This correlation between TUE and local temperature offers immediate insight into the power efficiency and thermal performance contributions of individual components, facilitating further development and optimization. Examples of WFD applications are presented under varying thermal loads and ambient conditions, demonstrating reasonable cooling strategies. Notably, the 3 U modular architecture maintains a consistent chip case temperature of 85°C, achieving a TUE of 1.016 at ambient temperature of 47°C, and a TUE of 1.026 at ambient temperature of 52°C. The WFD methodology provides an efficient, holistic, and streamlined framework for DC thermal management architecture assessment and enables design optimization which is important for addressing the thermal-fluidic energy challenges of current and next-generation DCs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Harnessing Quantum Computing for Energy Materials: Opportunities and Challenges

Developing high-performance materials is critical for diverse energy applications to increase efficiency, improve sustainability and reduce costs. Classical computational methods have enabled important breakthroughs in energy materials development, but they face scaling and time-complexity limitations, particularly for high-dimensional or strongly correlated material systems. Quantum computing (QC) promises to offer a paradigm shift by exploiting quantum bits with their superposition and entanglement to address challenging problems intractable for classical approaches. This Perspective discusses the opportunities in leveraging QC to advance energy materials research and the challenges QC faces in solving complex and high-dimensional problems. We present cases on how QC, when combined with classical computing methods, can be used for the design and simulation of practical energy materials. We also outline the outlook for error-corrected, fault-tolerant QC capable of achieving predictive accuracy and quantum advantage for complex material systems.

Algorithms↗

Seed-Mediated Colloidal Synthesis of Multimetallic and High-Entropy Alloy Nanocrystal Libraries with Enhanced Catalytic Performance

Engineering colloidally stable multimetallic nanocrystals offers many benefits in a wide range of applications and allows manipulation of physical, chemical, and electronic properties of materials at the nanoscale. Synthesis routes are challenged by the chemical complexity required to temporally and spatially coordinate the reduction and alloying of multiple metal species, which has hampered the development of tunable libraries of colloidal materials to date. Here, in this work, we demonstrate a seed-mediated synthesis method to incorporate five or more metal elements into uniform, colloidally stable nanocrystals. By integrating machine learning-accelerated simulations, the synthesis of shortlisted high-entropy alloy nanocrystals was demonstrated. Multiple seed materials can be used, leading to a library of multimetallic nanocrystals with tunable electronic, physical, and alloy structures. The advantage of this synthetic protocol is highlighted in the preparation of catalytic materials that showed 2 orders of magnitude higher reaction rates than monometallic catalysts and outstanding thermal stability, thus highlighting the promise of this approach for high-performance materials in many areas.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Computational Thermal Hydraulics of a High-Performance Low-Enriched-Uranium Annular Target for HFIR Irradiation

Molybdenum-99 has historically been generated via isolation from fissioned highly enriched uranium (HEU) targets. Here, this isotope is in high demand due to its daily use across the world in radiopharmaceutical medical procedures. The primary objective of this work was to design and analyze an experimental target assembly containing one low-enriched-uranium (LEU) annular target for irradiation at the High Flux Isotope Reactor (HFIR). Efforts included incorporating spatially dependent energy sources from neutron and gamma interactions, quantifying thermal contact conductance at material interfaces, performing grid-independent studies, comparing turbulence models, and simulating various steady-state and transient scenarios relevant for irradiation qualification and eventual insertion. These models provide velocity, pressure, and temperature distributions in both space and time. Such results enable the selection of an appropriate irradiation location, fission rate density, and flow-limiting orifice size and demonstrate compliance with HFIR safety requirements such that insertion into the reactor can be approved. This analysis shows that across all scenarios, wetted surface temperatures remain below the coolant saturation temperature with no net vapor formation in the coolant. In every scenario, all components stay below 30% of the aluminum 6061 melting temperature. Computational fluid dynamics and system-level models predict peak target temperatures that agree within 4%, though the predicted axial location of the peak differs by about 10% of the heated length due to differences in flow development length. These results de-risk the irradiation of LEU (annular targets) and strengthen a domestic, HEU-independent 99 Mo supply by providing important fuel performance data to form the foundation for a robust licensing basis.

Molybdenum-99↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗