Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain decomposition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Long-time integration of parametric evolution equations with physics-informed DeepONets

Ordinary and partial differential equations (ODEs/PDEs) play a paramount role in analyzing and simulating complex dynamic processes across all corners of science and engineering. In recent years machine learning tools are aspiring to introduce new effective ways of simulating such equations, however existing approaches are not able to reliably return stable and accurate predictions across long temporal horizons. We aim to address this challenge by introducing an effective framework for learning evolution operators that map random initial conditions to associated ODE/PDE solutions within a short time interval. Such operators can be parametrized by deep neural networks that are trained in an entirely self-supervised manner without requiring one to generate any paired input-output observations. Global long-time predictions across a range of initial conditions can be then obtained by iteratively evaluating the trained model using each prediction as the initial condition for the next evaluation step. Here, this introduces a new approach to temporal domain decomposition that is shown to be effective in performing accurate long-time simulations for a wide range of parametric ODE and PDE systems, from wave propagation, to reaction-diffusion dynamics and stiff chemical kinetics, introducing a new way of rapidly emulating non-equilibrium processes in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the spatial self-shielding impact for TRISO-based nuclear fuel depletion

Reactor physics analyses of nuclear cores with nuclear fuel concepts containing tristructural isotropic (TRISO) particles, such as pebbles or compact fuel elements, rely on various degrees of simplification to keep these highly heterogeneous problems computationally tractable. One such limitation regards the level of spatial discretization employed during burnup calculations, where traditionally only a limited number of spatial zones are modeled at the full core level and assume that the spectrum is constant within the fuel elements and TRISO particles in this depletion zone. This type of assumption neglects the impact of spatial self-shielding effect within the kernels (microscale level) as well as within the compact or pebbles (mesoscale level). Furthermore, the Monte Carlo code Serpent 2 contains many relevant features for efficiently modeling this type of geometry, including a collision-based domain decomposition intended for very large burnup calculations, which we leveraged for this work to quantify the impact of capturing neutron flux variations occurring at the micro- and mesoscale level on a series of high-temperature gas-cooled reactor fuel element depletion problems. While spatial self-shielding is observed at both scales, with differences from a volume-averaged burnup of ±7% within the kernels and ±2% between TRISO particles within the fuel element, the conjugated effect on nuclide inventories and multiplication factor are negligible, hence confirming that assuming a single average spectrum value may be sufficient for most applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Scalable freeform optimization of wide-aperture 3D metalenses by zoned discrete axisymmetry

We introduce a novel framework for design and optimization of 3D freeform metalenses that attains nearly linear scaling of computational cost with diameter, by breaking the lens into a sequence of radial “zones” with 𝑛-fold discrete axisymmetry, where 𝑛 increases with radius. This allows vastly more design freedom than imposing continuous axisymmetry, while avoiding the compromises of the locally periodic approximation (LPA) or scalar diffraction theory. Using a GPU-accelerated finite-difference time-domain (FDTD) solver in cylindrical coordinates, we perform full-wave simulation and topology optimization within each supra-wavelength zone. We validate our approach by designing millimeter and centimeter-scale, poly-achromatic, 3D freeform metalenses which outperform the state of the art. By demonstrating the scalability and resulting optical performance enabled by our “zoned discrete axisymmetry” (ZDA) and supra-wavelength domain decomposition, we highlight the potential of our framework to advance large-scale meta-optics and next-generation photonic technologies.

Sun, Mengdi [Wesleyan University]↗

Status of GPU capabilities within the Shift Monte Carlo radiation transport code

Shift is a general-purpose Monte Carlo (MC) radiation transport code for fission, fusion, and national security applications. Shift has been adapted to efficiently run on GPUs in order to leverage leadership-class supercomputers. This work presents Shift’s current GPU capabilities. These include core radiation transport capabilities for eigenvalue and fixed-source simulations, and support for non-uniform domain decomposition, Doppler broadening, free-gas elastic scattering, general-purpose geometry, hybrid MC/deterministic transport, and depletion. Transport results demonstrate a 2–5× GPU-to-CPU speedup on a per-node basis for an eigenvalue problem on the Frontier supercomputer and a 28× speedup for a fixed-source problem on the Summit supercomputer.

Biondo, Elliott [ORNL] (ORCID:0000000290881360)↗

Preparation of the SU(3) lattice Yang-Mills vacuum with variational quantum methods

Studying QCD and other gauge theories on quantum hardware requires the preparation of physically interesting states. The variational quantum eigensolver provides a way of performing vacuum state preparation on quantum hardware. Here in this work, variational quantum eigensolver is applied to pure SU(3) lattice Yang-Mills on a single plaquette and one dimensional plaquette chains. Bayesian optimization and gradient descent were investigated for performing the classical optimization. Ansatz states for plaquette chains are constructed in a scalable manner from smaller systems using domain decomposition and a stitching procedure analogous to the density matrix renormalization group. Small examples are performed on IBM’s superconducting Manila processor.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

pdas-experiments

SAND2025-04589O pdas-experiments automates computational experiments of fluid flow simulations. It uses the pressio-demoapps-schwarz package as a basis to break down complex simulations into smaller, manageable parts. This application is an extension of the Sandia Pressio software which uses domain decomposition to work with complex simulations more efficiently. Users can test different simulation setups, while keeping a detailed record of their experiments so they can be reproduced later. The software includes a C++ program that runs individual experiments based on user-defined settings in a YAML file, as well as a Python script that can manage multiple simulations at once. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tezaur, Irina [Sandia National Lab. (SNL-CA), Live↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression

Handling large-scale scientific data in high-performance computing (HPC) environments poses significant challenges, including excessive I/O, high storage costs, and slow query performance. Traditional approaches often require full data decompression and scans, making them impractical for real-time or interactive analysis. To address these limitations, we introduce Eureka, a unified data-index co-compression framework that enables fine-grained access and efficient range queries on compressed scientific datasets. Eureka integrates spatial domain decomposition with block-wise error-bounded lossy compression to support selective decompression. It constructs a hierarchical AVL-tree index during compression to capture block-level value ranges, enabling fast pruning during query execution. To reduce metadata overhead, the index itself is also compressed while ensuring recall-preserving results. Experiments on six diverse HPC simulation datasets show that Eureka achieves up to 25x data compression and over 300x index compression, surpassing state-of-the-art compressors such as SZ3 and ZFP in rate-distortion performance. Additionally, Eureka delivers over 30x speedup for low-selectivity range queries, making it a scalable and efficient solution for modern scientific data analysis.

Yan, Ning↗

DRDMannTurb: A Python package for scalable, data-driven synthetic turbulence

Synthetic turbulence models (STMs) are used in wind engineering to generate realistic flow fields and are employed as inputs to industrial wind simulations. Examples include prescribing inlet conditions in large eddy simulations that model loads on wind turbines and tall buildings. We are interested in STMs capable of generating fluctuations based on prescribed second-moment statistics since such models can simulate environmental conditions that closely resemble on-site observations. To this end, the widely used Mann model (see Mann, 1994, 1998) is the inspiration for DRDMannTurb. The Mann model is described by three physical parameters: a magnitude parameter influencing the global variance of the wind field and corresponding to the Kolmogorov constant multiplied by the rate of viscous dissipation of the turbulent kinetic energy to the two-thirds, αϵ 2/3 , a turbulence length scale parameter L, and a nondimensional parameter Γ related to the lifetime of the eddies. A number of studies, as well as international standards (e.g., those by the International Electrotechnical Commission (IEC)), include recommended values for these three parameters with the goal of standardizing wind simulations according to observed energy spectra. Yet, having only three parameters, the Mann model faces limitations in accurately representing the diversity of observable spectra. This Python package enables users to extend the Mann model and more accurately fit field measurements through flexible neural network models of the eddy lifetime function. Following Keith et al. (2021), we refer to this class of models as Deep Rapid Distortion (DRD) models. DRDMannTurb also includes a general module implementing an efficient method for synthetic turbulence generation based on a domain decomposition technique. This technique is also described in Keith et al. (2021).

17 WIND ENERGY↗

The Schwarz Alternating Method for the Seamless Coupling of Nonlinear Reduced Order Models and Full Order Models

Projection-based model order reduction allows for the parsimonious representation of full order models (FOMs), typically obtained through the discretization of a set of partial differential equations (PDEs) using conventional techniques (e.g., finite element, finite volume, finite difference methods) where the discretization may contain a very large number of degrees of freedom. As a result of this more compact representation, the resulting projection-based reduced order models (ROMs) can achieve considerable computational speedups, which are especially useful in real-time or multi-query analyses. One known deficiency of projection-based ROMs is that they can suffer from a lack of robustness, stability and accuracy, especially in the predictive regime, which ultimately limits their useful application. Another research gap that has prevented the widespread adoption of ROMs within the modeling and simulation community is the lack of theoretical and algorithmic foundations necessary for the “plug-and-play” integration of these models into existing multi-scale and multi-physics frameworks. This paper describes a new methodology that has the potential to address both of the aforementioned deficiencies by coupling projection-based ROMs with each other as well as with conventional FOMs by means of the Schwarz alternating method [41]. Leveraging recent work that adapted the Schwarz alternating method to enable consistent and concurrent multiscale coupling of finite element FOMs in solid mechanics [35, 36], we present a new extension of the Schwarz framework that enables FOM-ROM and ROM-ROM coupling, following a domain decomposition of the physical geometry on which a PDE is posed. In order to maintain efficiency and achieve computation speed-ups, we employ hyper-reduction via the Energy-Conserving Sampling and Weighting (ECSW) approach [13]. We evaluate the proposed coupling approach in the reproductive as well as in the predictive regime on a canonical test case that involves the dynamic propagation of a traveling wave in a nonlinear hyper-elastic material.

97 MATHEMATICS AND COMPUTING↗

Physics Informed Neural Networks as Computational Physics Emulators

This report is a brief overview and evaluation of Physics Informed Neural Networks (PINNs). Karniadakis and co-workers, e.g., Karniadakis et al. (2021) assert that the PINNs approach integrates seamlessly both data and mathematical physics models, even in partially understood, uncertain and high-dimensional contexts. They further claim that PINNs are effective and efficient for ill-posed and inverse problems, and when combined with domain decomposition, are scalable to large problems and a tool to discover hidden physics. While they demonstrate the capabilities in specific academic instances, their overarching claims about PINNs seem to be an overstatement, at least at the current time. We have briefly considered a few of the limitations of PINNs in this investigation. It is not clear to us if the PINNs approach can ever be competitive with approaches that use specialized algorithms to achieve high-accuracy solutions of governing equations and other techniques that can combine observational data with such solutions. As an example of the latter, consistent with the principles of Bayesian inference, data assimilation, or more generally data-model fusion, is a process that fuses observational data typically with a computational model that respects certain constraints such as conservation laws. For example, improvements in observational network combined with data assimilation have been key in improving weather predictions over the past four decades Kalnay (2003).

97 MATHEMATICS AND COMPUTING↗

Cold Plasma Measurements

We have continued the simulation campaign in support of our ongoing magnetospheric cold plasma research project. This project aims to develop the next-generation particle instruments to measure the properties of the cold particle populations in the Earth’s magnetosphere. For this purpose, simulations have been performed with a Particle-In-Cell (PIC) code called the Curvilinear PIC (CPIC). The code is formulated in curvilinear geometry and couples the standard PIC algorithm with algorithms for the generation and adaptation of the underlaying computational mesh. It conforms to complex objects like spacecraft and it can place more grid points in regions where higher resolution is needed. The code also features a scalable solver based on the multigrid algorithm and it is fully parallelized via domain decomposition and MPI.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multi-physics Preconditioning for Thermally Activated Batteries

Thermal batteries, also known as molten-salt batteries, are single-use reserve power systems activated by pyrotechnic heat generation, which transitions the solid electrolyte into a molten state. The simulation of these batteries relies on multiphysics modeling to evaluate performance and behavior under various conditions. This paper presents advancements in scalable preconditioning strategies for the Thermally Activated Battery Simulator (TABS) tool, enabling efficient solutions to the coupled electrochemical systems that dominate computational costs in thermal battery simulations. We propose a hierarchical block Gauss-Seidel preconditioner implemented through the Teko package in Trilinos, which effectively addresses the challenges posed by tightly coupled physics, including charge transport, porous flow, and species diffusion. The preconditioner leverages scalable subblock solvers, including smoothed aggregation algebraic multigrid (SA-AMG) methods and domain-decomposition techniques, to achieve robust convergence and parallel scalability. Strong and weak scaling studies demonstrate the solver’s ability to handle problem sizes up to 51.3 million degrees of freedom on 2048 processors, achieving near sub-second setup and solve times for the end-to-end electrochemical solve. These advancements significantly improve the computational efficiency and turnaround time of thermal battery simulations, paving the way for higher-resolution models and enabling the transition from 2D axisymmetric to full 3D simulations.

25 ENERGY STORAGE↗

MPAS-Seaice (v1.0.0): sea-ice dynamics on unstructured Voronoi meshes

Abstract. We present MPAS-Seaice, a sea-ice model which uses the Model for Prediction Across Scales (MPAS) framework and spherical centroidal Voronoi tessellation (SCVT) unstructured meshes. As well as SCVT meshes, MPAS-Seaice can run on the traditional quadrilateral grids used by sea-ice models such as CICE. The MPAS-Seaice velocity solver uses the elastic–viscous–plastic (EVP) rheology and the variational discretization of the internal stress divergence operator used by CICE, but adapted for the polygonal cells of MPAS meshes, or alternatively an integral (“finite-volume”) formulation of the stress divergence operator. An incremental remapping advection scheme is used for mass and tracer transport. We validate these formulations with idealized test cases, both planar and on the sphere. The variational scheme displays lower errors than the finite-volume formulation for the strain rate operator but higher errors for the stress divergence operator. The variational stress divergence operator displays increased errors around the pentagonal cells of a quasi-uniform mesh, which is ameliorated with an alternate formulation for the operator. MPAS-Seaice shares the sophisticated column physics and biogeochemistry of CICE and when used with quadrilateral meshes can reproduce the results of CICE. We have used global simulations with realistic forcing to validate MPAS-Seaice against similar simulations with CICE and against observations. We find very similar results compared to CICE, with differences explained by minor differences in implementation such as with interpolation between the primary and dual meshes at coastlines. We have assessed the computational performance of the model, which, because it is unstructured, runs with 70 % of the throughput of CICE for a comparison quadrilateral simulation. The SCVT meshes used by MPAS-Seaice allow removal of equatorial model cells and flexibility in domain decomposition, improving model performance. MPAS-Seaice is the current sea-ice component of the Energy Exascale Earth System Model (E3SM).

58 GEOSCIENCES↗

Extension of the high-resolution thermal-hydraulics code ESCOT to hexagonal core geometries for multi-physics calculations

The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

NeuroSEM: A hybrid framework for simulating multiphysics problems by coupling PINNs and spectral elements

Multiphysics problems that are characterized by complex interactions among fluid dynamics, heat transfer, structural mechanics, and electromagnetics, are inherently challenging due to their coupled nature. While experimental data on certain state variables may be available, integrating these data with numerical solvers remains a significant challenge. Physics-informed neural networks (PINNs) have shown promising results in various engineering disciplines, particularly in handling noisy data and solving inverse problems in partial differential equations (PDEs). However, their effectiveness in forecasting nonlinear phenomena in multiphysics regimes, particularly involving turbulence, is yet to be fully established. Here, this study introduces NeuroSEM, a hybrid framework integrating PINNs with the highfidelity Spectral Element Method (SEM) solver, Nektar++. NeuroSEM leverages the strengths of both PINNs and SEM, providing robust solutions for multiphysics problems. PINNs are trained to assimilate data and model physical phenomena in specific subdomains, which are then integrated into the Nektar++ solver. We demonstrate the efficiency and accuracy of NeuroSEM for thermal convection in cavity flow and flow past a cylinder. The framework effectively handles data assimilation by addressing those subdomains and state variables where the data is available. We applied NeuroSEM to the Rayleigh-B´enard convection system, including cases with missing thermal boundary conditions and noisy datasets. Finally, we applied the proposed NeuroSEM framework to real particle image velocimetry (PIV) data to capture flow patterns characterized by horseshoe vortical structures. Our results indicate that NeuroSEM accurately models the physical phenomena and assimilates the data within the specified subdomains. The framework’s plug-and-play nature facilitates its extension to other multiphysics or multiscale problems. Furthermore, NeuroSEM is optimized for efficient execution on emerging integrated GPU-CPU architectures. This hybrid approach enhances the accuracy and efficiency of simulations, making it a powerful tool for tackling complex engineering challenges in various scientific domains.

42 ENGINEERING↗