Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “direct solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Review of Lattice-Boltzmann Models Coupled with Geochemical Modeling Applied for Simulation of Advanced Waterflooding and Enhanced Oil Recovery Processes

To maintain economic profit and improve the oil production efficiency after the primary and secondary production phase, advanced waterflooding techniques such as low salinity waterflooding in carbonate reservoirs have been investigated in numerical simulations, laboratory experiments, and field pilot tests. Multiple underlying mechanisms have been proposed based on these studies, and they are still under debate. Various numerical modeling approaches are introduced, but there exists a lack of a pore-scale comprehensive modeling scheme to fully understand the processes. Lattice-Boltzmann method (LBM) is a type of numerical fluid flow modeling technique that shows capabilities and flexibilities in modeling pore-scale fluid flow to integrate physical–chemical processes within complex structures. The intrinsic feature of LBM makes it a promising framework for simulating advanced waterflooding due to its flexibility, accuracy, and parallel efficiency. LBM works either by itself for solving reactive transport problems or by coupling with a third-party reaction solver. This review mainly introduces the LBM fluid flow and reactive transport capabilities and the concept and modeling approaches to simulate advanced waterflooding techniques. Meanwhile, an evaluation of the coupled LBM models for enhanced oil recovery (EOR) simulations is discussed with future research challenges and directions concluded.

02 PETROLEUM↗

SEAS Communication Engine: An Extensible, Flexible Wrapper for Co-Simulation Agents

When modeling and analyzing the power grid and other large scale systems, researchers often express scenarios as optimization problems and feed them into advanced software solvers. In order to allow multiple solvers to communicate with each other and share data from different domains, the National Renewable Energy Laboratory (NREL) and associated Department of Energy (DOE) labs have developed a software framework called the Hierarchical Engine for Large-scale Infrastructure Co-Simulation (HELICS). HELICS allows cosimulation via a collection of client libraries for different languages that can be called from the appropriate optimization software. However, these client libraries do not provide a higher level of abstraction beyond reading and writing data off of the shared HELICS bus. In this paper, we describe a new software library called the SEAS Communication Engine that exposes a higher-level API for running cosimulation problems. The SEAS Engine provides a class-based abstraction on top of the Python HELICS client, in order to allow users to implement their domain-specific cosimulations without needing to interact with core HELICS primitives. This will make adoption of HELICS and cosimulation in general easier, by exposing a simpler API. In the second part of the paper, we validate our library on a collection of different simulation examples, including the canonical IEEE 13 Bus Feeder. Lastly, we demonstrate using the SEAS Engine to directly call domain-specific code written in the Julia programming language. Our hope is that this will serve as a template for easily calling software in different programming languages via the SEAS Engine, thereby avoiding code duplication and complexity.

co-simulation↗

Coupling between TOUGH3 and FLAC3D

Sandia National Laboratories has hired Itasca Consulting Group, Inc., the authors of the FLAC3D geomechanics software, to couple FLAC3D with TOUGH3, the porous media flow solver. The work is being done to enable a coupled mechanical-thermal-hydraulic analysis of a potential criticality event in a dual purpose cannister (DPC). The U.S. Department of Energy Office of Spent Fuel and Waste Science & Technology is investigating the performance of DPCs for direct geological disposal of spent nuclear fuel. Post closure criticality control is an important aspect of this investigation. Over geological timescales, it is envisioned that the canister and canister overpack will develop fractures due to stress corrosion processes. A breach in the canister could allow groundwater to fill the canister. Fresh water is a neutron moderator; thus, if the canister internals and fuel assemblies have been sufficiently degraded, a criticality event could occur. Such an event would release enough energy to boil the water between the fuel rods and pressurize the cannister. This internal pressurization may cause the initial fractures in the canister and overpack to grow. It is important to understand the change in hydraulic transmissivity between the canister and surroundings for two reasons: first, because it may control the potential for and frequency of subsequent criticality events; second, because it will control the release of radionuclides from the canister. The motivation for this work is to better understand the potential for periodic criticality events, cannister damage, and release of radionuclides during a criticality event in a DPC.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

SOMAFOAM: An OpenFOAM based solver for continuum simulations of low-temperature plasmas

Here, we report the development of SOMAFOAM, a finite volume framework for performing continuum simulations of low-temperature plasmas. The primary goal of this work is to discuss the features of SOMAFOAM along with representative results provided as examples for a range of operating conditions and geometries. This includes plasma and plasma–dielectric systems operating in direct current, radio frequency, and microwave regimes from pressures as low as 100 mTorr to atmospheric pressure. The code has several useful features including the ability to run massively parallel simulations using arbitrary geometries, structured/unstructured meshes, choice of models such as drift–diffusion/full-momentum at runtime, and species-dependent timesteps to name a few. The verification/validation studies presented include comparison with previously published continuum simulations (low-pressure direct current and radio frequency plasma), with experiments (Gaseous Electronics Conference Reference Cell and microwave microplasma ignited in a split ring resonator), and previously published kinetic simulations (low-pressure radio frequency plasma). Other examples provided include a direct current atmospheric pressure microplasma bounded by dielectric sidewalls and a helium–nitrogen plasma ignited using a needle electrode facing a dielectric. The performance of the code is also discussed with serial and distributed memory parallel runs demonstrated up to 512 cores. The design and implementation of the code in a modular object-oriented framework allows for easy extension and seamless coupling with other codes and can be expected to play an important role in both academia and industry.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

EchemAMR (electro-chemical microsctructure scale models with adaptive meshing) [SWR-23-111]

A 3D microstructure resolving electrochemical transport and interfacial chemistry solver. Electrode microstructure plays an important role in determining the performance of an electrochemical system, e.g. lithium ion battery. EchemAMR is a microstructure scale model that solves the governing equations for ion transport, electrical current continuity, interfacial chemistry and structural mechanics. Complex microstructure geometries from imaging can be directly imported into EchemAMR. A volume fraction based description of the geometry on Cartesian grid with an immersed interface formulation enables simplified meshing and large-scale simulations with millions of degrees of freedom. EchemAMR has been tested against systems with analytic solutions for numerical convergence and highly resolved lithium ion battery microstructures. EchemAMR demonstrates excellent mass conversation and efficient scaling on heterogenous High-Performance Computing (HPC) with central and graphics processing units.

Sitaraman, Hariswaran↗

Kinetic Plasma Simulation Capabilities in the MOOSE Framework: Verification of Particle-Particle Collisions

High-fidelity simulations of complex plasma systems allow researchers to gain key insights into and understanding of these systems. To facilitate massively parallel high-fidelity plasma simulations, finite-element-based particle-in-cell capabilities are being developed within the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) based framework called Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering & Research (SALAMANDER). While SALAMANDER’s primary objective is modeling edge plasmas and plasma-facing components in fusion devices, the particle-in-cell capabilities being developed are general and will support modeling low-temperature plasmas as well. Previously, collisionless magnetostatic simulation capabilities have been verified with the two-stream and Dorey-Guest-Harris instabilities, and single particle motion. Collisions were implemented using the direct simulation Monte Carlo method, and verification of this capability will be presented here several verification problems: relaxation of a randomly initialized gas to a Maxwellian distribution, Fourier heat flow, and comparison of reaction rates to both analytic calculations and those calculated using a multi-term Boltzmann solver.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload imbalance across the GPUs. This is particularly the case for \textit{Sparse Triangular Solver (SpTRSV)} which introduces additional two-dimensional computation dependencies among subsequent computation steps. Dependency information is exchanged and shared among GPUs, thus warrant for efficient memory allocation, data partitioning, and workload distribution as well as fine-grained communication and synchronization support. In this work, we demonstrate that directly adopting unified memory can adversely affect the performance of SpTRSV on multi-GPU architectures, despite linking via fast interconnect like NVLinks and NVSwitches. Alternatively, we employ the latest NVSHMEM technology based on Partitioned Global Address Space programming model to enable efficient fine-grained communication and drastic synchronization overhead reduction. Furthermore, to handle workload imbalance, we propose a malleable task-pool execution model which can further enhance the utilization of GPUs. By applying these techniques, our experiments on the NVIDIA multi-GPU supernode V100-DGX-1 and DGX-2 systems demonstrate that our design can achieve on average 3.53x (up to 9.86x) speedup on a DGX-1 system and 3.66x (up to 9.64x) speedup on a DGX-2 system with 4-GPUs over the Unified-Memory design. The comprehensive sensitivity and scalability studies also show that the proposed zero-copy SpTRSV is able to fully utilize the computing and communication resources of the multi-GPU system.

Xie, Chenhao↗

(Towards) DNS of a Laboratory Lean CH4/H2 Low-Swirl Flame Impinging on an Inclined Wall

Due to downsizing trends, flame-wall interaction (FWI) is increasingly prominent in gas turbines (GTs). FWI has direct consequences on flame stabilization and pollutant emissions, but it is not well understood in turbulent flows representative of GTs. We present results from a direct numerical simulation (DNS) of a turbulent CH4/H2 model GT low-swirl laboratory-scale flame interacting with an inclined wall. The results from the laboratory flame include simultaneous measurements of velocity using stereo particle imaging velocimetry and OHxCH2O planar laser induced fluorescence. The adaptive-mesh refinement solver PeleLMeX is used, with 24-species, 105-reaction reduced Aramco chemical kinetics mechanism. The premixed fuel-air mixture consists of hydrogen-enriched methane with 70% hydrogen volume fraction and 0.4 equivalence ratio. The inflow is prescribed to match experimental measurements at the burner exit. Karlovitz and turbulent Reynolds numbers are 300 and 400, respectively. The simulation and experimental results show excellent agreement. The flame features a bowl-shape stabilization, with a corrugated, continuous flame front at the leading edge, followed by fragmented reaction zones downstream. A large diffuse cloud of CH2O is formed downstream of the quenching point. The simulation results indicate that the cloud of CH2O is the result of incomplete methane combustion, with CH2O "leaking" from the locally quenched reaction zones.The DNS provides fine-grain resolution of turbulence-flame-wall interaction that cannot be captured with experimental measurements. With access to the entire solution vector at each cell of the computational domain, the local quenching.

flame-wall interactions↗

Component-Level Inverse Design of Transmon Qubits Using Neural Networks

Designing a superconducting qubit to realize specific Hamiltonian parameters typically requires iterating through a time and compute-intensive forward loop in which the designer chooses a layout geometry, simulates it, extracts circuit parameters such as capacitances, and refines the geometry. We study the inverse version of this task using a neural-network workflow that maps target Hamiltonian parameters directly to component-level layout parameters, which we subsequently demonstrate on a planar transmon layout. During training, we pair the inverse model with a frozen forward surrogate model and evaluate the loss in Hamiltonian space rather than in layout-parameter space. In validation against a conventional EM solver, 97% of generated designs produce usable geometries, and the inverse-plus-surrogate pipeline reaches mean percent errors of 0.73% for qubit frequency and 1.58% for anharmonicity, comparable to or below the fabrication and simulation-to-measurement uncertainty expected for academic-process transmon devices of this type. A single pipeline query takes ~60 ms on CPU, versus ~2 min for a conventional EM capacitance extraction on the same hardware, a speedup of approximately 2,000x. Batching minimizes the AI model inference overhead, reducing the runtime to 3.1 microseconds per sample on CPU and 2.6 microseconds per sample on GPU at a batch size of 2048, resulting in speedups of 3.9 x 10^7 and 4.6 x 10^7, respectively, relative to a single conventional CPU EM extraction. Our results indicate that component-level inverse design usefully extends and complements conventional EM simulation, including for small datasets on the order of 1,000 samples.

Seidel, Olivia [Fermilab; Texas U., Arlington]↗

Virtual Flow Solver - Geophysics: A 3D Incompressible Navier-Stokes Solver

Virtual Flow Solver - Geophysics (VFS-Geophysics) is a three-dimensional (3D) incompressible Navier-Stokes solver based on the Curvilinear Immersed Boundary (CURVIB) method. The CURVIB is a sharp interface type of immersed boundary (IB) method that enables the simulation of fluid flows in the presence of geometrically complex moving bodies. The CURVIB method can be applied to wind/MHK turbine simulations and energy applications. VFS-Geophysics is the result of many years of research work by several graduate students, post-docs, and research associates that have been involved in the Computational Hydrodynamics and Biofluids Laboratory directed by Professor Fotis Sotiropoulos. The preparation of the present manual has been supported by the U.S. Department of Energy (DE-EE 0005482).

16 TIDAL AND WAVE POWER↗

Cholla-MHD: An Exascale-capable Magnetohydrodynamic Extension to the Cholla Astrophysical Simulation Code

Abstract We present an extension of the massively parallel, GPU native, astrophysical hydrodynamics code Cholla to magnetohydrodynamics (MHD). Cholla solves the ideal MHD equations in their Eulerian form on a static Cartesian mesh utilizing the Van Leer + constrained transport integrator, the HLLD Riemann solver, and reconstruction methods at second and third order. Cholla’s MHD module can perform ≈260 million cell updates per GPU-second on an NVIDIA A100 while using the HLLD Riemann solver and second order reconstruction. The inherently parallel nature of GPUs combined with increased memory in new hardware allows Cholla’s MHD module to perform simulations with resolutions ∼500 3 cells on a single high-end GPU (e.g., an NVIDIA A100 with 80 GB of memory). We employ GPU direct Message Passing Interface to attain excellent weak scaling on the exascale supercomputer Frontier, while using 74,088 GPUs and simulating a total grid size of over 7.2 trillion cells. A suite of test problems highlights the accuracy of Cholla’s MHD module and demonstrates that zero magnetic divergence in solutions is maintained to round off error. We also present new testing and CI tools using GoogleTest, GitHub Actions, and Jenkins that have made development more robust and accurate and ensure reliability in the future.

Astronomy & Astrophysics↗

Effect of off-diagonal elements in the Wannier Hamiltonian on DFT + DMFT for low-symmetry materials: Study of Li 2 MnO 3

Here, we study the effect of the off-diagonal elements of the Wannier Hamiltonian on the electronic structure of the low-symmetry material Li 2 MnO 3 ( C2/m ), using dynamical mean field theory calculations with a continuous-time quantum Monte Carlo impurity solver. The presence of significant off-diagonal elements leads to a pronounced suppression of the energy gap. The off-diagonal elements are largest when the Wannier projection is used based on the global coordinate, and they remain substantial even with the projection using the local coordinate close to the direction of Mn-O bonds. We show that the energy gap is enhanced by the diagonalization of the Mn d block in the full p-d Hamiltonian with the application of a unitary rotation matrix. Additionally, the inclusion of small double counting energy is crucial for achieving the experimental gap by reducing p-d hybridization. Furthermore, we establish the efficiency of a low-energy (d-only basis) model for studying the electronic structure of Li 2 MnO 3 , as the Wannier basis represents a hybridized state of Mn d and O p orbitals. These findings suggest an appropriate approach for investigating low-symmetry materials using the density functional theory plus dynamical mean field theory (DFT + DMFT) method. We also find that the antiferromagnetic ground state $\Gamma$ 2u is stable with U ≤ 2 eV within density functional theory+$U$ calculations, which is much smaller than the widely used U = 5 eV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Multimodal X-ray Microscopy: The Kinetics of Cu in CdTe Absorbers

The behavior of solar cells is very often limited by inhomogeneously distributed nanoscale defects. This is the case throughout the entire lifecycle of the solar cell, from the distribution of elements and defects during solar cell growth as well as the charge-collection and recombination during operation, to degradation and failure mechanisms due to impurity diffusion, crack formation, and irradiation- and heat-induced cell damage This has been known for a while in the field of crystalline silicon, but inhomogeneities are far more abundant in polycrystalline materials, and are the limiting factor in thin-film solar cells where grain sizes are often on the order of the diffusion length. We showed that the high penetration of hard X-rays combined with the high sensitivity to elemental distribution, structure, and spatial resolution offers a unique avenue for highly correlative studies at the nanoscale. We presented results on Cu-doped CdTe, where carrier collection is directly correlated to the kinetics of Cu inside the absorber and the particular Cu-phases at the ZnTe/CdTe interface. We complemented the results with modeling using PVRD-FASP and PyCDTS, where the X-ray beam induced profile can be reproduced by the equilibrium defect concentrations proposed by the solvers.

Bertoni, Mariana↗

Evaluation of Spray and Combustion Models for Simulating Dilute Combustion in a Direct-Injection Spark-Ignition Engine

Dilute combustion in spark-ignition engines has the potential to improve thermal efficiency by mitigating knock and by reducing throttling and wall heat losses. However, ignition and combustion processes can become unstable for dilute operation due to a lowered laminar flame speed, resulting in excessive cycle-to-cycle variability (CCV) of the combustion process. To compensate for the slower combustion in less reactive mixtures, a modified intake port geometry can be employed to generate a strong tumble flow in the cylinder and elevate turbulence levels around the spark plug, thereby promoting a faster transition to turbulent deflagration. Consequently, optimizing combustion chamber geometry and operating strategy is crucial to maximizing the benefits of using dilute combustion with enhanced in-cylinder turbulence across a wide range of operating conditions. Computational fluid dynamics (CFD) simulations can be utilized for virtual engine optimization tasks, but this would require the models to be truly predictive regarding the impact of changes to the engine design and operational parameters.In this study, multicycle large-eddy simulations (LES) are performed for a direct-injection spark-ignition engine to investigate the model performance in predicting engine combustion characteristics with respect to changes in the intake configuration. A tumble plate that blocks the lower part of the intake port inlet is used to vary the tumble. A set of CFD models that have been recently developed are employed, which takes into account the drag of nonspherical droplets, flash-boiling behavior of liquid sprays, spray-wall interaction, surrogate formulation of a research-grade E10 gasoline, and fast chemical kinetic solvers. Simulation results are compared to experimental engine data in terms of cylinder pressure, apparent heat release rate, mass fraction burned timing, and flame images. It is found that LES employing the state-of-the-art CFD models are capable of properly predicting the spray processes and reproducing the measured mean cylinder pressure for the case with the tumble plate. On the other hand, the LES over-predicts the combustion rate during the early combustion stage and under-estimates the CCV, and these discrepancies become larger when the tumble plate is removed.

computational fluid dynamics simulation↗

A variational framework for residual-based adaptivity in neural PDE solvers and operator learning

Residual-based adaptive strategies are widely used in scientific machine learning yet remain largely heuristic. We introduce a variational framework that formalizes these methods through convex transformations of the residual, where different transformations correspond to distinct objective functionals. For instance, exponential weights target uniform error minimization, while linear weights recover quadratic error minimization. This perspective reveals adaptive weighting as a means of selecting sampling distributions that optimize a primal objective, directly linking discretization choices to error metrics. This principled approach yields three key benefits: it enables systematic design of adaptive schemes, reduces discretization error by lowering estimator variance, and enhances learning dynamics by improving gradient signal-to-noise ratio. Extending the framework to operator learning, we demonstrate substantial performance gains across diverse optimizers and architectures. Our results provide a theoretical perspective for residual-based adaptivity and establish a foundation for principled discretization and training.

97 MATHEMATICS AND COMPUTING↗

Cardinal: A Lower-Length-Scale Multiphysics Simulator for Pebble-Bed Reactors

This paper demonstrates a multiphysics solver for pebble-bed reactors, in particular, for Berkeley’s pebble-bed -fluoride-salt-cooled high-temperature reactor (PB-FHR) (Mark I design). The FHR is a class of advanced nuclear reactors that combines the robust coated particle fuel form from high-temperature gas-cooled reactors, the direct reactor auxiliary cooling system passive decay removal of liquid-metal fast reactors, and the transparent, high-volumetric heat capacitance liquid-fluoride salt working fluids (e.g., FLiBe) from molten salt reactors. This fuel and coolant combination enables FHRs to operate in a high-temperature, low-pressure design space that has beneficial safety and economic implications. The PB-FHR relies on a pebble-bed approach, and pebble-bed reactors are, in a sense, the poster child for multiscale analysis. Relying heavily on the MultiApp capability of the Multiphysics Object-Oriented Simulation Environment (MOOSE), we have developed Cardinal, a new platform for lower-length-scale simulation of pebble-bed cores. The lower-length-scale simulator comprises three physics: neutronics (OpenMC), thermal fluids (Nek5000/NekRS), and fuel performance (BISON). Cardinal tightly couples all three physics and leverages advances in MOOSE, such as the MultiApp system and the concept of MOOSE-wrapped applications. Moreover, Cardinal can utilize graphics processing units for accelerating solutions. In this paper, we discuss the development of Cardinal and the verification and validation and demonstration simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Deterministic-Monte Carlo Hybrid Methods for Eigenvalue Sensitivity Coefficient Calculations

The TSUNAMI suite within the SCALE code package includes several methods for generating sensitivity data, including multigroup (MG) and continuous-energy (CE) capabilities. For generating sensitivities with CE data, three methods are available in SCALE 6.3.0: (1) the iterated fission probability (IFP) method with the KENO Monte Carlo transport solver, (2) IFP with the Shift Monte Carlo transport solver, and (3) the Contributon-Linked eigenvalue sensitivity/Uncertainty estimation via Tracklength importance Characterization (CLUTCH) with the KENO Monte Carlo transport solver. Currently, it is difficult to generate accurate sensitivities with large reflectors when using the CLUTCH method, specifically with fissionable and hydrogenous materials. To address this issue, the work presented herein examines a methodology to calculate the adjoint flux externally with the 3D deterministic SN transport code DENOVO in SCALE; the result is then read directly into the CLUTCH-TSUNAMI sequence. This hybridization method replaces the Monte Carlo F*(r) calculation in CLUTCH while still utilizing the forward calculation. The critical benchmark HEU-MET-FAST-028-001 is used to generate sensitivities based on the inability of CLUTCH to generate accurate sensitivities. Results from the hybrid method appear to generate sensitivity values that are in excellent agreement with direct perturbations. Although further testing is needed, the method provides promising results for the development and utility of a hybrid method for use in TSUNAMI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗