Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Effect of non-uniform void distributions on the yielding of metals

High-throughput (several thousand) calculations have been carried out to investigate the yield behavior of porous materials with randomly distributed pores, porosity levels over four orders of magnitude and up to a hundred pores per simulation box. To this end, a Galerkin based fast Fourier transform (FFT) formulation was enhanced to deal with high phase contrast materials. In addition, GPU parallelization was employed in solving the governing equation for strain fluctuations using a Krylov iterative solver. Emphasis is laid on the conditions under which percolation of plastically non-deforming zones through the porous network emerge, a regime termed unhomogeneous yielding. By way of contrast, the regime where the plastic strain fluctuations (associated with the heterogeneous void-matrix aggregate) fall below the percolation threshold is defined as homogeneous yielding. Here, we find that nonuniform pore distributions only affect unhomogeneous yielding and have a universal softening effect. The extent of this distribution softening is analyzed as a function of porosity, cell size and number of realizations. Whether the uncovered universal distribution softening has direct implications on failure resistance of porous materials is discussed.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A thermodynamic model of integrated liquid-to-liquid thermoelectric heat pump systems

Thermoelectric (TE) heat pumps (TEHPs) are advantageous for heating and cooling in various applications because of their modularity and simple design. A TEHP system includes the TE modules with p- and n-type materials bonded to substrates, plus heat exchangers, thermal interfaces to the heat exchangers, and heat transfer fluids. Although modeling an individual TE module has been extensively studied, limited studies have reported performance at the larger system-level. Furthermore, no prior study has addressed the impact of temperature-dependent TE material properties (e.g., electric resistivity, thermal conductivity, and Seebeck coefficient) on overall heat-pump-system-level performance. Here, this work presents a mathematical model for TEHP system performance based on Goldsmid's approach for TE material performance, “effective” TE material properties, Gnielinski's correlation for convective heat transfer, and thermal balance theory for a heat exchange network. This combined approach provides an accurate model of the liquid-to-liquid TEHP system. Three different approaches—one empirical, one based on the manufacturer's specifications, and one drawn from the literature—were then used to determine values for TE material properties. The first two methods treated properties as constants, while the last approach treated properties as surface-temperature-based functions. Finally, experimental TEHP data was used to validate the models, all with relative absolute deviations of approximately 10% when predicting heating capacity and 10%–25% when forecasting cooling capacity up to a 30 K surface temperature lift. The results demonstrated that, at the TEHP system level, the TE material properties could be treated as constants, avoiding solver iterations and reducing the performance uncertainty by up to 95%.

42 ENGINEERING↗

NSFnets (Navier-Stokes flow nets): Physics-informed neural networks for the incompressible Navier-Stokes equations

In the last 50 years there has been a tremendous progress in solving numerically the Navier-Stokes equations using finite differences, finite elements, spectral, and even meshless methods. Yet, in many real cases, we still cannot incorporate seamlessly (multi-fidelity) data into existing algorithms, and for industrial-complexity applications the mesh generation is time consuming and still an art. Moreover, solving ill-posed problems (e.g., lacking boundary conditions) or inverse problems is often prohibitively expensive and requires different formulations and new computer codes. Here, we employ physics-informed neural networks (PINNs), encoding the governing equations directly into the deep neural network via automatic differentiation, to overcome some of the aforementioned limitations for simulating incompressible laminar and turbulent flows. We develop the Navier-Stokes flow nets (NSFnets) by considering two different mathematical formulations of the Navier-Stokes equations: the velocity-pressure (VP) formulation and the vorticity-velocity (VV) formulation. Since this is a new approach, we first select some standard benchmark problems to assess the accuracy, convergence rate, computational cost and flexibility of NSFnets; analytical solutions and direct numerical simulation (DNS) databases provide proper initial and boundary conditions for the NSFnet simulations. The spatial and temporal coordinates are the inputs of the NSFnets, while the instantaneous velocity and pressure fields are the outputs for the VP-NSFnet, and the instantaneous velocity and vorticity fields are the outputs for the VV-NSFnet. This is unsupervised learning and, hence, no labeled data are required beyond boundary and initial conditions and the fluid properties. The residuals of the VP or VV governing equations, together with the initial and boundary conditions, are embedded into the loss function of the NSFnets. No data is provided for the pressure to the VP-NSFnet, which is a hidden state and is obtained via the incompressibility constraint without extra computational cost. Unlike the traditional numerical methods, NSFnets inherit the properties of neural networks (NNs), hence the total error is composed of the approximation, the optimization, and the generalization errors. Here, we empirically attempt to quantify these errors by varying the sampling (“residual”) points, the iterative solvers, and the size of the NN architecture. For the laminar flow solutions, we show that both the VP and the VV formulations are comparable in accuracy but their best performance corresponds to different NN architectures. The initial convergence rate is fast but the error eventually saturates to a plateau due to the dominance of the optimization error. For the turbulent channel flow, we show that NSFnets can sustain turbulence at , but due to expensive training we only consider part of the channel domain and enforce velocity boundary conditions on the subdomain boundaries provided by the DNS data base. We also perform a systematic study on the weights used in the loss function for balancing the data and physics components, and investigate a new way of computing the weights dynamically to accelerate training and enhance accuracy. In the last part, we demonstrate how NSFnets should be used in practice, namely for ill-posed problems with incomplete or noisy boundary conditions as well as for inverse problems. We obtain reasonably accurate solutions for such cases as well without the need to change the NSFnets and at the same computational cost as in the forward well-posed problems. As a result, we also present a simple example of transfer learning that will aid in accelerating the training of NSFnets for different parameter settings.

97 MATHEMATICS AND COMPUTING↗

DG-IMEX method for a two-moment model for radiation transport in the $\mathscr{O}$($v$/$c$) limit

Here, we consider neutral particle systems described by moments of a phase-space density and propose a realizability-preserving numerical method to evolve a spectral two-moment model for particles interacting with a background fluid moving with nonrelativistic velocities. The system of nonlinear moment equations, with special relativistic corrections to $\mathscr{O}$($v$/$c$), expresses a balance between phase-space advection and collisions and includes velocity-dependent terms that account for spatial advection, Doppler shift, and angular aberration. The model is conservative for the correct $\mathscr{O}$($v$/$c$) Eulerian-frame number density and is consistent, to $\mathscr{O}$($v$/$c$), with Eulerian-frame energy and momentum conservation. This model is closely related to the one promoted by Lowrie et al. and similar to models currently used to study transport phenomena in large-scale simulations of astrophysical environments. The proposed numerical method is designed to preserve moment realizability, which guarantees that the moments correspond to a nonnegative phase-space density. The realizability-preserving scheme consists of the following key components: (i) a strong stability-preserving implicit-explicit (IMEX) time-integration method; (ii) a discontinuous Galerkin (DG) phase-space discretization with carefully constructed numerical uxes; (iii) a realizability-preserving implicit collision update; and(iv) a realizability-enforcing limiter. In time integration, nonlinearity of the moment model necessitates solution of nonlinear equations, which we formulate as fixed-point problems and solve with tailored iterative solvers that preserve moment realizability with guaranteed global convergence. We also analyze the simultaneous Eulerian-frame number and energy conservation properties of the semi-discrete DG scheme and propose a "spectral redistribution" scheme that promotes Eulerian-frame energy conservation. Through numerical experiments, we demonstrate the accuracy and robustness of this DG-IMEX method and investigate its Eulerian-frame energy conservation properties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simulating self-powered neutron detector responses to infer burnup-induced power distribution perturbations in next-generation light water reactors

Understanding how 3D power distribution will be monitored throughout reactor core volumetric space in next-generation nuclear power reactors is crucial to the design, deployment, and licensing of these reactors. Although numerous techniques exist for 3D power distribution monitoring based on the response of both in situ and ex situ sensors currently implemented or proposed for use in the US reactor fleet, crucial details about these techniques are often unclear. The publicly available documentation does not include information such as how well these techniques are characterized and optimized in their implementations and the levels of uncertainty in the inferred 3D power distribution. The work described herein investigated a recently developed 3D power distribution inferencing method as applied to two next-generation reactor simulations: (1) the NuScale small modular reactor design and (2) the Westinghouse AP1000 design, both of which contain in-core strings of vanadium self-powered neutron detectors (SPNDs). This investigation considered a range of SPND string sensor densities, as well as a range of 3D power distribution axial segment sizes. In this work, SPND response simulation is informed by neutron flux calculations in representative homogenized cores. For the different sensor densities and power distribution axial segment sizes in these simulations, the average solution error, solver iterations, and run time were tracked to parameterize the sensor-core configuration.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Robust scalable initialization for Bayesian variational inference with multi-modal Laplace approximations

Predictive modeling typically relies on Bayesian model calibration to provide uncertainty quantification. Variational inference utilizing fully independent (“mean-field”) Gaussian distributions are often used as approximate probability density functions. This simplification is attractive since the number of variational parameters grows only linearly with the number of unknown model parameters. However, the resulting diagonal covariance structure and unimodal behavior can be too restrictive to provide useful approximations of intractable Bayesian posteriors that exhibit highly non-Gaussian behavior, including multimodality. High-fidelity surrogate posteriors for these problems can be obtained by considering the family of Gaussian mixtures. Gaussian mixtures are capable of capturing multiple modes and approximating any distribution to an arbitrary degree of accuracy, while maintaining some analytical tractability. Unfortunately, variational inference using Gaussian mixtures with full-covariance structures suffers from a quadratic growth in variational parameters with the number of model parameters. The existence of multiple local minima due to strong nonconvex trends in the loss functions often associated with variational inference present additional complications, These challenges motivate the need for robust initialization procedures to improve the performance and computational scalability of variational inference with mixture models. In this work, we propose a method for constructing an initial Gaussian mixture model approximation that can be used to warm-start the iterative solvers for variational inference. The procedure begins with a global optimization stage in model parameter space. In this step, local gradient-based optimization, globalized through multistart, is used to determine a set of local maxima, which we take to approximate the mixture component centers. Around each mode, a local Gaussian approximation is constructed via the Laplace approximation. Finally, the mixture weights are determined through constrained least squares regression. The robustness and scalability of the proposed methodology is demonstrated through application to an ensemble of synthetic tests using high-dimensional, multimodal probability density functions. Here, the practical aspects of the approach are demonstrated with inversion problems in structural dynamics.

97 MATHEMATICS AND COMPUTING↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An unconditionally stable, time-implicit algorithm for solving the one-dimensional Vlasov–Poisson system

The development of an implicit, unconditionally stable, numerical method for solving the Vlasov–Poisson system in one dimension using a phase-space grid is presented. The algorithm uses the Crank–Nicolson discretization scheme and operator splitting allowing for direct solution of the finite difference equations. This method exactly conserves particle number, enstrophy and momentum. A variant of the algorithm which does not use splitting also exactly conserves energy but requires the use of iterative solvers. This algorithm has no dissipation and thus fine-scale variations can lead to oscillations and the production of negative values of the distribution function. We find that overall, the effects of negative values of the distribution function are relatively benign. We consider a variety of test cases that have been used extensively in the literature where numerical results can be compared with analytical solutions or growth rates. We examine higher-order differencing and construct higher-order temporal updates using standard composition methods.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Shadow molecular dynamics for flexible multipole models

Shadow molecular dynamics provide an efficient and stable atomistic simulation framework for flexible charge models with long-range electrostatic interactions. Shadow molecular dynamics simulations are driven by approximate “shadow” Born–Oppenheimer potentials for which the exact charges and forces are directly accessible without relying on costly (and approximate) iterative solvers. While previous implementations have been limited to atomic monopole charge distributions, we extend this approach to flexible multipole models. We derive detailed expressions for the shadow energy functions, potentials, and force terms, explicitly incorporating monopole–monopole, dipole–monopole, and dipole–dipole interactions. In our formulation, both atomic monopoles and atomic dipoles are treated as extended dynamical variables alongside the propagation of the nuclear degrees of freedom. We demonstrate that introducing the additional dipole degrees of freedom preserves the stability and accuracy previously seen in monopole-only shadow molecular dynamics simulations. In addition, we present a shadow molecular dynamics scheme where the monopole charges are held fixed while the dipoles remain flexible. Our extended shadow dynamics provide a framework for stable, computationally efficient, and versatile molecular dynamics simulations involving long-range interactions between flexible multipoles. This is of particular current interest in combination with machine-learned interatomic potentials, including long-range electrostatic interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Particle-Tracking Proton Computed Tomography—Data Acquisition, Preprocessing, and Preconditioning

Proton CT (pCT) is a promising new imaging technique that can reconstruct relative stopping power (RSP) more accurately than x-ray CT in each cubic millimeter voxel of the patient. This, in turn, will result in better proton range accuracy and, therefore, smaller planned tumor volumes (PTV). The hardware description and some reconstructed images have previously been reported. In a series of two contributions, we focus on presenting the software algorithms that convert pCT detector data to the final reconstructed pCT images for application in proton treatment planning. There were several options on how to accomplish this, and we will describe our solutions at each stage of the data processing chain. In the first paper of this series, we present the data acquisition with the pCT tracking and energy-range detectors and how the data are preprocessed, including the conversion to the well-formatted track information from tracking data and water-equivalent path length from the data of a calibrated multi-stage energy-range detector. These preprocessed data are then used for the initial image formation with an FDK cone-beam CT algorithm. The output of data acquisition, preprocessing, and FDK reconstruction is presented along with illustrative imaging results for two phantoms, including a pediatric head phantom. The second paper in this series will demonstrate the use of iterative solvers in conjunction with the superiorization methodology to further improve the images resulting from the upfront FDK image reconstruction and the implementation of these algorithms on a hybrid CPU/GPU computer cluster.

42 ENGINEERING↗

Energy Exascale Computational Fluid Dynamics Simulations With the Spectral Element Method

Development and application of the open-source GPU-based fluid-thermal simulation code, NekRS, are described. Time advancement is based on an efficient kth-order accurate timesplit formulation coupled with scalable iterative solvers. Spatial discretization is based on the high-order spectral element method (SEM), which affords the use of fast, low-memory, matrix-free operator evaluation. Further, recent developments include support for nonconforming meshes using overset grids and for GPU-based Lagrangian particle tracking. Results of large-eddy simulations of atmospheric boundary layers for wind-energy applications as well as extensive nuclear energy applications are presented.

42 ENGINEERING↗

SDE_quark

This program provides a fixed-grid quadrature algorithm to compute integrals in the self-energy of the quark propagator within the Maris-Tandy model. The quark propagator both on t he spacelike real axis and at complex-valued momenta is determined from its Schwinger-Dyson equation (SDE). We first apply an iterative solver to find the quark propagator on the spacelike real axis. The propagator at complex-valued momenta is then computed from its self-energy based on this solution, where demanding integrals are encountered. In order to compute of these integrals, we apply customized variable transformations for the radial integral after subtracting the asymptotics. We subsequently apply an optional compound of quadrature rules for the angular integral.

Jia, Shaoyang↗

Graph-based Reversible Evaluation and Tangents Library

GRETL is a C++ library for evaluation, re-evaluation and algorithmic differentiation of functional operations on an arbitrary computational graph with limited memory usage. Similar to popular machine learning frameworks in Python, like PyTorch and JAX, it tracks and stores both operations and output data as functions are evaluated. Once this composition of functions is built up, the entire chain of operations can be back propagated to compute sensitivities of the final result with respect to any number of inputs. In contrast to most machine learning applications, memory usage becomes the bottleneck for back propagation in many physics applications, especially for time-dependent PDEs. Dynamic check pointing becomes essential. An important distinguishing feature of GRETL is its ability to limit the maximum memory usage by automatically dynamic checkpointing the data output for each graph operation (see Wang, Moin, Iaccarino, 2009). During backpropagation, parts of the graph that are no longer in memory are automatically re-evaluated from upstream checkpointed states as needed for derivative sensitivity calculations (or more precisely, for vector-Jacobian products). GRETL is particularly beneficial for applications, such as coupled multi-physics, where deriving adjoint-based sensitivities and managing checkpoint memory across modules becomes onerous. Cases which can be readily handled by the GRETL library include: different time-integration algorithms per physics (e.g., coupled predictor-corrector algorithms, IMEX, etc.), sub-cycling, asynchronous integrators, state dependent timestep sizes, iterative solvers and coupling algorithms, controller algorithms, and more.

Tupek, MichaelR [Lawrence Livermore National Labor↗

Scientific Core Library Stack (SCLS) v2026

SCLS (Scientific Core Library Stack) is an opinionated build and packaging system for scientific computing libraries developed at Lawrence Berkeley National Laboratory. It produces a coherent, reproducible stack of numerical libraries — including BLAS/LAPACK, MPI, sparse direct and iterative solvers, graph partitioners, and parallel I/O libraries (e.g., PETSc, SLEPc, HDF5, NetCDF, MUMPS, OpenBLAS) — that work together without manual repair by downstream scientific software. From a single recipe-and-flavor model, SCLS produces native RPM packages for RHEL-family Linux, DEB packages for Debian/Ubuntu, direct Unix-style prefix installs for HPC and locked-down environments, and native macOS builds. Multiple build "flavors" (e.g., GCC+OpenBLAS, GCC+MKL, Intel+MKL, debug) coexist in distinct prefixes on the same host. Compared to general-purpose meta-build frameworks, SCLS is deliberately curated rather than infinitely configurable. It enforces deterministic, audit-friendly behavior: explicit build dependencies, no silent feature autodetection, a clear open-source license policy, and rpath-based runtime linkage so installs integrate cleanly with standard package-manager workflows.

Messe, Christian [Lawrence Berkeley National Labor↗

CUDO: closed-form universal dwell-time optimization for computer-controlled optical surfacing

Precision optical figuring demands fast and accurate dwell time optimization to reach nanometer- and sub-nanometer-level accuracy in next-generation optical systems. We introduce CUDO (closed-form universal dwell-time optimization), the first, to the best of our knowledge, unified closed-form analytical framework that supports both function-form and matrix-form dwell time models in computer-controlled optical surfacing (CCOS). In contrast to traditional methods, which rely on iterative optimization and hyperparameter tuning, our framework derives direct analytical solutions with no adjustable parameters. This approach unifies the solution principles of existing methods within a single mathematical model, delivering three key advantages: (1) accuracy on par with, or superior to, iterative solvers, (2) substantial reduction in computation time, and (3) numerical robustness. Comparative studies with prior art confirm that closed-form solutions achieve equivalent residual error while removing runtime bottlenecks. By simplifying the implementation and enabling real-time, scalable deployment, CUDO establishes a practical foundation for future deterministic fabrication of large-aperture and high-performance optics.

36 MATERIALS SCIENCE↗

Sierra/SolidMechanics 4.58 User's Guide

Sierra / SolidMechanics (Sierra / SM) is a Lagrangian, three-dimensional code for finite element analysis of solids and structures. It provides capabilities for explicit dynamic, implicit quasistatic and dynamic analyses. The explicit dynamics capabilities allow for the efficient and robust solution of models with extensive contact subjected to large, suddenly applied loads. For implicit problems, Sierra / SM uses a multi-level iterative solver, which enables it to effectively solve problems with large deformations, nonlinear material behavior, and contact. Sierra / SM has a versatile library of continuum and structural elements, an d a large library of material models. The code is written for parallel computing environments enabling scalable solutions of extremely large problems for both implicit and explicit analyses. It is built on the SIERRA Framework, which facilitates coupling with other SIERRA mechanics codes . This document describes the functionality and input syntax for Sierra / SM.

36 MATERIALS SCIENCE↗

Sierra/SolidMechanics 5.0 User's Guide

Sierra/SolidMechanics (Sierra/SM) is a Lagrangian, three-dimensional code for finite element analysis of solids and structures. It provides capabilities for explicit dynamic, implicit quasistatic and dynamic analyses. The explicit dynamics capabilities allow for the efficient and robust solution of models with extensive contact subjected to large, suddenly applied loads. For implicit problems, Sierra/SM uses a multi-level iterative solver, which enables it to effectively solve problems with large deformations, nonlinear material behavior, and contact. Sierra/SM has a versatile library of continuum and structural elements, and a large library of material models. The code is written for parallel computing environments enabling scalable solutions of extremely large problems for both implicit and explicit analyses. It is built on the SIERRA Framework, which facilitates coupling with other SIERRA mechanics codes. This document describes the functionality and input syntax for Sierra/SM.

97 MATHEMATICS AND COMPUTING↗

Mixed Precision Numerical Linear Algebra (Final Report)

The objective of this subcontract was to identify opportunities for the use of mixed precision within iterative solvers and to explore these opportunities both theoretically and experimentally. All quarterly milestones were achieved. We summarize the achievements each quarter in sections below.

97 MATHEMATICS AND COMPUTING↗