Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Shock-capturing parabolized Navier-Stokes model /SCIPVIS/ for the analysis of turbulent underexpanded jets

A new computational model, SCIPVIS, has been developed to predict the multiple-cell wave/shock structure in under or over-expanded turbulent jets. SCIPVIS solves the parabolized Navier-Stokes jet mixing equations utilizing a shock-capturing approach in supersonic regions of the jet and a pressure-split approach in subsonic regions. Turbulence processes are represented by the solution of compressibility corrected two-equation turbulence models. The formation of Mach discs in the jet and the interactive turbulent mixing process occurring behind the disc are handled in a detailed fashion. SCIPVIS presently analyzes jets exhausting into a quiescent or supersonic external stream for which a single-pass spatial marching solution can be obtained. The iterative coupling of SCIPVIS with a potential flow solver for the analysis of subsonic/transonic external streams is under development.

Dash, S. M.↗

Improving the convergence rate to steady state of parabolic ADI methods

The present, residuals' L(2)-norms analysis of the rate of convergence to steady state for parabolic ADI solvers allows the prediction of the number of iterations required for convergence, as a function of the Courant number alpha. A modification of current ADI codes is presented which significantly improves the convergence rate and is insensitive to the Courant number over a large range of alpha. This corrected algorithm is tested for the cases of Dirichlet problems for uniform grids of many mesh sizes, mixed Dirichlet-Neumann problems, and problems defined on stretched grids and/or problems with variable coefficients.

Abarbanel, Saul S.↗

Optimal model reduction and frequency-weighted extension

In this paper the quadratically optimal model reduction problem for single-input, single-output systems is considered. The reduced order model is determined by minimizing the integral of the magnitude-squared of the transfer function error. It is shown that the numerator coefficients of the optimal approximant satisfy a weighted least squares problem and, on this basis, a two-step iterative algorithm is developed combining a least squares solver with a gradient minimizer. The existence of globally optimal stable solutions to the optimization problem is established, and convergence of the algorithm to stationary values of the cost function is proved. The formulation is extended to handle the frequency-weighted optimal model reduction problem. Three examples demonstrate the optimization algorithm.

Spanos, J. T.↗

A new algorithm for L2 optimal model reduction

In this paper the quadratically optimal model reduction problem for single-input, single-output systems is considered. The reduced order model is determined by minimizing the integral of the magnitude-squared of the transfer function error. It is shown that the numerator coefficients of the optimal approximant satisfy a weighted least squares problem and, on this basis, a two-step iterative algorithm is developed combining a least squares solver with a gradient minimizer. Convergence of the proposed algorithm to stationary values of the quadratic cost function is proved. The formulation is extended to handle the frequency-weighted optimal model reduction problem. Three examples demonstrate the optimization algorithm.

Spanos, J. T.↗

An Adaptive Newton-Based Free-Boundary Grad–Shafranov Solver

Equilibria in magnetic confinement devices result from force balancing between the Lorentz force and the plasma pressure gradient. In an axisymmetric configuration like a tokamak, such an equilibrium is described by an elliptic equation for the poloidal magnetic flux, commonly known as the Grad–Shafranov equation. It is challenging to develop a scalable and accurate free-boundary Grad–Shafranov solver, since it is a fully nonlinear optimization problem that simultaneously solves for the magnetic field coil current outside the plasma to control the plasma shape. In this work, we develop a Newton-based free-boundary Grad–Shafranov solver using adaptive finite elements and preconditioning strategies. The free-boundary interaction leads to the evaluation of a domain-dependent nonlinear form of which its contribution to the Jacobian matrix is achieved through shape calculus. The optimization problem aims to minimize the distance between the plasma boundary and specified control points while satisfying two nontrivial constraints, which correspond to the nonlinear finite element discretization of the Grad–Shafranov equation and a constraint on the total plasma current involving a nonlocal coupling term. The linear system is solved by a block factorization, and AMG is called for subblock elliptic operators. The unique contributions of this work include the treatment of a global constraint, preconditioning strategies, nonlocal reformulation, and the implementation of adaptive finite elements. Furthermore, it is found that the resulting Newton solver is robust, successfully reducing the nonlinear residual to 1e-6 and lower in a small handful of iterations while addressing the challenging case to find a Taylor state equilibrium where conventional Picard-based solvers fail to converge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

An installed nacelle design code using a multiblock Euler solver. Volume 2: User guide

This is a user manual for the general multiblock Euler design (GMBEDS) code. The code is for the design of a nacelle installed on a geometrically complex configuration such as a complete airplane with wing/body/nacelle/pylon. It consists of two major building blocks: a design module developed by LaRC using directive iterative surface curvature (DISC); and a general multiblock Euler (GMBE) flow solver. The flow field surrounding a complex configuration is divided into a number of topologically simple blocks to facilitate surface-fitted grid generation and improve flow solution efficiency. This user guide provides input data formats along with examples of input files and a Unix script for program execution in the UNICOS environment.

Chen, H. C.↗

Parallel Implicit Algorithms for CFD

The main goal of this project was efficient distributed parallel and workstation cluster implementations of Newton-Krylov-Schwarz (NKS) solvers for implicit Computational Fluid Dynamics (CFD.) "Newton" refers to a quadratically convergent nonlinear iteration using gradient information based on the true residual, "Krylov" to an inner linear iteration that accesses the Jacobian matrix only through highly parallelizable sparse matrix-vector products, and "Schwarz" to a domain decomposition form of preconditioning the inner Krylov iterations with primarily neighbor-only exchange of data between the processors. Prior experience has established that Newton-Krylov methods are competitive solvers in the CFD context and that Krylov-Schwarz methods port well to distributed memory computers. The combination of the techniques into Newton-Krylov-Schwarz was implemented on 2D and 3D unstructured Euler codes on the parallel testbeds that used to be at LaRC and on several other parallel computers operated by other agencies or made available by the vendors. Early implementations were made directly in Massively Parallel Integration (MPI) with parallel solvers we adapted from legacy NASA codes and enhanced for full NKS functionality. Later implementations were made in the framework of the PETSC library from Argonne National Laboratory, which now includes pseudo-transient continuation Newton-Krylov-Schwarz solver capability (as a result of demands we made upon PETSC during our early porting experiences). A secondary project pursued with funding from this contract was parallel implicit solvers in acoustics, specifically in the Helmholtz formulation. A 2D acoustic inverse problem has been solved in parallel within the PETSC framework.

Keyes, David E.↗

Squash-Box Feasibility Driven Differential Dynamic Programming

Recently, Differential Dynamic Programming (DDP) and other similar algorithms have become the solvers of choice when performing non-linear Model Predictive Control (nMPC) with modern robotic devices. The reason is that they have a lower computational cost per iteration when compared with off-the-shelf Non-Linear Programming (NLP) solvers, which enables its online operation. However, they cannot handle constraints, and are known to have poor convergence capabilities. In this paper, we propose a method to solve the optimal control problem with control bounds through a squashing function (i.e., a sigmoid, which is bounded by construction). It has been shown that a naive use of squashing functions damage the convergence rate. To tackle this, we first propose to add a quadratic barrier that avoids the difficulty of the plateau produced by the sigmoid. Second, we add an outer loop that adapts both the sigmoid and the barrier; it makes the optimal control problem with the squashing function converge to the original control-bounded problem. To validate our method, we present simulation results for different types of platforms including a multi-rotor, a biped, a quadruped and a humanoid robot.

Navarro, Angel Santamaria↗

WARP3D-Release 10.8: Dynamic Nonlinear Analysis of Solids using a Preconditioned Conjugate Gradient Software Architecture

This report describes theoretical background material and commands necessary to use the WARP3D finite element code. WARP3D is under continuing development as a research code for the solution of very large-scale, 3-D solid models subjected to static and dynamic loads. Specific features in the code oriented toward the investigation of ductile fracture in metals include a robust finite strain formulation, a general J-integral computation facility (with inertia, face loading), an element extinction facility to model crack growth, nonlinear material models including viscoplastic effects, and the Gurson-Tver-gaard dilatant plasticity model for void growth. The nonlinear, dynamic equilibrium equations are solved using an incremental-iterative, implicit formulation with full Newton iterations to eliminate residual nodal forces. The history integration of the nonlinear equations of motion is accomplished with Newmarks Beta method. A central feature of WARP3D involves the use of a linear-preconditioned conjugate gradient (LPCG) solver implemented in an element-by-element format to replace a conventional direct linear equation solver. This software architecture dramatically reduces both the memory requirements and CPU time for very large, nonlinear solid models since formation of the assembled (dynamic) stiffness matrix is avoided. Analyses thus exhibit the numerical stability for large time (load) steps provided by the implicit formulation coupled with the low memory requirements characteristic of an explicit code. In addition to the much lower memory requirements of the LPCG solver, the CPU time required for solution of the linear equations during each Newton iteration is generally one-half or less of the CPU time required for a traditional direct solver. All other computational aspects of the code (element stiffnesses, element strains, stress updating, element internal forces) are implemented in the element-by- element, blocked architecture. This greatly improves vectorization of the code on uni-processor hardware and enables straightforward parallel-vector processing of element blocks on multi-processor hardware.

Koppenhoefer, Kyle C.↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Carbon Organisms Rhizosphere and Protection in Soil Environment model script and input data for soil moisture-respiration responses in tropical forests

Objectives: Climatic drying is predicted for many tropical forests, yet models remain poorly parameterized for tropical forests, hampering predictions of forest-climate feedbacks. We applied an integrated model–experiment approach, parameterizing an ecosystem model Carbon Organisms Rhizosphere and Protection in the Soil Environment (CORPSE) with tropical forest observational data, and comparing model predictions with a field drying manipulation. We hypothesized that drying would suppress soil CO2 fluxes (i.e., respiration) in already-drier tropical forests, but increases CO2 fluxes in wetter tropical forests by alleviating anaerobiosis. We measured soil CO2 fluxes, soil moisture, soil temperature, and forest floor biomass during wet-dry cycles (2015 – 2022) in four Panamanian forests that vary in rainfall and soil fertility. We used the field data to parameterize and run tests in the model.Results: Measured CO2 fluxes declined in the dry season and peaked in the early wet season ahead of peak soil moisture, resulting in a lower soil moisture optimum for respiration than previously modeled. We used this data to parameterize the model, which then predicted increased soil CO2 fluxes in wetter and fertile forests with drying, and decreased fluxes in drier, infertile forests. In contrast to model predictions, a chronic throughfall exclusion experiment in the forests initially suppressed soil CO2 fluxes across forests, with sustained suppression after four years in the wettest forest only (-28 ± 4% during the dry season), but elevated soil CO2 fluxes in a fertile forest after four years (+75 ± 28% during the late wet season), as predicted by the model. The unexpected negative drying effect in the wettest, most infertile forest could have resulted from reduced vertical flushing of nutrients into soils. Including hydro-nutrient interactions in ecosystem models could improve predictions of tropical forest-climate feedbacks (results presented in Cusack et al. 2023). Datasets included: Code files:CORPSE_array.py: Defines the equations of the CORPSE modelCORPSE_solvers: Functions for running the CORPSE model using either iterative or ordinary differential equation (ODE) solversrun_Panama_sims.py: Read in datasets and run the model simulations for this studyInput data:PanamaGradientEcosystemChem_BT_CPools_20152016CO2_DC_20190615.xlsx: Plot characteristics used in running model simulationsLiCor compiled surface flux only to 2020_03 DC_20200825.xlsx: Surface gas exchange fluxes used in model-data comparisonsPARCHED litterfall data for Ben Sulman LD 20200902.xlsx: Litterfall data used to drive model simulationsInitialization data:state_500y_20190823.csv: Initial state of model pools based on previous spinup runsOutput data:Outputs/prev_moisture_response.csv: Simulations of multiple sites using original model moisture response function.Outputs/updated_moisture_response.csv: Simulations of multiple sites using updated model moisture response function.Outputs/dry15_prev_moisture_response.csv: Simulations with soil moisture reduced by 15%, using original moisture response function.Outputs/dry15_updated_moisture_response.csv: Simulations with soil moisture reduced by 15%, using updated moisture response function.Outputs/dry30_prev_moisture_response.csv: Simulations with soil moisture reduced by 30%, using original moisture response function.Outputs/dry30_updated_moisture_response.csv: Simulations with soil moisture reduced by 30%, using updated moisture response function.Outputs/latestart_prev_moisture_response.csv: Simulations with extended dry season, using original moisture response function.Outputs/latestart_updated_moisture_response.csv: Simulations with extended dry season, using updated moisture response function.Outputs/[site name]_oneyear.csv: One-year simulation for each site in expanded site list using original moisture response function.Outputs/[site name]_oneyear_dried.csv: One-year simulation for each site in expanded site list using original moisture response function, with soil moisture reduced by 25%.Outputs/[site name]_oneyear_updated_moisture_response.csv: One-year simulation for each site in expanded site list using updated moisture response function.Outputs/[site name]_oneyear_updated_moisture_response_dried.csv: One-year simulation for each site in expanded site list using updated moisture response function, with soil moisture reduced by 25%.Field plot location data:There is also a .kml file that includes coordinates for all 32 plots included in the study of four forests (n = 4 throughfall reduction and n = 4 control plots per site).

54 ENVIRONMENTAL SCIENCES↗

ULTRA-SHARP nonoscillatory convection schemes for high-speed steady multidimensional flow

For convection-dominated flows, classical second-order methods are notoriously oscillatory and often unstable. For this reason, many computational fluid dynamicists have adopted various forms of (inherently stable) first-order upwinding over the past few decades. Although it is now well known that first-order convection schemes suffer from serious inaccuracies attributable to artificial viscosity or numerical diffusion under high convection conditions, these methods continue to enjoy widespread popularity for numerical heat transfer calculations, apparently due to a perceived lack of viable high accuracy alternatives. But alternatives are available. For example, nonoscillatory methods used in gasdynamics, including currently popular TVD schemes, can be easily adapted to multidimensional incompressible flow and convective transport. This, in itself, would be a major advance for numerical convective heat transfer, for example. But, as is shown, second-order TVD schemes form only a small, overly restrictive, subclass of a much more universal, and extremely simple, nonoscillatory flux-limiting strategy which can be applied to convection schemes of arbitrarily high order accuracy, while requiring only a simple tridiagonal ADI line-solver, as used in the majority of general purpose iterative codes for incompressible flow and numerical heat transfer. The new universal limiter and associated solution procedures form the so-called ULTRA-SHARP alternative for high resolution nonoscillatory multidimensional steady state high speed convective modelling.

Leonard, B. P.↗

Some Remarks on GMRES for Transport Theory

We review some work on the application of GMRES to the solution of the discrete ordinates transport equation in one-dimension. We note that GMRES can be applied directly to the angular flux vector, or it can be applied to only a vector of flux moments as needed to compute the scattering operator of the transport equation. In the former case we illustrate both the delights and defects of ILU right-preconditioners for problems with anisotropic scatter and for problems with upscatter. When working with flux moments we note that GMRES can be used as an accelerator for any existing transport code whose solver is based on a stationary fixed-point iteration, including transport sweeps and DSA transport sweeps. We also provide some numerical illustrations of this idea. We finally show how space can be traded for speed by taking multiple transport sweeps per GMRES iteration. Key Words: transport equation, GMRES, Krylov subspace

Patton, Bruce W.↗

Procedure for Determining One-Dimensional Flow Distributions in Arbitrarily Connected Passages Without the Influence of Pumping

A calculation procedure is presented which allows the one-dimensional determination of flow distributions in arbitrarily connected (branching) flow passages having multiple inlets and exits. The procedure uses an adaptation of the finite element technique, iteratively coupled with an accurate one-dimensional flow solver. The procedure eliminates the usual restrictions inherent with finite element flow calculations. Unlike existing one-dimensional methods, which require simplifications to the flow equations (uncoupling the momentum and energy equations), to allow for arbitrary branching and multiple inlets and exits, the only limitation of the described methodology is that, at present, it can only accommodate non-rotating configurations (no pumping effects). The calculation procedure is robust, and will always converge for physically possible flow. The procedure is described, and its use is illustrated by an example.

Meitner, Peter L.↗

Face- and Cell-Averaged Nodal-Gradient Approach to Cell-Centered Finite-Volume Method on Mixed Grids

In this paper, the averaged nodal-gradient approach previously developed for triangular grids is extended to mixed triangular-quadrilateral grids. It is shown that the face- averaged approach leads to deteriorated iterative convergence on quadrilateral grids. To develop a convergent solver, we consider cell-averaging instead of face-averaging for quadri- lateral cells. We show that the cell-averaged approach leads to a convergent solver and can be efficiently combined with the face-averaged approach on mixed grids. The method is demonstrated for various inviscid and viscous problems from low to high Mach numbers on two-dimensional mixed grids.

Nishikawa, Hiroaki↗