Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Algebraic Multigrid with Filtering: An Efficient Preconditioner for Interior Point Methods in Large-Scale Contact Mechanics Optimization

Large-scale contact mechanics simulations are crucial in many engineering fields such as structural design and manufacturing. In the frictionless case, contact can be modeled by minimizing an energy functional; however, these problems are often nonlinear, nonconvex, and increasingly difficult to solve as mesh resolution increases. In this work, we employ a Newton-based interior-point (IP) filter line-search method, an effective approach for large-scale constrained optimization. While this method converges rapidly, each iteration requires solving a large saddle-point linear system that becomes ill-conditioned as the optimization process converges, largely due to IP treatment of the contact constraints. Such ill-conditioning can hinder solver scalability and increase iteration counts with mesh refinement. Here, to address this, we introduce a novel preconditioner, algebraic multigrid with filtering (AMGF), tailored to the Schur complement of the saddle-point system. Building on the classical AMG solver, commonly used for elasticity, we augment it with a specialized subspace correction that filters near null space components introduced by contact interface constraints. Through theoretical analysis and numerical experiments on a range of linear and nonlinear contact problems, we demonstrate that the proposed solver achieves mesh independent convergence and maintains robustness against the ill-conditioning that notoriously plagues IP methods. These results indicate that AMGF makes contact mechanics simulations more tractable and broadens the applicability of Newton-based IP methods in challenging engineering scenarios. More broadly, AMGF is well suited for problems, optimization or otherwise, where solver performance is limited by a low-dimensional subspace, such as those arising from localized constraints, interface conditions, or model heterogeneities. This makes the method widely applicable beyond contact mechanics and constrained optimization.

Mathematics and Computing↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

ALESQP: An Augmented Lagrangian Equality-Constrained SQP Method for Optimization with General Constraints

Here we present a new algorithm for infinite-dimensional optimization with general constraints, called ALESQP. In short, ALESQP is an augmented Lagrangian method that penalizes inequality constraints and solves equality-constrained nonlinear optimization subproblems at every iteration. The subproblems are solved using a matrix-free trust-region sequential quadratic programming (SQP) method that takes advantage of iterative, i.e., inexact linear solvers, and is suitable for large-scale applications. A key feature of ALESQP is a constraint decomposition strategy that allows it to exploit problem-specific variable scalings and inner products. We analyze convergence of ALESQP under different assumptions. We show that strong accumulation points are stationary. Consequently, in finite dimensions ALESQP converges to a stationary point. In infinite dimensions we establish that weak accumulation points are feasible in many practical situations. Under additional assumptions we show that weak accumulation points are stationary. We present several infinite-dimensional examples where ALESQP shows remarkable discretization-independent performance in all of its iterative components, requiring a modest number of iterations to meet constraint tolerances at the level of machine precision. Also, we demonstrate a fully matrix-free solution of an infinite-dimensional problem with nonlinear inequality constraints.

97 MATHEMATICS AND COMPUTING↗

Delayed Neutron Precursor Group Parameter and Spectra Generation from Fast Fission of 235U in SCALE

Delayed neutron precursor (DNP) group data are important for modeling reactor dynamics. Although the data for individual DNPs have been developed over time, the DNP group data present in the Evaluated Nuclear Data Files (ENDF) have not been updated in the past 20 years. In this work, we use SCALE to recreate the Godiva experiment that was used to generate the original DNP group structure for fast fission of 235U. However, each DNP is modeled using up-to-date data, and the results are then converted into a newly updated group structure. This conversion uses an iterative linear least squares solver to minimize chi-squared. The approaches used in this work also enable energy spectrum generation and uncertainty tracking. The method used in this paper for fast 235U fission DNP group structure updating can be applied to different energies and fissile nuclides. Demonstration of the uncertainty tracking in reactor kinetics and dynamics simulations is shown using point reactor kinetics simulations. Results show that there are data discrepancies between the International Atomic Energy Agency database and data used in ORIGEN, which are currently being fixed. Results also show that the proposed method for group spectra generation performs well.

Seifert, Luke [University of Illinois]↗

In situ multi-tier auto-ignition detection applied to dual-fuel combustion simulations

Here we use an anomaly detection methodology that is centered on analyzing fourth-order joint moments (co-kurtosis), particularly focusing on its application in auto-ignition of combustion problems with large numbers of species. Unsupervised anomaly detection is challenging to generalize across problem types and domains. A recent technique, centered on analyzing information in the fourth-order joint moment co-kurtosis, has shown promise, especially for high-dimensional scientific data. In this work we present developments to the co-kurtosis based anomaly detection method needed to make it effective and scalable for large-scale distributed scientific data, such as those generated by massively parallel simulations. An in situ co-kurtosis algorithm is employed as the anomaly detection method for identifying ignition kernels in simulations of turbulent combustion. Here, we extend an existing methodology which identifies regions of the domain where anomalies are present, and add another tier of anomaly detection where the individual samples contributing to the anomaly are identified. We apply this algorithm on-the-fly to a variety of turbulent reacting flow problems and compare it to the widely used (but significantly more expensive) chemical explosive mode analysis (CEMA). We demonstrate the ability of the method to detect and identify the onset of low and high temperature ignition which can be used for computational steering, as chemical and combustion anomalies occur intermittently at spatio-temporal locations unknown a priori. Finally, we apply our lightweight in situ algorithm to an exascale high-fidelity simulation with a total of 2.4 Trillion degrees of freedom, performed using an adaptive mesh refinement solver. Furthermore, through a scalability analysis, we show that the relative computational cost of this in-situ anomaly detection algorithm compared to an iteration of the reacting flow solver is negligible.

97 MATHEMATICS AND COMPUTING↗

Simulating energetic ions and enhanced fusion rates from ion-cyclotron resonance heating with a full-wave/Fokker–Planck model

Reproducing fast-ion enhanced fusion rates from ion-cyclotron resonance heating (ICRH) in tokamaks requires the self-consistent coupling of a full-wave solver and a Fokker–Planck solver, which evolves multiple simultaneously resonant ion species. We introduce a new self-consistent model that iterates the TORIC full-wave solver with the CQL3D Fokker–Planck solver using the integrated plasma simulator (IPS). This model evolves the bounce-averaged ion distribution functions in both parallel and perpendicular velocity-space with a quasilinear radio frequency (RF) diffusion operator valid in the ion finite Larmor radius (FLR) limit and the RF electric fields with the resultant non-Maxwellian FLR dielectric tensor. This produces non-Maxwellian ICRH simulations that are fully self-consistent, fast, and interoperable with integrated modeling frameworks, such as TRANSP/GACODE/IPS-FASTRAN. We demonstrate our model's capabilities by validating it against experimental data in Alcator C-Mod. We then perform the first RF heating simulations of SPARC using self-consistent non-Maxwellian ion distributions to investigate the potential to enhance fusion rates using ion cyclotron resonance heating generated fast ions.

Physics↗

Accelerated coupled Monte Carlo-Thermal hydraulic calculations using a hybrid GTF-diffusion-based prediction block: first results

Accurate predictions of spatial power and temperature distributions require the coupling of a neutron transport solver with a thermal-hydraulic (TH) feedback. Nowadays, Monte Carlo (MC) codes are widely coupled to TH solvers, typically via a Picard iteration (PI) method, due to the higher fidelity that such frameworks can produce. To speed up a PI, a prediction step can produce an improved initial guess for a source distribution and feed it to the MC code. Recent work investigated a prediction step that uses generalized transfer functions (GTFs) to predict the macroscopic cross sections' variations following a perturbation in TH properties, such as coolant density. The previous method also relied on first order perturbation (FOP) theory to predict perturbed power profiles, rather than using an expensive MC iterate. The implemented FOP method relied on generating a fission matrix from which the forward and adjoint Eigenmodes were extracted and later used to for power calculations. The generation of the fission matrix can introduce a significant computational overhead, therefore undermining the performance of the proposed hybrid technique when applied to high-dimensional problems, e.g., full core calculations. This work attempts to improve the GTF-FOP prediction step by replacing the FOP solver with a nodal diffusion solver, thus eliminating the need to calculate a fission matrix. The GTF-diffusion step was tested for various moderator density perturbations. In each case, the predicted power distribution showed good agreement with the reference case. The latter is attributed to the generally good prediction of most spatially distributed macroscopic cross sections, except the transport cross section, which will become the focus of future work. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Accelerated Coupled Monte Carlo-Thermal Hydraulic Calculations using a Hybrid GTF-Diffusion-based Prediction Block: First Results

Accurate predictions of spatial power and temperature distributions require the coupling of a neutron transport solver with a thermal-hydraulic (TH) feedback. Nowadays, Monte Carlo (MC) codes are widely coupled to TH solvers, typically via a Picard iteration (PI) method, due to the higher fidelity that such frameworks can produce. To speed up a PI, a prediction step can produce an improved initial guess for a source distribution and feed it to the MC code. Recent work [1, 2] investigated a prediction step that uses generalized transfer functions (GTFs) to predict the macroscopic cross sections’ variations following a perturbation in TH properties, such as coolant density. The previous method also relied on first order perturbation (FOP) theory to predict perturbed power profiles, rather than using an expensive MC iterate. The implemented FOP method relied on generating a fission matrix from which the forward and adjoint eigenmodes were extracted and later used to for power calculations. The generation of the fission matrix can introduce a significant computational overhead, therefore undermining the performance of the proposed hybrid technique when applied to high-dimensional problems, e.g., full core calculations. This work attempts to improve the GTF-FOP prediction step by replacing the FOP solver with a nodal diffusion solver, thus eliminating the need to calculate a fission matrix. The GTF-diffusion step was tested for various moderator density perturbations. In each case, the predicted power distribution showed good agreement with the reference case. The latter is attributed to the generally good prediction of most spatially distributed macroscopic cross sections, except the transport cross section, which will become the focus of future work.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Fast meta-solvers for 3D complex-shape scatterers using neural operators trained on a non-scattering problem

Three-dimensional target identification using scattering techniques requires high accuracy solutions and very fast computations for real-time predictions in some critical applications. We first train a deep neural operator (DeepONet) to solve wave propagation problems described by the Helmholtz equation in a domain without scatterers but at different wavenumbers and with a complex absorbing boundary condition. We then design two classes of fast meta-solvers by combining DeepONet with either relaxation methods, such as Jacobi and Gauss-Seidel, or with Krylov methods, such as GMRES and BiCGStab, using the trunk basis of DeepONet as a coarse-scale preconditioner. We leverage the spectral bias of neural networks to account for the lower part of the spectrum in the error distribution while the upper part is handled inexpensively using relaxation methods or fine-scale preconditioners. The meta-solvers are then applied to solve scattering problems with different shape of scatterers, at no extra training cost. We first demonstrate that the resulting meta-solvers are shape-agnostic, fast, and robust, whereas the standard standalone solvers may even fail to converge without the DeepONet. We then apply both classes of meta-solvers to scattering from a submarine, a complex three-dimensional problem. We achieve very fast solutions, especially with the DeepONet-Krylov methods, which require orders of magnitude fewer iterations than any of the standalone solvers.

97 MATHEMATICS AND COMPUTING↗

Data transfers for full core heterogeneous reactor high- fidelity multiphysics studies

Multiphysics simulations for nuclear reactor analysis are usually performed by resorting to operator splitting and fixed point iterations between single-physics solvers. This enables the separate solution of each physics, such as neutronics, fuel performance, and thermal hydraulics, on meshes tailored to the requirements of the respective numerical discretizations of the equations. As the equations are coupled, several fields must be transferred between single-physics solves. Projecting fields between meshes while preserving order of accuracy, conservation properties, and mapping non-overlapping geometries is a complex endeavor. This conference paper will present the transfers as implemented in MOOSE, which can handle arbitrary meshes, arbitrary mappings, conservation of integral quantities, and are made to scale with distributed simulations on both ends of the transfers. Their adequacy for advanced nuclear reactor multiphysics coupling is shown through examples and numerical studies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multigrid Optimization for Large-Scale Ptychographic Phase Retrieval

Ptychography is a popular imaging technique that combines diffractive imaging with scanning microscopy. The technique consists of a coherent beam that is scanned across an object in a series of overlapping positions, leading to reliable and improved reconstructions. Ptychographic microscopes allow for large fields to be imaged at high resolution at the cost of additional computational expense. Herein, we propose a multigrid-based optimization framework to reduce the computational burdens of large-scale ptychographic phase retrieval. Our proposed method exploits the inherent hierarchical structures in ptychography through tailored restriction and prolongation operators for the object and data domains. Our numerical results show that our proposed scheme accelerates the convergence of its underlying solver and outperforms the ptychographical iterative engine, a workhorse in the optics community.

97 MATHEMATICS AND COMPUTING↗

PyAMG: Algebraic Multigrid Solvers in Python

PyAMG is a Python package of algebraic multigrid (AMG) solvers and supporting tools for approximating the solution to large, sparse linear systems of algebraic equations, Ax = b, where A is an n × n sparse matrix. Sparse linear systems arise in a range of problems in science, from fluid flows to solid mechanics to data analysis. While the direct solvers available in SciPy’s sparse linear algebra package (scipy.sparse.linalg) are highly efficient, in many cases iterative methods are preferred due to overall complexity. However, the iterative methods in SciPy, such as CG and GMRES, often require an efficient preconditioner in order to achieve a lower complexity. Preconditioning is a powerful tool whereby the conditioning of the linear system and convergence rate of the iterative method are both dramatically improved. PyAMG constructs multigrid solvers for use as a preconditioner in this setting. A summary of multigrid and algebraic multigrid solvers can be found in Olson (2015a), in Olson (2015b), and in Falgout (2006); a detailed description can be found in Briggs et al. (2000) and Trottenberg et al. (2001).

97 MATHEMATICS AND COMPUTING↗

A Performance and Energy Study of GPU-Resident Preconditioners for Conjugate Gradient Solvers: In the Context of Existing and Novel Approaches

Optimizing a particular subprogram out of the set of Basic (sparse) Linear Algebra Subprograms (BLAS) for a given architecture is a common topic of research. In applications, however, these BLAS functions rarely appear in isolation; usually, many of them are used together, in various combinations and with varying inputs. As the need to solve a large, sparse linear system is ubiquitous throughout HPC applications, linear solvers constitute a realistic, sufficiently complex and well-defined representative use case for composite BLAS routines. To this end, based on a representative set of matrices drawn from a diverse set of fields, we present a framework to study, from the performance and energy perspective, the efficacy of GPU- resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process. The development of this preconditioner was motivated by solving large graph Laplacian linear systems, for which the existing preconditioners either perform slow on GPU-based platforms or are not applicable. We compare the performance of these preconditioners on different hardware accelerator architectures, i.e., AMD MI250X, MI100, Nvidia A100, V100, and Jetson. Our experiments reveal performance trade-offs and provide information on how to select the best strategy for the given linear system, dictated by its properties, and the platform of interest. We demonstrate the application of our novel preconditioner for solving CG and graph Laplacian systems. Overall, the framework can be utilized as a benchmark to guide informed decisions in choosing a specific preconditioner, i.e., whether it is better to rely on the performance of a triangular solver or on the performance of sparse matrix-vector product. Finally, by considering power consumption to solve the linear systems, we report the energy footprint for the solvers.

Preconditioned Conjugate Gradient, GPUs, iterative↗

Quantum microgrid state estimation

This paper investigates the feasibility and efficiency of quantum-circuit-based algorithms for microgrid state estimation. Here, our new contributions include: (1) a general quantum state estimation (GQSE) formulation is devised for swing-bus-contained microgrids through the quantized Gaussian–Newton iteration, (2) a preconditioned quantum linear solver (PQLS) is developed for tackling the ill-conditioned GQSE with limited quantum resources, and (3) an enhanced quantum state estimation (EQSE) algorithm is further established for hierarchical-control-based microgrids with exogenous disturbances. Extensive case studies demonstrate the correctness of GQSE, PQLS and EQSE in two typical microgrids. The robustness and convergence performance of EQSE are also verified.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Jacobian-free Newton–Krylov method for the simulation of non-thermal plasma discharges with high-order time integration and physics-based preconditioning

A preconditioning framework for the numerical simulation of non-thermal streamer discharges is developed using the Jacobian-free Newton-Krylov (JFNK) method. A reduced plasma fluid model is considered, consisting of electrons, one positive ion, one negative ion, and the electrostatic potential. Here, the plasma kinetics model includes ionization, electron-ion recombination, electron attachment, electron detachment, and ion-ion recombination. The governing equations are made dimensionless, discretized in space with finite differences, and integrated in time with a fully implicit method based on high-order backward differentiation formulas. The preconditioning framework is based on a linearized form of the governing equations and physics-based operator splitting. The efficiency of the preconditioning strategy is assessed through two test cases: streamer propagation between parallel plates and an axisymmetric pin-to-pin discharge. The fully implicit approach overcomes traditional restrictions in the time step size due to processes such as electron drift, electron diffusion, and dielectric relaxation. Excellent performance is observed through relevant statistics of the JFNK solver, although the number of linear iterations increases for the pin-to-pin discharge when nonlinear numerical boundary conditions are imposed at the electrodes. Performance studies show scalability with O(100-1000) processors for O(10M) unknowns with ample room for optimization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Predicting Flow in Fracture Networks With Quantum Algorithms

Uncertainty quantification plays a crucial role in the modeling of subsurface flow. For instance, uncertainties in the properties of geologic fracture networks significantly impact flow, requiring numerous simulations to accurately estimate quantities of interest. However, each simulation is computationally expensive because it requires solving a large linear system to capture features that involve both small and large fractures. An example is in percolation, where the interaction of many small fractures (which cumulatively can have a large surface area) with the rock matrix must be modeled precisely. Quantum computing is an emerging tool with the potential to address this issue. Quantum algorithms offer a significant speedup in solving linear systems, achieving efficiencies that are challenging to match with classical approaches. These classical approaches include direct solvers, such as LU decomposition, and iterative methods, notably preconditioned conjugate gradient, commonly used in subsurface modeling to solve large sparse systems. However, applying quantum algorithms to geologic fracture flow requires careful attention to algorithmic and problem-specific constraints to fully realize this quantum advantage. In this work we describe a quantum algorithm for generalized Monte Carlo applications with a quadratic speedup over the classical approaches which can be combined with the quantum speedup, currently under investigation, for solving quantum linear systems for subsurface flow. We show that for quantum algorithms the computational cost of estimating a quantity of interest for a statistical ensemble of networks is roughly the same as that of a single realization, essentially implying that one can get uncertainty quantification for free.

58 GEOSCIENCES↗

Game theoretic modeling and optimization of competition and collaboration in dual channel electronic waste supply chains

The rapid growth of electronic waste (e-waste) presents critical challenges for sustainable resource recovery and environmental protection. This study develops a dual-channel closed-loop supply chain (CLSC) model formulated as a hierarchical Stackelberg game, that integrates dynamic pricing and cost-sharing mechanisms to optimize both economic and environmental outcomes. The model explicitly captures strategic interactions between manufacturer-led and third-party recycling channels, accounting for consumer behavior, regulatory incentives, and market competition. Numerical simulations conducted (implemented over a four-iteration horizon using a commercial optimization solver) show that, relative to the baseline equilibrium, manufacturer profit increases from 11.6 thousand USD to 37.9 thousand USD (+226.8%), total recycled volume rises from 7,848 to 7,942 units (+1.2%), and collector profit nearly doubles under cost-sharing, enabling more equitable profit distribution. Furthermore, scenario-based simulations across Sub-Saharan Africa, high-income economies, and emerging Asian industrial countries reveal that infrastructure quality, policy intensity, and labor costs critically shape recycling efficiency and profit allocation. These findings demonstrate that subsidies alone are insufficient to ensure system efficiency. Instead, coordinated strategies that integrate internal incentive alignment with context-sensitive policy support are required. Overall, this study offers a robust framework for designing resilient, efficient, and regionally adaptable e-waste management systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Discrete modeling of a transformer with ALEGRA

We report progress on a task to model transformers in ALEGRA using the “Transient Magnetics” option. We specifically evaluate limits of the approach resolving individual coil wires. There are practical limits to the number of turns in a coil that can be numerically modeled, but calculated inductance can be scaled to the correct number of turns in a simple way. Our testing essentially confirmed this “turns scaling” hypothesis. We developed a conceptual transformer design, representative of practical designs of interest, and that focused our analysis. That design includes three coils wrapped around a rectangular ferromagnetic core. The secondary and tertiary coils have multiple layers. The tertiary has three layers of 13 turns each; the secondary has five layers of 44 turns; the primary has one layer of 20 turns. We validated the turns scaling of inductance for simple (one-layer) coils in air (no core) by comparison to available independent calculations for simple rectangular coils. These comparisons quantified the errors versus reduced number of turns modeled. For more than 3 turns, the errors are <5%. The magnetic field solver failed to converge (within 5000 iterations) for >10 turns. Including the core introduced some complications. It was necessary to capture the core surfaces in thin grid sheaths to minimize errors in computed magnetic energy. We do not yet have quantitative benchmarks with which to compare, but calculated results are qualitatively reasonable.

42 ENGINEERING↗