Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “precondition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

BISON Robustness and Performance Improvements

BISON is a modern finite-element based nuclear fuel performance code that has been under development at the Idaho National Laboratory (USA) since 2009 [1]. The code is applicable to both steady and transient fuel behavior and can be used to analyze 1D (spherically symmetric), 2D (axisymmetric and generalized plane strain) or 3D geometries. BISON is the fuel performance code used within CASL for LWR fuel under both normal operating and accident conditions. BISON is built using the INL Multiphysics ObjectOriented Simulation Environment, or MOOSE [2, 3]. MOOSE is a massively parallel, finite element-based framework to solve systems of coupled non-linear partial differential equations using the Jacobian-Free Newton Krylov (JFNK) method [4]. This enables investigation of computationally large problems, for example a full stack of discrete pellets in a LWR fuel rod, or every rod in a full reactor core. MOOSE supports the use of complex two and three-dimensional meshes and uses implicit time integration, important for the widely varied time scale in nuclear fuel simulation. An object-oriented architecture is employed which greatly minimizes the programming effort required to add new material and behavioral models. The flexibility of the implicit and fully coupled multiphysics approach comes with a need for constructing suitable approximations for the Jacobian matrix of the coupled system used for either preconditioning a Krylov solve or in a direct Newton solve. Preconditioning options for Bison problems need to be revisited with new preconditioning methods becoming available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Preconditioners for multiphase poromechanics with strong capillarity

This paper aims to enhance the performance of Newton–Krylov solvers for coupled poromechanical problems with two-phase flow. In particular, we investigate the impact of capillary pressure on preconditioning strategies. Capillarity complicates the coupling between the solid deformation and fluid pressure degrees of freedom, as well as increases the nonlinearity of the system. Depending on the capillary pressure relation used in the constitutive formulation, the flow equations may exhibit a spectrum of advection-dominated to diffusion-dominated behavior. We propose preconditioning approaches that account for this behavior and lead to robust numerical performance within a broad range of regimes.

42 ENGINEERING↗

On a fully-implicit VMS-stabilized FE formulation for low Mach number compressible resistive MHD with application to MCF

This study presents the development and evaluation of a fully-implicit variational multiscale (VMS) stabilized unstructured finite element (FE) formulation for compressible magnetohydrodynamics (MHD) model, at low Mach number regime. The model describes the dynamics of a compressible conducting fluid in the low Mach number limit in the presence of electromagnetic fields and can be used to study aspects of astrophysical phenomena, important science and technology applications, and basic plasma physics phenomena. The specific applications that motivate this study are macroscopic simulations of the longer time-scale stability and disruptions of magnetic confinement fusion (MCF) devices, specifically the ITER tokamak. The discussion considers the development of the VMS FE representation, the structure of the stabilizing terms that deal with significant convective flows, the stabilization of the nearly incompressible response of the fluid flow, and the stabilization of the constraint that enforces the solenoidal involution on the magnetic field. The nonlinear discretized system is solved with scalable preconditioned Newton–Krylov iterative methods, which employs a multiphysics block preconditioning method based on approximate block factorizations and Schur complements. The study presents an evaluation of the VMS method on a 2D cartesian tearing mode instability, and illustrates the scalability of the solvers on MCF relevant problems. A set of results are also presented for longer time-scale stability and disruptions for the ITER tokamak. These include a vertical displacement event (VDE), and a (1,1) internal kink mode. Here, the formulation is demonstrated to be scalable and also reasonably robust with respect to the Lundquist number scaling.

42 ENGINEERING↗

Tusas: A fully implicit parallel approach for coupled phase-field equations

In this study, we develop a fully-coupled, fully-implicit approach for phase-field modeling of solidification in metals and alloys. Predictive simulation of solidification in pure metals and metal alloys remains a significant challenge in the field of materials science, as microstructure formation during the solidification process plays a critical role in the properties and performance of the solid material. Our simulation approach consists of a finite element spatial discretization of the fully-coupled nonlinear system of partial differential equations at the microscale, which is treated implicitly in time with a preconditioned Jacobian-free Newton-Krylov method. The approach is algorithmically scalable as well as efficient due to an effective preconditioning strategy based on algebraic multigrid and block factorization. We implement this approach in the open-source Tusas framework, which is a general, flexible tool developed in C++ for solving coupled systems of nonlinear partial differential equations. The performance of our approach is analyzed in terms of algorithmic scalability and efficiency, while the computational performance of Tusas is presented in terms of parallel scalability and efficiency on emerging heterogeneous architectures. We demonstrate that modern algorithms, discretizations, and computational science, and heterogeneous hardware provide a robust route for predictive phase-field simulation of microstructure evolution during additive manufacturing.

97 MATHEMATICS AND COMPUTING↗

An Adaptive Newton-Based Free-Boundary Grad–Shafranov Solver

Equilibria in magnetic confinement devices result from force balancing between the Lorentz force and the plasma pressure gradient. In an axisymmetric configuration like a tokamak, such an equilibrium is described by an elliptic equation for the poloidal magnetic flux, commonly known as the Grad–Shafranov equation. It is challenging to develop a scalable and accurate free-boundary Grad–Shafranov solver, since it is a fully nonlinear optimization problem that simultaneously solves for the magnetic field coil current outside the plasma to control the plasma shape. In this work, we develop a Newton-based free-boundary Grad–Shafranov solver using adaptive finite elements and preconditioning strategies. The free-boundary interaction leads to the evaluation of a domain-dependent nonlinear form of which its contribution to the Jacobian matrix is achieved through shape calculus. The optimization problem aims to minimize the distance between the plasma boundary and specified control points while satisfying two nontrivial constraints, which correspond to the nonlinear finite element discretization of the Grad–Shafranov equation and a constraint on the total plasma current involving a nonlocal coupling term. The linear system is solved by a block factorization, and AMG is called for subblock elliptic operators. The unique contributions of this work include the treatment of a global constraint, preconditioning strategies, nonlocal reformulation, and the implementation of adaptive finite elements. Furthermore, it is found that the resulting Newton solver is robust, successfully reducing the nonlinear residual to 1e-6 and lower in a small handful of iterations while addressing the challenging case to find a Taylor state equilibrium where conventional Picard-based solvers fail to converge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding Electric Vehicle Range and Charging Needs: Interactions Between Ambient Temperature, Commute Patterns, and State-of-Charge Usage

Electric vehicle (EV) performance can vary substantially under real-world operating conditions, particularly due to ambient temperature effects on energy consumption, battery behavior, and thermal management requirements. This study quantifies how weather conditions, daily driving patterns, and State-of-Charge (SOC) usage strategies jointly influence EV driving range, charging frequency, and overall energy efficiency. A detailed and experimentally validated Autonomie vehicle model is developed, integrating a powertrain, a mono-zonal cabin model, and a battery electro-thermal model. Three battery sizes (200-, 300-, and 400-mile homologated ranges) are assessed across five commute profiles (20–200 miles) and six ambient temperatures (−18 °C to 50 °C), including scenarios with and without preconditioning. Results show that extreme temperatures could significantly decrease the maximum achievable range by up to 55% in cold conditions (−18 °C) and 40% in hot conditions (50 °C), relative to moderate conditions. Larger battery packs retain a greater fraction of their nominal range under thermal stress, while smaller packs experience sharper relative penalties due to the higher contribution of thermal loads to total energy demand. The analysis further demonstrates that limiting operation to partial SOC windows (e.g., 80–20%), a common real-world practice, significantly reduces achievable range and increases charging frequency, particularly in cold weather. Thermal preconditioning while plugged in is shown to mitigate these effects for short trips, reducing energy consumption by up to 31% in hot conditions and 7% in cold conditions. The findings demonstrate how climate, SOC usage behavior, and thermal management jointly shape the practical driving capability of EVs, highlighting the importance of efficient thermal management and realistic user charging strategies for ensuring reliable EV operation across diverse climatic scenarios.

33 ADVANCED PROPULSION SYSTEMS↗

An adaptive Hessian approximated stochastic gradient MCMC method

Bayesian approaches have been successfully integrated into training deep neural networks. One popular family is stochastic gradient Markov chain Monte Carlo methods (SG-MCMC), which have gained increasing interest due to their ability to handle large datasets and the potential to avoid overfitting. Although standard SG-MCMC methods have shown great performance in a variety of problems, they may be inefficient when the random variables in the target posterior densities have scale differences or are highly correlated. Here, we present an adaptive Hessian approximated stochastic gradient MCMC method to incorporate local geometric information while sampling from the posterior. The idea is to apply stochastic approximation (SA) to sequentially update a preconditioning matrix at each iteration. The preconditioner possesses second-order information and can guide the random walk of a sampler efficiently. Instead of computing and saving the full Hessian of the log posterior, we use limited memory of the samples and their stochastic gradients to approximate the inverse Hessian-vector multiplication in the updating formula. Moreover, by smoothly optimizing the preconditioning matrix via SA, our proposed algorithm can asymptotically converge to the target distribution with a controllable bias under mild conditions. To reduce the training and testing computational burden, we adopt a magnitude-based weight pruning method to enforce the sparsity of the network. Our method is user-friendly and demonstrates better learning results compared to standard SG-MCMC updating rules. The approximation of inverse Hessian alleviates storage and computational complexities for large dimensional models. Numerical experiments are performed on several problems, including sampling from 2D correlated distribution, synthetic regression problems, and learning the numerical solutions of heterogeneous elliptic PDE. The numerical results demonstrate great improvement in both the convergence rate and accuracy.

97 MATHEMATICS AND COMPUTING↗

A Comparison of Linear Solvers for Resolving Flow in Three-Dimensional Discrete Fracture Networks

We compare various methods for resolving steady flow within three-dimensional discrete fracture networks, including direct methods, Krylov subspace methods with and without preconditioning, and multi-grid methods. We compared the performance of the methods based on compute times and scaling of the solution as a function of the number of grid nodes and log-variance of the hydraulic aperture. The methods are applied to three test cases: (a) variable density of networks with a truncated power-law distribution of fracture lengths, (b) a fixed network composed of monodisperse fracture sizes but varied permeability/aperture heterogeneity, (c) and a network based on field site in Nevada, US. We chose these cases to allow us to study the impact of the mesh size and flow properties, as well as to demonstrate our conclusions on a large-scale, realistic problem (more than 40 million mesh nodes). A direct solution using Cholesky factorization outperformed other methods for every example but was closely followed in performance by some algebraic multigrid (AMG) preconditioned Krylov subspace methods. Among the Krylov methods, conjugate gradients (CG) with an AMG preconditioner performs the best. Generally, Cholesky factorization is recommended, but CG with an AMG preconditioner may be suitable for very large problems beyond 40 million nodes where the entire linear system cannot reside in memory.

58 GEOSCIENCES↗

Kinetic Model for Moisture-Controlled CO 2 Sorption

The understanding of the sorption/desorption kinetics is essential for practical applications of moisture-controlled CO 2 sorption. We introduce an analytic model of the kinetics of moisture-controlled CO 2 sorption and its interpretation in two limiting cases. In one case, chemical reaction kinetics on pore surfaces dominates, in the other case, diffusive transport through the sorbent defines the kinetics. Here, we show that reaction kinetics, which is dominant in the first case, can be expressed as a linear combination of 1st and 2nd order kinetics in agreement with the static isotherm equation derived and validated in a previous paper. The interior transport kinetics can be described by non-linear diffusion equations. By combining all carbon species into a single equation, we can eliminate — in certain limits — the source terms associated with chemical reactions. In this case, the governing equation is ∂θ/∂t = –∇ · (–D eff ∇ θ ). For a sorbent in a form of a flat sheet or a membrane, one can maintain the same functional form of a diffusion equation by introducing a generalized effective diffusivity D M that combines contributions from both surface chemical reaction kinetics and interior diffusive transport kinetics. Experimental data of transient CO 2 flux in a preconditioned commercial anion exchange membrane fit well to the 1st order model as long as very dry states are avoided, validating the theory. The observed DM for a preconditioned commercial anion exchange membrane ranges from 6.6× 10 -14 to 7.1× 10 -14 m 2 s -1 at 35°C. These small values compared to typical ionic diffusivities imply a very slow kinetics, which will be the largest issue that needs to be addressed for practical application. The collected transient CO 2 flux data are used to predict the magnitude of a continuous CO 2 pumping flux in an active membrane that transports CO 2 against a CO 2 concentration gradient. The pumped CO 2 flux is supported by water flux due to a water concentration gradient.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parametric Study of Used Nuclear Oxide Fuel Constituent Dissolution in Molten LiCl-KCl-UCl 3

Prior work identified dissolution of used nuclear oxide fuel constituents from a uranium oxide matrix into molten LiCl-KCl-UCl 3 at 500°C, prompting a subsequent series of three progressive studies (including an initial scoping study, an electrolytic dissolution study, and a chemical-seeded dissolution study) to further investigate associated parameters and mechanisms. Thermodynamic calculations were performed to identify possible reaction mechanisms and their propensities in used oxide fuel constituent dissolution. Used nuclear oxide fuels with varying preconditions from fast and thermal test reactors were separately immersed in the subject salt system to assess fuel constituent migration from the bulk fuel matrix to the salt phase in an initial scoping study. Dissolution of expected fuel constituents, including alkali, alkaline earth, lanthanide, and transuranium oxides, into the chloride salt phase varied widely, ranging from 12% to 99% in the initial study. Uranium isotope blending between the salt phase and bulk fuel matrix was also observed, which was attributed to reducing conditions in the fuel matrix. Electrolytic and chemical-seeded dissolution studies were subsequently performed to effect reducing conditions in the fuel. Other parameters, including temperature (at 500°C, 650°C, 725°C, and 800°C) and uranium trichloride concentrations (at 6, 9, and 19 wt% uranium), were investigated in the latter two studies, resulting in fuel constituent dissolution above 90%. Extents of dissolution were based on initial and final fuel constituent concentrations in the oxide fuels following operations in the salt and subsequent removal of the salt via distillation. Finally, in this series of progressive studies, oxide fuel preconditioning and in situ reducing conditions, along with elevated temperature and uranium trichloride concentrations, were the primary parameters promoting used nuclear oxide fuel constituent dissolution in accordance with identified reaction mechanisms.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Matrix-Free High-Performance Saddle-Point Solvers for High-Order Problems in \(\boldsymbol{H}(\operatorname{\textbf{div}})\)

Here, this work describes the development of matrix-free GPU-accelerated solvers for high-order finite element problems in H(div). The solvers are applicable to grad-div and Darcy problems in saddle-point formulation, and have applications in radiation diffusion and porous media flow problems, among others. Using the interpolation–histopolation basis, efficient matrix-free preconditioners can be constructed for the (1, 1)-block and Schur complement of the block system. With these approximations, block-preconditioned MINRES converges in a number of iterations that is independent of the mesh size and polynomial degree. The approximate Schur complement takes the form of an M-matrix graph Laplacian and therefore can be well-preconditioned by highly scalable algebraic multigrid methods. High-performance GPU-accelerated algorithms for all components of the solution algorithm are developed, discussed, and benchmarked. Numerical results are presented on a number of challenging test cases, including the “crooked pipe” grad-div problem, the SPE10 reservoir modeling benchmark problem, and a nonlinear radiation diffusion test case.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

Performance portable ice-sheet modeling with MALI

High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.

54 ENVIRONMENTAL SCIENCES↗

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

High-order algorithmic developments and optimizations for large-scale GPU-accelerated simulations (Milestone CEED-MS36)

The goal of this milestone was to improve the high-order software ecosystem for CEED-enabled ECP applications by making progress on efficient matrix-free kernels targeting forthcoming ECP architectures. These kernels included matrix-free preconditioning and the development of new set of CEED solver bake-off problems. As part of this milestone, we also released the next version of the CEED software stack, CEED-4.0, reported on results from several application collaborations, and documented the efforts of porting to AMD GPUs for Frontier and other modern architectures, such as Fugaku. The specific tasks addressed in this milestone were: (1) Port and run CEED benchmarks/miniapps on Frontier EA systems; (2) Demonstrate performant libCEED integration in MFEM, Nek and applications; (3) Matrix-free preconditioning of high-order operators; (4) Benchmark problems for fast high-order solvers on GPU platforms; and (5) Public release of CEED-4.0. The artifacts delivered include the next version of the CEED software stack, CEED-4.0, the next libCEED release, libCEED-0.8, and a number of developments integrated within applications to improve their GPU and CPU performance and capabilities. See the CEED website, https://ceed.exascaleproject.org and the CEED GitHub organization, https://github.com/ceed for more details.

97 MATHEMATICS AND COMPUTING↗

An Analog Preconditioner for Solving Linear Systems [Slides]

This presentation concludes in situ computation enables new approaches to linear algebra problems which can be both more effective and more efficient as compared to conventional digital systems. Preconditioning is well-suited to analog computation due to the tolerance for approximate solutions. When combined with prior work on in situ MVM for scientific computing, analog preconditioning can enable significant speedups for important linear algebra applications.

97 MATHEMATICS AND COMPUTING↗

Ensemble Simulation Techniques and Fast Randomized Algorithms

The major goals of the project were to develop and analyze new ensemble simulation techniques, including trajectory stratification and preconditioned MCMC techniques, as well as develop fast numerical linear algebra techniques closely related to ensemble simulation ideas. The trajectory stratification techniques involve simulating in parallel short trajectory fragments of a Markov process confined to a specific region of space‐time and then patching together the statistics gathered to assemble estimates of very general dynamical properties. We have also developed this approach for rare event simulation and extended the techniques to applications requiring a more general framework (such as electronic structure calculations). The preconditioned MCMC techniques involve simulating multiple Markov chains in parallel and then using information from the ensemble to speed the mixing of each individual chain. The fast randomized linear algebra methods are motivated by the diffusion Monte Carlo technique, but are applicable to finding the dominant eigenvalue of (almost) general matrices. For most non‐negative matrices, the schemes result in an error (compared to the power method) that is constant in the dimension of the problem. For more general matrices, we see a very clear sublinear cost trend in computational tests.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗