Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Krylov solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Scalable Multiphysics Block Preconditioning for Low Mach Number Compressible Resistive MHD with Application to Magnetic Confinement Fusion

This study investigates multiphysics block preconditioners that are critical in devising scalable Newton–Krylov iterative solvers for longer time-scale fully implicit fluid plasma models. The specific model of interest is the visco-resistive, low Mach number, compressible magnetohydrodynamics (MHD) model. This model describes the dynamics of conducting fluids in the presence of electromagnetic fields and can be used to study aspects of astrophysical phenomena, important science and technology applications, and basic plasma physics. The specific application of interest that motivates this study is the macroscopic simulation of longer time-scale stability and disruptions of magnetic confinement fusion devices, specifically the ITER Tokamak. The computational solution of the governing balance equations for mass, momentum, heat transfer, and magnetic induction for resistive MHD systems can be extremely challenging. These difficulties arise from both the strong nonlinear, nonsymmetric coupling of fluid and electromagnetic phenomena as well as the significant range of time and length scales that the interactions of these physical mechanisms produce. To handle the range of time and spatial scales of interest, a fully implicit unstructured variational multiscale finite element formulation is employed. For the scalable solution of the Newton linearized systems, fully coupled block preconditioners are designed to leverage algebraic multigrid subsolves. In conclusion, results are presented for the strong and weak scaling of the method as well as the robustness of these techniques for a large range of Lundquist numbers.

97 MATHEMATICS AND COMPUTING↗

Reproduced Computational Results Report for “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing”

The article titled “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing” by Anzt et al. presents a modern, linear operator centric, C++ library for sparse linear algebra. Experimental results in the article demonstrate that Ginkgo is a flexible and user-friendly framework capable of achieving high-performance on state-of-the-art GPU architectures. In this report, the Ginkgo library is installed and a subset of the experimental results are reproduced. Specifically, the experiment that shows the achieved memory bandwidth of the Ginkgo Krylov linear solvers on NVIDIA A100 and AMD MI100 GPUs is redone and the results are compared to what presented in the published article. Upon completion of the comparison, the published results are deemed reproducible.

97 MATHEMATICS AND COMPUTING↗

Globalized Newton-Krylov-Schwarz Algorithms and Software for Parallel Implicit CFD

Implicit solution methods are important in applications modeled by PDEs with disparate temporal and spatial scales. Because such applications require high resolution with reasonable turnaround, "routine" parallelization is essential. The pseudo-transient matrix-free Newton-Krylov-Schwarz (Psi-NKS) algorithmic framework is presented as an answer. We show that, for the classical problem of three-dimensional transonic Euler flow about an M6 wing, Psi-NKS can simultaneously deliver: globalized, asymptotically rapid convergence through adaptive pseudo- transient continuation and Newton's method-, reasonable parallelizability for an implicit method through deferred synchronization and favorable communication-to-computation scaling in the Krylov linear solver; and high per- processor performance through attention to distributed memory and cache locality, especially through the Schwarz preconditioner. Two discouraging features of Psi-NKS methods are their sensitivity to the coding of the underlying PDE discretization and the large number of parameters that must be selected to govern convergence. We therefore distill several recommendations from our experience and from our reading of the literature on various algorithmic components of Psi-NKS, and we describe a freely available, MPI-based portable parallel software implementation of the solver employed here.

Gropp, W. D.↗

Investigation of a Smooth Local Correlation-based Transition Model in a Discrete-Adjoint Aerodynamic Shape Optimization Algorithm

A smooth local correlation-based transition model is fully coupled to a RANS-based Newton-Krylov flow solver and discrete-adjoint gradient-based optimization algorithm. The free-transition optimization framework is evaluated using lift-constrained drag minimizations of airfoils at design conditions ranging from light to single-aisle aircraft and an infinite swept wing at design conditions representative of a transonic strut-braced wing aircraft. The impact of the streamwise grid resolution on the ability of the optimization algorithm to delay boundary-layer transition is investigated, with the results demonstrating that streamwise grid resolution requirements increase as the transition length decreases with increasing Reynolds number. The optimization problem at the light aircraft design conditions is demonstrated to be multi-modal, with the optimization algorithm producing two distinct designs: one with a thin, reflexed trailing edge and steep pressure recovery regions, the other with increased aft loading, with the latter design outperforming the former. A drag minimization of an airfoil at transonic design conditions demonstrates that the optimization algorithm successfully trades a decrease in viscous drag by delaying boundary-layer transition with an increase in wave drag, while the drag minimization of an infinite swept wing demonstrates the capability of the optimizational gorithm to delay both Tollmien-Schlichting and stationary crossflow instabilities.

AATT↗

Efficient Solution of Three-Dimensional Problems of Acoustic and Electromagnetic Scattering by Open Surfaces

We present a computational methodology (a novel Nystrom approach based on use of a non-overlapping patch technique and Chebyshev discretizations) for efficient solution of problems of acoustic and electromagnetic scattering by open surfaces. Our integral equation formulations (1) Incorporate, as ansatz, the singular nature of open-surface integral-equation solutions, and (2) For the Electric Field Integral Equation (EFIE), use analytical regularizes that effectively reduce the number of iterations required by iterative linear-algebra solution based on Krylov-subspace iterative solvers.

sound-soft acoustic scattering↗

PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU Applications

Iterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS.

Zhang, Lingqi↗

Incompressible Navier-Stokes, GMRES, and Multi-Element Airfoils

The talk will discuss current work in developing and applying the INS2D and INS3D flow solvers to problems in high-lift aerodynamics, primarily multi-element airfoils. These flow solvers solve the incompressible Navier-Stokes equations using the method of artificial compressibility. High-lift system flowfields are perhaps the most challenging aeronautical configurations for CFD. Difficulties with convergence for multi-element configurations using fine-grid overset meshes has led to research in using Krylov-space iterative matrix solvers. In particular, a Generalized Minimum Residual (GMRES) solver has found to be a dramatic improvement over a previously used Gauss-Seidal line-relaxation solver.

Rogers, Stuart↗

Improvements to the Griffin Transport Solvers

Griffin is a Multiphysics Object-Oriented Simulation Environment (MOOSE) based reactor multiphysics analysis application jointly developed by Idaho National Laboratory and Argonne National Laboratory. The code includes a variety of steady-state solvers for fixed-source, k-eigenvalue, adjoint, and subcritical multiplication, as well as transient solvers for point-kinetics, improved quasi-static, and spatial dynamics. This document summarizes the transport solver development efforts pursued during Fiscal Year 2022. We added the multiphysics transient capability for the coarse-mesh finite difference accelerated Richardson iteration for discontinuous finite element method discrete ordinates (DFEM-SN) scheme to support high-order heterogeneous transport simulations. HFEM (hybrid finite element method) - PN (spherical harmonics expansion) was completed and red-black iteration was added for solving the HFEM-PN system with both preconditioned Jacobian-free Newton Krylov and Richardson iteration solvers. The HFEM-PN scheme, as one of the low-order transport schemes, is expected for supporting routine design simulations. Pin power reconstruction capability was also designed and implemented with the Griffin ISOXML module to enhance all the low-order transport solvers for more accurate multiphysics simulations. Numerical results are presented for demonstrating the capabilities and verifying their performance, and future works are discussed.

97 MATHEMATICS AND COMPUTING↗

FEOTS v0.0.0: a new offline code for the fast equilibration of tracers in the ocean

Abstract. In this paper we introduce a new software framework for the offline calculation of tracer transport in the ocean. The Fast Equilibration of Ocean Tracers Software (FEOTS) is an end-to-end set of tools to efficiently calculate tracer distributions on a global or regional sub-domain using transport operators diagnosed from a comprehensive ocean model. To the best of our knowledge, this is the first application of a transport matrix model to an eddying ocean state. While a Newton–Krylov-based equilibration capability is still under development and not presented here, we demonstrate in this paper the transient modeling capabilities of FEOTS in an application focused on the Argentine Basin, where intense eddy activity and the Zapiola Anticyclone lead to strong mixing of water masses. The demonstration illustrates progress in developing offline passive tracer simulation capabilities, while highlighting the challenges of the impulse response functions approach in capturing tracer transports by a non-linear advection scheme. Our future work will focus on improving the computational efficiency of the code to reduce time-to-solution, using different basis functions to better represent non-linear advection operators, applying FEOTS to a parent model with unstructured grids (Ocean Model for Prediction Across Scales, MPAS-Ocean), and fully implementing a Newton–Krylov steady-state solver.

54 ENVIRONMENTAL SCIENCES↗

Towards an Automated Unstructured Grid Adaptation Workflow with VULCAN

Early work is presented for an unstructured grid adaptation workflow with VULCAN and refine. Anisotropic simplex grids are iteratively adapted to match a Riemannian metric tensor field describing desired mesh spacing. The Riemannian metric tensor field is obtained from Hessians of CFD solution output scalar sensor fields; both Mach number and static temperature sensor fields are explored. In addition, we describe a Newton-method-based solver recently implemented in VULCAN utilizing Jacobian-Free-Newton-Krylov that can be used to increase flow solver automation on early grids in the adadptation process. Hypersonic flow solutions are presented on a high Reynolds number flat plate and wall heat flux is compared against a highly resolved structured solution. Additionally, complex shock boundary-layer interaction is explored in a high Mach number compression corner and complex 3D flow phenomena are evaluated on the Boundary Layer Transition (BOLT) vehicle.

Matthew O'Connell↗

Krylov methods preconditioned with incompletely factored matrices on the CM-2

The performance is measured of the components of the key interative kernel of a preconditioned Krylov space interative linear system solver. In some sense, these numbers can be regarded as best case timings for these kernels. Sweeps were timed over meshes, sparse triangular solves, and inner products on a large 3-D model problem over a cube shaped domain discretized with a seven point template. The performance of the CM-2 is highly dependent on the use of very specialized programs. These programs mapped a regular problem domain onto the processor topology in a careful manner and used the optimized local NEWS communications network. The rather dramatic deterioration in performance was documented when these ideal conditions no longer apply. A synthetic workload generator was developed to produce and solve a parameterized family of increasingly irregular problems.

Berryman, Harry↗

eddy Users Manual

eddy is a collection of tools - nonlinear solvers, meshing, post-processing, visualization, optimization, etc. - for performing scale-resolving simulations of multi-physics applications. The framework is designed to enable advanced R&D on a variety of topics by leveraging a mature capability for scale resolving simulations, and simultaneously be an appropriate tool for application analysis and support. Currently, eddy is at a relatively low technical readiness level (TRL), and users and developers should maintain appropriate expectations. The technical details behind eddy are outlined in several publications which can be consulted for more information [1–10]. The solvers are built around an unstructured high-order capability, and heavily utilize the tensor product sum-factorization approach for efficiency. The unsteady formulation utilizes a fully implicit space-time approach with a matrix-free Newton- Krylov method. A primitive steady-state solver is available for testing purposes, but is not expected to converge for all but simple verification cases. The Navier-Stokes fluid solvers do not support either RANS or hybrid-RANS capability, only LES and wall-modeled LES approaches. All of the solvers within eddy support three modes of operation: a primal solve of the full nonlinear problem, and two linearization approaches of the primal solve - the ad joint and the tangent solution. Details on how to select and use these three modes are outlined in Sec. 3.

Murman, Scott M.↗

An implicit particle code with exact energy and charge conservation for electromagnetic studies of dense plasmas

A collisional particle code based on implicit energy- and charge-conserving methods is presented. A modified version of the particle-suppressed Jacobian-Free Newton-Krylov method that can enhance the solver efficiency is introduced. Mathematically, it is shown that this new approach can be viewed as a fixed-point iteration method for the particle positions. The model can exactly conserve global energy and local charge and can efficiently use time steps larger than the plasma period. In conclusion, the algorithm's ability to simulate dense plasmas accurately and efficiently is quantified by simulating the dynamic compression of a plasma slab via a magnetic piston in 1D planar geometry.

97 MATHEMATICS AND COMPUTING↗

Implicit Preconditioning for Explicit Multigrid Solvers on Cut-Cell Cartesian Meshes

This work assesses the effectiveness of linearized implicit Euler preconditioning for multigrid solvers using an unpreconditioned, Jacobian-free Newton Krylov method to converge the linear system of equations. Multigrid convergence rates improve to approximately 0.75 across the cases tested including a Mach 2 supersonic wedge, transonic NACA 0012 airfoil, and ONERA M6 wing. While larger Krylov subspaces increase the convergence rate, they also increase the computational cost, such that 4-8 Krylov vectors often offers the fastest turnaround. Further reductions in computational cost are achieved with a sequential hybrid preconditioner that begins with the explicit multigrid solver before transitioning to the preconditioned algorithm later on. In addition, a novel implementation of dual time stepping is extended to include both common BDF methods as well as high-order implicit Runge-Kutta schemes. This particular formulation, which uses A −1 preconditioning, is amenable to matrix-free solvers, and the L-stable methods are especially suited for meshes with arbitrarily small cut-cells. Asymptotic order of convergence is demonstrated for BDF1, BDF2, SDIRK2, and 3rd-order Radau IIA time integration with unsteady 2D vortex simulations.

ARMD↗

A Scalable Interior‐Point Gauss–Newton Method for PDE‐Constrained Optimization With Bound Constraints

Here, we present a scalable approach to solve a class of partial differential equation (PDE)‐constrained optimization problems with bound constraints. This approach utilizes a robust full‐space interior‐point (IP)‐Gauss–Newton optimization method. To cope with the poorly‐conditioned IP‐Gauss–Newton saddle‐point linear systems that need to be solved approximately, once per optimization step, we propose two spectrally related preconditioners. These preconditioners leverage the limited informativeness of data in regularized PDE‐constrained optimization problems. A block Gauss–Seidel preconditioner is proposed for the GMRES‐based solution of the IP‐Gauss–Newton linear systems. It is shown, for a large‐class of PDE‐ and bound‐constrained optimization problems, that the spectrum of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix is asymptotically independent of discretization and is not impacted by the ill‐conditioning that notoriously plagues interior‐point methods. We exploit symmetry of the IP‐Gauss–Newton linear systems and propose a regularization and log‐barrier Hessian preconditioner for the preconditioned conjugate gradient (PCG)‐based solution of the equivalent IP‐Gauss–Newton–Schur complement linear systems. The eigenvalues of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix, that are not equal to one, are identical to the eigenvalues of the regularization and log‐barrier Hessian preconditioned Schur complement matrix. The scalability of the approach is demonstrated on two example problems. The numerical solution of these optimization problems is shown to require a discretization independent number of IP‐Gauss–Newton linear solves. Furthermore, the linear systems are solved in a discretization and IP ill‐conditioning independent number of preconditioned Krylov subspace iterations. The parallel scalability of the preconditioner, achieved via algebraic multigrid component solvers when applicable, and the aforementioned algorithmic scalability permits a parallel scalable means to compute solutions of a large class of PDE‐ and bound‐constrained problems.

PDE-constrained optimization↗

Application of the TRANAIR rectangular grid approach to the aerodynamic analysis of complex configurations

A numerical method is described which uses a rectangular grid to solve the nonlinear full potential equation about complex configurations. The grid is locally refined to resolve high velocity gradients arising from leading edge expansions or shock waves. The grid penetrates the boundary (described by networks of quadrilateral panels) and is generated automatically. Discrete operators are constructed using the finite element method. The system of nonlinear discrete equations is solved iteratively using a Krylov subspace method preconditioned by an exterior Poisson solver and a direct sparse solver. The primary emphasis is to provide design engineers with an aerodynamic analysis tool (the TRANAIR code) which is accurate, reliable, economical, and flexible to use. Computational results for many interesting configurations are presented.

Johnson, Forrester T.↗

Fast meta-solvers for 3D complex-shape scatterers using neural operators trained on a non-scattering problem

Three-dimensional target identification using scattering techniques requires high accuracy solutions and very fast computations for real-time predictions in some critical applications. We first train a deep neural operator (DeepONet) to solve wave propagation problems described by the Helmholtz equation in a domain without scatterers but at different wavenumbers and with a complex absorbing boundary condition. We then design two classes of fast meta-solvers by combining DeepONet with either relaxation methods, such as Jacobi and Gauss-Seidel, or with Krylov methods, such as GMRES and BiCGStab, using the trunk basis of DeepONet as a coarse-scale preconditioner. We leverage the spectral bias of neural networks to account for the lower part of the spectrum in the error distribution while the upper part is handled inexpensively using relaxation methods or fine-scale preconditioners. The meta-solvers are then applied to solve scattering problems with different shape of scatterers, at no extra training cost. We first demonstrate that the resulting meta-solvers are shape-agnostic, fast, and robust, whereas the standard standalone solvers may even fail to converge without the DeepONet. We then apply both classes of meta-solvers to scattering from a submarine, a complex three-dimensional problem. We achieve very fast solutions, especially with the DeepONet-Krylov methods, which require orders of magnitude fewer iterations than any of the standalone solvers.

97 MATHEMATICS AND COMPUTING↗

Sylvester-preconditioned adaptive-rank implicit time integrators for advection-diffusion equations with variable coefficients

Here, we consider the adaptive-rank integration of multi-dimensional time-dependent advection-diffusion partial differential equations (PDEs) with variable coefficients. We employ a standard finite-difference method for spatial discretization coupled with high-order diagonally implicit Runge-Kutta temporal schemes. The discrete equation is a generalized Sylvester equation (GSE), which we solve with a projection-based adaptive-rank algorithm structured around two key strategies: (i) constructing dimension-wise subspaces using a novel atypical extended Krylov strategy, and (ii) efficiently solving the basis coefficient matrix with a preconditioned GMRES solver. The low-rank decomposition is performed in 2D using SVD and with high-order SVD (HOSVD) in 3D to represent the tensor in a compressed Tucker format. For d-dimensional problems (here, d = 2 or 3), the computational complexity and memory storage of the approach are found numerically to scale as and $\mathscr{O}(Nr^2) + \mathscr{O} (r^{d+1})$ and $\mathscr{O}(Nr) + \mathscr{O} (r^{d})$, respectively, with the one-dimensional resolution and the maximal rank during the Krylov iteration (which we find to be largely independent of on our numerical examples). We present numerical examples that illustrate the advertised properties of the algorithm.

97 MATHEMATICS AND COMPUTING↗