Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “iterative solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

xSDK-batched Subcontract - Ginkgo Batched Iterative Solver Development (Final Report)

Iterative solvers are fundamentally different from direct solvers in terms of execution as they generally do not execute a pre-defined sequence of operations or steps, but adapt the number of iterations to the specific problem and the preset solution quality. Generally, the adaptation of the iteration count to the problem is realized by monitoring the solver convergence and stopping the iteration process once the monitored metric, e.g., the residual norm, hits a pre-defined threshold. When addressing a set of problems with different properties, it is necessary to monitor the threshold for each problem individually and break up the SIMD execution style to avoid excess iterations for “easier” problems. Ginkgo integrates a simple but customizable stopping criterion for the residual norm and generally uses a pre-defined (relative or absolute) residual norm as the stopping criterion. In order to avoid the overhead of launching a kernel at every iteration, the iteration convergence and iteration control is part of the solver kernel. Each thread maintains its own copy of the iteration count.

97 MATHEMATICS AND COMPUTING↗

A family of independent Variable Eddington Factor methods with efficient preconditioned iterative solvers

We present a family of discretizations for the Variable Eddington Factor (VEF) equations that have high-order accuracy on curved meshes and efficient preconditioned iterative solvers. The VEF discretizations are combined with the Discontinuous Galerkin transport discretization from to form effective high-order, linear transport methods. The VEF discretizations are derived by extending the unified analysis of Discontinuous Galerkin methods for elliptic problems presented by Arnold et al. to the VEF equations. This framework is used to define analogs of the interior penalty, second method of Bassi and Rebay, minimal dissipation local Discontinuous Galerkin, and continuous finite element methods. The analysis of subspace correction preconditioners, which use a continuous operator to iteratively precondition the discontinuous discretization, is extended to the case of the non-symmetric VEF system. Numerical results demonstrate that the VEF discretizations have arbitrary-order accuracy on curved meshes, preserve the thick diffusion limit, and are effective on a proxy problem from thermal radiative transfer in both outer transport iterations and inner preconditioned linear solver iterations. We demonstrate that the VEF solution converges to the S N transport solution as the mesh is refined on both problems with smooth and non-smooth behavior in angle. Parallel performance studies show that the interior penalty VEF discretization's linear solve weak scales out to 1024 processors and strong scales well on a single node. Particular attention is paid to the parallel performance of the VEF algorithm when used in combination with a parallel block Jacobi transport sweep.

97 MATHEMATICS AND COMPUTING↗

Fast nonlinear iterative solver for an implicit, energy-conserving, asymptotic-preserving charged-particle orbit integrator

Here recently, an asymptotic-preserving (AP) particle orbit integrator has been proposed with remarkable properties including exact energy conservation, the ability to capture of all first-order drifts (including the ∇B-drift), the ability to capture trapped-passing boundaries with parallel velocity extremely close to the critical velocity, and the ability to transition from strongly to weakly magnetized spatial regions. The new AP orbit integrator is implicit, employing a Crank-Nicolson (CN) temporal discretization to ensure exact energy conservation. This, in turn, requires a local nonlinear iteration involving particle velocities and positions, and the local electromagnetic fields, to obtain the new-time solution. Ref. [1] did not attempt to provide an efficient solver for this system, and employed a brute-force GMRES-driven Jacobian-free Newton-Krylov (JFNK) solver to invert the particle orbit equations at every timestep for expediency. While JFNK is robust and reliable, it is also expensive and very intrusive for practical implementations of the method (it requires having the JFNK machinery available and solving a 6 x 6 Jacobian system iteratively once per iteration per particle).

97 MATHEMATICS AND COMPUTING↗

Quarter 4 Report: Report on Final Findings and Opportunities for Future Work in the Use of Mixed Precision in Iterative Solvers

The fourth quarter of the project was spent developing an error analysis of the s-step Lanczos and CG algorithms. Our theoretical bounds and numerical experiments show that the numerical behavior of the algorithm can be significantly improved by using extra precision in a small part of the computation related to the computation and application of the Gram matrix. We have published a technical report which includes all steps of the analysis [8]; a shortened version for journal submission is in preparation. We plan to submit this paper in the following weeks. Activities related to this also include a collaboration with Ichitaro Yamazaki on gathering performance results for these new mixed precision s-step Krylov subspace methods using single/double precision on GPUs. Namely, we would like to obtain performance results that show that the performance overhead of using double the working precision in these select computations is minimal. Other activities include attending biweekly xSDK meetings and presenting a pitch talk on this work to the group on February 25, 2021. In the remainder of the document, we summarize our findings on the potential for mixed precision in classical Krylov subspace methods and s-step Krylov subspace methods, as well as key opportunities for future work.

97 MATHEMATICS AND COMPUTING↗

Handling Iterative Solvers in an Algorithmic Differentiation Framework Using Implicit Methods

Differentiable programming is a powerful concept as it enables the seemly propagation of gradients through functions, algorithms, and/or whole physics simulations. These gradients are useful for a wide variety of applications, including sensitivity studies and machine learning, but one of particular interest is optimization. Gradient-based optimization, enabled through automatic/algorithmic differentiation (AD), can be used on predictive physical models to efficiently optimize a set of design variables. AD methods are a particularly promising approach to complex physics simulations because they can be shown to scale well with an increasing number of design variables; however, care must be taken when coupling between different models or different states of a single model.

algorithmic differentiation↗

An implicit barotropic mode solver for MPAS-ocean using a modern Fortran solver interface

Here, we demonstrate use of a modern Fortran solver interface to manage solver algorithms for an implicit barotropic mode solver in the Model for Predictions Across Scales-Ocean (MPAS-O). ForTrilinos, a Fortran interface to Trilinos that contains a large collection of solver capabilities written in C++, has been implemented in MPAS-O to provide access to a suite of linear solver options. By virtue of the simplified wrapper and interface generator (SWIG) automation tool that generates modern Fortran interfaces to C++ code, we were able to implement the Fortran solver interface in MPAS-O using a familiar Fortran coding style while minimizing performance degradation. The ForTrilinos solver interface is written within MPAS-O’s time stepping modules as a subroutine in conjunction with MPAS-O code. Applied to an idealized ocean and a high-resolution realistic ocean test case, parallel performance of ForTrilinos solvers is examined. It is found that parallel scalability of the ForTrilinos solvers is highly dependent on the number of global synchronization points per solver iteration in each iterative solver algorithm. ForTrilinos solvers perform best compared to the Fortran hand-crafted (FHC) solver when the amount of work per processor is large enough. However, parallel scalability is better with the FHC solver and so when the work per core is modest FHC outperforms ForTrilinos. The intercomparison between the ForTrilinos and FHC solvers reveals that this performance hit in the ForTrilinos solver mostly comes from the global synchronization process, while suggesting that the matrix-vector multiplication process in the FHC solver needs to be optimized for better performance.

97 MATHEMATICS AND COMPUTING↗

Large-scale harmonic balance simulations with Krylov subspace and preconditioner recycling

The multi-harmonic balance method combined with numerical continuation provides an efficient framework to compute a family of time-periodic solutions, or response curves, for large-scale, nonlinear mechanical systems. The predictor and corrector steps repeatedly solve a sequence of linear systems that scale by the model size and number of harmonics in the assumed Fourier series approximation. In this paper, a novel Newton–Krylov iterative method is embedded within the multi-harmonic balance and continuation algorithm to efficiently compute the approximate solutions from the sequence of linear systems that arise during the prediction and correction steps. Further, the method recycles, or reuses, both the preconditioner and the Krylov subspace generated by previous linear systems in the solution sequence. A delayed frequency preconditioner refactorizes the preconditioner only when the performance of the iterative solver deteriorates. The GCRO-DR iterative solver recycles a subset of harmonic Ritz vectors to initialize the solution subspace for the next linear system in the sequence. The performance of the iterative solver is demonstrated on two exemplars with contact-type nonlinearities and benchmarked against a direct solver with traditional Newton–Raphson iterations.

97 MATHEMATICS AND COMPUTING↗

Comparative investigation of iterative solutions of the all-frequency stable formulation and its vector-potential-only variation

An all-frequency stable formulation has been proposed and applied to solve low-frequency and multiscale electromagnetic problems to avoid the low-frequency breakdown catastrophe of the traditional vector wave equation. Based on the potential representation of the fields, this formulation involves vector and scalar potentials, as well as an auxiliary potential to enforce the gauge condition. In this paper, such a formulation is first simplified to involve vector potential only. Their iterative solutions are then sought by using an incomplete LU preconditioner and an iterative solver. The iterative performance of both formulations are investigated in a comparative study.

Mekonnen, Minyechil↗

Analysis of a New Implicit Solver for a Semiconductor Model

Here, we present and analyze a new iterative solver for implicit discretizations of a simplified Boltzmann--Poisson system. The algorithm builds on recent work that incorporated a sweeping algorithm for the Vlasov--Poisson equations as part of nested inner-outer iterative solvers for the Boltzmann--Poisson equations. The new method eliminates the need for nesting and requires only one transport sweep per iteration. It arises as a new fixed-point formulation of the discretized system which we prove to be contractive for a given electric potential. We also derive an accelerator to improve the convergence rate for systems in the drift-diffusion regime. We numerically compare the efficiency of the new solver, with and without acceleration, against a recently developed nested iterative solver.

97 MATHEMATICS AND COMPUTING↗

A Scalable Semi-Implicit Barotropic Mode Solver for the MPAS-Ocean

A scalable semi-implicit barotropic mode solver for the ocean component of the model for prediction across scales has been implemented as a competitor to an existing explicit-subcycling scheme to allow faster and more stable simulations while not sacrificing accuracy. The semi-implicit solver adopts the pipelined preconditioned bi-conjugate gradient stabilization algorithm as an iterative solver in conjunction with the restricted additive Schwarz preconditioner that accelerates the convergence rate of the iterative solver. The preconditioner is constructed from a linearized barotropic system that also reorders the system for optimal performance, while the semi-implicit solver deals with the fully nonlinear barotropic system that requires reassembly of the coefficient matrix for every time step. Several numerical experiments, from simple one-dimensional tests to three-dimensional real-world tests, demonstrate that the semi-implicit solver has almost the same accuracy and better parallel scalability compared with the existing scheme while allowing faster and more stable simulations. Furthermore, the semi-implicit solver accelerates the barotropic mode up to 2.9 times faster than the existing scheme on 16,320 processors, leading to an overall runtime speedup of 1.9.

97 MATHEMATICS AND COMPUTING↗

Simulation of electron Bernstein waves using FullWave with a 2D non-local hot plasma model

Hot plasma wave simulation capability is expanded in the FullWave code by updating the hybrid iterative solver in the code with a semi-implicit time stepping method. The new approach is used to simulate Electron Bernstein Wave (EBW) heating in over-dense spherical tokamak plasmas. The code’s hybrid iterative solver circumvents the prohibitive memory cost of direct methods by combining a time evolution of Maxwell’s equations with frequency-domain relaxation, while the conductivity kernel, calculated via 3D particle tracking, captures the essential non-local wave–particle interactions. One-dimensional EBW simulations verify the algorithm’s accuracy by demonstrating mode conversion from X-mode wave to EBW at the upper hybrid resonance and a strong cyclotron damping near the plasma core. Two-dimensional simulation reproduces the predicted short EBW wavelength and quantitatively matches the hot-plasma dispersion relation. This study demonstrates the fidelity of the hybrid solver for the electron cyclotron frequency range.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A DG-IMEX Method for Two-moment Neutrino Transport: Nonlinear Solvers for Neutrino–Matter Coupling

Neutrino-matter interactions play an important role in core-collapse supernova (CCSN) explosions as they contribute to both lepton number and/or four-momentum exchange between neutrinos and matter, and thus act as the agent for neutrino-driven explosions. Due to the multiscale nature of neutrino transport in CCSN simulations, an implicit treatment of neutrino-matter interactions is desired, which requires solutions of coupled nonlinear systems in each step of the time integration scheme. In this paper we design and compare nonlinear iterative solvers for implicit systems with energy coupling neutrino-matter interactions commonly used in CCSN simulations. Specifically, we consider electron neutrinos and antineutrinos, which interact with static matter configurations through the Bruenn 85 opacity set. The implicit systems arise from the discretization of a nonrelativistic two-moment model for neutrino transport, which employs the discontinuous Galerkin (DG) method for phase-space discretization and an implicit-explicit (IMEX) time integration scheme. In the context of this DG-IMEX scheme, we propose two approaches to formulate the nonlinear systems — a coupled approach and a nested approach. For each approach, the resulting systems are solved with Anderson-accelerated fixedpoint iteration and Newton’s method. The performance of these four iterative solvers has been compared on relaxation problems with various degree of collisionality, as well as proto-neutron star deleptonization problems with several matter profiles adopted from spherically symmetric CCSN simulations. Here, numerical results suggest that the nested Anderson-accelerated fixed-point solver is more efficient than other tested solvers for solving implicit nonlinear systems with energy coupling neutrino-matter interactions.

79 ASTRONOMY AND ASTROPHYSICS↗

A robust solver for wavefunction-based density functional theory calculations

A new iterative solver is proposed to efficiently calculate the ground state electronic structure in density functional theory calculations. This algorithm is particularly useful for simulating physical systems considered difficult to converge by standard solvers, in particular metallic systems. Here, the effectiveness of the proposed algorithm is demonstrated on various applications.

42 ENGINEERING↗