Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

Implementation of hybrid finite element method based transport solver in GRIFFIN

A new transport solver option based on the hybrid FEM (HFEM) was implemented in GRIFFIN, the MOOSE-based reactor analysis code, as an effort to support routine core design calculations for advanced reactor applications. The HFEM formulation with P{sub N} (spherical harmonics expansion), akin to the variational nodal method, is effective for solving a spatially homogenized problem with strong transport effect. The residual and Jacobian evaluations of the HFEM weak form were derived and successfully implemented in GRIFFIN, having the diffusion and the PN options available in the new HFEM based transport solver. The performance was tested with the simplified ABTR benchmark problems. The results indicate that the HFEM-based transport solver is a feasible option for solving problems with spatially homogenized and strong streaming by providing superior accuracy with a proper p-refinement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of Surface Tension on a Reacting Flow Solver, PeleLM: Preprint

In liquid rocket engines, the fuel is supplied to the combustion chamber in the liquid state though injectors. Such fuel undergoes atomization, vaporization, and combustion processes. To design reliable and efficient injectors, it is required to understand the full processes. This research is part of an effort to develop a full atomization-vaporization-combustion solver from first principles. As an initial step to tackle the atomization process, a multiphase flow solver is under development. For the development, a library of the volume of fluid scheme for multiphase, IRL is coupled with a reacting Navier-Stokes equation solver, PeleLM. Furthermore, as the surface tension has considerable effects on spray breakup. surface tension is implemented in the momentum equation using the continuum surface force model and the improved height function technique.

height function↗

A fully implicit, scalable, conservative nonlinear relativistic Fokker–Planck 0D-2P solver for runaway electrons

Upon application of a sufficiently strong electric field, electrons break away from thermal equilibrium and approach relativistic speeds. These highly energetic ‘runaway’ electrons (~ MeV) play a significant role in tokamak disruption physics, and therefore their accurate understanding is essential to develop reliable mitigation strategies. As such, we have developed a fully implicit solver for the 0D-2P (i.e., including two momenta coordinates) relativistic nonlinear Fokker–Planck equation (rFP). As in earlier implicit rFP studies (NORSE, CQL3D), electron–ion interactions are modeled using the Lorentz operator, and synchrotron damping using the Abraham–Lorentz–Dirac reaction term. However, our implementation improves on these earlier studies by (1) ensuring exact conservation properties for electron collisions, (2) strictly preserving positivity, and (3) being scalable algorithmically and in parallel. Key to our proposed approach is an efficient multigrid preconditioner for the linearized rFP equation, a multigrid elliptic solver for the Braams–Karney potentials, and a novel adaptive technique to determine the associated boundary values. We verify the accuracy and efficiency of the proposed scheme with numerical results ranging from small electric-field electrical conductivity measurements to the accurate reproduction of runaway tail dynamics when strong electric fields are applied.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Matrix-Free High-Performance Saddle-Point Solvers for High-Order Problems in \(\boldsymbol{H}(\operatorname{\textbf{div}})\)

Here, this work describes the development of matrix-free GPU-accelerated solvers for high-order finite element problems in H(div). The solvers are applicable to grad-div and Darcy problems in saddle-point formulation, and have applications in radiation diffusion and porous media flow problems, among others. Using the interpolation–histopolation basis, efficient matrix-free preconditioners can be constructed for the (1, 1)-block and Schur complement of the block system. With these approximations, block-preconditioned MINRES converges in a number of iterations that is independent of the mesh size and polynomial degree. The approximate Schur complement takes the form of an M-matrix graph Laplacian and therefore can be well-preconditioned by highly scalable algebraic multigrid methods. High-performance GPU-accelerated algorithms for all components of the solution algorithm are developed, discussed, and benchmarked. Numerical results are presented on a number of challenging test cases, including the “crooked pipe” grad-div problem, the SPE10 reservoir modeling benchmark problem, and a nonlinear radiation diffusion test case.

97 MATHEMATICS AND COMPUTING↗

An efficient, conservative, time-implicit solver for the fully kinetic arbitrary-species 1D-2V Vlasov-Ampère system

In this paper, we consider the solution of the fully kinetic (including electrons) Vlasov-Ampère system in a one-dimensional physical space and two-dimensional velocity space (1D-2V) for an arbitrary number of species with a time-implicit Eulerian algorithm. The problem of velocity-space meshing for disparate thermal and bulk velocities is dealt with by an adaptive coordinate transformation of the Vlasov equation for each species, which is then discretized, including the resulting inertial terms. Mass, momentum, and energy are conserved, and Gauss's law is enforced to within the nonlinear convergence tolerance of the iterative solver through a set of nonlinear constraint functions while permitting significant flexibility in choosing discretizations in time, configuration, and velocity space. We mitigate the temporal stiffness introduced by, e.g., the plasma frequency through the use of high-order/low-order (HOLO) acceleration of the iterative implicit solver. We present several numerical results for canonical problems of varying degrees of complexity, including the multiscale ion-acoustic shock wave problem, which demonstrate the efficacy, accuracy, and efficiency of the scheme.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast nonlinear iterative solver for an implicit, energy-conserving, asymptotic-preserving charged-particle orbit integrator

Here recently, an asymptotic-preserving (AP) particle orbit integrator has been proposed with remarkable properties including exact energy conservation, the ability to capture of all first-order drifts (including the ∇B-drift), the ability to capture trapped-passing boundaries with parallel velocity extremely close to the critical velocity, and the ability to transition from strongly to weakly magnetized spatial regions. The new AP orbit integrator is implicit, employing a Crank-Nicolson (CN) temporal discretization to ensure exact energy conservation. This, in turn, requires a local nonlinear iteration involving particle velocities and positions, and the local electromagnetic fields, to obtain the new-time solution. Ref. [1] did not attempt to provide an efficient solver for this system, and employed a brute-force GMRES-driven Jacobian-free Newton-Krylov (JFNK) solver to invert the particle orbit equations at every timestep for expediency. While JFNK is robust and reliable, it is also expensive and very intrusive for practical implementations of the method (it requires having the JFNK machinery available and solving a 6 x 6 Jacobian system iteratively once per iteration per particle).

97 MATHEMATICS AND COMPUTING↗

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗

A robust solver for wavefunction-based density functional theory calculations

A new iterative solver is proposed to efficiently calculate the ground state electronic structure in density functional theory calculations. This algorithm is particularly useful for simulating physical systems considered difficult to converge by standard solvers, in particular metallic systems. Here, the effectiveness of the proposed algorithm is demonstrated on various applications.

42 ENGINEERING↗

Newly Released Capabilities in the Distributed-Memory SuperLU Sparse Direct Solver

We present the new features available in the recent release of SuperLU_DIST, Version 8.1.1. SuperLU_DIST is a distributed-memory parallel sparse direct solver. The new features include (1) a 3D communication-avoiding algorithm framework that trades off inter-process communication for selective memory duplication, (2) multi-GPU support for both NVIDIA GPUs and AMD GPUs, and (3) mixed-precision routines that perform single-precision LU factorization and double-precision iterative refinement. Apart from the algorithm improvements, we also modernized the software build system to use CMake and Spack package installation tools to simplify the installation procedure. Throughout the article, we describe in detail the pertinent performance-sensitive parameters associated with each new algorithmic feature, show how they are exposed to the users, and give general guidance of how to set these parameters. We illustrate that the solver’s performance both in time and memory can be greatly improved after systematic tuning of the parameters, depending on the input sparse matrix and underlying hardware.

97 MATHEMATICS AND COMPUTING↗

A Tutorial for Using an Open-Source Solver for the Regional Energy Deployment System (ReEDS) Model

ReEDS is a publicly available model developed at NREL that can be used to analyze the potential evolution of the U.S. electric power system into the future. ReEDS is formulated as a linear program, written in GAMS, and solved using a linear programming solver (e.g., CPLEX). This tutorial provides context and understanding for the implications of using an open-source solver for ReEDS.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

High-order matrix-free incompressible flow solvers with GPU acceleration and low-order refined preconditioners

In this work, we present a matrix-free flow solver for high-order finite element discretizations of the incompressible Navier-Stokes and Stokes equations with GPU acceleration. For high polynomial degrees, assembling the matrix for the linear systems resulting from the finite element discretization can be prohibitively expensive, both in terms of computational complexity and memory. For this reason, it is necessary to develop matrix-free operators and preconditioners, which can be used to efficiently solve these linear systems without access to the matrix entries themselves. The matrix-free operator evaluations utilize GPU-accelerated sum-factorization techniques to minimize memory movement and maximize throughput. The preconditioners developed in this work are based on a low-order refined methodology with parallel subspace corrections, as described for diffusion problems in [1]. The saddle-point Stokes system is solved using block-preconditioning techniques, which are robust in mesh size, polynomial degree, time step, and viscosity. For the incompressible Navier-Stokes equations, we make use of projection (fractional step) methods, which require Helmholtz and Poisson solves at each time step. The performance of our flow solvers is assessed on several benchmark problems in two and three spatial dimensions.

97 MATHEMATICS AND COMPUTING↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

Validation of NSFsim as a Grad-Shafranov equilibrium solver at DIII-D

Plasma shape is a significant factor that must be considered for any Fusion Pilot Plant (FPP) as it has significant consequences for plasma stability and core confinement. A new simulator, NSFsim, has been developed based on a historically successful code, DINA [1], offering tools to simulate both transport and plasma shape. Specifically, NSFsim is a free boundary equilibrium and transport solver and has been configured to match the properties of the DIII-D tokamak. This paper is focused on validating the Grad-Shafranov (GS) solver of NSFsim by analyzing its ability to recreate the plasma shape, the poloidal flux distribution, and the measurements of the simulated diagnostic signals originating from flux loops and magnetic probes in DIII-D. Five different plasma shapes are simulated to show the robustness of NSFsim to different plasma conditions; these shapes are Lower Single Null (LSN), Upper Single Null (USN), Double Null (DN), Inner Wall Limited (IWL), and Negative Triangularity (NT). The NSFsim results are compared against real measured signals, magnetic profile fits from EFIT [2], and another plasma equilibrium simulator, GSevolve [3]. EFIT reconstructions of shots are readily available at DIII-D, but GSevolve was manually ran by us to provide simulation data to compare against.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exact and locally implicit source term solvers for multifluid-Maxwell systems

Recently, a family of models that couple multifluid systems to the full Maxwell equations have been used in laboratory, space, and astrophysical plasma modeling. These models are more complete descriptions of the plasma than reduced models like magnetohydrodynamic (MHD) since they are derived more closely from the full kinetic Vlasov-Maxwell system, without assumptions like quasi-neutrality, negligible electron mass, etc. Thus these models naturally retain non-ideal MHD effects like electron inertia, Hall term, pressure anisotropy/nongyrotropy, displacement current, among others. One obstacle to broader application of these model is that an explicit treatment of their source terms leads to the need to resolve rapid processes like plasma oscillation and electron cyclotron motion, even when these are not important. In this paper, we suggest two ways to address this issue. First, we derive the analytic solutions to the source update equations, which can be implemented as a practical, but less generic solver. We then develop a time-centered, locally implicit algorithm to update the source terms, allowing stepping over the fast kinetic time-scales. For a plasma with S species, the locally implicit algorithm involves inverting a local (3 S + 3) × (3 S + 3) matrix only, thus is very efficient. The performance can be further increased by using the direct update formulas to skip null calculations. In this paper, we present benchmarks illustrating the exact energy-conservation of the locally implicit solver, as well as its efficiency and robustness for both small-scale, idealized problems and largescale, complex systems. The locally implicit algorithm can be also easily extended to include other local sources, like collisions and ionization, which are difficult to solve analytically.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗