Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Edge-Based Viscous Method for Mixed-Element Node-Centered Finite-Volume Solvers

A novel, efficient, edge-based viscous (EBV) discretization method has been recently developed, implemented in a practical, unstructured-grid, node-centered, finite-volume flow solver, and applied to viscous-kernel computations that include evaluations of meanflow viscous fluxes, turbulence-model and chemistry-model diffusion terms, and the corresponding Jacobian contributions. Initially, the EBV method had been implemented for tetrahedral grids and demonstrated multifold acceleration of all viscous-kernel computations. This paper presents an extension of the EBV method for mixed-element grids. In addition to the primal edges of a given mixed-element grid, virtual edges are introduced to connect cell nodes that are not connected by a primal edge. The EBV method uses an efficient loop over all (primal and virtual) edges and features a compact discretization stencil based on the nearest neighbors. This study verifies the EBV method and assesses its efficiency on mixed-element grids by comparing the EBV solution accuracy and iterative convergence with those of well-established solutions obtained using a cell-based viscous (CBV) discretization method. The EBV solver’s memory footprint is optimized and often smaller than the memory footprint of the CBV solver. A multifold speedup is demonstrated for all viscous-kernel computations resulting in significant reduction of the time to solutions for several benchmark mixed-element-grid computations, including simulations of a flow around NASA’s juncture-flow model and a hypersonic, chemically reacting flow around a blunt body.

CFD↗

Edge-Based Viscous Method for Mixed-Element Node-Centered Finite-Volume Solvers

A novel, efficient, edge-based viscous (EBV) discretization method has been recently developed, implemented in a practical, unstructured-grid, node-centered, finite-volume flow solver, and applied to viscous-kernel computations that include evaluations of meanflow viscous fluxes, turbulence-model and chemistry-model diffusion terms, and the corresponding Jacobian contributions. Initially, the EBV method had been implemented for tetrahedral grids and demonstrated multifold acceleration of all viscous-kernel computations. This paper presents an extension of the EBV method for mixed-element grids. In addition to the primal edges of a given mixed-element grid, virtual edges are introduced to connect cell nodes that are not connected by a primal edge. The EBV method uses an efficient loop over all (primal and virtual) edges and features a compact discretization stencil based on the nearest neighbors. This study verifies the EBV method and assesses its efficiency on mixed-element grids by comparing the EBV solution accuracy and iterative convergence with those of well-established solutions obtained using a cell-based viscous (CBV) discretization method. The EBV solver’s memory footprint is optimized and often smaller than the memory footprint of the CBV solver. A multifold speedup is demonstrated for all viscous-kernel computations resulting in significant reduction of the time to solutions for several benchmark mixed-element-grid computations, including simulations of a flow around NASA’s juncture-flow model and a hypersonic, chemically reacting flow around a blunt body.

Edge-based viscous method↗

Computational Analysis of a Boundary-Layer Ingesting Tailcone Thruster Configuration Within the National Transonic Facility Using the LAVA Solver

Boundary-layer ingesting (BLI) propulsion systems are one of the many technologies currently under investigation within the aerospace community to meet NASA’s Advanced Air Transport Technology project goal of a sustainable future in aviation. To that end, an experimental campaign was conducted in the National Transonic Facility (NTF) to study the propulsion-airframe integration effects of the Boundary-Layer Ingesting Tailcone System installed in the aft portion of a 2.7% scale design of the NASA Common Research Model (CRM). This test was accompanied by a numerical simulation effort using the Launch, Ascent, and Vehicle Aerodynamics (LAVA) solver, to validate the applicability of current best-practice Reynolds-averaged Navier Stokes (RANS) models in accurately predicting the relevant flow-physics in free-air. A total of 205 steady RANS simulations were performed using LAVA, covering the range of conditions present in the NTF test matrix. Flow conditions ranged from 5- to 15-million Reynolds number, Mach numbers of 0.75, 0.80 and 0.85, and angles-of-attack between –3° and 4°. Four different mass-flow plugs were also tested to assess propulsor operating condition effects. Sensitivity to these flow conditions are analyzed in detail, especially in terms of the nacelle flow distortion which was the main subject of the study. Aerodynamic load coefficients, aftbody boundary-layer measurements and fuselage pressure tap results are also discussed, providing a wide range of validation results from this experimental campaign. This validation effort shows that current best-practices using RANS models within the LAVA solver flow solver are well suited for predicting the inlet flow distortion characteristics of novel aircraft configurations employing BLI at the aft end of the fuselage.

AATT↗

Implicit Preconditioning for Explicit Multigrid Solvers on Cut-Cell Cartesian Meshes

This work assesses the effectiveness of linearized implicit Euler preconditioning for multigrid solvers using an unpreconditioned, Jacobian-free Newton Krylov method to converge the linear system of equations. Multigrid convergence rates improve to approximately 0.75 across the cases tested including a Mach 2 supersonic wedge, transonic NACA 0012 airfoil, and ONERA M6 wing. While larger Krylov subspaces increase the convergence rate, they also increase the computational cost, such that 4-8 Krylov vectors often offers the fastest turnaround. Further reductions in computational cost are achieved with a sequential hybrid preconditioner that begins with the explicit multigrid solver before transitioning to the preconditioned algorithm later on. In addition, a novel implementation of dual time stepping is extended to include both common BDF methods as well as high-order implicit Runge-Kutta schemes. This particular formulation, which uses A −1 preconditioning, is amenable to matrix-free solvers, and the L-stable methods are especially suited for meshes with arbitrarily small cut-cells. Asymptotic order of convergence is demonstrated for BDF1, BDF2, SDIRK2, and 3rd-order Radau IIA time integration with unsteady 2D vortex simulations.

ARMD↗

Computational Analysis of a Boundary Layer Ingesting Tailcone Thruster Configuration Using the LAVA Curvilinear Solver

Boundary layer ingesting (BLI) propulsion systems are one of the many technologies currently under investigation within the aerospace community to meet NASA's Advanced Air Transport Technology project goal of a sustainable future in aviation. To that end, an experimental campaign was conducted in the National Transonic Facility (NTF) wind tunnel to study the propulsion-airframe integration effects of the Boundary Layer Ingesting Tailcone System installed in the aft portion of a 2.7% scale design of the NASA Common Research Model (CRM). This test was accompanied by a numerical simulation effort using the Launch, Ascent, and Vehicle Aerodynamics (LAVA) curvilinear solver to validate the applicability of current best-practice Reynolds-Averaged Navier-Stokes (RANS) models in accurately predicting the relevant flow physics in free-air. A total of 205 steady RANS simulations were performed using LAVA, covering the range of conditions present in the NTF test matrix. Flow conditions ranged from 5- to 15-million Reynolds number, Mach numbers of 0.75, 0.80 and 0.85, and angles of attack between -3 and 4 degrees. Four different mass-flow plugs were also tested to assess propulsor operating condition effects. Sensitivity to these flow conditions is analyzed in detail, especially in terms of the nacelle flow distortion which was the main subject of the study. Aerodynamic load coefficients, aft body boundary layer measurements and fuselage pressure tap results are also discussed, providing a wide range of validation results from this experimental campaign. This validation effort shows that current best-practices using RANS models within the LAVA curvilinear solver flow solver are well suited for predicting the inlet flow distortion characteristics of novel aircraft configurations employing BLI at the aft end of the fuselage. This talk will cover the combined experimental and numerical efforts that resulted from this BLI-focused investigation.

ARMD↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

Implementation of hybrid finite element method based transport solver in GRIFFIN

A new transport solver option based on the hybrid FEM (HFEM) was implemented in GRIFFIN, the MOOSE-based reactor analysis code, as an effort to support routine core design calculations for advanced reactor applications. The HFEM formulation with P{sub N} (spherical harmonics expansion), akin to the variational nodal method, is effective for solving a spatially homogenized problem with strong transport effect. The residual and Jacobian evaluations of the HFEM weak form were derived and successfully implemented in GRIFFIN, having the diffusion and the PN options available in the new HFEM based transport solver. The performance was tested with the simplified ABTR benchmark problems. The results indicate that the HFEM-based transport solver is a feasible option for solving problems with spatially homogenized and strong streaming by providing superior accuracy with a proper p-refinement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of Surface Tension on a Reacting Flow Solver, PeleLM: Preprint

In liquid rocket engines, the fuel is supplied to the combustion chamber in the liquid state though injectors. Such fuel undergoes atomization, vaporization, and combustion processes. To design reliable and efficient injectors, it is required to understand the full processes. This research is part of an effort to develop a full atomization-vaporization-combustion solver from first principles. As an initial step to tackle the atomization process, a multiphase flow solver is under development. For the development, a library of the volume of fluid scheme for multiphase, IRL is coupled with a reacting Navier-Stokes equation solver, PeleLM. Furthermore, as the surface tension has considerable effects on spray breakup. surface tension is implemented in the momentum equation using the continuum surface force model and the improved height function technique.

height function↗

Matrix-Free High-Performance Saddle-Point Solvers for High-Order Problems in \(\boldsymbol{H}(\operatorname{\textbf{div}})\)

Here, this work describes the development of matrix-free GPU-accelerated solvers for high-order finite element problems in H(div). The solvers are applicable to grad-div and Darcy problems in saddle-point formulation, and have applications in radiation diffusion and porous media flow problems, among others. Using the interpolation–histopolation basis, efficient matrix-free preconditioners can be constructed for the (1, 1)-block and Schur complement of the block system. With these approximations, block-preconditioned MINRES converges in a number of iterations that is independent of the mesh size and polynomial degree. The approximate Schur complement takes the form of an M-matrix graph Laplacian and therefore can be well-preconditioned by highly scalable algebraic multigrid methods. High-performance GPU-accelerated algorithms for all components of the solution algorithm are developed, discussed, and benchmarked. Numerical results are presented on a number of challenging test cases, including the “crooked pipe” grad-div problem, the SPE10 reservoir modeling benchmark problem, and a nonlinear radiation diffusion test case.

97 MATHEMATICS AND COMPUTING↗

Fast nonlinear iterative solver for an implicit, energy-conserving, asymptotic-preserving charged-particle orbit integrator

Here recently, an asymptotic-preserving (AP) particle orbit integrator has been proposed with remarkable properties including exact energy conservation, the ability to capture of all first-order drifts (including the ∇B-drift), the ability to capture trapped-passing boundaries with parallel velocity extremely close to the critical velocity, and the ability to transition from strongly to weakly magnetized spatial regions. The new AP orbit integrator is implicit, employing a Crank-Nicolson (CN) temporal discretization to ensure exact energy conservation. This, in turn, requires a local nonlinear iteration involving particle velocities and positions, and the local electromagnetic fields, to obtain the new-time solution. Ref. [1] did not attempt to provide an efficient solver for this system, and employed a brute-force GMRES-driven Jacobian-free Newton-Krylov (JFNK) solver to invert the particle orbit equations at every timestep for expediency. While JFNK is robust and reliable, it is also expensive and very intrusive for practical implementations of the method (it requires having the JFNK machinery available and solving a 6 x 6 Jacobian system iteratively once per iteration per particle).

97 MATHEMATICS AND COMPUTING↗

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗

A robust solver for wavefunction-based density functional theory calculations

A new iterative solver is proposed to efficiently calculate the ground state electronic structure in density functional theory calculations. This algorithm is particularly useful for simulating physical systems considered difficult to converge by standard solvers, in particular metallic systems. Here, the effectiveness of the proposed algorithm is demonstrated on various applications.

42 ENGINEERING↗

Newly Released Capabilities in the Distributed-Memory SuperLU Sparse Direct Solver

We present the new features available in the recent release of SuperLU_DIST, Version 8.1.1. SuperLU_DIST is a distributed-memory parallel sparse direct solver. The new features include (1) a 3D communication-avoiding algorithm framework that trades off inter-process communication for selective memory duplication, (2) multi-GPU support for both NVIDIA GPUs and AMD GPUs, and (3) mixed-precision routines that perform single-precision LU factorization and double-precision iterative refinement. Apart from the algorithm improvements, we also modernized the software build system to use CMake and Spack package installation tools to simplify the installation procedure. Throughout the article, we describe in detail the pertinent performance-sensitive parameters associated with each new algorithmic feature, show how they are exposed to the users, and give general guidance of how to set these parameters. We illustrate that the solver’s performance both in time and memory can be greatly improved after systematic tuning of the parameters, depending on the input sparse matrix and underlying hardware.

97 MATHEMATICS AND COMPUTING↗

A Tutorial for Using an Open-Source Solver for the Regional Energy Deployment System (ReEDS) Model

ReEDS is a publicly available model developed at NREL that can be used to analyze the potential evolution of the U.S. electric power system into the future. ReEDS is formulated as a linear program, written in GAMS, and solved using a linear programming solver (e.g., CPLEX). This tutorial provides context and understanding for the implications of using an open-source solver for ReEDS.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗