Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Two-Stage Gauss-Seidel Preconditioners and Smoothers for Krylov Solvers on a GPU Cluster: Preprint

Gauss-Seidel (GS) relaxation is often employed as a preconditioner for a Krylov solver or as a smoother for Algebraic Multigrid (AMG). However, the requisite sparse triangular solve is difficult to parallelize on many-core architectures such as graphics processing units (GPUs). In the present study, the performance of the sequential GS relaxation based on a triangular solve is compared with two-stage variants, replacing the direct triangular solve with a fixed number of inner Jacobi-Richardson (JR) iterations. When a small number of inner iterations is sufficient to maintain the Krylov convergence rate, the two-stage GS (GS2) often outperforms the sequential algorithm on many-core architectures. The GS2 algorithm is also compared with JR. When they perform the same number of ops for SpMV (e.g. three JR sweeps compared to two GS sweeps with one inner JR sweep), the GS2 iterations, and the Krylov solver preconditioned with GS2, may converge faster than the JR iterations. Moreover, for some problems (e.g. elasticity), it was found that JR may diverge with a damping factor of one, whereas two-stage GS may improve the convergence with more inner iterations. Finally, to study the performance of the two-stage smoother and preconditioner for a practical problem, these were applied to incompressible uid ow simulations on GPUs.

algebraic multigrid↗

Discrete-Element and Material-Point Method (DEM and MPM) Based Solvers for Sustainable Technologies

We present the use of discrete element method (DEM) and material point method (MPM) in three relevant green technology applications that include biomass feedstock handling, lithium-ion battery manufacturing, and high-pressure reverse osmosis. Our open-source DEM and MPM solvers are developed using performance portable grid and particle management library, AMReX, thus enabling superior performance on NVIDIA and AMD GPUs with > 100 million particles. Our DEM solver resolves the motion of individual particles in a granular system and includes a bonded sphere method for modeling non-spherical particles along with Hertzian and liquid bridge-based contact models. We simulate highly variable biomass feedstock flows in large-scale hoppers for biofuel production and electrode calendering in battery manufacturing using DEM. Our simulations predict flow blockage in large scale biomass hoppers and electrode microstructure variations, thus providing valuable information for biofuel and battery manufacturers, respectively. The second half of the talk will be on MPM and its application towards pore resolved simulations of reverse osmosis membranes under compressive loads. We present a validation study of our MPM simulations with membrane microscopy imaging thus providing useful insights on membrane stability under high pressure conditions. We also present a spectral stability analysis of using linear hat, quadratic and cubic spline basis in MPM indicating regions of numerical stability.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

Investigation of grid-based vorticity-velocity large eddy simulation off-body solvers for application to overset CFD

Accurately predicting unsteady wakes and vortex-dominated flows is essential to a wide range of engineering applications, including aircraft, rotorcraft, shipboard operations, bio-inspired unsteady flight and propulsion, wind turbines, and urban flows. While current CFD software can model the complete flow field and wake system, the computational costs incurred in high Reynolds number unsteady turbulent flow simulations often remain prohibitive for routine engineering use, particularly for applications involving moving components. Prior work has demonstrated that by adopting a vorticity-velocity formulation in a grid-based off-body flow solver (VorTran-M and VorTran-M2) one can lower these costs by several orders of magnitude when compared to conventional approaches. This paper describes the extensions made to VorTran-M2 to support turbulent flows, and associated benchmarking activity to assess its performance for problems involving strong stretching and diffusion processes, whose competing contributions to the vorticity field are core drivers of turbulent flow evolution. Predictions are presented for: (i) the Kida-Pelz problem whose inviscid form is of mathematical interest due to its apparent formation of singular flow in finite time; and (ii) the Taylor Green vortex arrangement, which has been extensively studied as a fundamental simulation challenge in the turbulent modeling community. Here, the results are used to evaluate the overall predictive ability and performance of two sub-grid scale models incorporated into VorTran-M2. Results indicate that the computational cost savings seen previously for inviscid and convection dominated problems extend to turbulent flow simulations supporting the viability of VorTran-M2 as a low cost means for accurately modeling the far-field and background flow, particularly when long duration vorticity evolution is of interest.

42 ENGINEERING↗

Benchmark of the KGMf with a coupled Boltzmann equation solver

The Kinetic Global Model framework (KGMf) is an open-source general-purpose global model (spatially averaged) simulation code developed to explore the reaction kinetics and pathways in plasma discharge systems. It contains species continuity and electron energy balance equations, with a time-dependent evaluated electron energy distribution function (EEDF) for electron impact reactions. The EEDF is utilized to determine the rate coefficients for electron impact reactions, which can have profound impact on the temporal evolutions of plasma parameters. Previously, the EEDF was commonly assumed as Maxwellian or an analytical function of the effective electron temperature. In this work, the KGMf is coupled with a Boltzmann equation (BE) solver to self-consistently compute the EEDF. The EEDF evolution frequency is determined based on relative changes of the reduced electric field. The KGMf is benchmarked with the ZDPlasKin code based on high-pressure low-temperature argon plasma discharge cases. Additionally, the temporal evolutions of reduced electric field, electron temperature, EEDF, reaction rates, and species densities, are obtained and compared under different discharge conditions, showing good agreement between the KGMf and the ZDPlasKin simulations. The application of the KGMf for predicting breakdown times in high power microwave discharges is also presented, which shows qualitative agreement with particle-in-cell simulations. The KGMf can be further applied for more complicated plasma discharge systems, where the reaction kinetics are more intricate (e.g., plasma-assisted combustion systems).

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Accelerated impurity solver for DMFT and its diagrammatic extensions

Here, we present ComCTQMC, a GPU accelerated quantum impurity solver. It uses the continuous-time quantum Monte Carlo (CTQMC) algorithm wherein the partition function is expanded in terms of the hybridisation function (CT-HYB). ComCTQMC supports both partition and worm-space measurements, and it uses improved estimators and the reduced density matrix to improve observable measurements whenever possible. ComCTQMC efficiently measures all one and two-particle Green's functions, all static observables which commute with the local Hamiltonian, and the occupation of each impurity orbital. ComCTQMC can solve complex-valued impurities with crystal fields that are hybridized to both fermionic and bosonic baths. Most importantly, ComCTQMC utilizes graphical processing units (GPUs), if available, to dramatically accelerate the CTQMC algorithm when the Hilbert space is sufficiently large. We demonstrate acceleration by a factor of over 600 (100) in a simulation of δ-Pu at 600 K with (without) crystal fields. In easier problems, the GPU offers less impressive acceleration or even decelerates the CTQMC. Here we describe the theory, algorithms, and structure used by ComCTQMC in order to achieve this set of features and level of acceleration.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

A numerical Poisson solver with improved radial solutions for a self-consistent locally scaled self-interaction correction method

Abstract The universal applicability of density functional approximations is limited by self-interaction error made by these functionals. Recently, a novel one-electron self-interaction-correction (SIC) method that uses an iso-orbital indicator to apply the SIC at each point in space by scaling the exchange-correlation and Coulomb energy densities was proposed. The locally scaled SIC (LSIC) method is exact for the one-electron densities, and unlike the well-known Perdew–Zunger SIC (PZSIC) method recovers the uniform electron gas limit of the uncorrected density functional approximation, and reduces to PZSIC method as a special case when isoorbital indicator is set to the unity. Here, we present a numerical scheme that we have adopted to evaluate the Coulomb potential of the electron density scaled by the iso-orbital indicator required for the self-consistent LSIC calculations. After analyzing the behavior of the finite difference method (FDM) and the green function solution to the radial part of the Poisson equation, we adopt a hybrid approach that uses the FDM for the Coulomb potential due to the monopole and the GF for all higher-order terms. The performance of the resultant hybrid method is assessed using a variety of systems. The results show improved accuracy than earlier numerical schemes. We also find that, even with a generic set of radial grid parameters, accurate energy differences can be obtained using a numerical Coulomb solver in standard density functional studies.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

A Semi-Algebraic Two Level Solver

We develop a simple semi-algebraic 2-level solver built on traditional multigrid ideas. It is designed to be easily incorporated into existing simulation software. It exhibits good convergence for many classes of challenging problems including discontinuous diffusion, convection- diffusion, and Helmholtz equations. It has built-in structure that makes it simple to generalize in several interesting directions.

97 MATHEMATICS AND COMPUTING↗

A multipurpose lifting-line flow solver for arbitrary wind energy concepts

Abstract. In this work, we extend the AeroDyn module of OpenFAST to support arbitrary collections of wings, rotors, and towers. The new standalone AeroDyn driver supports arbitrary motions of the lifting surfaces and complex turbulent inflows. Aerodynamics and inflow are assembled into one module that can be readily coupled with an elastic solver. We describe the features and updates necessary for the implementation of the new AeroDyn driver. We present different case studies of the driver to illustrate its application to concepts such as multirotors, kites, or vertical-axis wind turbines. We perform verification and validation of some of the new features using the following test cases: elliptical wings, horizontal-axis wind turbines, and 2D and 3D vertical-axis wind turbines. The wind turbine simulations are compared to existing tools and field measurements. We use this opportunity to describe some limitations of current models and to highlight areas that we think should be the focus of future research in wind turbine aerodynamics.

17 WIND ENERGY↗

A compressible Navier-Stokes solver with two-equation and Reynolds stress turbulence closure models

This report outlines the development of a general purpose aerodynamic solver for compressible turbulent flows. Turbulent closure is achieved using either two equation or Reynolds stress transportation equations. The applicable equation set consists of Favre-averaged conservation equations for the mass, momentum and total energy, and transport equations for the turbulent stresses and turbulent dissipation rate. In order to develop a scheme with good shock capturing capabilities, good accuracy and general geometric capabilities, a multi-block cell centered finite volume approach is used. Viscous fluxes are discretized using a finite volume representation of a central difference operator and the source terms are treated as an integral over the control volume. The methodology is validated by testing the algorithm on both two and three dimensional flows. Both the two equation and Reynolds stress models are used on a two dimensional 10 degree compression ramp at Mach 3, and the two equation model is used on the three dimensional flow over a cone at angle of attack at Mach 3.5. With the development of this algorithm, it is now possible to compute complex, compressible high speed flow fields using both two equation and Reynolds stress turbulent closure models, with the capability of eventually evaluating their predictive performance.

Navier-Stoke solver↗

Using parallel banded linear system solvers in generalized eigenvalue problems

Subspace iteration is a reliable and cost effective method for solving positive definite banded symmetric generalized eigenproblems, especially in the case of large scale problems. This paper discusses an algorithm that makes use of two parallel banded solvers in subspace iteration. A shift is introduced to decompose the banded linear systems into relatively independent subsystems and to accelerate the iterations. With this shift, an eigenproblem is mapped efficiently into the memories of a multiprocessor and a high speedup is obtained for parallel implementations. An optimal shift is a shift that balances total computation and communication costs. Under certain conditions, we show how to estimate an optimal shift analytically using the decay rate for the inverse of a banded matrix, and how to improve this estimate. Computational results on iPSC/2 and iPSC/860 multiprocessors are presented.

DISTRIBUTED MEMORY MULTIPROCES↗

Development of a One-Domain Volume-Averaged Navier–Stokes Solver

The interaction between a high-enthalpy flow and a thermal protection material is inherently multiscale and multiphysics. In conventional aerothermal analyses, the external flow and material response are generally modeled using separate computational domains coupled through boundary conditions at the material surface. Although this approach has supported many practical applications, it requires assumptions about the location and behavior of the interface and may become difficult to apply when material decomposition, internal reactions, and surface recession substantially alter the porous structure. This report presents the development of a one-domain formulation in which the free-fluid and porous-material regions are represented within a single computational domain. The formulation is based on the volume-averaged Navier–Stokes (VANS) equations, derived from the governing equations for reacting, compressible flow and condensed material. Volume averaging transfers the influence of the unresolved material microstructure to the macroscale equations through effective transport properties, interfacial source terms, and dispersion fluxes. Particular attention is given to regions in which porosity and permeability vary rapidly, including the diffuse transition between a porous material and the surrounding fluid. The resulting equations are implemented in the Porous-material Analysis Toolbox based on OpenFOAM (PATO). The report describes the pressure–velocity coupling strategy used by the solver, examines spatial filtering techniques for deriving effective properties, and evaluates the influence of a smoothly varying interface permeability. Numerical demonstrations include canonical porous-flow configurations, a flow-tube configuration representative of FiberForm® permeability experiments, and the oxidation of a porous carbon material. The purpose of this work is to establish a mathematical and computational foundation for a unified treatment of flow and thermal protection material response. The present formulation is intended to support the progressive inclusion of additional physical processes, including multicomponent transport, finite-rate gas–surface chemistry, pyrolysis, internal oxidation, and material recession. It also provides a framework for connecting pore-scale simulations and microstructural characterization with macroscale aerothermal-response calculations. This report is intended for researchers and engineers working in computational fluid dynamics, porous-media transport, material response, and thermal protection system modeling. It documents both the theoretical development and the initial numerical assessment of the one-domain approach, while identifying the closure of effective and dispersion terms as an important subject for continued investigation.

Ablation↗

Using Griffin's Transmutation Solver to Calculate Radiation Damage

Displacement Radiation Damage originates from all nuclides, not just those that are naturally occurring. Currently only damage from naturally-occurring nuclides, or sometimes damage from one transmutation product is considered. It is proposed that the transmutation solvers implemented in many codes be used to calculate this radiation damage. This would explicitly treat damage from all sources without additionally burdening the user. A proof-of-concept implementation was created in Griffin. The implementation showed that only minor code modifications are necessary to add this feature. When compared against an analytical benchmark the results Griffin could calculate were very accurate.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗