Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The alpha(3) Scheme - A Fourth-Order Neutrally Stable CESE Solver

The conservation element and solution element (CESE) development is driven by a belief that a solver should (i) enforce conservation laws in both space and time, and (ii) be built from a non-dissipative (i.e., neutrally stable) core scheme so that the numerical dissipation can be controlled effectively. To provide a solid foundation for a systematic CESE development of high order schemes, in this paper we describe a new 4th-order neutrally stable CESE solver of the advection equation Theta u/Theta + alpha Theta u/Theta x = 0. The space-time stencil of this two-level explicit scheme is formed by one point at the upper time level and three points at the lower time level. Because it is associated with three independent mesh variables u(sup n) (sub j), (u(sub x))(sup n) (sub j) , and (uxz)(sup n) (sub j) (the numerical analogues of u, Theta u/Theta x, and Theta(exp 2)u/Theta x(exp 2), respectively) and four equations per mesh point, the new scheme is referred to as the alpha(3) scheme. As in the case of other similar CESE neutrally stable solvers, the alpha(3) scheme enforces conservation laws in space-time locally and globally, and it has the basic, forward marching, and backward marching forms. These forms are equivalent and satisfy a space-time inversion (STI) invariant property which is shared by the advection equation. Based on the concept of STI invariance, a set of algebraic relations is developed and used to prove that the alpha(3) scheme must be neutrally stable when it is stable. Moreover it is proved rigorously that all three amplification factors of the alpha(3) scheme are of unit magnitude for all phase angles if |v| <= 1/2 (v = alpha delta t/delta x). This theoretical result is consistent with the numerical stability condition |v| <= 1/2. Through numerical experiments, it is established that the alpha(3) scheme generally is (i) 4th-order accurate for the mesh variables u(sup n) (sub j) and (ux)(sup n) (sub j); and 2nd-order accurate for (uxx)(sup n) (sub j). However, in some exceptional cases, the scheme can achieve perfect accuracy aside from round-off errors.

Chang, Sin-Chung↗

The a(4) Scheme-A High Order Neutrally Stable CESE Solver

The CESE development is driven by a belief that a solver should (i) enforce conservation laws in both space and time, and (ii) be built from a nondissipative (i.e., neutrally stable) core scheme so that the numerical dissipation can be controlled effectively. To provide a solid foundation for a systematic CESE development of high order schemes, in this paper we describe a new high order (4-5th order) and neutrally stable CESE solver of a 1D advection equation with a constant advection speed a. The space-time stencil of this two-level explicit scheme is formed by one point at the upper time level and two points at the lower time level. Because it is associated with four independent mesh variables (the numerical analogues of the dependent variable and its first, second, and third-order spatial derivatives) and four equations per mesh point, the new scheme is referred to as the a(4) scheme. As in the case of other similar CESE neutrally stable solvers, the a(4) scheme enforces conservation laws in space-time locally and globally, and it has the basic, forward marching, and backward marching forms. Except for a singular case, these forms are equivalent and satisfy a space-time inversion (STI) invariant property which is shared by the advection equation. Based on the concept of STI invariance, a set of algebraic relations is developed and used to prove the a(4) scheme must be neutrally stable when it is stable. Numerically, it has been established that the scheme is stable if the value of the Courant number is less than 1/3

Chang, Sin-Chung↗

Overset Techniques for Hypersonic Multibody Configurations with the DPLR Solver

Three unit problems in shock-shock/shock-boundary layer interactions are considered in the evaluation overset techniques with the Data Parallel Line Relaxation (DPLR) computational fluid dynamics solver, a three dimensional Navier-Stokes solver . The unit problems considered are those of two stacked hemispherical cylinders (of different diameters and lengths, and at various orientations relative to each other or relative to the nozzle axis) tested in a hypersonic wind tunnel. These problems are taken as representative of a Two-Stage-To-Orbit design. The objective of the present presentation would be to discuss the techniques used to develop suitable overset grid systems and then evaluate their respective solutions by comparing to corresponding point matched grid solutions and experimental data. Both successful and unsuccessful techniques would be discussed. All solutions would be calculated using the DPLR solver and SUGGAR will be used to develop the domain connectivity information.

Hyatt, Andrew James↗

Modularization and Validation of FUN3D as a CREATE-AV Helios Near-Body Solver

Under a recent collaborative effort between the US Army Aeroflightdynamics Directorate (AFDD) and NASA Langley, NASA's general unstructured CFD solver, FUN3D, was modularized as a CREATE-AV Helios near-body unstructured grid solver. The strategies adopted in Helios/FUN3D integration effort are described. A validation study of the new capability is performed for rotorcraft cases spanning hover prediction, airloads prediction, coupling with computational structural dynamics, counter-rotating dual-rotor configurations, and free-flight trim. The integration of FUN3D, along with the previously integrated NASA OVERFLOW solver, lays the ground for future interaction opportunities where capabilities of one component could be leveraged with those of others in a relatively seamless fashion within CREATE-AV Helios.

Jain, Rohit↗

A Survey of Solver-Related Geometry and Meshing Issues

There is a concern in the computational fluid dynamics community that mesh generation is a significant bottleneck in the CFD workflow. This is one of several papers that will help set the stage for a moderated panel discussion addressing this issue. Although certain general "rules of thumb" and a priori mesh metrics can be used to ensure that some base level of mesh quality is achieved, inadequate consideration is often given to the type of solver or particular flow regime on which the mesh will be utilized. This paper explores how an analyst may want to think differently about a mesh based on considerations such as if a flow is compressible vs. incompressible or hypersonic vs. subsonic or if the solver is node-centered vs. cell-centered. This paper is a high-level investigation intended to provide general insight into how considering the nature of the solver or flow when performing mesh generation has the potential to increase the accuracy and/or robustness of the solution and drive the mesh generation process to a state where it is no longer a hindrance to the analysis process.

Masters, James↗

An Optimized Multicolor Point-Implicit Solver for Unstructured Grid Applications on Graphics Processing Units

In the field of computational fluid dynamics, the Navier-Stokes equations are often solved using an unstructuredgrid approach to accommodate geometric complexity. Implicit solution methodologies for such spatial discretizations generally require frequent solution of large tightly-coupled systems of block-sparse linear equations. The multicolor point-implicit solver used in the current work typically requires a significant fraction of the overall application run time. In this work, an efficient implementation of the solver for graphics processing units is proposed. Several factors present unique challenges to achieving an efficient implementation in this environment. These include the variable amount of parallelism available in different kernel calls, indirect memory access patterns, low arithmetic intensity, and the requirement to support variable block sizes. In this work, the solver is reformulated to use standard sparse and dense Basic Linear Algebra Subprograms (BLAS) functions. However, numerical experiments show that the performance of the BLAS functions available in existing CUDA libraries is suboptimal for matrices representative of those encountered in actual simulations. Instead, optimized versions of these functions are developed. Depending on block size, the new implementations show performance gains of up to 7x over the existing CUDA library functions.

Zubair, Mohammad↗

Optimization of a Solver for Computational Materials and Structures Problems on NVIDIA Volta and AMD Instinct GPUs

The Scalable Implementation of Finite Elements by NASA (ScIFEN) is a software package developed to solve complex computational materials and structures problems using the finite element method (FEM). In this paper, we describe optimization techniques to speed up the linear solver computation that occurs within the ScIFEN application. We consider GPUs from two different vendors, NVIDIA and AMD as our target platforms for optimization and highlight differences in performance and optimization techniques. The NVIDIA GPU Volta V100 is used in the Summit system deployed at Oak Ridge National Laboratory, and the new exascale system, Frontier, will be using AMD Radeon Instinct GPU. We evaluated the performance of various optimization techniques on test matrices, ranging in size from100K to 4M, that are representative of ScIFEN applications. The linear solver computation is memory-bound on both GPUs. Our experiments show that on the NVIDIA GPU we obtained up to79%of the theoretical peak bandwidth, while the AMD GPU achieved 59%. Overall, the NVIDIA V100 GPU outperforms the AMD MI 25 GPU1. We observed an overall speedup of up to37X on an NVIDIA V100 compared to an Intel Skylake 12-coremachine. The solver for a 4M degree of freedom system took under 2.5 seconds.

Mohammad Zubair↗

Performance and Portability of a Linear Solver Across Emerging Architectures

A linear solver algorithm used by a large-scale unstructured-grid computational fluid dynamics application is examined for a broad range of familiar and emerging architectures. Efficient implementation of a linear solver is challenging on recent CPUs offering vector architectures. Vector loads and stores are essential to effectively utilize available memory bandwidth on CPUs, and maintaining performance across different CPUs can be difficult in the face of varying vector lengths offered by each. A similar challenge occurs on GPU architectures, where it is essential to have coalesced memory accesses to utilize memory bandwidth effectively. In this work, we demonstrate that restructuring a computation, and possibly data layout, with regard to architecture is essential to achieve optimal performance by establishing a performance benchmark for each target architecture in a low level language such as vector intrinsics or CUDA. In doing so, we demonstrate how a linear solver kernel can be mapped to Intel® Xeon™ and Xeon Phi™, Marvell® ThunderX2®, NEC® SX-Aurora™ TSUBASA Vector Engine, and NVIDIA® and AMD® GPUs. We further demonstrate that the required code restructuring can be achieved in higher level programming environments such as OpenACC, OCCA, and Intel® OneAPI™/SYCL, and that each generally results in optimal performance on the target architecture. Relative performance metrics for all implementations are shown, and subjective ratings for ease of implementation and optimization are suggested.

Programming models↗

Scalability of Cohesive Fatigue Analyses Using Explicit Solvers

A cohesive fatigue law has been integrated into a constitutive material model compatible with an explicit finite element solver. The cohesive fatigue model response is based on engineering approximations of the endurance limit and the Goodman diagram. This approach can predict stress-life diagrams for crack initiation, the Paris law regime, and transient effects of crack initiation and stable tearing. Simplified cyclic loading is utilized so that the applied load(or displacement) corresponds to the peak load of a fatigue cycle. Loads are held constant during fatigue while damage develops with increasing solution increments. An automatically-calculated ratio of fatigue cycles per solution increment controls the rate of damage growth, ensuring that damage growth is modeled with a sufficient minimum number of increments and damage growth advances to a minimum desired extent within the explicit analysis step time. The compatibility with an explicit finite element solver enables the analysis of structures that are computationally intractable for implicit finite element solvers. Scalability studies are conducted for geometrically nonlinear problems involving fiber-reinforced composite structures that exhibit fatigue damage growth of interacting matrix cracks and delaminations

Frank A Leone↗

Improvements to a Batch Pentadiagonal Solver on NVIDIA GPUs

This poster presents the recent work in OVERFLOW to port the batched pentadiagonal solver to NVIDIA GPUs. There are five pentadiagonal systems for each pencil in the grid but three of these systems share the same LHS. Our first simple approach for porting the pentadiagonal solver to the GPUs was to take advantage of the shared LHS by assigning three threads to the three LHS of each pencil. We demonstrated that this custom solver was 92% faster than the NVIDIA batched pentadiagonal library implementation on a V100 GPU due to the lower memory bandwidth requirements. The second approach treated each pentadiagonal system as a 2x2 block tridiagonal system and used a variant of the parallel cyclic reduction algorithm to solve the problem. One benefit of this approach is that it does not require interleaving the data between each system. We demonstrated that this algorithm is 2.18x faster than the NVIDIA library implementation for the same amount of work. If we take advantage of our shared LHS, this approach is 2.58x faster than the library implementation on a V100 GPU.

GPU Programming↗

Efficient Preconditioning of a High-Order Solver for Multiple Physics

This work addresses preconditioning approaches for an implicit high-order solver frame-work applied to multiple physics. The solver is based on a space-time spectral element method and matrix-free Newton-Krylov solver developed at NASA over the recent years. Within this context, most preconditioning methods are impractical, as the computational time and memory requirements scale poorly with increasing polynomial orders. To improve computational efficiency, we first describe a novel entity-based Block Jacobi preconditioner for the continuous-Galerkin solution of the linear-elasticity and linear-shell equations. Second, we introduce a multigrid algorithm to further reduce time-to-solution on stiff cases arising from continuous-and discontinuous-Galerkin discretizations. Results obtained on relevant single-physics reference solutions, demonstrate the feasibility of the methods, paving the way for high-order solutions of fully coupled multi-physics problems.

STMD↗

Performance of Coupled Physics Solvers for Multidisciplinary Hypersonic Flow Simulations on Several Classes of Computer Architectures

The application of hypersonic flow simulation tools to realistic flight scenarios will require the coupling of multiple physical effects to the baseline fluid dynamics. Such multiphysics effects can include the aerooelastic response of the airframe or engine components, dynamic transport of atmospheric particles, the deformation of solid-fluid interfaces that can ablate, pyrolyze, or erode, as well as a host of other processes, all of which are governed by unique sets of physical equations and models. Coupling multiple (and potentially disparate) physics solvers to a robust compressible flow solver poses additional challenges related to the stability, performance and scalability of the combined solver. The choices made during the software design process can therefore lead to a variation in simulation efficiency across different computer architectures. In this paper, we will consider two representative multiphysics hypersonic flow scenarios: the interaction of solid particulates with the flow field created by a hypersonic lifting body and the aerooelastic deformation of a model airframe under high-Mach-number flow conditions. For these simulations we explore the behavior of several hypersonic simulation tools, including Kestrel, FUN3D, US3D, and JENRE multiphysics framework, on several high performance computing systems containing various CPU and GPU architectures.

architecture↗

Edge-Based Viscous Method for Mixed-Element Node-Centered Finite-Volume Solvers

A novel, efficient, edge-based viscous (EBV) discretization method has been recently developed, implemented in a practical, unstructured-grid, node-centered, finite-volume flow solver, and applied to viscous-kernel computations that include evaluations of meanflow viscous fluxes, turbulence-model and chemistry-model diffusion terms, and the corresponding Jacobian contributions. Initially, the EBV method had been implemented for tetrahedral grids and demonstrated multifold acceleration of all viscous-kernel computations. This paper presents an extension of the EBV method for mixed-element grids. In addition to the primal edges of a given mixed-element grid, virtual edges are introduced to connect cell nodes that are not connected by a primal edge. The EBV method uses an efficient loop over all (primal and virtual) edges and features a compact discretization stencil based on the nearest neighbors. This study verifies the EBV method and assesses its efficiency on mixed-element grids by comparing the EBV solution accuracy and iterative convergence with those of well-established solutions obtained using a cell-based viscous (CBV) discretization method. The EBV solver’s memory footprint is optimized and often smaller than the memory footprint of the CBV solver. A multifold speedup is demonstrated for all viscous-kernel computations resulting in significant reduction of the time to solutions for several benchmark mixed-element-grid computations, including simulations of a flow around NASA’s juncture-flow model and a hypersonic, chemically reacting flow around a blunt body.

CFD↗

Edge-Based Viscous Method for Mixed-Element Node-Centered Finite-Volume Solvers

A novel, efficient, edge-based viscous (EBV) discretization method has been recently developed, implemented in a practical, unstructured-grid, node-centered, finite-volume flow solver, and applied to viscous-kernel computations that include evaluations of meanflow viscous fluxes, turbulence-model and chemistry-model diffusion terms, and the corresponding Jacobian contributions. Initially, the EBV method had been implemented for tetrahedral grids and demonstrated multifold acceleration of all viscous-kernel computations. This paper presents an extension of the EBV method for mixed-element grids. In addition to the primal edges of a given mixed-element grid, virtual edges are introduced to connect cell nodes that are not connected by a primal edge. The EBV method uses an efficient loop over all (primal and virtual) edges and features a compact discretization stencil based on the nearest neighbors. This study verifies the EBV method and assesses its efficiency on mixed-element grids by comparing the EBV solution accuracy and iterative convergence with those of well-established solutions obtained using a cell-based viscous (CBV) discretization method. The EBV solver’s memory footprint is optimized and often smaller than the memory footprint of the CBV solver. A multifold speedup is demonstrated for all viscous-kernel computations resulting in significant reduction of the time to solutions for several benchmark mixed-element-grid computations, including simulations of a flow around NASA’s juncture-flow model and a hypersonic, chemically reacting flow around a blunt body.

Edge-based viscous method↗

Computational Analysis of a Boundary-Layer Ingesting Tailcone Thruster Configuration Within the National Transonic Facility Using the LAVA Solver

Boundary-layer ingesting (BLI) propulsion systems are one of the many technologies currently under investigation within the aerospace community to meet NASA’s Advanced Air Transport Technology project goal of a sustainable future in aviation. To that end, an experimental campaign was conducted in the National Transonic Facility (NTF) to study the propulsion-airframe integration effects of the Boundary-Layer Ingesting Tailcone System installed in the aft portion of a 2.7% scale design of the NASA Common Research Model (CRM). This test was accompanied by a numerical simulation effort using the Launch, Ascent, and Vehicle Aerodynamics (LAVA) solver, to validate the applicability of current best-practice Reynolds-averaged Navier Stokes (RANS) models in accurately predicting the relevant flow-physics in free-air. A total of 205 steady RANS simulations were performed using LAVA, covering the range of conditions present in the NTF test matrix. Flow conditions ranged from 5- to 15-million Reynolds number, Mach numbers of 0.75, 0.80 and 0.85, and angles-of-attack between –3° and 4°. Four different mass-flow plugs were also tested to assess propulsor operating condition effects. Sensitivity to these flow conditions are analyzed in detail, especially in terms of the nacelle flow distortion which was the main subject of the study. Aerodynamic load coefficients, aftbody boundary-layer measurements and fuselage pressure tap results are also discussed, providing a wide range of validation results from this experimental campaign. This validation effort shows that current best-practices using RANS models within the LAVA solver flow solver are well suited for predicting the inlet flow distortion characteristics of novel aircraft configurations employing BLI at the aft end of the fuselage.

AATT↗

Implicit Preconditioning for Explicit Multigrid Solvers on Cut-Cell Cartesian Meshes

This work assesses the effectiveness of linearized implicit Euler preconditioning for multigrid solvers using an unpreconditioned, Jacobian-free Newton Krylov method to converge the linear system of equations. Multigrid convergence rates improve to approximately 0.75 across the cases tested including a Mach 2 supersonic wedge, transonic NACA 0012 airfoil, and ONERA M6 wing. While larger Krylov subspaces increase the convergence rate, they also increase the computational cost, such that 4-8 Krylov vectors often offers the fastest turnaround. Further reductions in computational cost are achieved with a sequential hybrid preconditioner that begins with the explicit multigrid solver before transitioning to the preconditioned algorithm later on. In addition, a novel implementation of dual time stepping is extended to include both common BDF methods as well as high-order implicit Runge-Kutta schemes. This particular formulation, which uses A −1 preconditioning, is amenable to matrix-free solvers, and the L-stable methods are especially suited for meshes with arbitrarily small cut-cells. Asymptotic order of convergence is demonstrated for BDF1, BDF2, SDIRK2, and 3rd-order Radau IIA time integration with unsteady 2D vortex simulations.

ARMD↗

Computational Analysis of a Boundary Layer Ingesting Tailcone Thruster Configuration Using the LAVA Curvilinear Solver

Boundary layer ingesting (BLI) propulsion systems are one of the many technologies currently under investigation within the aerospace community to meet NASA's Advanced Air Transport Technology project goal of a sustainable future in aviation. To that end, an experimental campaign was conducted in the National Transonic Facility (NTF) wind tunnel to study the propulsion-airframe integration effects of the Boundary Layer Ingesting Tailcone System installed in the aft portion of a 2.7% scale design of the NASA Common Research Model (CRM). This test was accompanied by a numerical simulation effort using the Launch, Ascent, and Vehicle Aerodynamics (LAVA) curvilinear solver to validate the applicability of current best-practice Reynolds-Averaged Navier-Stokes (RANS) models in accurately predicting the relevant flow physics in free-air. A total of 205 steady RANS simulations were performed using LAVA, covering the range of conditions present in the NTF test matrix. Flow conditions ranged from 5- to 15-million Reynolds number, Mach numbers of 0.75, 0.80 and 0.85, and angles of attack between -3 and 4 degrees. Four different mass-flow plugs were also tested to assess propulsor operating condition effects. Sensitivity to these flow conditions is analyzed in detail, especially in terms of the nacelle flow distortion which was the main subject of the study. Aerodynamic load coefficients, aft body boundary layer measurements and fuselage pressure tap results are also discussed, providing a wide range of validation results from this experimental campaign. This validation effort shows that current best-practices using RANS models within the LAVA curvilinear solver flow solver are well suited for predicting the inlet flow distortion characteristics of novel aircraft configurations employing BLI at the aft end of the fuselage. This talk will cover the combined experimental and numerical efforts that resulted from this BLI-focused investigation.

ARMD↗

Approximate Inverse Chain Preconditioner: Iteration Count Case Study for Spectral Support Solvers

As the growing availability of computational power slows, there has been an increasing reliance on algorithmic advances. However, faster algorithms alone will not necessarily bridge the gap in allowing computational scientists to study problems at the edge of scientific discovery in the next several decades. Often, it is necessary to simplify or precondition solvers to accelerate the study of large systems of linear equations commonly seen in a number of scientific fields. Preconditioning a problem to increase efficiency is often seen as the best approach; yet, preconditioners which are fast, smart, and efficient do not always exist. Following the progress of [1], we present a new preconditioner for symmetric diagonally dominant (SDD) systems of linear equations. These systems are common in certain PDEs, network science, and supervised learning among others. Based on spectral support graph theory, this new preconditioner builds off of the work of [2], computing and applying a V-cycle chain of approximate inverse matrices. This preconditioner approach is both algebraic in nature as well as hierarchically-constrained depending on the condition number of the system to be solved. Due to its generation of an Approximate Inverse Chain of matrices, we refer to this as the AIC preconditioner. We further accelerate the AIC preconditioner by utilizing precomputations to simplify setup and multiplications in the con-text of an iterative Krylov-subspace solver. While these iterative solvers can greatly reduce solution time, the number of iterations can grow large quickly in the absence of good preconditioners. Initial results for the AIC preconditioner have shown a very large reduction in iteration counts for SDD systems as compared to standard preconditioners such as Incomplete Cholesky (ICC) and Multigrid (MG). We further show significant reduction in iteration counts against the more advanced Combinatorial Multigrid (CMG) preconditioner. We have further developed no-fill sparsification techniques to ensure that the computational cost of applying the AIC preconditioner does not grow prohibitively large as the depth of the V-cycle grows for systems with larger condition numbers. Our numerical results have shown that these sparsifiers maintain the sparsity structure of our system while also displaying significant reductions in iteration counts.1 2

97 MATHEMATICS AND COMPUTING↗