Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stencil computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Interference Lattice-based Loop Nest Tilings for Stencil Computations

A common method for improving performance of stencil operations on structured multi-dimensional discretization grids is loop tiling. Tile shapes and sizes are usually determined heuristically, based on the size of the primary data cache. We provide a lower bound on the numbers of cache misses that must be incurred by any tiling, and a close achievable bound using a particular tiling based on the grid interference lattice. The latter tiling is used to derive highly efficient loop orderings. The total number of cache misses of a code is the sum of (necessary) cold misses and misses caused by elements being dropped from the cache between successive loads (replacement misses). Maximizing temporal locality is equivalent to minimizing replacement misses. Temporal locality of loop nests implementing stencil operations is optimized by tilings that avoid data conflicts. We divide the loop nest iteration space into conflict-free tiles, derived from the cache miss equation. The tiling involves the definition of the grid interference lattice an equivalence class of grid points whose images in main memory map to the same location in the cache-and the construction of a special basis for the lattice. Conflicts only occur on the boundaries of the tiles, unless the tiles are too thin. We show that the surface area of the tiles is bounded for grids of any dimensionality, and for caches of any associativity, provided the eccentricity of the fundamental parallelepiped (the tile spanned by the basis) of the lattice is bounded. Eccentricity is determined by two factors, aspect ratio and skewness. The aspect ratio of the parallelepiped can be bounded by appropriate array padding. The skewness can be bounded by the choice of a proper basis. Combining these two strategies ensures that pathologically thin tiles are avoided. They do not, however, minimize replacement misses per se. The reason is that tile visitation order influences the number of data conflicts on the tile boundaries. If two adjacent tiles are visited successively, there will be no replacement misses on the shared boundary. The iteration space may be covered with pencils larger than the size of the cache while avoiding data conflicts if the pencils are traversed by a scanning-face method. Replacement misses are incurred only on the boundaries of the pencils, and the number of misses is minimized by maximizing the volume of the scanning face, not the volume of the tile. We present an algorithm for constructing the most efficient scanning face for a given grid and stencil operator. In two dimensions it is based on a continued fraction algorithm. In three dimensions it follows Voronoi's successive minima algorithm. We show experimental results of using the scanning face, and compare with canonical loop orderings.

VanderWijngaart, Rob F.↗

Compact finite difference schemes with spectral-like resolution

The present finite-difference schemes for the evaluation of first-order, second-order, and higher-order derivatives yield improved representation of a range of scales and may be used on nonuniform meshes. Various boundary conditions may be invoked, and both accurate interpolation and spectral-like filtering can be accomplished by means of schemes for derivatives at mid-cell locations. This family of schemes reduces to the Pade schemes when the maximal formal accuracy constraint is imposed with a specific computational stencil. Attention is given to illustrative applications of these schemes in fluid dynamics.

Lele, Sanjiva K.↗

An Immersed Boundary Method for Solving the Compressible Navier-Stokes Equations with Fluid Structure Interaction

An immersed boundary method for the compressible Navier-Stokes equation and the additional infrastructure that is needed to solve moving boundary problems and fully coupled fluid-structure interaction is described. All the methods described in this paper were implemented in NASA's LAVA solver framework. The underlying immersed boundary method is based on the locally stabilized immersed boundary method that was previously introduced by the authors. In the present paper this method is extended to account for all aspects that are involved for fluid structure interaction simulations, such as fast geometry queries and stencil computations, the treatment of freshly cleared cells, and the coupling of the computational fluid dynamics solver with a linear structural finite element method. The current approach is validated for moving boundary problems with prescribed body motion and fully coupled fluid structure interaction problems in 2D and 3D. As part of the validation procedure, results from the second AIAA aeroelastic prediction workshop are also presented. The current paper is regarded as a proof of concept study, while more advanced methods for fluid structure interaction are currently being investigated, such as geometric and material nonlinearities, and advanced coupling approaches.

Equations↗

High-Order Methods for Computational Fluid Dynamics: A Brief Review of Compact Differential Formulations on Unstructured Grids

Popular high-order schemes with compact stencils for Computational Fluid Dynamics (CFD) include Discontinuous Galerkin (DG), Spectral Difference (SD), and Spectral Volume (SV) methods. The recently proposed Flux Reconstruction (FR) approach or Correction Procedure using Reconstruction (CPR) is based on a differential formulation and provides a unifying framework for these high-order schemes. Here we present a brief review of recent developments for the FR/CPR schemes as well as some pacing items.

Huynh, H. T.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngaart, Rob F.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngarrt, Rob F.↗

Edge Based Viscous Method for Node-Centered Formulations

This paper presents a novel, efficient, conservative, edge-based method for evaluation of mean flow viscous fluxes and turbulence-model diffusion terms of the Reynolds-averaged Navier-Stokes equations on tetrahedral grids. The new method is implemented in a practical, node-centered, finite-volume computational fluid dynamics solver. The baseline finite-volume scheme that is equivalent to a second-order accurate finite-element Galerkin approximation of viscous stresses is reformulated. The order of operations to compute the cell-based Green-Gauss gradients is changed to combine the operations by edge, which leads to an equivalent formulation on tetrahedral grids, improves efficiency, and preserves the compact discretization stencil based on the nearest neighbors. The computational results presented in this paper verify the implementation of this edge-based method by comparing its accuracy and iterative convergence with those of the well verified and validated baseline formulation. Efficiency gains for residual and Jacobian evaluations result in significant reduction of time to solution. This novel edge-based formulation on tetrahedra can be seamlessly combined with the baseline formulation on cells of other types for computing solutions on mixed-element grids.

Edge Based↗

Large Eddy Simulation (LES) of Particle-Laden Temporal Mixing Layers

High-fidelity models of plume-regolith interaction are difficult to develop because of the widely disparate flow conditions that exist in this process. The gas in the core of a rocket plume can often be modeled as a time-dependent, high-temperature, turbulent, reacting continuum flow. However, due to the vacuum conditions on the lunar surface, the mean molecular path in the outer parts of the plume is too long for the continuum assumption to remain valid. Molecular methods are better suited to model this region of the flow. Finally, granular and multiphase flow models must be employed to describe the dust and debris that are displaced from the surface, as well as how a crater is formed in the regolith. At present, standard commercial CFD (computational fluid dynamics) software is not capable of coupling each of these flow regimes to provide an accurate representation of this flow process, necessitating the development of custom software. This software solves the fluid-flow-governing equations in an Eulerian framework, coupled with the particle transport equations that are solved in a Lagrangian framework. It uses a fourth-order explicit Runge-Kutta scheme for temporal integration, an eighth-order central finite differencing scheme for spatial discretization. The non-linear terms in the governing equations are recast in cubic skew symmetric form to reduce aliasing error. The second derivative viscous terms are computed using eighth-order narrow stencils that provide better diffusion for the highest resolved wave numbers. A fourth-order Lagrange interpolation procedure is used to obtain gas-phase variable values at the particle locations.

Bellan, Josette↗

Weighted Least-Squares Cell-Average Gradient Construction Methods for the VULCAN-CFD Second-Order Accurate Unstructured-Grid Cell-Centered Finite-Volume Solver

The ability to solve the equations governing the hypersonic turbulent flow of a real gas on unstructured grids using a spatially-elliptic, 2nd-order accurate, cell-centered, finite-volume method has been recently implemented in the VULCAN-CFD code. The construction of cell-average gradients using a weighted linear least-squares method and the use of these gradients in the construction of the inviscid fluxes is the focus of this paper. A comparison of least-squares stencil construction methodologies is presented and approaches to augment the number of cells participating in the stencil while preserving accuracy are explored. Due to our interest in hypersonic flow, a robust multidimensional cell-average gradient limiter procedure that is consistent with the stencil used to construct the cell-average gradients is described and investigated. Canonical problems are computed to illustrate the challenges and investigate the accuracy, robustness and convergence behavior of the cell-average gradient methods on unstructured cell-centered finite-volume grids. Finally, thermally perfect, chemically frozen, Mach 8 turbulent flow of air around a blunt wedge is computed to demonstrate the robustness and convergence behavior of the new method for constructing stencils of use in a weighted linear least-squares gradient method for a hypersonic flow.

White, Jeffery A.↗

A New Approach for Constructing Highly Stable High Order CESE Schemes

A new approach is devised to construct high order CESE schemes which would avoid the common shortcomings of traditional high order schemes including: (a) susceptibility to computational instabilities; (b) computational inefficiency due to their local implicit nature (i.e., at each mesh points, need to solve a system of linear/nonlinear equations involving all the mesh variables associated with this mesh point); (c) use of large and elaborate stencils which complicates boundary treatments and also makes efficient parallel computing much harder; (d) difficulties in applications involving complex geometries; and (e) use of problem-specific techniques which are needed to overcome stability problems but often cause undesirable side effects. In fact it will be shown that, with the aid of a conceptual leap, one can build from a given 2nd-order CESE scheme its 4th-, 6th-, 8th-,... order versions which have the same stencil and same stability conditions of the 2nd-order scheme, and also retain all other advantages of the latter scheme. A sketch of multidimensional extensions will also be provided.

Chang, Sin-Chung↗

Practical aspects of spatially high accurate methods

The computational qualities of high order spatially accurate methods for the finite volume solution of the Euler equations are presented. Two dimensional essentially non-oscillatory (ENO), k-exact, and 'dimension by dimension' ENO reconstruction operators are discussed and compared in terms of reconstruction and solution accuracy, computational cost and oscillatory behavior in supersonic flows with shocks. Inherent steady state convergence difficulties are demonstrated for adaptive stencil algorithms. An exact solution to the heat equation is used to determine reconstruction error, and the computational intensity is reflected in operation counts. Standard MUSCL differencing is included for comparison. Numerical experiments presented include the Ringleb flow for numerical accuracy and a shock reflection problem. A vortex-shock interaction demonstrates the ability of the ENO scheme to excel in simulating unsteady high-frequency flow physics.

Godfrey, Andrew G.↗

Steady and Unsteady Nozzle Simulations Using the Conservation Element and Solution Element Method

This paper presents results from computational fluid dynamic (CFD) simulations of a three-stream plug nozzle. Time-accurate, Euler, quasi-1D and 2D-axisymmetric simulations were performed as part of an effort to provide a CFD-based approach to modeling nozzle dynamics. The CFD code used for the simulations is based on the space-time Conservation Element and Solution Element (CESE) method. Steady-state results were validated using the Wind-US code and a code utilizing the MacCormack method while the unsteady results were partially validated via an aeroacoustic benchmark problem. The CESE steady-state flow field solutions showed excellent agreement with solutions derived from the other methods and codes while preliminary unsteady results for the three-stream plug nozzle are also shown. Additionally, a study was performed to explore the sensitivity of gross thrust computations to the control surface definition. The results showed that most of the sensitivity while computing the gross thrust is attributed to the control surface stencil resolution and choice of stencil end points and not to the control surface definition itself.Finally, comparisons between the quasi-1D and 2D-axisymetric solutions were performed in order to gain insight on whether a quasi-1D solution can capture the steady and unsteady nozzle phenomena without the cost of a 2D-axisymmetric simulation. Initial results show that while the quasi-1D solutions are similar to the 2D-axisymmetric solutions, the inability of the quasi-1D simulations to predict two dimensional phenomena limits its accuracy.

Nozzle gross thrust calculations↗

Time-Accurate Local Time Stepping and High-Order Time CESE Methods for Multi-Dimensional Flows Using Unstructured Meshes

With the wide availability of affordable multiple-core parallel supercomputers, next generation numerical simulations of flow physics are being focused on unsteady computations for problems involving multiple time scales and multiple physics. These simulations require higher solution accuracy than most algorithms and computational fluid dynamics codes currently available. This paper focuses on the developmental effort for high-fidelity multi-dimensional, unstructured-mesh flow solvers using the space-time conservation element, solution element (CESE) framework. Two approaches have been investigated in this research in order to provide high-accuracy, cross-cutting numerical simulations for a variety of flow regimes: 1) time-accurate local time stepping and 2) highorder CESE method. The first approach utilizes consistent numerical formulations in the space-time flux integration to preserve temporal conservation across the cells with different marching time steps. Such approach relieves the stringent time step constraint associated with the smallest time step in the computational domain while preserving temporal accuracy for all the cells. For flows involving multiple scales, both numerical accuracy and efficiency can be significantly enhanced. The second approach extends the current CESE solver to higher-order accuracy. Unlike other existing explicit high-order methods for unstructured meshes, the CESE framework maintains a CFL condition of one for arbitrarily high-order formulations while retaining the same compact stencil as its second-order counterpart. For large-scale unsteady computations, this feature substantially enhances numerical efficiency. Numerical formulations and validations using benchmark problems are discussed in this paper along with realistic examples.

Chang, Chau-Lyan↗

Efficient and Robust Weighted Least-Squares Cell-Average Gradient Construction Methods for the Simulation of Scramjet Flows

The ability to solve the equations governing the hypersonic turbulent flow of a real gas on unstructured grids using a spatially-elliptic, 2nd-order accurate, cell-centered, finite-volume method has been recently implemented in the VULCAN-CFD code. The construction of cell-average gradients using a weighted linear least-squares method and the use of these gradients in the construction of the inviscid fluxes is the focus of this paper. A comparison of least-squares stencil construction methodologies is presented and approaches designed to minimize the number of cells used to augment/stabilize the least-squares stencil while preserving accuracy are explored. Due to our interest in hypersonic flow, a robust multidimensional cell-average gradient limiter procedure that is consistent with the stencil used to construct the cellaverage gradients is described. Canonical problems are computed to illustrate the challenges and investigate the accuracy, robustness and convergence behavior of the cell-average gradient methods on unstructured cell-centered finite-volume grids. Finally, thermally perfect, chemically frozen, Mach 7.8 turbulent flow of air through a scramjet engine flowpath is computed and compared with experimental data to demonstrate the robustness, accuracy and convergence behavior of the preferred gradient method for a realistic 3-D geometry on a non-hex-dominant grid.

White, Jeffery A.↗

Direct Replacement of Arbitrary Grid-Overlapping by Non-Structured Grid

A new approach that uses nonstructured mesh to replace the arbitrarily overlapped structured regions of embedded grids is presented. The present methodology uses the Chimera composite overlapping mesh system so that the physical domain of the flowfield is subdivided into regions which can accommodate easily-generated grid for complex configuration. In addition, a Delaunay triangulation technique generates nonstructured triangular mesh which wraps over the interconnecting region of embedded grids. It is designed that the present approach, termed DRAGON grid, has three important advantages: eliminating some difficulties of the Chimera scheme, such as the orphan points and/or bad quality of interpolation stencils; making grid communication in a fully conservative way; and implementation into three dimensions is straightforward. A computer code based on a time accurate, finite volume, high resolution scheme for solving the compressible Navier-Stokes equations has been further developed to include both the Chimera overset grid and the nonstructured mesh schemes. For steady state problems, the local time stepping accelerates convergence based on a Courant - Friedrichs - Leury (CFL) number near the local stability limit. Numerical tests on representative steady and unsteady supersonic inviscid flows with strong shock waves are demonstrated.

Kao, Kai-Hsiung↗

Optimized Combined Compact Differencing Methods for Computational Aeroacoustics

The sixth order Combined Compact Differencing (CCD) scheme of Chu and Fan is optimized, resulting in exceptional performance. Results indicate that the best of these optimized schemes can accurately propagate linear simple harmonic waves with fewer than three points per wavelength, while using a three point differencing stencil. Stable and accurate boundary treatments for these schemes are given, and the performance of these schemes is assessed using a benchmark problem.

Computational Aeroacoustics↗

A New Method for the Accurate Calculation of Derivatives at Structured Multiblock Grid Singularities

The new Structured Mapping At Reduced dimension Topology elements (SMART) method for the accurate calculation of derivatives at multiblock structured grid block interfaces is presented. In the SMART method, a ’local’ curvilinear mapping is defined at each interface point, giving a consistent multidimensional definition of the derivatives. The resulting system of equations is overdetermined, and a least squares method is used to determine the value of the derivatives in the local coordinate system. In the special case of a nonsingular interface point, the SMART method automatically recovers the spatial differencing stencils used at the interior points in the neighboring grid blocks.

Computational Aeroacoustics↗

Fresh look at floating shock fitting

A fast implicit upwind procedure for the two-dimensional Euler equations is described that allows accurate computations of shocked flows on nonadapted meshes. Away from shocks, the second-order accurate upwinding is based on the split-coefficient-matrix (SCM) method. In the presence of shocks, the difference stencils are modified using a floating shock fitting technique. Rapid convergence to steady-state solutions is attained with a diagonalized approximate factorization (AF) algorithm. Results are presented for Riemann's problem, for a regular shock reflection at an inviscid wall, for supersonic flow past a cylinder, and for a transonic airfoil. All computed shocks are ideally sharp and in excellent agreement with other numerical results or 'exact' solutions. Most importantly, this has been accomplished on unusually crude meshes without any attempt to align grid lines with shock fronts or to cluster grid lines around shocks.

Hartwich, PETER-M.↗