Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stencil computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Fresh look at floating shock fitting

A fast implicit upwind procedure for the two-dimensional Euler equations is described that allows accurate computations of shocked flows on nonadapted meshes. Away from shocks, the second-order accurate upwinding is based on the split-coefficient-matrix (SCM) method. In the presence of shocks, the difference stencils are modified using a floating shock fitting technique. Rapid convergence to steady-state solutions is attained with a diagonalized approximate factorization (AF) algorithm. Results are presented for Riemann's problem, for a regular shock reflection at an inviscid wall, for supersonic flow past a cylinder, and for a transonic airfoil. All computed shocks are ideally sharp and in excellent agreement with other numerical results or 'exact' solutions. Most importantly, this has been accomplished on unusually crude meshes without any attempt to align grid lines with shock fronts or to cluster grid lines around shocks.

Hartwich, PETER-M.↗

Patchy nanoparticles by atomic stencilling

Stencilling, in which patterns are created by painting over masks, has ubiquitous applications in art, architecture and manufacturing. Modern, top-down microfabrication methods have succeeded in reducing mask sizes to under 10 nm, enabling ever smaller microdevices as today’s fastest computer chips. Meanwhile, bottom-up masking using chemical bonds or physical interactions has remained largely unexplored, despite its advantages of low cost, solution-processability, scalability and high compatibility with complex, curved and three-dimensional (3D) surfaces. Here we report atomic stencilling to make patchy nanoparticles (NPs), using surface-adsorbed iodide submonolayers to create the mask and ligand-mediated grafted polymers onto unmasked regions as ‘paint’. We use this approach to synthesize more than 20 different types of NP coated with polymer patches in high yield. Polymer scaling theory and molecular dynamics (MD) simulation show that stencilling, along with the interplay of enthalpic and entropic effects of polymers, generates patchy particle morphologies not reported previously. These polymer-patched NPs self-assemble into extended crystals owing to highly uniform patches, including different non-closely packed superlattices. We propose that atomic stencilling opens new avenues in patterning NPs and other substrates at the nanometre length scale, leading to precise control of their chemistry, reactivity and interactions for a wide range of applications, such as targeted delivery, catalysis, microelectronics, integrated metamaterials and tissue engineering.

36 MATERIALS SCIENCE↗

A 3-D Nodal-Averaged Gradient Approach for Unstructured-Grid Cell-Centered Finite-Volume Methods for Application to Turbulent Hypersonic Flow

A 2-D nodal weighted least-squares gradient method and a related face-averaged nodal gradient approach that were developed for use with triangular grids are extended to 3-D for use with tetrahedral grids. In addition, a method, developed in 2-D, to stabilize the iterative convergence of these methods on quadrilateral cells is described and extended to 3-D and remedies are investigated to determine the nodal gradient averaging approach most suitable for use with grids made up of hexahedral, prismatic, pyramidal and tetrahedral cells. Moreover, due to an interest in hypersonic flow, a robust multidimensional gradient limiter procedure that is consistent with the stencil used to construct the nodal gradients is described. Finally, we demonstrate that the resulting 3-D methods are sufficiently robust for use in scramjet computations through the solution of three canonical turbulent hypersonic flow problems as well as a physically realistic 3-D scramjet inlet geometry.

Jeffery A White↗

A Numerical and Experimental Study of Coflow Laminar Diffusion Flames: Effects of Gravity and Inlet Velocity

In this work, the influence of gravity, fuel dilution, and inlet velocity on the structure, stabilization, and sooting behavior of laminar coflow methane-air diffusion flames was investigated both computationally and experimentally. A series of flames measured in the Structure and Liftoff in Combustion Experiment (SLICE) was assessed numerically under microgravity and normal gravity conditions with the fuel stream CH4 mole fraction ranging from 0.4 to 1.0. Computationally, the MC-Smooth vorticity-velocity formulation of the governing equations was employed to describe the reactive gaseous mixture; the soot evolution process was considered as a classical aerosol dynamics problem and was represented by the sectional aerosol equations. Since each flame is axisymmetric, a two-dimensional computational domain was employed, where the grid on the axisymmetric domain was a nonuniform tensor product mesh. The governing equations and boundary conditions were discretized on the mesh by a nine-point finite difference stencil, with the convective terms approximated by a monotonic upwind scheme and all other derivatives approximated by centered differences. The resulting set of fully coupled, strongly nonlinear equations was solved simultaneously using a damped, modified Newton's method and a nested Bi-CGSTAB linear algebra solver. Experimentally, the flame shape, size, lift-off height, and soot temperature were determined by flame emission images recorded by a digital camera, and the soot volume fraction was quantified through an absolute light calibration using a thermocouple. For a broad spectrum of flames in microgravity and normal gravity, the computed and measured flame quantities (e.g., temperature profile, flame shape, lift-off height, and soot volume fraction) were first compared to assess the accuracy of the numerical model. After its validity was established, the influence of gravity, fuel dilution, and inlet velocity on the structure, stabilization, and sooting tendency of laminar coflow methane-air diffusion flames was explored further by examining quantities derived from the computational results.

microgravity↗

GPU Implementation of the OVERFLOW CFD Code

The high-performance computing (HPC) landscape is quickly changing to systems where most of the performance comes from specialized chips, specifically graphics processing units (GPUs). Such GPU systems are throughput machines, where efficient use of the GPU often requires code refactoring to expose a few orders of magnitude more fine grain parallelism than was previously used on the CPU. Recent modifications to OVERFLOW, an overset, structured grid, computational fluid dynamics flow solver, written in Fortran will be presented. These modifications include both code modernization efforts and algorithmic changes to enable OVERFLOW to efficiently utilize GPUs. Many of these algorithmic changes would likely also be applicable for other structured grid, stencil-based codes wanting to utilize GPUs. The capabilities that have been ported to run on the GPUs are presented, along with the performance gains of the GPU version relative the CPU version of OVERFLOW.

GPU Programming↗

GPU Implementation of the OVERFLOW CFD Code

The high-performance computing (HPC) landscape is quickly changing to systems where most of the performance comes from specialized chips, specifically graphics processing units (GPUs). Such GPU systems are throughput machines, where efficient use of the GPU often requires code refactoring to expose a few orders of magnitude more fine grain parallelism than was previously used on the CPU. Recent modifications to OVERFLOW, an overset, structured grid, computational fluid dynamics flow solver, written in Fortran will be presented. These modifications include both code modernization efforts and algorithmic changes to enable OVERFLOW to efficiently utilize GPUs. Many of these algorithmic changes would likely also be applicable for other structured grid, stencil-based codes wanting to utilize GPUs. The capabilities that have been ported to run on the GPUs are presented, along with the performance gains of the GPU version relative the CPU version of OVERFLOW.

GPU Programming↗

An Automated Approach to Very High Order Aeroacoustic Computations in Complex Geometries

Computational aeroacoustics requires efficient, high-resolution simulation tools. And for smooth problems, this is best accomplished with very high order in space and time methods on small stencils. But the complexity of highly accurate numerical methods can inhibit their practical application, especially in irregular geometries. This complexity is reduced by using a special form of Hermite divided-difference spatial interpolation on Cartesian grids, and a Cauchy-Kowalewslci recursion procedure for time advancement. In addition, a stencil constraint tree reduces the complexity of interpolating grid points that are located near wall boundaries. These procedures are used to automatically develop and implement very high order methods (>15) for solving the linearized Euler equations that can achieve less than one grid point per wavelength resolution away from boundaries by including spatial derivatives of the primitive variables at each grid point. The accuracy of stable surface treatments is currently limited to 11th order for grid aligned boundaries and to 2nd order for irregular boundaries.

Dyson, Rodger W.↗

Automated Approach to Very High-Order Aeroacoustic Computations

Computational aeroacoustics requires efficient, high-resolution simulation tools. For smooth problems, this is best accomplished with very high-order in space and time methods on small stencils. However, the complexity of highly accurate numerical methods can inhibit their practical application, especially in irregular geometries. This complexity is reduced by using a special form of Hermite divided-difference spatial interpolation on Cartesian grids, and a Cauchy-Kowalewski recursion procedure for time advancement. In addition, a stencil constraint tree reduces the complexity of interpolating grid points that am located near wall boundaries. These procedures are used to develop automatically and to implement very high-order methods (> 15) for solving the linearized Euler equations that can achieve less than one grid point per wavelength resolution away from boundaries by including spatial derivatives of the primitive variables at each grid point. The accuracy of stable surface treatments is currently limited to 11th order for grid aligned boundaries and to 2nd order for irregular boundaries.

Dyson, Rodger W.↗

Efficient Cache use for Stencil Operations on Structured Discretization Grids

We derive tight bounds on the cache misses for evaluation of explicit stencil operators on structured grids. Our lower bound is based on the isoperimetrical property of the discrete octahedron. Our upper bound is based on a good surface to volume ratio of a parallelepiped spanned by a reduced basis of the interference lattice of a grid. Measurements show that our algorithm typically reduces the number of cache misses by a factor of three, relative to a compiler optimized code. We show that stencil calculations on grids whose interference lattice have a short vector feature abnormally high numbers of cache misses. We call such grids unfavorable and suggest to avoid these in computations by appropriate padding. By direct measurements on a MIPS R10000 processor we show a good correlation between abnormally high numbers of cache misses and unfavorable three-dimensional grids.

Frumkin, Michael↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

Three-Dimensional High-Order Spectral Volume Method for Solving Maxwell's Equations on Unstructured Grids

A three-dimensional, high-order, conservative, and efficient discontinuous spectral volume (SV) method for the solutions of Maxwell's equations on unstructured grids is presented. The concept of discontinuous 2nd high-order loca1 representations to achieve conservation and high accuracy is utilized in a manner similar to the Discontinuous Galerkin (DG) method, but instead of using a Galerkin finite-element formulation, the SV method is based on a finite-volume approach to attain a simpler formulation. Conventional unstructured finite-volume methods require data reconstruction based on the least-squares formulation using neighboring cell data. Since each unknown employs a different stencil, one must repeat the least-squares inversion for every cell at each time step, or to store the inversion coefficients. In a high-order, three-dimensional computation, the former would involve impractically large CPU time, while for the latter the memory requirement becomes prohibitive. In the SV method, one starts with a relatively coarse grid of triangles or tetrahedra, called spectral volumes (SVs), and partition each SV into a number of structured subcells, called control volumes (CVs), that support a polynomial expansion of a desired degree of precision. The unknowns are cell averages over CVs. If all the SVs are partitioned in a geometrically similar manner, the reconstruction becomes universal as a weighted sum of unknowns, and only a few universal coefficients need to be stored for the surface integrals over CV faces. Since the solution is discontinuous across the SV boundaries, a Riemann solver is thus necessary to maintain conservation. In the paper, multi-parameter and symmetric SV partitions, up to quartic for triangle and cubic for tetrahedron, are first presented. The corresponding weight coefficients for CV face integrals in terms of CV cell averages for each partition are analytically determined. These discretization formulas are then applied to the integral form of the Maxwell equations. All numerical procedures for outer boundary, material interface, zonal interface, and interior SV face are unified with a single characteristic formulation. The load balancing in a massive parallel computing environment is therefore easier to achieve. A parameter is introduced in the Riemann solver to control the strength of the smoothing term. Important aspects of the data structure and its effects to communication and the optimum use of cache memory are discussed. Results will be presented for plane TE and TM waves incident on a perfectly conducting cylinder for up to fifth order of accuracy, and a plane wave incident on a perfectly conducting sphere for up to fourth order of accuracy. Comparisons are made with exact solutions for these cases.

Liu, Yen↗

Assessment of Machine Learning Wall Modeling Approaches for Large Eddy Simulation of Gas Turbine Film Cooling Flows: An a Priori Study

Here, in this work, a priori analysis of machine learning (ML) strategies is carried out with the goal of data-driven wall modeling for large eddy simulation (LES) of gas turbine film cooling flows. High-fidelity flow datasets are extracted from wall-resolved LES (WRLES) of flow over a flat plate interacting with the coolant flow supplied by a single row of 7-7-7 shaped cooling holes inclined at 30 degrees with the flat plate at different blowing ratios (BR). The WRLES are performed using the high-order Nek5000 spectral element computational fluid dynamics (CFD) solver. Light gradient boosting machine (LightGBM) is employed as the ML algorithm for the data-driven wall model. Parametric tests are conducted to systematically assess the influence of a wide range of input flow features (velocity components, velocity gradients, pressure gradients, and fluid properties) on the accuracy of ML wall model with respect to prediction of wall shear stress. In addition, the use of spatial stencil and time delay is also explored within the ML wall modeling framework. It is shown that features associated with gradients of the streamwise and spanwise velocity components have a major impact on the prediction fidelity of wall model, while the effect of gradients of wall-normal velocity component is found to be negligible. Moreover, adding flow feature information from an x-y-z spatial stencil significantly improves the ML model accuracy and generalizability compared to just using local flow features from the matching location. Overall, highest prediction accuracy is achieved when both spatial stencil and time delay features are incorporated within the data-driven wall modeling paradigm.

33 ADVANCED PROPULSION SYSTEMS↗

High resolution upwind schemes for the three-dimensional incompressible Navier-Stokes equations

Based on flux-difference splitting, implicit high resolution schemes are constructed for efficient computations of steady-state solutions to the three-dimensional, incompressible Navier-Stokes equations in curvilinear coordinates. These schemes use first-order accurate Euler backward-time differencing and second-order central differencing for the viscous shear fluxes. Up to third-order accurate upwind differencing is achieved through a reconstruction of the solution from its cell averages. The reconstruction is accomplished by linear interpolation, where the node stencils are selected such that in regions of smooth solution the flow is highly resolved while spurious oscillations in regions of rapid changes in gradient are still suppressed. Fairly rapid convergence to steady-state solutions is attained with a completely vectorizable hybrid time-marching method. Flows around a sharp-edged delta wing are computed with the maximum accuracy of the upwind-differencing restricted to first-, second-, and third-order, to illustrate the effect of accuracy on the global and on the local vortical flow fields. The results are validated with experimental data.

Hartwich, PETER-M.↗

A New Semistructured Algebraic Multigrid Method

Multigrid methods are well suited to large massively parallel computer architectures because they are mathematically optimal and display good parallelization properties. Since current architecture trends are favoring regular compute patterns to achieve high performance, the ability to express structure has become much more important. The hypre software library provides high-performance multigrid preconditioners and solvers through conceptual interfaces, including a semistructured interface that describes matrices primarily in terms of stencils and logically structured grids. This paper presents a new semistructured algebraic multigrid (SSAMG) method built on this interface. The numerical convergence and performance of a CPU implementation of this method are evaluated for a set of semistructured problems. In conclusion, SSAMG achieves significantly better setup times than hypre’s unstructured AMG solvers and comparable convergence. In addition, the new method is capable of solving more complex problems than hypre’s structured solvers.

97 MATHEMATICS AND COMPUTING↗

Computing unsteady shock waves for aeroacoustic applications

The computation of unsteady shock waves, which contribute significantly to noise generation in supersonic jet flows, is investigated. This paper focuses on the difficulties of computing slowly moving shock waves. Numerical error is found to manifest itself principally as a spurious entropy wave. Calculations presented are performed using a third order essentially nonoscillatory scheme. The effect of stencil biasing parameters and of two versions of numerical flux formulas on the magnitude of spurious entropy are investigated. The level of numerical error introduced in the calculation in quantified as a function of shock pressure ratio, shock speed, Courant number, and mesh density. The spurious entropy relative to the entropy jump across a static shock decreases with increasing shock strength and shock velocity relative to the grid, but is insensitive to Courant number. The structure of the spurious entropy wave is affected by the choice of flux formulas and algorithm biasing parameters. The effect of the spurious numerical waves on the calculation of sound amplification by a shock wave is investigated. For this class of problem, the acoustic pressure waves are relatively unaffected by the spurious numerical phenomena.

Meadows,, Kristine r.↗

Discontinuous Spectral Difference Method for Conservation Laws on Unstructured Grids

A new, high-order, conservative, and efficient discontinuous spectral finite difference (SD) method for conservation laws on unstructured grids is developed. The concept of discontinuous and high-order local representations to achieve conservation and high accuracy is utilized in a manner similar to the Discontinuous Galerkin (DG) and the Spectral Volume (SV) methods, but while these methods are based on the integrated forms of the equations, the new method is based on the differential form to attain a simpler formulation and higher efficiency. Conventional unstructured finite-difference and finite-volume methods require data reconstruction based on the least-squares formulation using neighboring point or cell data. Since each unknown employs a different stencil, one must repeat the least-squares inversion for every point or cell at each time step, or to store the inversion coefficients. In a high-order, three-dimensional computation, the former would involve impractically large CPU time, while for the latter the memory requirement becomes prohibitive. In addition, the finite-difference method does not satisfy the integral conservation in general. By contrast, the DG and SV methods employ a local, universal reconstruction of a given order of accuracy in each cell in terms of internally defined conservative unknowns. Since the solution is discontinuous across cell boundaries, a Riemann solver is necessary to evaluate boundary flux terms and maintain conservation. In the DG method, a Galerkin finite-element method is employed to update the nodal unknowns within each cell. This requires the inversion of a mass matrix, and the use of quadratures of twice the order of accuracy of the reconstruction to evaluate the surface integrals and additional volume integrals for nonlinear flux functions. In the SV method, the integral conservation law is used to update volume averages over subcells defined by a geometrically similar partition of each grid cell. As the order of accuracy increases, the partitioning for 3D requires the introduction of a large number of parameters, whose optimization to achieve convergence becomes increasingly more difficult. Also, the number of interior facets required to subdivide non-planar faces, and the additional increase in the number of quadrature points for each facet, increases the computational cost greatly.

Liu, Yen↗