Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stencil computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A New Approach for Constructing Highly Stable High Order CESE Schemes

A new approach is devised to construct high order CESE schemes which would avoid the common shortcomings of traditional high order schemes including: (a) susceptibility to computational instabilities; (b) computational inefficiency due to their local implicit nature (i.e., at each mesh points, need to solve a system of linear/nonlinear equations involving all the mesh variables associated with this mesh point); (c) use of large and elaborate stencils which complicates boundary treatments and also makes efficient parallel computing much harder; (d) difficulties in applications involving complex geometries; and (e) use of problem-specific techniques which are needed to overcome stability problems but often cause undesirable side effects. In fact it will be shown that, with the aid of a conceptual leap, one can build from a given 2nd-order CESE scheme its 4th-, 6th-, 8th-,... order versions which have the same stencil and same stability conditions of the 2nd-order scheme, and also retain all other advantages of the latter scheme. A sketch of multidimensional extensions will also be provided.

Chang, Sin-Chung↗

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Practical aspects of spatially high accurate methods

The computational qualities of high order spatially accurate methods for the finite volume solution of the Euler equations are presented. Two dimensional essentially non-oscillatory (ENO), k-exact, and 'dimension by dimension' ENO reconstruction operators are discussed and compared in terms of reconstruction and solution accuracy, computational cost and oscillatory behavior in supersonic flows with shocks. Inherent steady state convergence difficulties are demonstrated for adaptive stencil algorithms. An exact solution to the heat equation is used to determine reconstruction error, and the computational intensity is reflected in operation counts. Standard MUSCL differencing is included for comparison. Numerical experiments presented include the Ringleb flow for numerical accuracy and a shock reflection problem. A vortex-shock interaction demonstrates the ability of the ENO scheme to excel in simulating unsteady high-frequency flow physics.

Godfrey, Andrew G.↗

A high order Cartesian grid, finite volume method for elliptic interface problems

We present a higher-order finite volume method for solving elliptic PDEs with jump conditions on interfaces embedded in a 2D Cartesian grid. Second, fourth, and sixth order accuracy is demonstrated on a variety of tests including problems with high-contrast and spatially varying coefficients, large discontinuities in the source term, and complex interface geometries. We include a generalized truncation error analysis based on cell-centered Taylor series expansions, which then define stencils in terms of local discrete solution data and geometric information. In the process, we develop a simple method based on Green's theorem for computing exact geometric moments directly from an implicit function definition of the embedded interface. This approach produces stencils with a simple bilinear representation, where spatially-varying coefficients and jump conditions can be easily included and finite volume conservation can be enforced.

97 MATHEMATICS AND COMPUTING↗

Steady and Unsteady Nozzle Simulations Using the Conservation Element and Solution Element Method

This paper presents results from computational fluid dynamic (CFD) simulations of a three-stream plug nozzle. Time-accurate, Euler, quasi-1D and 2D-axisymmetric simulations were performed as part of an effort to provide a CFD-based approach to modeling nozzle dynamics. The CFD code used for the simulations is based on the space-time Conservation Element and Solution Element (CESE) method. Steady-state results were validated using the Wind-US code and a code utilizing the MacCormack method while the unsteady results were partially validated via an aeroacoustic benchmark problem. The CESE steady-state flow field solutions showed excellent agreement with solutions derived from the other methods and codes while preliminary unsteady results for the three-stream plug nozzle are also shown. Additionally, a study was performed to explore the sensitivity of gross thrust computations to the control surface definition. The results showed that most of the sensitivity while computing the gross thrust is attributed to the control surface stencil resolution and choice of stencil end points and not to the control surface definition itself.Finally, comparisons between the quasi-1D and 2D-axisymetric solutions were performed in order to gain insight on whether a quasi-1D solution can capture the steady and unsteady nozzle phenomena without the cost of a 2D-axisymmetric simulation. Initial results show that while the quasi-1D solutions are similar to the 2D-axisymmetric solutions, the inability of the quasi-1D simulations to predict two dimensional phenomena limits its accuracy.

Nozzle gross thrust calculations↗

Time-Accurate Local Time Stepping and High-Order Time CESE Methods for Multi-Dimensional Flows Using Unstructured Meshes

With the wide availability of affordable multiple-core parallel supercomputers, next generation numerical simulations of flow physics are being focused on unsteady computations for problems involving multiple time scales and multiple physics. These simulations require higher solution accuracy than most algorithms and computational fluid dynamics codes currently available. This paper focuses on the developmental effort for high-fidelity multi-dimensional, unstructured-mesh flow solvers using the space-time conservation element, solution element (CESE) framework. Two approaches have been investigated in this research in order to provide high-accuracy, cross-cutting numerical simulations for a variety of flow regimes: 1) time-accurate local time stepping and 2) highorder CESE method. The first approach utilizes consistent numerical formulations in the space-time flux integration to preserve temporal conservation across the cells with different marching time steps. Such approach relieves the stringent time step constraint associated with the smallest time step in the computational domain while preserving temporal accuracy for all the cells. For flows involving multiple scales, both numerical accuracy and efficiency can be significantly enhanced. The second approach extends the current CESE solver to higher-order accuracy. Unlike other existing explicit high-order methods for unstructured meshes, the CESE framework maintains a CFL condition of one for arbitrarily high-order formulations while retaining the same compact stencil as its second-order counterpart. For large-scale unsteady computations, this feature substantially enhances numerical efficiency. Numerical formulations and validations using benchmark problems are discussed in this paper along with realistic examples.

Chang, Chau-Lyan↗

Julia as a unifying end-to-end workflow language on the Frontier exascale system

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.

Godoy, William↗

Efficient and Robust Weighted Least-Squares Cell-Average Gradient Construction Methods for the Simulation of Scramjet Flows

The ability to solve the equations governing the hypersonic turbulent flow of a real gas on unstructured grids using a spatially-elliptic, 2nd-order accurate, cell-centered, finite-volume method has been recently implemented in the VULCAN-CFD code. The construction of cell-average gradients using a weighted linear least-squares method and the use of these gradients in the construction of the inviscid fluxes is the focus of this paper. A comparison of least-squares stencil construction methodologies is presented and approaches designed to minimize the number of cells used to augment/stabilize the least-squares stencil while preserving accuracy are explored. Due to our interest in hypersonic flow, a robust multidimensional cell-average gradient limiter procedure that is consistent with the stencil used to construct the cellaverage gradients is described. Canonical problems are computed to illustrate the challenges and investigate the accuracy, robustness and convergence behavior of the cell-average gradient methods on unstructured cell-centered finite-volume grids. Finally, thermally perfect, chemically frozen, Mach 7.8 turbulent flow of air through a scramjet engine flowpath is computed and compared with experimental data to demonstrate the robustness, accuracy and convergence behavior of the preferred gradient method for a realistic 3-D geometry on a non-hex-dominant grid.

White, Jeffery A.↗

A MultiGPU Performance-Portable Solution for Array Programming Based on Kokkos

Today, multiGPU nodes are widely used in high-performance computing and data centers. However, current programming models do not provide simple, transparent, and portable support for automatically targeting multiple GPUs within a node on application areas of array programming. In this paper, we describe a new application programming interface based on the Kokkos programming model to enable array computation on multiple GPUs in a transparent and portable way across both NVIDIA and AMD GPUs. We implement different variations of this technique to accommodate the exchange of stencils (array boundaries) among different GPU memory spaces, and we provide autotuning to select the proper number of GPUs, depending on the computational cost of the operations to be computed on arrays, that is completely transparent to the programmer. We evaluate our multiGPU extension on Summit (#5 TOP500), with six NVIDIA V100 Volta GPUs per node, and Crusher that contains identical hardware/software as Frontier (#1 TOP500), with four AMD MI250X GPUs, each with 2 Graphics Compute Dies (GCDs)for a total of 8 GCDs per node. We also compare the performance of this solution against the use of MPI + Kokkos, which is the cur-rent de facto solution for multiple GPUs in Kokkos. Our evaluation shows that the new Kokkos solution provides good scalability for many GPUs and a faster and simpler solution (from a programming productivity perspective) than MPI + Kokkos.

Valero Lara, Pedro↗

A scalable exponential-DG approach for nonlinear conservation laws: With application to Burger and Euler equations

In this work, we propose an Exponential DG framework for partial differential equations. We decompose 7 governing equations into linear and nonlinear parts to which we apply the discontinuous Galerkin 8 (DG) spatial discretization. In particular, we construct the linear part using Jacobian that effectively 9 capture stiff characteristics in the system. The former is integrated analytically, whereas the latter 10 is approximated. This approach i) is stable with a large Courant number (Cr > 1); ii) supports 11 high-order solutions both in time and space; iii) is computationally favorable compared to IMEX 12 DG methods with no preconditioner; iv) becomes comparable to explicit RKDG methods on uniform 13 mesh and beneficial on non-uniform grid for Euler equations; v) is scalable in a modern massively 14 parallel computing architecture due to its explicit nature of exponential time integrators and com15 pact communication stencil of DG method. Numerical results demonstrate the performance of our 16 proposed methods through various examples. We also discuss the stability and convergence analysis 17 for our exponential DG scheme in the context of Burgers equation.

42 ENGINEERING↗

Towards a Verifiable Domain-Specific Language for Hardware-Accelerated Stencils

Defining a domain-specific language (DSL) that supports vector-calculus abstractions eases the porting of partial differential equation (PDE) solvers to specialized architectures. Sufficiently high-level abstractions empower users to express universal laws with sufficient generality that the laws must always hold true within their domain of validity. A broad class of PDE solvers employs stencil-based algorithms, the target domain of Berkeley Lab's stencil accelerator chip co-design project. First released as open-source in January 2026, the Formal software framework lays a foundation for defining an embedded DSL based on composable operators that implement mimetic numerical methods -- stencil algorithms that guarantee satisfaction of discrete versions of important vector calculus theorems. The Formal DSL will be the frontend to a new class of stencil-PDE accelerators developed jointly by LBNL, UHCL, and UC Berkeley through the DOE Competitive Portfolios for Computer Science Project. This offers the potential of an order of magnitude acceleration for this important category of computational methods to serve the DOE mission. Future work on the Formal DSL will facilitate software verification via type-safe templates that enable problem-specific correctness proofs relying upon generic function theory and carefully crafted unit tests.

Rouson, Damian↗

Direct Replacement of Arbitrary Grid-Overlapping by Non-Structured Grid

A new approach that uses nonstructured mesh to replace the arbitrarily overlapped structured regions of embedded grids is presented. The present methodology uses the Chimera composite overlapping mesh system so that the physical domain of the flowfield is subdivided into regions which can accommodate easily-generated grid for complex configuration. In addition, a Delaunay triangulation technique generates nonstructured triangular mesh which wraps over the interconnecting region of embedded grids. It is designed that the present approach, termed DRAGON grid, has three important advantages: eliminating some difficulties of the Chimera scheme, such as the orphan points and/or bad quality of interpolation stencils; making grid communication in a fully conservative way; and implementation into three dimensions is straightforward. A computer code based on a time accurate, finite volume, high resolution scheme for solving the compressible Navier-Stokes equations has been further developed to include both the Chimera overset grid and the nonstructured mesh schemes. For steady state problems, the local time stepping accelerates convergence based on a Courant - Friedrichs - Leury (CFL) number near the local stability limit. Numerical tests on representative steady and unsteady supersonic inviscid flows with strong shock waves are demonstrated.

Kao, Kai-Hsiung↗

Optimized Combined Compact Differencing Methods for Computational Aeroacoustics

The sixth order Combined Compact Differencing (CCD) scheme of Chu and Fan is optimized, resulting in exceptional performance. Results indicate that the best of these optimized schemes can accurately propagate linear simple harmonic waves with fewer than three points per wavelength, while using a three point differencing stencil. Stable and accurate boundary treatments for these schemes are given, and the performance of these schemes is assessed using a benchmark problem.

Computational Aeroacoustics↗

Implicit fast sweeping method for hyperbolic systems of conservation laws

Implicit time-accurate methods are often used to integrate stiff problems where explicit schemes impose severe time step restrictions. This paper presents an efficient numerical framework based on the Fast Sweeping Method (FSM) for solving linear and nonlinear hyperbolic systems of conservation laws. The solution at each discrete location is computed by sweeping the numerical domain in several predetermined directions that follow the causality of the characteristic families. The use of a fractional step strategy eliminates the need for a solution selection criterion while one-sided stencils limit the number of sweeps to at most 2 d for d space dimensions. This work focuses on the first-order implicit upwind method since it constitutes the building block for high-order conservative schemes. For problems where the degree of stiffness evolves over time, implicit-explicit hybridization can be accomplished with the same algorithm by simply switching the stencil at each time level. As opposed to traditional implicit solvers, the sweeping method does not require a local time linearization of the fluxes thereby preserving the nonlinear stability properties of the original implicit scheme. It also avoids the large computational and memory requirements associated with solving large block-diagonal systems of equations. Here, a series of one- and two-dimensional test cases are presented for the inviscid Burgers' equation and the reactive Euler equations. The results indicate that the implicit FSM can allow a major reduction in the number of time steps even in the presence of discontinuous solution profiles.

74 ATOMIC AND MOLECULAR PHYSICS↗

A New Method for the Accurate Calculation of Derivatives at Structured Multiblock Grid Singularities

The new Structured Mapping At Reduced dimension Topology elements (SMART) method for the accurate calculation of derivatives at multiblock structured grid block interfaces is presented. In the SMART method, a ’local’ curvilinear mapping is defined at each interface point, giving a consistent multidimensional definition of the derivatives. The resulting system of equations is overdetermined, and a least squares method is used to determine the value of the derivatives in the local coordinate system. In the special case of a nonsingular interface point, the SMART method automatically recovers the spatial differencing stencils used at the interior points in the neighboring grid blocks.

Computational Aeroacoustics↗

Fresh look at floating shock fitting

A fast implicit upwind procedure for the two-dimensional Euler equations is described that allows accurate computations of shocked flows on nonadapted meshes. Away from shocks, the second-order accurate upwinding is based on the split-coefficient-matrix (SCM) method. In the presence of shocks, the difference stencils are modified using a floating shock fitting technique. Rapid convergence to steady-state solutions is attained with a diagonalized approximate factorization (AF) algorithm. Results are presented for Riemann's problem, for a regular shock reflection at an inviscid wall, for supersonic flow past a cylinder, and for a transonic airfoil. All computed shocks are ideally sharp and in excellent agreement with other numerical results or 'exact' solutions. Most importantly, this has been accomplished on unusually crude meshes without any attempt to align grid lines with shock fronts or to cluster grid lines around shocks.

Hartwich, PETER-M.↗