Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stencil computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Algorithmic Extensions of Low-Dispersion Scheme and Modeling Effects for Acoustic Wave Simulation

Accurate computation of acoustic wave propagation may be more efficiently performed when their dispersion relations are considered. Consequently, computational algorithms which attempt to preserve these relations have been gaining popularity in recent years. In the present paper, the extensions to one such scheme are discussed. By solving the linearized, 2-D Euler and Navier-Stokes equations with such a method for the acoustic wave propagation, several issues were investigated. Among them were higher-order accuracy, choice of boundary conditions and differencing stencils, effects of viscosity, low-storage time integration, generalized curvilinear coordinates, periodic series, their reflections and interference patterns from a flat wall and scattering from a circular cylinder. The results were found to be promising en route to the aeroacoustic simulations of realistic engineering problems.

Kaushik, Dinesh K.↗

F-ANG+: A 3-D Augmented-Stencil Face-Averaged Nodal-Gradient Cell-Centered Finite-Volume Method for Hypersonic Flows

We describe the extension of a 2-D simplified face-averaged nodal-gradient (F-ANG) method to 3-D and demonstrate that the 3-D simplified F-ANG method is accomplished by augmenting the nodecentered gradient least squares stencil. This augmented stencil F-ANG method is shown to result in advection and diffusion schemes that are stable for hexahedral, prismatic, pyramidal and tetrahedral cells without having to resort to cell-averaged nodal gradients. In addition, we describe the modifications to the augmented stencil required to support the use of wall function boundary conditions. Finally we describe a consistent, face-stencil based multi-dimensional limiter procedure (MLP), and show it to be fully consistent and compatible with the linearity-preserving unstructured- MUSCL (LP-U-MUSCL) scheme for all values of kappa. These methods and schema are implemented in the cell-centered finite-volume code VULCAN-CFD, which is then used to investigate whether the robustness improvements demonstrated in 2-D carry over to 3-D by computing hypersonic flows using mixed-element grids as well as highly adapted tetrahedral grids.

Weighted Least-Squares↗

A High-Order Finite Spectral Volume Method for Conservation Laws on Unstructured Grids

A time accurate, high-order, conservative, yet efficient method named Finite Spectral Volume (FSV) is developed for conservation laws on unstructured grids. The concept of a 'spectral volume' is introduced to achieve high-order accuracy in an efficient manner similar to spectral element and multi-domain spectral methods. In addition, each spectral volume is further sub-divided into control volumes (CVs), and cell-averaged data from these control volumes is used to reconstruct a high-order approximation in the spectral volume. Riemann solvers are used to compute the fluxes at spectral volume boundaries. Then cell-averaged state variables in the control volumes are updated independently. Furthermore, TVD (Total Variation Diminishing) and TVB (Total Variation Bounded) limiters are introduced in the FSV method to remove/reduce spurious oscillations near discontinuities. A very desirable feature of the FSV method is that the reconstruction is carried out only once, and analytically, and is the same for all cells of the same type, and that the reconstruction stencil is always non-singular, in contrast to the memory and CPU-intensive reconstruction in a high-order finite volume (FV) method. Discussions are made concerning why the FSV method is significantly more efficient than high-order finite volume and the Discontinuous Galerkin (DG) methods. Fundamental properties of the FSV method are studied and high-order accuracy is demonstrated for several model problems with and without discontinuities.

Wang, Z. J.↗

A multilevel adaptive projection method for unsteady incompressible flow

There are two main requirements for practical simulation of unsteady flow at high Reynolds number: the algorithm must accurately propagate discontinuous flow fields without excessive artificial viscosity, and it must have some adaptive capability to concentrate computational effort where it is most needed. We satisfy the first of these requirements with a second-order Godunov method similar to those used for high-speed flows with shocks, and the second with a grid-based refinement scheme which avoids some of the drawbacks associated with unstructured meshes. These two features of our algorithm place certain constraints on the projection method used to enforce incompressibility. Velocities are cell-based, leading to a Laplacian stencil for the projection which decouples adjacent grid points. We discuss features of the multigrid and multilevel iteration schemes required for solution of the resulting decoupled problem. Variable-density flows require use of a modified projection operator--we have found a multigrid method for this modified projection that successfully handles density jumps of thousands to one. Numerical results are shown for the 2D adaptive and 3D variable-density algorithms.

Howell, Louis H.↗

Application of traditional CFD methods to nonlinear computational aeroacoustics problems

This paper describes an implementation of a high order finite difference technique and its application to the category 2 problems of the ICASE/LaRC Workshop on Computational Aeroacoustics (CAA). Essentially, a popular Computational Fluid Dynamics (CFD) approach (central differencing, Runge-Kutta time integration and artificial dissipation) is modified to handle aeroacoustic problems. The changes include increasing the order of the spatial differencing to sixth order and modifying the artificial dissipation so that it does not significantly contaminate the wave solution. All of the results were obtained from the CM5 located at the Numerical Aerodynamic Simulation Laboratory. lt was coded in CMFortran (very similar to HPF), using programming techniques developed for communication intensive large stencils, and ran very efficiently.

Chyczewski, Thomas S.↗

Implementing Connected Component Labeling as a User Defined Operator for SciDB

We have implemented a flexible User Defined Operator (UDO) for labeling connected components of a binary mask expressed as an array in SciDB, a parallel distributed database management system based on the array data model. This UDO is able to process very large multidimensional arrays by exploiting SciDB's memory management mechanism that efficiently manipulates arrays whose memory requirements far exceed available physical memory. The UDO takes as primary inputs a binary mask array and a binary stencil array that specifies the connectivity of a given cell to its neighbors. The UDO returns an array of the same shape as the input mask array with each foreground cell containing the label of the component it belongs to. By default, dimensions are treated as non-periodic, but the UDO also accepts optional input parameters to specify periodicity in any of the array dimensions. The UDO requires four stages to completely label connected components. In the first stage, labels are computed for each subarray or chunk of the mask array in parallel across SciDB instances using the weighted quick union (WQU) with half-path compression algorithm. In the second stage, labels around chunk boundaries from the first stage are stored in a temporary SciDB array that is then replicated across all SciDB instances. Equivalences are resolved by again applying the WQU algorithm to these boundary labels. In the third stage, relabeling is done for each chunk using the resolved equivalences. In the fourth stage, the resolved labels, which so far are "flattened" coordinates of the original binary mask array, are renamed with sequential integers for legibility. The UDO is demonstrated on a 3-D mask of O(1011) elements, with O(108) foreground cells and O(106) connected components. The operator completes in 19 minutes using 84 SciDB instances.

UDO↗

TEMPI: An Interposed MPI Library with Canonical Representation of MPI Datatypes [Slides]

These points are covered in this presentation: Distributed GPU stencil, non-contiguous data; Equivalence of strided datatypes and minimal representation; GPU communication methods; Deploying on managed systems; Large messages and MPI datatypes; Translation and canonicalization; Automatic model-driven transfer method selection; and Interposed library implementation.

97 MATHEMATICS AND COMPUTING↗

A Two-Dimensional Linear Bicharacteristic Scheme for Electromagnetics

The upwind leapfrog or Linear Bicharacteristic Scheme (LBS) has previously been implemented and demonstrated on one-dimensional electromagnetic wave propagation problems. This memorandum extends the Linear Bicharacteristic Scheme for computational electromagnetics to model lossy dielectric and magnetic materials and perfect electrical conductors in two dimensions. This is accomplished by proper implementation of the LBS for homogeneous lossy dielectric and magnetic media and for perfect electrical conductors. Both the Transverse Electric and Transverse Magnetic polarizations are considered. Computational requirements and a Fourier analysis are also discussed. Heterogeneous media are modeled through implementation of surface boundary conditions and no special extrapolations or interpolations at dielectric material boundaries are required. Results are presented for two-dimensional model problems on uniform grids, and the Finite Difference Time Domain (FDTD) algorithm is chosen as a convenient reference algorithm for comparison. The results demonstrate that the two-dimensional explicit LBS is a dissipation-free, second-order accurate algorithm which uses a smaller stencil than the FDTD algorithm, yet it has less phase velocity error.

Beggs, John H.↗

High-order essentially non-oscillatory methods for computational aeroacoustics

The desire to obtain acoustic information from the numerical solution of a nonlinear system of equations is a demanding proposition for a computational algorithm. High-order accuracy is required for the propagation of high-frequency, low-amplitude waves. In addition, it is desirable to highly resolve discontinuities that can develop in the solutions of the Euler or Navier-Stokes equations. The class of essentially non-oscillatory (ENO) shock-capturing schemes has been designed to have both of these properties. The dual capacity of ENO schemes for high-order accuracy and non-oscillatory shock-capturing is achieved through the use of adaptive stenciling, which makes these schemes highly nonlinear. These schemes are briefly described and referenced herein. A fourth-order algorithm is then applied to the solution of an acoustic wave in a quasi-one-dimensional converging-diverging nozzle.

Casper, Jay↗

Distributed Relaxation for Conservative Discretizations

A multigrid method is defined as having textbook multigrid efficiency (TME) if the solutions to the governing system of equations are attained in a computational work that is a small (less than 10) multiple of the operation count in one target-grid residual evaluation. The way to achieve this efficiency is the distributed relaxation approach. TME solvers employing distributed relaxation have already been demonstrated for nonconservative formulations of high-Reynolds-number viscous incompressible and subsonic compressible flow regimes. The purpose of this paper is to provide foundations for applications of distributed relaxation to conservative discretizations. A direct correspondence between the primitive variable interpolations for calculating fluxes in conservative finite-volume discretizations and stencils of the discretized derivatives in the nonconservative formulation has been established. Based on this correspondence, one can arrive at a conservative discretization which is very efficiently solved with a nonconservative relaxation scheme and this is demonstrated for conservative discretization of the quasi one-dimensional Euler equations. Formulations for both staggered and collocated grid arrangements are considered and extensions of the general procedure to multiple dimensions are discussed.

Diskin, Boris↗

Preserving Superconvergence of Spectral Elements for Curved Domains

Spectral element methods (SEM), extensions of finite element methods (FEM), have emerged as significant techniques for solving partial differential equations in physics and engineering. SEM can potentially deliver superior accuracy due to the potential superconvergence in nodal solutions for well-shaped tensor-product elements. However, the accuracy of SEM often degrades in complex geometries due to geometric inaccuracies near curved boundaries and the loss of superconvergence with simplicial or non-tensor-product elements. To overcome the first issue, we propose using geometric refinement, which both refines the mesh near high-curvature regions and increases the degree of geometric basis functions. We show that when using mixed-element meshes with tensor-product elements in the interior of the domain, curvature-based geometric refinement near boundaries can improve the accuracy of the interior elements by reducing pollution errors and preserving the superconvergence in nodal solutions. To address the second issue, we introduce ApSEM, a post-processing technique using the adaptive extended stencil finite element method (AES-FEM) to recover the accuracy near the curved boundaries. The combination of curvature-based geometric refinement and accurate post-processing offers an effective and easier-to-implement alternative to methods reliant on exact geometries. We demonstrate our techniques by solving the convection-diffusion equation in 2D and 3D and show up to two orders of magnitude of improvement in the solution accuracy, even when the elements are poorly shaped near boundaries. We also show the efficiency of ApSEM as it can recover superconvergence in nodal solutions without drastically increasing the computational cost.

97 MATHEMATICS AND COMPUTING↗

The Linear Bicharacteristic Scheme for Electromagnetics

The upwind leapfrog or Linear Bicharacteristic Scheme (LBS) has previously been implemented and demonstrated on electromagnetic wave propagation problems. This paper extends the Linear Bicharacteristic Scheme for computational electromagnetics to model lossy dielectric and magnetic materials and perfect electrical conductors. This is accomplished by proper implementation of the LBS for homogeneous lossy dielectric and magnetic media and for perfect electrical conductors. Heterogeneous media are modeled through implementation of surface boundary conditions and no special extrapolations or interpolations at dielectric material boundaries are required. Results are presented for one-dimensional model problems on both uniform and nonuniform grids, and the FDTD algorithm is chosen as a convenient reference algorithm for comparison. The results demonstrate that the explicit LBS is a dissipation-free, second-order accurate algorithm which uses a smaller stencil than the FDTD algorithm, yet it has approximately one-third the phase velocity error. The LBS is also more accurate on nonuniform grids.

Beggs, John H.↗

A 6th Order Mehrstellen Finite Volume Discretization of Poisson's Equation in Three Dimensions

We discuss the derivation of a new, sixth-order finite volume scheme for Poisson’s equation on 3D Cartesian equispaced grids. The scheme is based on a discretization of the Laplace operator with a compact (Mehrstellen) 27-point stencil. To achieve sixth order convergence the right hand side of the equation is replaced with a discrete operator that involves the discrete Laplace and Biharmonic operators and the sum of discrete fourth-order cross derivatives applied to the charge function. Numerical tests demonstrate the superiority of the proposed method compared to the well known schemes associated with the 7-point and 19-point discretizations of the Laplacian.

97 MATHEMATICS AND COMPUTING↗

A Dynamic Amplitude-Correcting Gradient Estimation Technique to Align X-ray Focusing Optics

High-brightness X-rays, as produced at synchrotrons and X-ray free electron laser (XFEL) facilities, are used to characterize materials in a variety of scientific experiments. In most cases, effective use of the high-energy light requires precisely-aligned focusing optics; one example being a compound refractive lens (CRL). To align a CRL, the position and rotation must be optimized along four axes. In practice, this is a labor-intensive, time-consuming manual process that can monopolize scarce experimental time at the necessary X-ray facilities. Models of the expected Xray transmission function suggest that this task can be automated; however, the temporally-varying intensity at X-ray free electron laser facilities preclude the direct use of standard implementations of optimization solvers such as steepest descent algorithms. In this paper, we propose a novel technique to estimate the gradient of noisy functions with temporally-varying amplitudes. We construct this dynamicamplitude correction by systematically sampling a fixed central location within the standard finite difference stencil, accounting for the observed changes in time, and normalizing the difference quotients against those fluctuations. In addition to a rigorous error analysis of the sampling technique, we demonstrate its efficacy in stochastic descent optimization methods. Further, we demonstrate how this approach may be implemented to optimize X-ray focusing optics at synchrotrons or XFEL facilities

97 MATHEMATICS AND COMPUTING↗

A compact solution to computational acoustics

This paper demonstrates that the linearized, dimensional Euler equations for acoustic computation can be accurately solved as a set of decoupled first-order wave equations, and that if ordered properly, this system of simple waves has unambiguous, easily implemented boundary conditions, allowing waves of same group speeds to pass through numerical boundaries or comply with wall conditions. Thus, the task of designing a complex multi-dimensional scheme with approximate far-field boundary conditions reduces to the design of higher order schemes for the one-dimensional simple wave equation. A compact finite-difference scheme and a characteristically exact but numerically n(th) order accurate boundary condition are introduced for solving the first order wave equation. Spanning a three-point two-level stencil, this low-dispersion implicit scheme has a third order spatial accuracy when used on nonuniform meshes, fourth order accurate on uniform meshes, and a temporal accuracy of second order due to the choice of trapezoidal integration for algorithmic simplicity. The robustness and accuracy of the scheme are demonstrated through a series of numerical experiments and comparisons with published results. When tested on the one-dimensional wave equation on a uniform grid, this scheme allows a Gaussian wave packet to pass through any finite domain with low numerical dispersion characteristic of a spatially fourth-order scheme and reflections at numerical boundaries maintained below truncation error. On highly stretched and irregular grids, only mild dispersions are found in the solution while solutions by other methods fail or are severely distorted. Yet, this scheme is no more sophisticated to solve or implement than the Crank-Nicolson scheme. This scheme has been tested on four categories of the ICASE/LaRC benchmark problems, which include propagation of acoustic and convective waves in Cartesian and cylindrical domains, reflection of acoustic wave at stationary/moving boundaries, and sound generation by gust-blade interaction.

Fung, K.-Y.↗

The Linear Bicharacteristic Scheme for Computational Electromagnetics

The upwind leapfrog or Linear Bicharacteristic Scheme (LBS) has previously been implemented and demonstrated on electromagnetic wave propagation problems. This paper extends the Linear Bicharacteristic Scheme for computational electromagnetics to treat lossy dielectric and magnetic materials and perfect electrical conductors. This is accomplished by proper implementation of the LBS for homogeneous lossy dielectric and magnetic media, and treatment of perfect electrical conductors (PECs) are shown to follow directly in the limit of high conductivity. Heterogeneous media are treated through implementation of surface boundary conditions and no special extrapolations or interpolations at dielectric material boundaries are required. Results are presented for one-dimensional model problems on both uniform and nonuniform grids, and the FDTD algorithm is chosen as a convenient reference algorithm for comparison. The results demonstrate that the explicit LBS is a dissipation-free, second-order accurate algorithm which uses a smaller stencil than the FDTD algorithm, yet it has approximately one-third the phase velocity error. The LBS is also more accurate on nonuniform grids.

Beggs, John H.↗

An Online Dynamic Amplitude-Correcting Gradient Estimation Technique to Align X-ray Focusing Optics

High-brightness X-ray pulses, as generated at synchrotrons and X-ray free electron lasers (XFELs), are used in a variety of scientific experiments. At these facilities, measurements often require optical equipment, e.g Compound Refractive Lenses (CRLs) to be precisely aligned and focused. The lateral alignment of CRLs to a beamline requires precise positioning along four axes: two translational, and the two rotational. At a synchrotron, alignment is often accomplished manually. However, XFEL beamlines present a beam brightness that fluctuates stochastically, making manual alignment a time-consuming endeavor. Automation using simplex or classic stochastic descent often fails, given the errant gradient estimates. Herein we present a dynamic-amplitude correction to the usual gradient based on the combination of a generalized finite difference stencil and a time-dependent sampling pattern. Intensity is recorded periodically, then used to normalize numerical derivatives against fluctuations. Error expectation is analyzed, and efficacy is demonstrated on classic benchmarks. We provide a proof of concept by laterally aligning optics on a simulated XFEL beamline using data recorded at both synchrotron and XFEL facilities.

97 MATHEMATICS AND COMPUTING↗