Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stencil computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Towards a Verifiable Domain-Specific Language for Hardware-Accelerated Stencils

Defining a domain-specific language (DSL) that supports vector-calculus abstractions eases the porting of partial differential equation (PDE) solvers to specialized architectures. Sufficiently high-level abstractions empower users to express universal laws with sufficient generality that the laws must always hold true within their domain of validity. A broad class of PDE solvers employs stencil-based algorithms, the target domain of Berkeley Lab's stencil accelerator chip co-design project. First released as open-source in January 2026, the Formal software framework lays a foundation for defining an embedded DSL based on composable operators that implement mimetic numerical methods -- stencil algorithms that guarantee satisfaction of discrete versions of important vector calculus theorems. The Formal DSL will be the frontend to a new class of stencil-PDE accelerators developed jointly by LBNL, UHCL, and UC Berkeley through the DOE Competitive Portfolios for Computer Science Project. This offers the potential of an order of magnitude acceleration for this important category of computational methods to serve the DOE mission. Future work on the Formal DSL will facilitate software verification via type-safe templates that enable problem-specific correctness proofs relying upon generic function theory and carefully crafted unit tests.

Rouson, Damian↗

Implicit fast sweeping method for hyperbolic systems of conservation laws

Implicit time-accurate methods are often used to integrate stiff problems where explicit schemes impose severe time step restrictions. This paper presents an efficient numerical framework based on the Fast Sweeping Method (FSM) for solving linear and nonlinear hyperbolic systems of conservation laws. The solution at each discrete location is computed by sweeping the numerical domain in several predetermined directions that follow the causality of the characteristic families. The use of a fractional step strategy eliminates the need for a solution selection criterion while one-sided stencils limit the number of sweeps to at most 2 d for d space dimensions. This work focuses on the first-order implicit upwind method since it constitutes the building block for high-order conservative schemes. For problems where the degree of stiffness evolves over time, implicit-explicit hybridization can be accomplished with the same algorithm by simply switching the stencil at each time level. As opposed to traditional implicit solvers, the sweeping method does not require a local time linearization of the fluxes thereby preserving the nonlinear stability properties of the original implicit scheme. It also avoids the large computational and memory requirements associated with solving large block-diagonal systems of equations. Here, a series of one- and two-dimensional test cases are presented for the inviscid Burgers' equation and the reactive Euler equations. The results indicate that the implicit FSM can allow a major reduction in the number of time steps even in the presence of discontinuous solution profiles.

74 ATOMIC AND MOLECULAR PHYSICS↗

Patchy nanoparticles by atomic stencilling

Stencilling, in which patterns are created by painting over masks, has ubiquitous applications in art, architecture and manufacturing. Modern, top-down microfabrication methods have succeeded in reducing mask sizes to under 10 nm, enabling ever smaller microdevices as today’s fastest computer chips. Meanwhile, bottom-up masking using chemical bonds or physical interactions has remained largely unexplored, despite its advantages of low cost, solution-processability, scalability and high compatibility with complex, curved and three-dimensional (3D) surfaces. Here we report atomic stencilling to make patchy nanoparticles (NPs), using surface-adsorbed iodide submonolayers to create the mask and ligand-mediated grafted polymers onto unmasked regions as ‘paint’. We use this approach to synthesize more than 20 different types of NP coated with polymer patches in high yield. Polymer scaling theory and molecular dynamics (MD) simulation show that stencilling, along with the interplay of enthalpic and entropic effects of polymers, generates patchy particle morphologies not reported previously. These polymer-patched NPs self-assemble into extended crystals owing to highly uniform patches, including different non-closely packed superlattices. We propose that atomic stencilling opens new avenues in patterning NPs and other substrates at the nanometre length scale, leading to precise control of their chemistry, reactivity and interactions for a wide range of applications, such as targeted delivery, catalysis, microelectronics, integrated metamaterials and tissue engineering.

36 MATERIALS SCIENCE↗

Boundary-consistent B-spline filtering schemes and application to high-fidelity simulations of turbulence

A filtering operation, based on B-spline discretizations, is introduced to target weakly growing mesh-scale oscillations that can arise in high-fidelity turbulence simulations. This is a spectral regularization that can be described using the singular values of a banded matrix operator, with the filtering strength set by a scalar- or vector-valued penalty parameter. The penalty parameter can be specified though it can also be advantageously selected to minimize the generalized cross validation (GCV) measure of distance between the pre- and post-filtered solutions. Efficient algorithms are developed to compute both the scalar and vector penalty parameters. The B-spline filter has a sharper localization to high-wavenumber than compact or explicit filters of the same stencil width and is demonstrated for solutions of the Burgers' equation, decaying Burgers' turbulence, and compressible Navier–Stokes turbulent channel flow. Furthermore, these simulations confirm the scheme's numerical stability and ability to narrowly target the high wavenumber components of numerical solutions. An advantage over finite-difference filters is that these B-spline filters are stable on bounded domains and even preserve formal order of accuracy.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

Assessment of Machine Learning Wall Modeling Approaches for Large Eddy Simulation of Gas Turbine Film Cooling Flows: An a Priori Study

Here, in this work, a priori analysis of machine learning (ML) strategies is carried out with the goal of data-driven wall modeling for large eddy simulation (LES) of gas turbine film cooling flows. High-fidelity flow datasets are extracted from wall-resolved LES (WRLES) of flow over a flat plate interacting with the coolant flow supplied by a single row of 7-7-7 shaped cooling holes inclined at 30 degrees with the flat plate at different blowing ratios (BR). The WRLES are performed using the high-order Nek5000 spectral element computational fluid dynamics (CFD) solver. Light gradient boosting machine (LightGBM) is employed as the ML algorithm for the data-driven wall model. Parametric tests are conducted to systematically assess the influence of a wide range of input flow features (velocity components, velocity gradients, pressure gradients, and fluid properties) on the accuracy of ML wall model with respect to prediction of wall shear stress. In addition, the use of spatial stencil and time delay is also explored within the ML wall modeling framework. It is shown that features associated with gradients of the streamwise and spanwise velocity components have a major impact on the prediction fidelity of wall model, while the effect of gradients of wall-normal velocity component is found to be negligible. Moreover, adding flow feature information from an x-y-z spatial stencil significantly improves the ML model accuracy and generalizability compared to just using local flow features from the matching location. Overall, highest prediction accuracy is achieved when both spatial stencil and time delay features are incorporated within the data-driven wall modeling paradigm.

33 ADVANCED PROPULSION SYSTEMS↗

A New Semistructured Algebraic Multigrid Method

Multigrid methods are well suited to large massively parallel computer architectures because they are mathematically optimal and display good parallelization properties. Since current architecture trends are favoring regular compute patterns to achieve high performance, the ability to express structure has become much more important. The hypre software library provides high-performance multigrid preconditioners and solvers through conceptual interfaces, including a semistructured interface that describes matrices primarily in terms of stencils and logically structured grids. This paper presents a new semistructured algebraic multigrid (SSAMG) method built on this interface. The numerical convergence and performance of a CPU implementation of this method are evaluated for a set of semistructured problems. In conclusion, SSAMG achieves significantly better setup times than hypre’s unstructured AMG solvers and comparable convergence. In addition, the new method is capable of solving more complex problems than hypre’s structured solvers.

97 MATHEMATICS AND COMPUTING↗

Regional surrogates for predictive control of digital twins

Digital twins of complex systems must involve a model that is fast, generalizable, and usable for real-time control. For example, high-fidelity nonlinear multiphysics simulations can capture laser-material interactions, but are too slow for optimization or model predictive control (MPC). Reduced-order models, used to accelerate such computation, frequently fail to generalize to unseen inputs or control states. We show theoretically that this failure is intrinsic, i.e., that a learned model is non-unique outside the sampled subspace when its low-rank structure arises from limited excitation and clustered eigenvalues, rather than from a user-imposed truncation alone. Motivated by this result, we propose a control-ready regional surrogate-construction framework for both autonomous and nonautonomous dynamics; it employs Koopman lifting to represent nonlinearities, while preserving spatial locality. We illustrate our approach by constructing a control-ready surrogate for the digital twin of a thermal component of additive-manufacturing process. Our surrogate, localized in space through a von Neumann stencil, is learned from noisy high-fidelity simulations that emulate thermal-camera images collected during the manufacturing. It is linear in thermo-physically augmented states so that MPC reduces to a convex quadratic program. The surrogate requires no online correction, generalizes to unseen scan paths and power profiles of the laser, and is more than three orders of magnitude faster than a finite-difference solver. Furthermore, when the MPC sequence computed on the digital twin is applied to this solver, closed-loop temperature regulation is recovered, showing that the surrogate preserves control-relevant input-output behavior.

Data-driven model↗

FFTX-IRIS: Towards Performance Portability and Heterogeneity for SPIRAL Generated Code

FFTX-IRIS is a dynamic system to efficiently utilize novel heterogeneous platforms. This system links two next-generation frameworks, FFTX and IRIS, to navigate the complexity of different hardware architectures. FFTX provides a runtime code generation framework for high-performance Fast Fourier Transform kernels. IRIS runtime provides portability and multi-device heterogeneity, allowing computation on any available compute resource. Together, FFTX-IRIS enables code generation, seamless portability, and performance without user involvement. We show the design of the FFTX-IRIS system along with an evaluation of various small FFT benchmarks. We also demonstrate multi-device heterogeneity of FFTX-IRIS with a larger stencil application.

Rao, Sanil↗

A second-order distributed memory parallel fast sweeping method for the Eikonal equation

The Eikonal equation is used to calculate wave propagation and distance fields, and due to its complexity requires numerical treatment for its solution. In this work, we present a second-order distributed memory parallel fast sweeping method. The second-order solution switches on a two-point stencil when two upwind points are available, and reverts to first-order otherwise. In all examples, the second-order method improves the solution over the first-order, allowing for significant savings in memory while achieving the same accuracy. Parallelization over distributed memory saw good weak scaling with optimal convergence. The computational time for second-order was approximately 2.5 times slower than first-order, where the largest amount of mesh points ran on 144 cores (512 GB) was ≈20 billion. The savings in memory from the second-order method combined with the distributed memory algorithm result in the ability to solve problems much larger than are possible with the serial first-order method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Preserving Superconvergence of Spectral Elements for Curved Domains [Slides]

Finite Element Methods (FEM) and Spectral Element Methods (SEM) are crucial for solving partial differential equations (PDEs) on complex geometries. SEM offers superior accuracy due to potential superconvergence for simple domains. Challenges persist for domains with curved boundaries, restricting SEM’s advantages in real-world applications. A proposed solution is the introduction of a novel strategy to enhance accuracy and maintain superconvergence of SEM in curved domains. The strategy includes a mesh-generation procedure with geometrically refined elements near curved boundaries and a post-processing phase using the Adaptive Extended Stencil Finite Element Method (AES-FEM). The method, named AES-FEM post-processed Spectral Element Method (ApSEM), aligns the accuracy of non-tensor-product elements with superconvergent spectral elements.

97 MATHEMATICS AND COMPUTING↗

Sharp front tracking with geometric interface reconstruction

Here, this paper presents a novel sharp front-tracking method designed to address limitations in classical front-tracking approaches, specifically their reliance on smooth interpolation kernels and extended stencils for coupling the front and fluid mesh. In contrast, the proposed method employs exclusively sharp, localized interpolation and spreading kernels, restricting the coupling to the interfacial fluid cells–those containing the interface/front. This localized coupling is achieved by integrating a divergence-preserving velocity interpolation method with a piecewise parabolic interface calculation (PPIC) and a polyhedron intersection algorithm to compute the indicator function and local interface curvature. Surface tension is computed using the Continuum Surface Force (CSF) method, maintaining consistency with the sharp representation. Additionally, we propose an efficient local roughness smoothing implementation to account for surface mesh undulations, which is easily applicable to any triangulated surface mesh. Building on our previous work, the primary innovation of this study lies in the localization of the coupling for both the indicator function and surface tension calculations. By reducing the interface thickness on the fluid mesh to a single cell, as opposed to the 4–5 cell spans typical in classical methods, the proposed sharp front-tracking method achieves a highly localized and accurate representation of the interface. This sharper representation mitigates parasitic currents and improves force balancing, making it particularly suitable for scenarios where the interface plays a critical role, such as microfluidics, fluid-fluid interactions, and fluid-structure interactions. The proposed method is comprehensively validated and tested on canonical interfacial flow problems, including stationary and translating Laplace equilibria, oscillating droplets, and rising bubbles. The presented results demonstrate that the sharp front-tracking method significantly outperforms the classical approach in terms of accuracy, stability, and computational efficiency. Notably, parasitic currents are reduced by approximately two orders of magnitude and stable results are obtained for parameter ranges where classical front tracking fails to converge.

42 ENGINEERING↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

TEMPI: An Interposed MPI Library with Canonical Representation of MPI Datatypes [Slides]

These points are covered in this presentation: Distributed GPU stencil, non-contiguous data; Equivalence of strided datatypes and minimal representation; GPU communication methods; Deploying on managed systems; Large messages and MPI datatypes; Translation and canonicalization; Automatic model-driven transfer method selection; and Interposed library implementation.

97 MATHEMATICS AND COMPUTING↗

Preserving Superconvergence of Spectral Elements for Curved Domains

Spectral element methods (SEM), extensions of finite element methods (FEM), have emerged as significant techniques for solving partial differential equations in physics and engineering. SEM can potentially deliver superior accuracy due to the potential superconvergence in nodal solutions for well-shaped tensor-product elements. However, the accuracy of SEM often degrades in complex geometries due to geometric inaccuracies near curved boundaries and the loss of superconvergence with simplicial or non-tensor-product elements. To overcome the first issue, we propose using geometric refinement, which both refines the mesh near high-curvature regions and increases the degree of geometric basis functions. We show that when using mixed-element meshes with tensor-product elements in the interior of the domain, curvature-based geometric refinement near boundaries can improve the accuracy of the interior elements by reducing pollution errors and preserving the superconvergence in nodal solutions. To address the second issue, we introduce ApSEM, a post-processing technique using the adaptive extended stencil finite element method (AES-FEM) to recover the accuracy near the curved boundaries. The combination of curvature-based geometric refinement and accurate post-processing offers an effective and easier-to-implement alternative to methods reliant on exact geometries. We demonstrate our techniques by solving the convection-diffusion equation in 2D and 3D and show up to two orders of magnitude of improvement in the solution accuracy, even when the elements are poorly shaped near boundaries. We also show the efficiency of ApSEM as it can recover superconvergence in nodal solutions without drastically increasing the computational cost.

97 MATHEMATICS AND COMPUTING↗

A 6th Order Mehrstellen Finite Volume Discretization of Poisson's Equation in Three Dimensions

We discuss the derivation of a new, sixth-order finite volume scheme for Poisson’s equation on 3D Cartesian equispaced grids. The scheme is based on a discretization of the Laplace operator with a compact (Mehrstellen) 27-point stencil. To achieve sixth order convergence the right hand side of the equation is replaced with a discrete operator that involves the discrete Laplace and Biharmonic operators and the sum of discrete fourth-order cross derivatives applied to the charge function. Numerical tests demonstrate the superiority of the proposed method compared to the well known schemes associated with the 7-point and 19-point discretizations of the Laplacian.

97 MATHEMATICS AND COMPUTING↗