Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

SPEL: Software tool for Porting E3SM Land Model with OpenACC in a Function Unit Test Framework

Most high-end computers adopt hybrid architecture, porting a large-scale scientific code onto accelerators is necessary. The paper presents a generic method for porting large-scale scientific code onto accelerators using compiler directives within a modularized function unit test platform. We have implemented the method and designed a software tool (SPEL) to port the E3SM Land Model (ELM) onto the GPUs in the Summit computer. SPEL automatically generates GPU-ready test modules for all ELM functions, such as CanopyFlux, SoilTemperature, and EcosystemDynamics. SPEL breaks the ELM into a collection of standalone unit test programs for easy code verification and further performance improvement. We further optimize several ELM test modules with advanced techniques, including memory reduction, reconstructed parallel loops, and asynchronous GPU kernel launch. We hope our study will inspire new toolkit developments that expedite large-scale scientific code porting with compiler directives.

Schwartz, Peter↗

Radiation Damage Mitigation in FeCrAl Alloy at Sub-Recrystallization Temperatures

Traditional defect recovery methods rely on high-temperature annealing, often exceeding 750 °C for FeCrAl. In this study, we introduce electron wind force (EWF)-assisted annealing as an alternative approach to mitigate irradiation-induced defects at significantly lower temperatures. FeCrAl samples irradiated with 5 MeV Zr 2+ ions at a dose of 10 14 cm −2 were annealed using EWF at 250 °C for 60 s. We demonstrate a remarkable transformation in the irradiated microstructure, where significant increases in kernel average misorientation (KAM) and low-angle grain boundaries (LAGBs) typically indicate heightened defect density; the use of EWF annealing reversed these effects. X-ray diffraction (XRD) confirmed these findings, showing substantial reductions in full width at half maximum (FWHM) values and a realignment of peak positions toward their original states, indicative of stress and defect recovery. To compare the effectiveness of EWF, we also conducted traditional thermal annealing at 250 °C for 7 h, which proved less effective in defect recovery as evidenced by less pronounced improvements in XRD FWHM values.

FeCrAl alloys↗

On finite-dimensional smoothed-particle Hamiltonian reductions of the Vlasov equation

The inclusion of spatial smoothing in finite-dimensional particle-based Hamiltonian reductions of the Vlasov equation and related models is considered. Here, this work investigates the underlying Hamiltonian structure of such smoothed particle-based methods for Hamiltonian systems and the small-scale regularization such methods implicitly make in approximating the continuum theory. In the context of the Vlasov–Poisson equation and other mean-field Lie–Poisson systems, of which Vlasov–Poisson is a special case, smoothing amounts to a convolutive regularization of the Hamiltonian. This regularization may be interpreted as a change of the inner product structure used to identify the dual space in the Lie–Poisson Hamiltonian formulation. In particular, the shape function used for spatial smoothing may be identified as the kernel function of a reproducing kernel Hilbert space whose inner product is used to define the Lie–Poisson Hamiltonian structure. It is likewise possible to introduce smoothing in the Vlasov–Maxwell system, but in this case the Poisson bracket must be modified rather than the Hamiltonian. The smoothing applied to the Vlasov–Maxwell system is incorporated by inserting smoothing in the map from canonical to kinematic coordinates. In the filtered system, the Lorentz force law and the current, the two terms coupling the Vlasov equation with Maxwell’s equations, are spatially smoothed.

Hamiltonian mechanics↗

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero↗

Real-space inversion and super-resolution of ultrafast scattering

Ultrafast scattering using x-rays or electrons is an emerging method to obtain structure dynamics at the atomic lengthscales and timescales. However, directly resolving in real-space atomic motions is inherently limited by the finite detector range and the probe energy. So, as a result, the time-resolved signal interpretation is mostly done in reciprocal space and relies on modeling and simulations of specific structures and processes. Here, we introduce a model-free approach to directly resolve scattering signals in real space, surpassing the diffraction limit, using scattering kernels and signal priors that naturally arise from the measurement constraints. We demonstrate the approach on simulated and experimental data, recover multiple atomic motions at sub-angstrom resolutions, and discuss the recovery accuracy and resolution limits versus signal fidelity. The approach offers a robust path to obtain high-resolution real-space information of atomic-scale structure dynamics using current time-resolved x-ray or electron scattering sources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Next-Cycle Optimal Fuel Control for Cycle-to-Cycle Variability Reduction in EGR-Diluted Combustion

In this simulation study, cycle-to-cycle fuel control was used to reduce CCV by injecting additional fuel in operating conditions with sporadic misfires and partial burns. An optimal control policy was proposed that utilizes 1) a physics-based model that tracks in-cylinder gas composition and 2) a one-step-ahead prediction of the combustion efficiency based on a kernel density estimator. The optimal solution, however, presents a tradeoff between the reduction in combustion CCV and the increase in fuel injection quantity required to stabilize the charge. Such a tradeoff can be ad- just by a single parameter embedded in the cost function.

Maldonado, BryanP. [Oak Ridge National Lab. (ORNL)↗

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

High-performance strategies for the recent MRSF-TDDFT in GAMESS

Multiple ERI (Electron Repulsion Integral) tensor contractions (METC) with several matrices are ubiquitous in quantum chemistry. In response theories, the contraction operation, rather than ERI computations, can be the major bottleneck, as its computational demands are proportional to the multiplicatively combined contributions of the number of excited states and the kernel pre-factors. Here, this paper presents several high-performance strategies for METC. Optimal approaches involve either the data layout reformations of interim density and Fock matrices, the introduction of intermediate ERI quartet buffer, and loop-reordering optimization for a higher cache hit rate. The combined strategies remarkably improve the performance of the MRSF (mixed reference spin flip)-TDDFT (time-dependent density functional theory) by nearly 300%. The results of this study are not limited to the MRSF-TDDFT method and can be applied to other METC scenarios.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient quadrature rules for finite element discretizations of nonlocal equations

In this paper we design efficient quadrature rules for finite element discretizations of nonlocal diffusion problems with compactly supported kernel functions. Two of the main challenges in nonlocal modeling and simulations are the prohibitive computational cost and the nontrivial implementation of discretization schemes, especially in three-dimensional settings. In this work we circumvent both challenges by introducing a parametrized mollifying function that improves the regularity of the integrand, utilizing an adaptive integration technique, and exploiting parallelization. We first showthat the “mollified” solution converges to the exact one as the mollifying parameter vanishes, then we illustrate the consistency and accuracy of the proposed method on several two- and three-dimensional test cases. Furthermore, we demonstrate the good scaling properties of the parallel implementation of the adaptive algorithm and we compare the proposed method with recently developed techniques for efficient finite element assembly.

97 MATHEMATICS AND COMPUTING↗

CSPlib - A Software Toolkit for the Analysis of Dynamical Systems and Chemical Kinetic Models

CSPlib is an open source software library for analyzing general ordinary differential equation (ODE) systems and detailed chemical kinetic ODE systems. It relies on the computational singular perturbation (CSP) method for the analysis of these systems. The software provides support for: General ODE models (gODE model class) for computing source terms and Jacobians for a generic ODE system; TChem model (ChemElemODETChem model class) for computing source term, Jacobian, other necessary chemical reaction data, as well as the rates of progress for a homogenous batch reactor using an elementary step detailed chemical kinetic reaction mechanism. This class relies on the TChem [2] library; A set of functions to compute essential elements of CSP analysis (Kernel class). This includes computations of the eigensolution of the Jacobian matrix, CSP basis vectors and co-vectors, time scales (reciprocals of the magnitudes of the Jacobian eigenvalues), mode amplitudes, CSP pointers, and the number of exhausted modes. This class relies on the Tines library; A set of functions to compute the eigensolution of the Jacobian matrix using Tines library GPU eigensolver; A set of functions to compute CSP indices (Index Class). This includes participation indices and both slow and fast importance indices.

97 MATHEMATICS AND COMPUTING↗

Efficient quadrature rules for finite element discretizations of nonlocal equations

Here, in this paper, we design efficient quadrature rules for finite element (FE) discretizations of nonlocal diffusion problems with compactly supported kernel functions. Two of the main challenges in nonlocal modeling and simulations are the prohibitive computational cost and the nontrivial implementation of discretization schemes, especially in three-dimensional settings. In this work, we circumvent both challenges by introducing a parametrized mollifying function that improves the regularity of the integrand, utilizing an adaptive integration technique, and exploiting parallelization. We first show that the “mollified” solution converges to the exact one as the mollifying parameter vanishes, then we illustrate the consistency and accuracy of the proposed method on several two- and three-dimensional test cases. Furthermore, we demonstrate the good scaling properties of the parallel implementation of the adaptive algorithm and we compare the proposed method with recently developed techniques for efficient FE assembly.

97 MATHEMATICS AND COMPUTING↗

Revisiting Temporal Blocking Stencil Optimizations

Iterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark.

Zhang, Lingqi↗

Scalable quantum processor noise characterization

Measurement fidelity matrices (MFMs) (also called error kernels) are a natural way to characterize state preparation and measurement errors in near-term quantum hardware. They can be employed in post processing to mitigate errors and substantially increase the effective accuracy of quantum hardware. However, the feasibility of using MFMs is currently limited as the experimental cost of determining the MFM for a device grows exponentially with the number of qubits. In this work we present a scalable way to construct approximate MFMs for many-qubit devices based on cumulant expansions. Our method can also be used to characterize various types of correlation error.

Hamilton, Kathleen↗

Atomistic and mesoscale simulations to determine effective diffusion coefficient of fission products in SiC

The silicon carbide (SiC) layer in tristructural isotropic (TRISO) particles serves as the barrier to prevent escape of fission products produced in the fuel kernel. Knowing the diffusion coefficient of fission products through SiC is critical to determining whether fission gas can escape from the particle. It has been observed in experiments that Ag accumulated in grain boundaries and triple junctions in SiC. It is hypothesized that grain boundary diffusion is the primary pathway by which fission products penetrate the SiC layer. In this report, the effective diffusion coefficient of the fission product Ag through the grain boundary network is calculated using a combination of atomistic and phase-field methods. The grain boundary diffusion coefficient is calculated using molecular dynamics simulations. The bulk diffusion coefficient is determined using a combination of density functional theory and nudged elastic band methods. An effective diffusion coefficient is calculated, accounting for the grain structure using a phase-field method. The effective diffusion coefficient will be incorporated into Bison and fission product release calculations are compared to available experimental data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

GPU Profiling and Optimizing xRAGE (Final Report)

Our project’s objective is to increase the efficiency of GPU-enabled kernels in xRAGE. To do so, we conduct GPU profiling with NSight Systems on xRAGE tests unsplit_sod_1d and unsplit_sedov_2d to identify bottlenecks and understand the behavior of the GPU during code execution. Next, we analyze these generated GPU profiles to locate the lines of code whose optimization have the most potential for improving runtime. We replicate the structure of the code in smaller test problems that are easier to understand, edit, and run quickly. Within these test problems, we implement two different methods of improving performance: transformation of nested loops into a single MDRangePolicy and hierarchical parallelization using teams of threads. Both methods show speedups in the test code, and after transferring them to xRAGE, they both show up to 30x speedups on various computing platforms. Profiling the edited versions of xRAGE reveals that the GPU successfully executed the bottlenecks with greater efficiency

97 MATHEMATICS AND COMPUTING↗

Enhanced relaxed physical factorization preconditioner for coupled poromechanics

The relaxed physical factorization (RPF) preconditioner is a recent algorithm allowing for the efficient and robust solution to the block linear systems arising from the three-field displacement-velocity-pressure formulation of coupled poromechanics. For its application, however, it is necessary to invert blocks with the algebraic form C^ = (C + βFF T ), where C is a symmetric positive definite matrix, FF T a rank-deficient term, and β a real non-negative coefficient. The inversion of C^, performed in an inexact way, can become unstable for large values of β, as it usually occurs at some stages of a full poromechanical simulation. In this work, we propose a family of algebraic techniques to stabilize the inexact solve with C^. This strategy can prove useful in other problems as well where such an issue might arise, such as augmented Lagrangian preconditioning techniques for Navier-Stokes or incompressible elasticity. First, we introduce an iterative scheme obtained by a natural splitting of matrix C^. Second, we develop a technique based on the use of a proper projection operator annihilating the near-kernel modes of C^. Both approaches give rise to a novel class of preconditioners denoted as Enhanced RPF (ERPF). Furthermore, effectiveness and robustness of the proposed algorithms are demonstrated in both theoretical benchmarks and real-world large-size applications, outperforming the native RPF preconditioner.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Scheme for Semilinear Nonlocal Diffusion Equations with Volume Constraints

This work presents a probabilistic scheme for solving semilinear nonlocal diffusion equations with volume constraints and integrable kernels. The nonlocal model of interest is defined by a time-dependent semilinear partial integro-differential equation (PIDE), in which the integro-differential operator consists of both local convection-diffusion and nonlocal diffusion operators. Here, our numerical scheme is based on the direct approximation of the nonlinear Feynman–Kac formula that establishes a link between nonlinear PIDEs and stochastic differential equations. The exploitation of the Feynman–Kac representation avoids solving dense linear systems arising from nonlocal operators. Compared with existing stochastic approaches, our method can achieve first-order convergence after balancing the temporal and spatial discretization errors, which is a significant improvement of existing probabilistic/stochastic methods for nonlocal diffusion problems. Error analysis of our numerical scheme is established. The effectiveness of our approach is shown in two numerical examples. The first example considers a three-dimensional nonlocal diffusion equation to numerically verify the error analysis results. The second example presents a physics problem motivated by the study of heat transport in magnetically confined fusion plasmas.

97 MATHEMATICS AND COMPUTING↗