Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “KERNEL FUNCTION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The total neutron cross section of liquid and solid ammonia

Ammonia is a material of interest for future neutron moderators at high-power sources due to its high hydrogen density, low melting point, and resistance to polymerization in an intense radiation field. Its performance in such applications cannot currently be calculated due to the absence of suitable computer models for the interaction of neutrons with ammonia under relevant conditions. In an effort to develop suitable scattering kernels for computer simulations of moderator performance, we have conducted a series of Density Functional Theory and Molecular Dynamics calculations of the molecular-level thermal properties of ammonia at various temperatures within both the solid and liquid phases. In this paper, we compare computer calculations for the energy-dependent total neutron cross section of ammonia, based on these models, to experimental measurements of those cross sections at temperatures of 221 K, 180 K, and 35 K. The experimental data were collected over an energy range from 0.1 meV to 10 eV using time-of-flight techniques at the Low Energy Neutron Source (LENS) facility at Indiana University. This comparison provides a first validation in the development of thermal scattering libraries for Monte Carlo source design simulations based on liquid and solid ammonia. In conclusion, we also provide some insights into where additional development of tools for creating such models may be needed.

Ammonia↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Global weak solutions for a nonlocal multispecies Fokker–Planck–Landau system

The global-in-time existence of weak solutions to a spatially homogeneous multispecies Fokker–Planck–Landau system for plasmas in the three-dimensional whole space is shown. The Fokker–Planck–Landau system is a simplification of the Landau equations assuming a linearized, velocity-independent, and isotropic kernel. The resulting equations depend nonlocally and nonlinearly on the moments of the distribution functions via the multispecies local Maxwellians. Furthermore, the existence proof is based on a three-level approximation scheme, energy and entropy estimates, as well as compactness results, and it holds for both soft and hard potentials.

97 MATHEMATICS AND COMPUTING↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

Polymer coatings on zirconia microspheres via rotating flow fluid dynamics

Conventional tristructural isotropic (TRISO) coatings for nuclear fuels require multi-step and expensive formation processes; any breakage of the coatings may lead to fission species release. Polymer derived ceramic (PDC) coatings can be a suitable alternative to address these issues. In this study, allylhydridopolycarbosilane (SMP-10) coatings were created on yttria stabilized zirconia (YSZ) microspheres using a Rotating Flow Fluid Dynamics (RFFD) coating method. The effects of curing temperature, rotation speed, and coating cycle/time on the coating were analyzed. Surface functionalization of YSZ microspheres with NaOH resulted in good adhesion between the polymer precursor and the YSZ kernel particles. Spectroscopy analysis revealed complete curing of SMP-10 coated YSZ at 160 °C. Rotation speed and coating time significantly affect the coating characteristics. After 3 cycles at lower rotation speed (50 rpm) for 10 min each, the coating obtained was uniform and homogeneous. In comparison, the coating showed almost 10 times less eccentricity at a high rotation speed of 300 rpm. Compared with the coatings prepared at 300 rpm condition, the coatings have higher sphericity and lower eccentricity at 50 rpm. Overall, this study provides a novel and effective route for fabricating and curing SMP-10 precursor coatings on YSZ microspheres.

Ravi, Nivetha [University of Alabama, Birmingham]↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection↗

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗

Training quantum neural networks using the quantum information bottleneck method

Abstract We provide in this paper a concrete method for training a quantum neural network to maximize the relevant information about a property that is transmitted through the network. This is significant because it gives an operationally well founded quantity to optimize when training autoencoders for problems where the inputs and outputs are fully quantum. We provide a rigorous algorithm for computing the value of the quantum information bottleneck quantity within error ε that requires O ( log 2 ⁡ ( 1 / ϵ ) + 1 / δ 2 ) queries to a purification of the input density operator if its spectrum is supported on { 0 } ⋃ [ δ , 1 − δ ] for δ > 0 and the kernels of the relevant density matrices are disjoint. We further provide algorithms for estimating the derivatives of the QIB function, showing that quantum neural networks can be trained efficiently using the QIB quantity given that the number of gradient steps required is polynomial.

Çatlı, Ahmet Burak (ORCID:0000000152294141)↗

QCD factorization with multihadron fragmentation functions

Important aspects of quantum chromodynamics (QCD) factorization theorems are the properties of the objects involved that can be identified as universal. One example is that the definitions of parton densities and fragmentation functions for different types of hadrons differ only in the identity of the nonperturbative states that form the matrix elements, but are otherwise the same. This leads to independence of perturbative calculations on nonperturbative details of external states. It also lends support to interpretations of correlation functions as encapsulations of intrinsic nonperturbative properties. These characteristics have usually been presumed to still hold true in fragmentation functions even when the observed nonperturbative state is a small-mass cluster of n hadrons rather than simply a single isolated hadron. However, the multidifferential aspect of cross sections that rely on these latter types of fragmentation functions complicates the treatment of kinematical approximations in factorization derivations. That has led to recent claims that the operator definitions for fragmentation functions need to be modified from the single hadron case with nonuniversal prefactors. With such concerns as our motivation, we retrace the steps for factorizing the unpolarized semi-inclusive e + e − annihilation cross section and confirm that they do apply without modification to the case of a small-mass multihadron observed in the final state. In particular, we verify that the standard operator definition from single hadron fragmentation, with its usual prefactor, remains equally valid for the small-mass n -hadron case with the same hard parts and evolution kernels, whereas the more recently proposed definitions with nonuniversal prefactors do not. Our results reaffirm the reliability of most past phenomenological applications of dihadron fragmentation functions. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Building workflows for an interactive human-in-the-loop automated experiment (hAE) in STEM-EELS

Exploring the structural, chemical, and physical properties of matter on the nano- and atomic scales has become possible with the recent advances in aberration-corrected electron energy-loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). However, the current paradigm of STEM-EELS relies on the classical rectangular grid sampling, in which all surface regions are assumed to be of equal a priori interest. However, this is typically not the case for real-world scenarios, where phenomena of interest are concentrated in a small number of spatial locations, such as interfaces, structural and topological defects, and multi-phase inclusions. One of the foundational problems is the discovery of nanometer- or atomic-scale structures having specific signatures in EELS spectra. Herein, we systematically explore the hyperparameters controlling deep kernel learning (DKL) discovery workflows for STEM-EELS and identify the role of the local structural descriptors and acquisition functions in experiment progression. In agreement with the actual experiment, we observe that for certain parameter combinations the experiment path can be trapped in the local minima. We demonstrate the approaches for monitoring the automated experiment in the real and feature space of the system and knowledge acquisition of the DKL model. Based on these, we construct intervention strategies defining the human-in-the-loop automated experiment (hAE). This approach can be further extended to other techniques including 4D STEM and other forms of spectroscopic imaging. The hAE library is available on Github at https://github.com/utkarshp1161/hAE/tree/main/hAE.

Pratiush, Utkarsh [Univ. of Tennessee, Knoxville, ↗

Microstructural Topology as a Prescriptor for Quantum Coherence: Towards A Unified Framework for Decoherence in Superconducting Qubits

In superconducting quantum circuits, decoherence improvements are frequently obtained through process interventions that simultaneously modify surface chemistry, microstructural topology, and device geometry, leaving mechanistic attribution structurally underdetermined. Predictive materials engineering requires measurable structural statistics to be separated from geometry-dependent coupling coefficients into independently testable factors. We introduce the concept of classical and quantum microstructure. In that context, we formulate a channel-wise separable framework for decoherence in superconducting transmon qubits in which each loss channel is described by a reduced prescriptor. Here, a channel-specific microstructural state variable is determined independently of device geometry, and a geometry-dependent coupling functional is computable from field solutions without reference to surface chemistry. We derive this product form from a spatially resolved kernel representation and establish a perturbative separability criterion that defines the regime where independent variation of the variables is valid. The framework specifies five prescriptor classes for dominant loss pathways in transmon-class devices. Falsifiability is operationalized through a pre-committed 2x2 experimental protocol in which the variables must satisfy independent ratio checks within propagated uncertainty. A Minimum-Dataset Specification standardizes reporting for cross-laboratory inference. Part I establishes the conceptual and mathematical architecture; coordinated experimental validation is reserved for Part II.

Dravid, Vinayak P. [Northwestern U.]↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Excited-State Densities from Time-Dependent Density Functional Response Theory

While the variational principle for excited-state energies leads to a route to obtaining excited-state densities from time-dependent density functional theory, relatively little attention has been paid to the quality of the resulting densities in real space obtained with different exchange-correlation functional approximations or how nonadiabatic approximations developed for energies of states of double-excitation character perform for their densities. Here we derive an expression directly in real space for the excited-state density, which includes the case of nonadiabatic kernels and consequently is able, for the first time, to yield densities of states of double-excitation character. Under some well-defined simplifications, we compare the performance of the local-density approximation and exact-exchange approximation, which are in a sense at the opposite extremes of the fundamental functional approximations, on local and charge-transfer excitations in one-dimensional model systems and show that the dressed Time-Dependent Density Functional Theory (TDDFT) approach gives good densities of double excitations.

approximation↗

Scatter and Blur Corrections for High-Energy X-Ray Radiography

High-energy X-ray radiography is useful as a highly penetrating method for imaging through dense materials. However, the primary modes of interaction of X-rays at these energies involve scattering or the production of secondary high-energy photons, which can interfere with the image. In addition, detector blurring, often resulting from scatter within the detector, can reduce image sharpness. Both of these processes can be mitigated with the use of convolution kernels, with the main challenge being that the proper kernel to use is not known, particularly for the scatter contribution. By radiographing solid slabs of uniform attenuation, we show that point spread functions and material-specific point scatter functions can be determined to significantly reduce the effect of detector blurring and object scatter. Constraining the fits to the slabs and uniform transmission within the slabs is sufficient to recover these functions. A functional form that reproduces the angular distribution of high-energy bremsstrahlung X-rays is presented for recovering point scatter functions. In conclusion, the method is applied to radiographs of objects from bremsstrahlung X-ray sources operating at 4- and 7.5-MV endpoint energies and a significant increase in sharpness is observed.

Blind deconvolution↗