Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Raman spectroscopy of uranium nitride kernels

Uranium nitride is an advanced fuel candidate for a wide variety of advanced nuclear reactors. This work summarizes the first characterization of UN kernels by Raman spectroscopy. First-principles density functional theory calculations were performed to predict the Raman spectra of uranium sesquinitride (U 2 N 3 ), uranium dinitride (UN 2 ), uranium mononitride (UN), uranium monocarbide (UC), as well as U-N-C (UN 1-x C x ) and a U-N-C-O mixture. Further, a core–shell structure was identified by scanning electron microscopy and Raman spectroscopy imaging. A signal at ~500 cm -1 was identified on the periphery of the core-shell structure, possibly corresponding to U 2 N 3 and/or UN 2 . This signal broadens and shifts to 470 cm -1 because of the formation of UNC, UNCO or U 2 N 3+x structures. The culmination of this work demonstrates the feasibility of using Raman spectroscopy to identify variations in composition and phases in UN kernels.

36 MATERIALS SCIENCE↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗

An Efficient Bayesian Approach to Learning Droplet Collision Kernels: Proof of Concept Using “Cloudy,” a New n -Moment Bulk Microphysics Scheme

The small-scale microphysical processes governing the formation of precipitation particles cannot be resolved explicitly by cloud resolving and climate models. Instead, they are represented by microphysics schemes that are based on a combination of theoretical knowledge, statistical assumptions, and fitting to data (“tuning”). Historically, tuning was done in an ad hoc fashion, leading to parameter choices that are not explainable or repeatable. Recent work has treated it as an inverse problem that can be solved by Bayesian inference. The posterior distribution of the parameters given the data—the solution of Bayesian inference—is found through computationally expensive sampling methods, which require over $\mathcal{O}$(10 5 ) evaluations of the forward model; this is prohibitive for many models. We present a proof of concept of Bayesian learning applied to a new bulk microphysics scheme named “Cloudy,” using the recently developed Calibrate-Emulate-Sample (CES) algorithm. Cloudy models collision-coalescence and collisional breakup of cloud droplets with an adjustable number of prognostic moments and with easily modifiable assumptions for the cloud droplet mass distribution and the collision kernel. The CES algorithm uses machine learning tools to accelerate Bayesian inference by reducing the number of forward evaluations needed to $\mathcal{O}$(10 2 ). It also exhibits a smoothing effect when forward evaluations are polluted by noise. In a suite of perfect-model experiments, we show that CES enables computationally efficient Bayesian inference of parameters in Cloudy from noisy observations of moments of the droplet mass distribution. In an additional imperfect-model experiment, a collision kernel parameter is successfully learned from output generated by a Lagrangian particle-based microphysics model.

54 ENVIRONMENTAL SCIENCES↗

Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels

Next generation High-Energy Physics (HEP) experiments are presented with significant computational challenges, both in terms of data volume and processing power. Using compute accelerators, such as GPUs, is one of the promising ways to provide the necessary computational power to meet the challenge. The current programming models for compute accelerators often involve using architecture-specific programming languages promoted by the hardware vendors and hence limit the set of platforms that the code can run on. Developing software with platform restrictions is especially unfeasible for HEP communities as it takes significant effort to convert typical HEP algorithms into ones that are efficient for compute accelerators. Multiple performance portability solutions have recently emerged and provide an alternative path for using compute accelerators, which allow the code to be executed on hardware from different vendors. We apply several portability solutions, such as Kokkos, SYCL, C++17 std::execution::par, Alpaka, and OpenMP/OpenACC, on two mini-apps extracted from the mkFit project: p2z and p2r. These apps include basic kernels for a Kalman filter track fit, such as propagation and update of track parameters, for detectors at a fixed z or fixed r position, respectively. The two mini-apps explore different memory layout formats. We report on the development experience with different portability solutions, as well as their performance on GPUs and many-core CPUs, measured as the throughput of the kernels from different GPU and CPU vendors such as NVIDIA, AMD and Intel.

Kwok, Ka Hei Martin↗

Discovering causal structure with reproducing-kernel Hilbert space ε -machines

We merge computational mechanics’ definition of causal states (predictively equivalent histories) with reproducing-kernel Hilbert space (RKHS) representation inference. The result is a widely applicable method that infers causal structure directly from observations of a system’s behaviors whether they are over discrete or continuous events or time. A structural representation—a finite- or infinite-state kernel ϵ-machine—is extracted by a reduced-dimension transform that gives an efficient representation of causal states and their topology. In this way, the system dynamics are represented by a stochastic (ordinary or partial) differential equation that acts on causal states. We introduce an algorithm to estimate the associated evolution operator. Paralleling the Fokker–Planck equation, it efficiently evolves causal-state distributions and makes predictions in the original data space via an RKHS functional mapping. We demonstrate these techniques, together with their predictive abilities, on discrete-time, discrete-value infinite Markov-order processes generated by finite-state hidden Markov models with (i) finite or (ii) uncountably infinite causal states and (iii) continuous-time, continuous-value processes generated by thermally driven chaotic flows. The method robustly estimates causal structure in the presence of varying external and measurement noise levels and for very high-dimensional data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hybrid programming-model strategies for GPU offloading of electronic structure calculation kernels

To address the challenge of performance portability and facilitate the implementation of electronic structure solvers, we developed the basic matrix library (BML) and Parallel, Rapid O(N), and Graph-based Recursive Electronic Structure Solver (PROGRESS) library. The BML implements linear algebra operations necessary for electronic structure kernels using a unified user interface for various matrix formats (dense and sparse) and architectures (CPUs and GPUs). Focusing on density functional theory and tight-binding models, PROGRESS implements several solvers for computing the single-particle density matrix and relies on BML. In this paper, we describe the general strategies used for these implementations on various computer architectures, using OpenMP target functionalities on GPUs, in conjunction with third-party libraries to handle performance critical numerical kernels. In this study, we demonstrate the portability of this approach and its performance in benchmark problems.

36 MATERIALS SCIENCE↗

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Collins-Soper kernel and reduced soft function in lattice QCD

We evaluate the Collins-Soper kernel and the reduced soft function in lattice QCD, incorporating 𝒪⁡(𝛼 𝑠 ) matching corrections. The calculation relies on the evaluation of the quasitransverse momentum–dependent wave function with asymmetric staple-shaped quark bilinear operators and four-point meson form factors. These quantities are computed nonperturbatively using two 𝑁 𝑓 =2 + 1 + 1 twisted-mass fermion ensembles with the same lattice spacing of 𝑎 = 0.093 fm: the first ensemble has a lattice size of 24 3 × 48 and a pion mass of 346 MeV, and the second one has a lattice size of 32 3 × 64 and a pion mass of 261 MeV. The Collins-Soper kernel and the soft function are needed for the determination of the transverse momentum–dependent parton distribution functions.

Lattice QCD↗

Determination of the Collins-Soper Kernel from Lattice QCD

This Letter presents a determination of the quark Collins-Soper kernel, which relates transverse-momentum-dependent parton distributions (TMDs) at different rapidity scales, using lattice quantum chromodynamics (QCD). This is the first such determination with systematic control of quark mass, operator mixing, and discretization effects. Next-to-next-to-leading logarithmic matching is used to match lattice-calculable distributions to the corresponding TMDs. The continuum-extrapolated lattice QCD results are consistent with several recent phenomenological parametrizations of the Collins-Soper kernel and are precise enough to disfavor other parametrizations. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Asymptotically Compatible Reproducing Kernel Collocation and Meshfree Integration for Nonlocal Diffusion

Reproducing kernel (RK) approximations are meshfree methods that construct shape functions from sets of scattered data. We present an asymptotically compatible (AC) RK collocation method for nonlocal diffusion models with Dirichlet boundary condition. The numerical scheme is shown to be convergent to both nonlocal diffusion and its corresponding local limit as nonlocal interaction vanishes. The analysis is carried out on a special family of rectilinear Cartesian grids for a linear RK method with designed kernel support. The key idea for the stability of the RK collocation scheme is to compare the collocation scheme with the standard Galerkin scheme, which is stable. In addition, assembling the stiffness matrix of the nonlocal problem requires costly computational resources because high-order Gaussian quadrature is necessary to evaluate the integral. We thus provide a remedy to the problem by introducing a quasi-discrete nonlocal diffusion operator for which no numerical quadrature is further needed after applying the RK collocation scheme. The quasi-discrete nonlocal diffusion operator combined with RK collocation is shown to be convergent to the correct local diffusion problem by taking the limits of nonlocal interaction and spatial resolution simultaneously. The theoretical results are then validated with numerical experiments. We additionally illustrate a connection between the proposed technique and an existing optimization based approach based on generalized moving least squares.

97 MATHEMATICS AND COMPUTING↗

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

97 MATHEMATICS AND COMPUTING↗

Temperature Effect of Gas Bubble Evolution in UCN Fuel Kernels Irradiated by Swift Xe Ions

UC1-xNx fuel kernels provided by Oak Ridge National Laboratory (ORNL) were irradiated by 84 MeV Xe ions at two different temperatures (450°C & 750°C) at the Argonne Tandem Linac Accelerator System (ATLAS) at Argonne National Laboratory, followed by post-irradiation examination. The primary goal of this study was to understand gas bubble formation (due to accumulation of Xe gas) and corresponding size evolution dependent upon net Xe deposition at the two different temperatures. From the post-irradiation examinations of the samples, it can be concluded at 750°C, with same amount of dose received, the Xe gas bubbles seems to coarsen much more easily compared to 450°C. The results generated for fission gas bubble evolution observed in this study can be used to support fuel performance models for UC1-xNx fuel kernels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Preliminary results from Low Pressure Steam Oxidation Testing of ALD ZrN and ZrO2 Coating Deposited over UCN Fuel Kernels

Steam oxidation testing was used to investigate as-developed 100 nm ALD ZrN and 1000 nm ZrO 2 coatings deposited over UC1-xNx fuel kernels. The results of these tests were compared against oxidation of uncoated kernels. This work was performed to support development of coatings which both provide high temperature Zr metal diffusion barrier and also provide resistance against oxidation, especially against high temperature steam and/or air. This work was performed in collaboration with Oak Ridge National Laboratory (ORNL), using ORNL-provided samples. The oxidation tests yielded valuable information, including relative differences in performance of different coating materials and coating thicknesses.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ExaSGD: 2021 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-214 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY21 the main objective of the Kernel Thrust was (i) provide robust optimization solver(s) that run efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations and (ii) provide coarse-grain parallel optimization capabilities that exploit the decomposition opportunities present in the SC-ACOPF challenge problems to provide exascale-capable solvers.

97 MATHEMATICS AND COMPUTING↗