Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Collins-Soper kernel and reduced soft function in lattice QCD

We evaluate the Collins-Soper kernel and the reduced soft function in lattice QCD, incorporating 𝒪⁡(𝛼 𝑠 ) matching corrections. The calculation relies on the evaluation of the quasitransverse momentum–dependent wave function with asymmetric staple-shaped quark bilinear operators and four-point meson form factors. These quantities are computed nonperturbatively using two 𝑁 𝑓 =2 + 1 + 1 twisted-mass fermion ensembles with the same lattice spacing of 𝑎 = 0.093 fm: the first ensemble has a lattice size of 24 3 × 48 and a pion mass of 346 MeV, and the second one has a lattice size of 32 3 × 64 and a pion mass of 261 MeV. The Collins-Soper kernel and the soft function are needed for the determination of the transverse momentum–dependent parton distribution functions.

Lattice QCD↗

Determination of the Collins-Soper Kernel from Lattice QCD

This Letter presents a determination of the quark Collins-Soper kernel, which relates transverse-momentum-dependent parton distributions (TMDs) at different rapidity scales, using lattice quantum chromodynamics (QCD). This is the first such determination with systematic control of quark mass, operator mixing, and discretization effects. Next-to-next-to-leading logarithmic matching is used to match lattice-calculable distributions to the corresponding TMDs. The continuum-extrapolated lattice QCD results are consistent with several recent phenomenological parametrizations of the Collins-Soper kernel and are precise enough to disfavor other parametrizations. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Asymptotically Compatible Reproducing Kernel Collocation and Meshfree Integration for Nonlocal Diffusion

Reproducing kernel (RK) approximations are meshfree methods that construct shape functions from sets of scattered data. We present an asymptotically compatible (AC) RK collocation method for nonlocal diffusion models with Dirichlet boundary condition. The numerical scheme is shown to be convergent to both nonlocal diffusion and its corresponding local limit as nonlocal interaction vanishes. The analysis is carried out on a special family of rectilinear Cartesian grids for a linear RK method with designed kernel support. The key idea for the stability of the RK collocation scheme is to compare the collocation scheme with the standard Galerkin scheme, which is stable. In addition, assembling the stiffness matrix of the nonlocal problem requires costly computational resources because high-order Gaussian quadrature is necessary to evaluate the integral. We thus provide a remedy to the problem by introducing a quasi-discrete nonlocal diffusion operator for which no numerical quadrature is further needed after applying the RK collocation scheme. The quasi-discrete nonlocal diffusion operator combined with RK collocation is shown to be convergent to the correct local diffusion problem by taking the limits of nonlocal interaction and spatial resolution simultaneously. The theoretical results are then validated with numerical experiments. We additionally illustrate a connection between the proposed technique and an existing optimization based approach based on generalized moving least squares.

97 MATHEMATICS AND COMPUTING↗

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

97 MATHEMATICS AND COMPUTING↗

Temperature Effect of Gas Bubble Evolution in UCN Fuel Kernels Irradiated by Swift Xe Ions

UC1-xNx fuel kernels provided by Oak Ridge National Laboratory (ORNL) were irradiated by 84 MeV Xe ions at two different temperatures (450°C & 750°C) at the Argonne Tandem Linac Accelerator System (ATLAS) at Argonne National Laboratory, followed by post-irradiation examination. The primary goal of this study was to understand gas bubble formation (due to accumulation of Xe gas) and corresponding size evolution dependent upon net Xe deposition at the two different temperatures. From the post-irradiation examinations of the samples, it can be concluded at 750°C, with same amount of dose received, the Xe gas bubbles seems to coarsen much more easily compared to 450°C. The results generated for fission gas bubble evolution observed in this study can be used to support fuel performance models for UC1-xNx fuel kernels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Preliminary results from Low Pressure Steam Oxidation Testing of ALD ZrN and ZrO2 Coating Deposited over UCN Fuel Kernels

Steam oxidation testing was used to investigate as-developed 100 nm ALD ZrN and 1000 nm ZrO 2 coatings deposited over UC1-xNx fuel kernels. The results of these tests were compared against oxidation of uncoated kernels. This work was performed to support development of coatings which both provide high temperature Zr metal diffusion barrier and also provide resistance against oxidation, especially against high temperature steam and/or air. This work was performed in collaboration with Oak Ridge National Laboratory (ORNL), using ORNL-provided samples. The oxidation tests yielded valuable information, including relative differences in performance of different coating materials and coating thicknesses.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ExaSGD: 2021 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-214 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY21 the main objective of the Kernel Thrust was (i) provide robust optimization solver(s) that run efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations and (ii) provide coarse-grain parallel optimization capabilities that exploit the decomposition opportunities present in the SC-ACOPF challenge problems to provide exascale-capable solvers.

97 MATHEMATICS AND COMPUTING↗

Technical note: Recommendations for diagnosing cloud feedbacks and rapid cloud adjustments using cloud radiative kernels

Abstract. The cloud radiative kernel method is a popular approach to quantify cloud feedbacks and rapid cloud adjustments to increased CO2 concentrations and to partition contributions from changes in cloud amount, altitude, and optical depth. However, because this method relies on cloud property histograms derived from passive satellite sensors or produced by passive satellite simulators in models, changes in obscuration of lower-level clouds by upper-level clouds can cause apparent low-cloud feedbacks and adjustments, even in the absence of changes in lower-level cloud properties. Here, we provide a methodology for properly diagnosing the impact of changing obscuration on cloud feedbacks and adjustments and quantify these effects across climate models. Averaged globally and across global climate models, properly accounting for obscuration leads to weaker positive feedbacks from lower-level clouds and stronger positive feedbacks from upper-level clouds while simultaneously removing a mostly artificial anti-correlation between them. Given that the methodology for diagnosing cloud feedbacks and adjustments using cloud radiative kernels has evolved over several papers, and obscuration effects have only occasionally been considered in recent papers, this paper serves to establish recommended best practices and to provide a corresponding code base for community use.

54 ENVIRONMENTAL SCIENCES↗

Bell nozzle kernel analysis program

Bell Nozzle Kernel Analysis Program computes and analyzes the supersonic flowfield in the kernel, or initial expansion region, of a bell or conical nozzle. It analyzes both plane and axisymmetric geometrices for specified gas properties, nozzle throat geometry and input line.

Elliot, J. J.↗

A kernel function method for computing steady and oscillatory supersonic aerodynamics with interference.

The method presented uses a collocation technique with the nonplanar kernel function to solve supersonic lifting surface problems with and without interference. A set of pressure functions are developed based on conical flow theory solutions which account for discontinuities in the supersonic pressure distributions. These functions permit faster solution convergence than is possible with conventional supersonic pressure functions. An improper integral of a 3/2 power singularity along the Mach hyperbola of the nonplanar supersonic kernel function is described and treated. The method is compared with other theories and experiment for a variety of cases.

Cunningham, A. M., Jr.↗

Oscillatory supersonic kernel function method for interfering surfaces

In the method presented in this paper, a collocation technique is used with the nonplanar supersonic kernel function to solve multiple lifting surface problems with interference in steady or oscillatory flow. The pressure functions used are based on conical flow theory solutions and provide faster solution convergence than is possible with conventional functions. In the application of the nonplanar supersonic kernel function, an improper integral of a 3/2 power singularity along the Mach hyperbola is described and treated. The method is compared with other theories and experiment for two wing-tail configurations in steady and oscillatory flow.

Cunningham, A. M., Jr.↗

A numerical solution for two-dimensional Fredholm integral equations of the second kind with kernels of the logarithmic potential form

Two dimensional Fredholm integral equations with logarithmic potential kernels are numerically solved. The explicit consequence of these solutions to their true solutions is demonstrated. The results are based on a previous work in which numerical solutions were obtained for Fredholm integral equations of the second kind with continuous kernels.

Gabrielsen, R. E.↗