Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Chemical Interaction of Palladium in Uranium Oxycarbide Nuclear Fuel Kernel

Chemical Interaction of Palladium in Uranium Oxycarbide Nuclear Fuel Kernel Jana Howard1,3, Guang Yang3, Haiyan Zhao1*, Patrick Warren4, Tiankai Yao3, Steven Cavazos4 Elizabeth Sooby Wood4*, Ching-heng Shiau2 1 University of Idaho, Environmental Science, Idaho Falls, Idaho, USA 2 Boise State University, Microscopy and Characterization Suite, Idaho Falls, Idaho, USA 3 Idaho National Laboratory, Idaho Falls, Idaho, USA 4 University of Texas San Antonio, Department of Physics and Astronomy, San Antonio, Texas, USA *Corresponding author: haiyanz@uidaho.edu; elizabeth.soobywood@utsa.edu Tri-structural Isotropic (TRISO) particle fuel is the preferred choice for the newest and next-generation high-temperature nuclear reactors due to its robust construction [1]. The fuel design effectively contains fission products inside the particle at extremely high temperatures [2]. However, the effect of transition metal fission products on fuel performance and structural integrity remains largely unknown, creating uncertainties in predicting fuel performance and potential failure mechanisms [2]. For example, palladium (Pd) has a strong tendency to form intermetallic phases with uranium (U). These new phases tend to be hard and brittle which could lead to fracture of fuel kernel during irradiation [3,4]. Pd is not normally produced in high concentrations in normal fission reactions however, driving the production of Pd into the locality of a Uranium Oxycarbide (UCO) nuclear fuel provides an opportunity to closely observe the diffusion and chemical interactions. This study focuses on detailed transmission electron microscopy characterization of the intermetallic phases formed by the interaction between a UCO kernel and a Pd bar after annealing at 1100 °C for 100 hours. Figure 1 provides an overview of the UCO kernel in contact with the Pd bar. The UCO-Pd sample surface was examined using a Focused Ion Beam (FIB) Quanta 3D FEG FEI in VCD detector mode at 10kV, 83pA to reveal surface features. Several key surface features in the interaction region were observed including: (1) a small area with dendritic microstructure, (2) color contrast near the UCO kernel and Pd boundary, and (3) darker and lighter regions throughout the sample surface. To analyze diffusion behavior, FIB was used to prepare lamellae from five different areas of the UCO-Pd sample, as well as one for the as-received sample. These lamellae were examined using a Scanning Transmission Electron Microscope ThermoFisher Spectra 300 (STEM-Spectra). Energy Dispersive Spectroscopy results confirmed the diffusion between the UCO fuel and Pd bar. Figure 2 shows the identified phases of UC, UO2, and Pd dispersed throughout the interaction region. It was also found that the Pd concentration decreases as the radial distance from the UCO-Pd interface increases. The Pd concentration was 36.93% atomic fraction 11µm from the interaction region then decreased to 7.51% atomic fraction at 257 µm away from the interaction region. Figure 3 shows how concentrations of U and Pd vary across the interaction zone, highlighting the extent of diffusion. This study confirms that Pd interacts with surrounding material to form U-Pd phase and diffusion zones. These diffusion zones vary in composition depending on its radial distance from the point of contact with the Pd bar. These findings contribute to a deeper understanding of the fission product behavior in high temperature nuclear fuels, aiding in the prediction and mitigation of potential fuel degradation mechanisms. A B C Fig. 1. SEM micrographs of UCO fuel kernel in contact with solid Pd bar and the lift out locations. Figure 1A shows UCO fuel kernel in contact with solid Pd bar. Figure 1B shows a higher magnification image of the UCO fuel-Pd interaction zone. Figure 1C shows lift out sites for the lamellae. A B C Fig 2. EDS maps reveal the distribution of U, O, and Pd in location 3. Figure 2A shows the distribution of uranium. Figure 2B shows the distribution of oxygen. Figure 2C shows the distribution of palladium. A B C Fig 3. Palladium and uranium concentration across interaction zone in location 2. Figure 3A shows EDS map of lift out number 2. Figure 3B shows a zoomed in EDS map taken from the interaction zone. Figure 3C shows a line graph of uranium and palladium concentrations across the interaction zone. References: 1. B.E. Wells, N.R. Phillips, K.J. Geelhood. Pacific West Laboratory. (2021). TRISO Fuel: Properties and Failure Modes. https://www.nrc.gov/docs/ML2117/ML21175A152.pdf (Accessed January 16, 2025). 2. American Nuclear Society. TRISO Fuel Development Progresses in INL, ORNL. https://www.gen-4.org/gif/upload/docs/application/pdf/2014-03/nov13nn_fuel_reprint.pdf (Accessed January 16, 2025) 3. Clark R.A., M.A. Con

36 - MATERIALS SCIENCE↗

Evaluation of CloudSat Radiative Kernels Using ARM and CERES Observations and ERA5 Reanalysis

Despite the widespread use of the radiative kernel technique for studying radiative feedbacks and radiative forcings, there has not been any systematic, observation-based validation of the radiative kernel method. Here, we utilize observed and reanalyzed radiative fluxes and atmospheric profiles from the Atmospheric Radiation Measurement (ARM) program and ERA5 reanalysis to assess a set of observation-based radiative kernels from CloudSat for six ARM sites. The CloudSat radiative kernels, convoluted with the ERA5 state variables, can almost perfectly reconstruct the monthly anomalies of shortwave (SW) and longwave (LW) radiative fluxes in ERA5 at the surface (SFC) and top-of-atmosphere (TOA) with correlations significantly being greater than 0.95. The biases of kernel-estimated flux anomalies calculated using the ARM-observed state variables can be more than twice as large when compared with the ARM-observed surface flux anomalies and Clouds and Earth’s Radiant Energy System (CERES) observed anomalies at the TOA. Generally, clouds contribute to most (>60%) of the variance of flux anomalies at Southern Great Plain (SGP), Tropical Western Pacific (TWP), and Eastern North Atlantic (ENA), and surface albedo dominates (>69%) the variance of SW flux anomalies at North Slope of Alaska (NSA). Furthermore, the radiative kernels exhibit the lowest correlation (r ~ [0.55,0.85]) when reconstructing SFC LW flux anomalies at SGP, TWP, and ENA, whose biases are related to the possibility that the kernels may not fully capture the characteristics associated with MJO and ENSO at TWP and the presence of clouds at SGP and ENA.

54 ENVIRONMENTAL SCIENCES↗

Importance of kernel bandwidth in quantum machine learning

Quantum kernel methods are considered a promising avenue for applying quantum computers to machine learning problems. Identifying hyperparameters controlling the inductive bias of quantum machine learning models is expected to be crucial given the central role hyperparameters play in determining the performance of classical machine learning methods. In this work we introduce the hyperparameter controlling the bandwidth of a quantum kernel and show that it controls the expressivity of the resulting model. We use extensive numerical experiments with multiple quantum kernels and classical data sets to show consistent change in the model behavior from underfitting (bandwidth too large) to overfitting (bandwidth too small), with optimal generalization in between. We draw a connection between the bandwidth of classical and quantum kernels and show analogous behavior in both cases. Furthermore, we show that optimizing the bandwidth can help mitigate the exponential decay of kernel values with qubit count, which is the cause behind recent observations that the performance of quantum kernel methods decreases with qubit count. Here, we reproduce these negative results and show that if the kernel bandwidth is optimized, the performance instead improves with growing qubit count and becomes competitive with the best classical methods.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Numerical evidence against advantage with quantum fidelity kernels on classical data

Quantum machine learning techniques are commonly considered one of the most promising candidates for demonstrating practical quantum advantage. In particular, quantum kernel methods have been demonstrated to be able to learn certain classically intractable functions efficiently if the kernel is well aligned with the target function. In the more general case, quantum kernels are known to suffer from exponential “flattening” of the spectrum as the number of qubits grows, preventing generalization and necessitating the control of the inductive bias by hyperparameters. We show that the general-purpose hyperparameter-tuning techniques proposed to improve the generalization of quantum kernels lead to the kernel becoming well approximated by a classical kernel, removing the possibility of quantum advantage. We provide extensive numerical evidence for this phenomenon utilizing multiple previously studied quantum feature maps and both synthetic and real data. Our results show that unless novel techniques are developed to control the inductive bias of quantum kernels, they are unlikely to provide a quantum advantage on classical data that lacks special structure.

Quantum Information, Science & Technology↗

First-principles wave-vector- and frequency-dependent exchange-correlation kernel for jellium at all densities

Here we propose a spatially and temporally nonlocal exchange correlation (XC) kernel for the spin-unpolarized fluid phase of ground-state jellium for use in time-dependent density functional and linear response calculations. The kernel is constructed to satisfy known properties of the exact XC kernel to accurately describe the correlation energies of bulk jellium and to satisfy frequency-moment sum rules at a wide range of bulk jellium densities, including those low densities that display strong correlation and symmetry breaking. These effects are easier to understand in the simple jellium model than in real systems. All exact constraints satisfied by the recent MCP07 kernel are maintained in the revised MCP07 (rMCP07) kernel, while others are added. The revision $f^{rMCP07}_{XC}$ (q, ω) differs from MCP07 only for nonzero frequencies ω. Only at densities much lower than those of real bulk metals is the frequency dependence of the kernel important for the correlation energy of jellium. As the wave vector q tends to zero, the kernel has a -4$πα(ω)/q^2$ divergence whose frequency-dependent ultranonlocality coefficient $α(ω)$ vanishes in jellium, and is predicted by rMCP07 to be extremely small for the real metals Al and Na.

36 MATERIALS SCIENCE↗

Non-Blind Deblurring for Fluorescence: A Deformable Latent Space Approach with Kernel Parameterization

We report N\non-blind deblurring (NBD) is a modeling method of the image deblurring problem in computer vision, where the blurring kernel is known or can be externally estimated. In this paper, we attempt to solve a parametric NBD problem, inspired by the simultaneous acquisition of ptychography and fluorescent imaging (FI). Ptychography is an imaging method that favors larger probes, i.e. convolutional kernels, while FI relies on a small probe for high resolution. Also, the kernel can be solved during ptychographic reconstruction. With Ptycho-FI using the same larger kernel, we can perform NBD on the blurred fluorescent images to achieve high-resolution FI, and thus speed up the experiments. To this end, we design a deep latent space deformation network that is directly parameterized by the kernel. The network consists of three components: encoder, deformer, and decoder, where the deformer is specifically meant to rectify the latent space representations of blurred images to a standard latent space, regardless of the kernel. The deformation network is trained with a two-stage training scheme. We conduct extensive experiments to confirm that our parametric model can adapt to drastically different blurring kernels and perform robust deblurring.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Decoupled Fréchet kernels based on a fractional viscoacoustic wave equation

We formulate the Fréchet kernel computation using the adjoint-state method based on a fractional viscoacoustic wave equation. We first numerically prove that both the 1/2- and the 3/2-order fractional Laplacian operators are self-adjoint. Using this property, we show that the adjoint wave propagator preserves the dispersion and compensates the amplitude, while the time-reversed adjoint wave propagator behaves identically as the forward propagator with the same dispersion and dissipation characters. Without introducing rheological mechanisms, this formulation adopts an explicit Q parameterization, which avoids the implicit Q in the conventional viscoacoustic/viscoelastic full waveform inversion (Q-FWI). In addition, because of the decoupling of operators in the wave equation, the viscoacoustic Fréchet kernel is separated into three distinct contributions with clear physical meanings: lossless propagation, dispersion, and dissipation. Here, we find that the lossless propagation kernel dominates the velocity kernel, while the dissipation kernel dominates the attenuation kernel over the dispersion kernel.

58 GEOSCIENCES↗

Temperature Distribution Within an Ignition Kernel Initiated by a Laser-Induced Plasma

The sizes of, and temperature distributions within, ignition kernels initiated by a Q-switched neodymium-doped yttrium aluminum garnet laser-induced plasma in an unconfined lean premixed hydrogen-air upward jet flow are investigated. The experiments involved a range of jet velocities and a range of deposited laser energies at a fixed height above the exit along the axis of a burner. The growth of, and the temperature distributions within, the ignition kernels, as affected by the size and the energy distribution of the laser-induced plasma, are monitored with an infrared camera. The initial ignition kernels’ areas are larger with higher laser pulse energies and remain unchanged up to [Formula: see text] and then increase by factors of up to 3 at [Formula: see text]. The change in the kernel area caused by the jet velocities is less than 1.5%. An increase of the bulk velocity by 190% decreases the ignition kernel temperature by 6%. This reduction in the ignition kernel temperatures is because of an increase in energy losses by a factor of 2 and decreases in heat releases by 2% at [Formula: see text] and by 11% at [Formula: see text]. The present contributions are: measurements of and insights into temperature distributions and kernel development rates during the laser-induced plasma ignition process at different deposited energies and flow velocities.

Engineering↗

Atomistic and cluster dynamics modeling of fission gas (Xe) diffusivity in TRISO fuel kernels

TRISO fuel particles are candidates for use in next generation reactors including gas reactors, fluoride salt-cooled high temperature reactors, and micro-reactors. The UCO fuel kernel consists of a uranium dioxide (UO) and uranium carbide mixture. The addition of UC helps suppress the formation of carbon monoxide gas, which led to failures during initial TRISO development. The addition of uranium carbide alters the chemistry of the UO kernel, which is known to influence performance parameters such as fission gas diffusivity, although the impact has not been quantified and no models exist that take the change in chemistry into account. Therefore, better understanding and more accurate models of the impact of chemistry on fuel performance are of high priority. In this paper, a first-principles density functional theory (DFT) and empirical potential based multi-scale study has been carried out to model the diffusivity of fission gas xenon (Xe) in UCO TRISO fuel kernels. The focus is on the UO component in the UCO fuel kernels, as that represents the largest volume fraction of the fuel kernels. The study relies on DFT and empirical potential calculations to determine Xe and point defect properties, which are then used in thermodynamic and kinetic models to predict diffusion for intrinsic conditions. In addition, the information is utilized in cluster dynamics simulations using the Centipede code to estimate the impact of irradiation on defect transport. Additionally, the presence of UC or UC in the UCO fuel kernels is shown to have a substantial impact on the UO non-stoichiometry by inducing oxygen vacancies and driving UO sub-stoichiometric, which causes much slower Xe diffusion in UCO compared to light water reactor UO fuel. The application of this model in fuel performance simulations using the Bison code is also demonstrated.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Progress towards understanding ultranonlocality through the wave-vector and frequency dependence of approximate exchange-correlation kernels

In the framework of time-dependent density functional theory (TDDFT), the exact exchange-correlation (xc) kernel f xc (n, q, ω) determines the ground-state energy, excited-state energies, lifetimes, and the time-dependent linear density response of any many-electron system. The recently developed MCP07 xc kernel f xc (n, q, ω) of A. Ruzsinszky et al. [Phys. Rev. B 101, 245135 (2020)] yields excellent uniform electron gas (UEG) ground-state energies and plausible plasmon lifetimes. As MCP07 is constructed to describe f xc of the UEG, it cannot capture optical properties of real materials. To verify this claim, we follow Nazarov et al. [Phys. Rev. Lett. 102, 113001 (2009)] to construct the long-range, dynamic xc kernel, lim q→0 f xc (n, q, ω) = -α(ω)e 2 /q 2 , of a weakly inhomogeneous electron gas, using MCP07 and other common xc kernels. The strong wavevector and frequency dependence of the “ultranonlocality” coefficient α(ω) is demonstrated for a variety of simple metals and semiconductors. We examine how imposing exact constraints on an approximate kernel shapes α(ω). Comparisons to kernels derived from correlated-wavefunction calculations are drawn.

36 MATERIALS SCIENCE↗

Advanced stationary and nonstationary kernel designs for domain-aware Gaussian processes

Gaussian process regression is a widely-applied method for function approximation and uncertainty quantification. The technique has gained popularity recently in the machine learning community due to its robustness and interpretability. The mathematical methods we discuss in this paper are an extension of the Gaussian-process framework. We are proposing advanced kernel designs that only allow for functions with certain desirable characteristics to be elements of the reproducing kernel Hilbert space (RKHS) that underlies all kernel methods and serves as the sample space for Gaussian process regression. These desirable characteristics reflect the underlying physics; two obvious examples are symmetry and periodicity constraints. In addition, non-stationary kernel designs can be defined in the same framework to yield flexible multi-task Gaussian processes. We will show the impact of advanced kernel designs on Gaussian processes using several synthetic and two scientific data sets. The results of our research show that including domain knowledge, communicated through advanced kernel designs, has a significant impact on the accuracy and relevance of the function approximation.

97 MATHEMATICS AND COMPUTING↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Simultaneous dissection of grain carotenoid levels and kernel color in biparental maize populations with yellow-to-orange grain

Maize enriched in provitamin A carotenoids could be key in combatting vitamin A deficiency in human populations relying on maize as a food staple. Consumer studies indicate that orange maize may be regarded as novel and preferred. This study identifies genes of relevance for grain carotenoid concentrations and kernel color, through simultaneous dissection of these traits in 10 families of the US maize nested association mapping panel that have yellow to orange grain. Quantitative trait loci were identified via joint-linkage analysis, with phenotypic variation explained for individual kernel color quantitative trait loci ranging from 2.4% to 17.5%. These quantitative trait loci were cross-analyzed with significant marker-trait associations in a genome-wide association study that utilized ~27 million variants. Nine genes were identified: four encoding activities upstream of the core carotenoid pathway, one at the pathway branchpoint, three within the α- or β-pathway branches, and one encoding a carotenoid cleavage dioxygenase. Of these, three exhibited significant pleiotropy between kernel color and one or more carotenoid traits. Kernel color exhibited moderate positive correlations with β-branch and total carotenoids and negligible correlations with α-branch carotenoids. These findings can be leveraged to simultaneously achieve desirable kernel color phenotypes and increase concentrations of provitamin A and other priority carotenoids.

59 BASIC BIOLOGICAL SCIENCES↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Inverse Mapping of the Collision Kernel and Wall Flux Scaling in a Tall Convection‐Cloud Chamber Using Local Sensors and Knowledge‐Informed Deep Learning

Droplet collision–coalescence is a crucial process in cloud physics, but accurately representing this process under different dynamical conditions remains challenging. A proposed future convective‐cloud chamber aims to investigate this key process, but the method for observing it remains unclear, even though it is theoretically established that collision‐coalescence will occur. This study serves as a proof‐of‐concept demonstration of how knowledge‐informed deep learning, combined with measurement data from local sensors in the chamber, can be used to estimate the collision kernels, which determine how the droplet size distribution evolves during collision‐coalescence. In addition to estimating the collision kernel, we also address wall fluxes, another uncertain but important process that acts as a source of heat and moisture in the chamber. Ensemble runs of large‐eddy simulations are conducted by scaling the wall fluxes and the collision kernel, while the measured flow and cloud properties are used as inputs for a neural network. Results indicate that this approach successfully maps the scaling of wall fluxes and the collision kernel with biases of approximately 1% or less relative to the range of the target data. This proof‐of‐concept lays the groundwork for future applications; when the real measurements are available, real sensor data combined with the trained model presented in this work will enable estimation of the actual wall fluxes and collision kernel.

cloud chamber↗

Exact Gaussian processes for massive datasets via non-stationary sparsity-discovering kernels

Abstract A Gaussian Process (GP) is a prominent mathematical framework for stochastic function approximation in science and engineering applications. Its success is largely attributed to the GP’s analytical tractability, robustness, and natural inclusion of uncertainty quantification. Unfortunately, the use of exact GPs is prohibitively expensive for large datasets due to their unfavorable numerical complexity of $$O(N^3)$$ O ( N 3 ) in computation and $$O(N^2)$$ O ( N 2 ) in storage. All existing methods addressing this issue utilize some form of approximation—usually considering subsets of the full dataset or finding representative pseudo-points that render the covariance matrix well-structured and sparse. These approximate methods can lead to inaccuracies in function approximations and often limit the user’s flexibility in designing expressive kernels. Instead of inducing sparsity via data-point geometry and structure, we propose to take advantage of naturally-occurring sparsity by allowing the kernel to discover—instead of induce—sparse structure. The premise of this paper is that the data sets and physical processes modeled by GPs often exhibit natural or implicit sparsities, but commonly-used kernels do not allow us to exploit such sparsity. The core concept of exact, and at the same time sparse GPs relies on kernel definitions that provide enough flexibility to learn and encode not only non-zero but also zero covariances. This principle of ultra-flexible, compactly-supported, and non-stationary kernels, combined with HPC and constrained optimization, lets us scale exact GPs well beyond 5 million data points.

97 MATHEMATICS AND COMPUTING↗

Time-dependent orbital-free density functional theory: Background and Pauli kernel approximations

Time-dependent orbital-free density functional theory (DFT) is an efficient method for calculating the dynamic properties of large-scale quantum systems due to the low computational cost compared to standard time-dependent DFT. In this work, we formalize this method by mapping the real system of interacting fermions onto a fictitious system of noninteracting bosons. The dynamic Pauli potential and associated kernel emerge as key ingredients of time-dependent orbital-free DFT. Using the uniform electron gas as a model system, we derive an approximate frequency-dependent Pauli kernel. Pilot calculations suggest that space nonlocality is a key feature for this kernel. Nonlocal terms arise already in the second-order expansion with respect to unitless frequency and reciprocal space variable ($\frac{ω}{qk_F}$ and $\frac{q}{2k_F}$, respectively). Given the encouraging performance of the proposed kernel, we expect it will lead to more accurate orbital-free DFT simulations of nanoscale systems out of equilibrium. Additionally, the proposed path to formulate nonadiabatic Pauli kernels presents several avenues for further improvements which can be exploited in future work to improve the results.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Understanding Performance Portability of SYCL Kernels: A Case Study with the All-Pairs Distance Calculation in Bioinformatics on GPUs

SYCL is a portable programming model. Toward the goal of a better understanding of performance portability of SYCL kernels on GPUs, we select a bioinformatics kernel for computing the all-pairs distance as a case study. After migrating the kernel from CUDA to HIP and SYCL, we evaluate the performance of the CUDA, HIP, and SYCL kernels on NVIDIA V100 and AMD MI210 GPUs. We analyze the GPU instructions from the kernels to explain performance gaps between SYCL and CUDA/HIP. We hope that the findings are valuable for improving performance portability of SYCL.

Jin, Zheming↗