Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “KERNELS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Initial Kernel Timing Using a Simple PIM Performance Model

This presentation will describe some initial results of paper-and-pencil studies of 4 or 5 application kernels applied to a processor-in-memory (PIM) system roughly similar to the Cascade Lightweight Processor (LWP). The application kernels are: * Linked list traversal * Sun of leaf nodes on a tree * Bitonic sort * Vector sum * Gaussian elimination The intent of this work is to guide and validate work on the Cascade project in the areas of compilers, simulators, and languages. We will first discuss the generic PIM structure. Then, we will explain the concepts needed to program a parallel PIM system (locality, threads, parcels). Next, we will present a simple PIM performance model that will be used in the remainder of the presentation. For each kernel, we will then present a set of codes, including codes for a single PIM node, and codes for multiple PIM nodes that move data to threads and move threads to data. These codes are written at a fairly low level, between assembly and C, but much closer to C than to assembly. For each code, we will present some hand-drafted timing forecasts, based on the simple PIM performance model. Finally, we will conclude by discussing what we have learned from this work, including what programming styles seem to work best, from the point-of-view of both expressiveness and performance.

BRIEFING CHARTS↗

Evaluating HPC Kernels for Processing in Memory

Memory subsystems contribute significantly to the performance and energy efficiency of high-performance computing (HPC) applications. Traditional memory technologies with conventional organization (e.g., DRAM) are struggling to keep up with the increasing memory requirements of modern applications. Techniques such as multilayer cache hierarchy and out-of-order execution are still falling short of mitigating the penalty incurred by memory accesses. Processing-in-memory (PIM), which involves moving memory-intensive kernels to memory for execution instead of bringing the data to the processing unit, is emerging as a promising technique. PIM has recently received traction among computer architecture researchers, and the increasing research activity surrounding this technique indicates its potential to alleviate main memory performance bottlenecks. In this paper, we characterize and identify memory-intensive HPC kernels, perform a first-order evaluation of the PIM technique for selected HPC kernels, quantify performance deviation, and analyze the key factors that affect PIM efficiency.

Asifuzzaman, Kazi↗

Kernel-based global sensitivity analysis obtained from a single data set

Results from global sensitivity analysis (GSA) often guide the understanding of complicated input–output systems. Kernel-based GSA methods have recently been proposed for their capability of treating a broad scope of complex systems. In this paper, we develop a new set of kernel GSA tools when only a single set of input–output data is available. Three key advances are made: (1) A new numerical estimator is proposed that demonstrates an empirical improvement over previous procedures. (2) A computational method for generating inner statistical functions from a single data set is presented. (3) A theoretical extension is made to define conditional sensitivity indices, which reveal the degree that the inputs carry shared information about the output when inherent input–input correlations are present. Utilizing these conditional sensitivity indices, a decomposition is derived for the output uncertainty based on what is called the optimal learning sequence of the input variables, which remains consistent when correlations exist between the input variables. Further, while these advances cover a range of GSA subjects, a common single data set numerical solution is provided by a technique known as the conditional mean embedding of distributions. The new methodology is implemented on benchmark systems to demonstrate the provided insights.

42 ENGINEERING↗

Digital Modeling on Large Kernel Metamaterial Neural Network

Deep neural networks (DNNs) utilized recently are physically deployed with computational units (e.g., CPUs and GPUs). Such a design might lead to a heavy computational burden, significant latency, and intensive power consumption, which are critical limitations in applications such as Internet of Things (IoT), edge computing, and usage of drones. Recent advances in optical computational units (e.g., metamaterial) have shed light on energy-free and light-speed neural networks. However, the digital design of the metamaterial neural network (MNN) is fundamentally limited by its physical limitations, such as precision, noise, and bandwidth during fabrication. Moreover, the unique advantages of MNN’s (e.g., light-speed computation) are not fully explored via standard 3×3 convolution kernels. In this paper, we propose a novel large kernel metamaterial neural network (LMNN) that maximizes the digital capacity of the state-of-the-art (SOTA) MNN with model re-parametrization and network compression, while also considering the optical limitation explicitly. The new digital learning scheme can maximize the learning capacity of MNN while modeling the physical restrictions of meta-optics. With the proposed LMNN, the computation cost of the convolutional front-end can be offloaded to fabricated optical hardware. The experimental results on two publicly available datasets demonstrate that the optimized hybrid design improved classification accuracy while reducing computational latency. In conclusion, the development of the proposed LMNN is a promising step towards the ultimate goal of energy-free and light-speed AI.

97 MATHEMATICS AND COMPUTING↗

Gaussian Kernel Methods for Seismic Fragility and Risk Assessment of Mid-Rise Buildings

Seismic fragility functions can be evaluated using the cloud analysis method with linear regression which makes three fundamental assumptions about the relation between structural response and seismic intensity: log-linear median relationship, constant standard deviation, and Gaussian distributed errors. While cloud analysis with linear regression is a popular method, the degree to which these individual and compounded assumptions affect the fragility and the risk of mid-rise buildings needs to be systematically studied. This paper conducts such a study considering three building archetypes that make up a bulk of the building stock: RC moment frame, steel moment frame, and wood shear wall. Gaussian kernel methods are employed to capture the data-driven variations in the median structural response and standard deviation and the distributions of residuals with the intensity level. With reference to the Gaussian kernels approach, it is found that while the linear regression assumptions may not affect the fragility functions of lower damage states, this conclusion does not hold for the higher damage states (such as the Complete state). In addition, the effects of linear regression assumptions on the seismic risk are evaluated. For predicting the demand hazard, it is found that the linear regression assumptions can impact the computed risk for larger structural response values. However, for predicting the loss hazard with downtime as the decision variable, linear regression can be considered adequate for all practical purposes.

58 GEOSCIENCES↗

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING↗

Multi-Kernel Adaptive Support Vector Machine for Scalable Predictive Maintenance

Application of data-driven solutions across an industry is challenging, since the data are often stored locally, and increasing privacy and security concerns restrict access to the data. In addition, it is highly unlikely that all potential data patterns are captured in a single data source. Because it is highly unlikely that all potential data patterns are captured in a single data source, machine learning (ML) models developed from a single source cannot be robust enough. An alternative is to train the ML model at each source and develop a distributed knowledge discovery and aggregation approach to build global knowledge. In this paper, we develop and demonstrate a distributed ML model, federated transfer learning (FTL), using a multi-kernel-based adaptive support vector machine (MK-A-SVM). For federated learning (FL), the multi-kernel (MK) approach enables feature-specific model aggregation under data heterogeneity; whereas for transfer learning (TL) the adaptive model enables utilization of an aggregated model from a different task. The proposed approach is validated using nuclear power plant (NPP) vertical motor-driven pump data to predict the health condition of vertical motor-driven pumps as an anomaly detection. The efficiency of the proposed approach is also quantified and compared with neural network.

42 ENGINEERING↗

Neural Network-Enhanced Reproducing Kernel Particle Method for Image-Based Multiphysics Damage Modeling of Energy Storage Materials

Energy storage materials undergo significant stresses during charge/discharge cycling, which makes understanding their reliability and durability fundamental in predicting performance and service life. Strong electrochemical-mechanical coupling and highly anisotropic material properties contribute to the formation and propagation of micro-cracking, largely along material interfaces and grain boundaries. With microstructural images supplied by the National Renewable Energy Laboratory (NREL), image-based modeling techniques are used to represent the complex material microstructures that dictate the coupled physics of these systems. Traditional electrochemical-mechanical models rely on mesh-based finite element methods, which can lead to difficulties in capturing crack propagation due to mesh dependency. Additionally, commonly used damage models, such as the continuous damage model and the cohesive zone model, often have steep tradeoffs between discontinuous field accuracy and computational expense. In this work, a neural network-enhanced reproducing kernel particle method (NN-RKPM) [1] is leveraged to accurately capture damage and crack propagation throughout the material by learning the location, orientation, and sharpness of discontinuity while allowing for a coarser nodal distribution than that necessary for capturing sharp solution transitions using traditional mesh-based methods. NN-RKPM is used to inform how crack opening and closure in turn affect the coupled chemical equations and material microstructure. Reference: [1] Baek, J., Chen, J. S., Susuki, K., "Neural Network enhanced Reproducing Kernel Particle Method for Modeling Localizations," International Journal for Numerical Methods in Engineering, Vol. 123, pp 4422-4454, https://doi.org/10.1002/nme.7040, 2022.

damage modeling↗

A note on thermal history kernel for unsteady heat transfer of a spherical particle

When a particle is subjected to an unsteady ambient flow, in terms of either time-dependent relative velocity or time-dependent temperature difference, the net heat transfer from the particle cannot be calculated based on the quasi-steady heat transfer correlation alone. Due to unsteady evolution of the thermal boundary layer, there is also a history contribution to heat transfer. The history contribution to heat transfer is expressed as a convolution integral of past evolution of temperature difference between the particle and the surrounding. While Basset history force and its finite Reynolds number extension have been well studied, similar understanding of unsteady heat transfer and thermal history kernel is lacking. Here, we use existing particle-resolved simulation results to develop a finite Peclet number thermal history kernel, which when used with the convolution integral is demonstrated to accurately predict unsteady heat transfer over a range of Peclet numbers and particle-to-fluid heat capacity ratio.

42 ENGINEERING↗

Method for measurement of TRISO kernel and layer volumes by X-ray computed tomography

Layer dimensions are key parameters for as-fabricated tristructural-isotropic (TRISO) particle fuel as well as for post-irradiation examination of particle performance. Layer thicknesses are typically measured by optical microscopy of the particle cross section near mid-plane, while layer volumes are estimated with serial sectioning and microscopy. This method for measuring layer volumes is limited due to the resolution limit imposed by slice thickness and the effect of the polishing process on delicate irradiated particle microstructures. In this study, image processing software has been developed to segregate the three-dimensional (3D) TRISO particle images provided by x-ray computed tomography (XCT) into high-resolution volumetric data for the kernel, all particle layers, and any internal voids introduced during irradiation. These data can be used to analyze key in-reactor behaviors of TRISO particles such as kernel swelling or buffer shrinkage, which are important inputs for fuel performance modeling.

36 MATERIALS SCIENCE↗

Postirradiation examination from separate effects irradiation testing of uranium nitride kernels and coated particles

An overview of postirradiation examination results for uranium nitride kernels and uranium nitride coated particles irradiated in the High Flux Isotope Reactor are presented. This is the first postirradiation examination of the MiniFuel irradiation vehicle that was recently developed to rapidly accumulate burnup during separate effects irradiation testing. In general, the burnup and fuel temperatures measured postirradiation were consistent with the design calculations. The burnup measured by mass spectrometry ranged from 5.9 to 10 MWd/kgU and was achieved after only 68 effective full-power days of irradiation. The dilatometric evaluation of passive silicon carbide thermometry indicated that the fuel was irradiated at temperatures ranging from 410 to 460 °C. Because the irradiation temperatures and burnup were low, the UN kernels showed minimal fission gas release that was within the range of the expected recoil (athermal) release. While it is possible to measure fuel swelling using x-ray computed tomography, the observed swelling was too small to quantify in this case. Extensive microstructural characterization of the irradiated fuel was performed in this study, and no significant irradiation induced changes were observed.

36 MATERIALS SCIENCE↗

Raman spectroscopy of uranium nitride kernels

Uranium nitride is an advanced fuel candidate for a wide variety of advanced nuclear reactors. This work summarizes the first characterization of UN kernels by Raman spectroscopy. First-principles density functional theory calculations were performed to predict the Raman spectra of uranium sesquinitride (U 2 N 3 ), uranium dinitride (UN 2 ), uranium mononitride (UN), uranium monocarbide (UC), as well as U-N-C (UN 1-x C x ) and a U-N-C-O mixture. Further, a core–shell structure was identified by scanning electron microscopy and Raman spectroscopy imaging. A signal at ~500 cm -1 was identified on the periphery of the core-shell structure, possibly corresponding to U 2 N 3 and/or UN 2 . This signal broadens and shifts to 470 cm -1 because of the formation of UNC, UNCO or U 2 N 3+x structures. The culmination of this work demonstrates the feasibility of using Raman spectroscopy to identify variations in composition and phases in UN kernels.

36 MATERIALS SCIENCE↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗

An Efficient Bayesian Approach to Learning Droplet Collision Kernels: Proof of Concept Using “Cloudy,” a New n -Moment Bulk Microphysics Scheme

The small-scale microphysical processes governing the formation of precipitation particles cannot be resolved explicitly by cloud resolving and climate models. Instead, they are represented by microphysics schemes that are based on a combination of theoretical knowledge, statistical assumptions, and fitting to data (“tuning”). Historically, tuning was done in an ad hoc fashion, leading to parameter choices that are not explainable or repeatable. Recent work has treated it as an inverse problem that can be solved by Bayesian inference. The posterior distribution of the parameters given the data—the solution of Bayesian inference—is found through computationally expensive sampling methods, which require over $\mathcal{O}$(10 5 ) evaluations of the forward model; this is prohibitive for many models. We present a proof of concept of Bayesian learning applied to a new bulk microphysics scheme named “Cloudy,” using the recently developed Calibrate-Emulate-Sample (CES) algorithm. Cloudy models collision-coalescence and collisional breakup of cloud droplets with an adjustable number of prognostic moments and with easily modifiable assumptions for the cloud droplet mass distribution and the collision kernel. The CES algorithm uses machine learning tools to accelerate Bayesian inference by reducing the number of forward evaluations needed to $\mathcal{O}$(10 2 ). It also exhibits a smoothing effect when forward evaluations are polluted by noise. In a suite of perfect-model experiments, we show that CES enables computationally efficient Bayesian inference of parameters in Cloudy from noisy observations of moments of the droplet mass distribution. In an additional imperfect-model experiment, a collision kernel parameter is successfully learned from output generated by a Lagrangian particle-based microphysics model.

54 ENVIRONMENTAL SCIENCES↗

Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels

Next generation High-Energy Physics (HEP) experiments are presented with significant computational challenges, both in terms of data volume and processing power. Using compute accelerators, such as GPUs, is one of the promising ways to provide the necessary computational power to meet the challenge. The current programming models for compute accelerators often involve using architecture-specific programming languages promoted by the hardware vendors and hence limit the set of platforms that the code can run on. Developing software with platform restrictions is especially unfeasible for HEP communities as it takes significant effort to convert typical HEP algorithms into ones that are efficient for compute accelerators. Multiple performance portability solutions have recently emerged and provide an alternative path for using compute accelerators, which allow the code to be executed on hardware from different vendors. We apply several portability solutions, such as Kokkos, SYCL, C++17 std::execution::par, Alpaka, and OpenMP/OpenACC, on two mini-apps extracted from the mkFit project: p2z and p2r. These apps include basic kernels for a Kalman filter track fit, such as propagation and update of track parameters, for detectors at a fixed z or fixed r position, respectively. The two mini-apps explore different memory layout formats. We report on the development experience with different portability solutions, as well as their performance on GPUs and many-core CPUs, measured as the throughput of the kernels from different GPU and CPU vendors such as NVIDIA, AMD and Intel.

Kwok, Ka Hei Martin↗

Discovering causal structure with reproducing-kernel Hilbert space ε -machines

We merge computational mechanics’ definition of causal states (predictively equivalent histories) with reproducing-kernel Hilbert space (RKHS) representation inference. The result is a widely applicable method that infers causal structure directly from observations of a system’s behaviors whether they are over discrete or continuous events or time. A structural representation—a finite- or infinite-state kernel ϵ-machine—is extracted by a reduced-dimension transform that gives an efficient representation of causal states and their topology. In this way, the system dynamics are represented by a stochastic (ordinary or partial) differential equation that acts on causal states. We introduce an algorithm to estimate the associated evolution operator. Paralleling the Fokker–Planck equation, it efficiently evolves causal-state distributions and makes predictions in the original data space via an RKHS functional mapping. We demonstrate these techniques, together with their predictive abilities, on discrete-time, discrete-value infinite Markov-order processes generated by finite-state hidden Markov models with (i) finite or (ii) uncountably infinite causal states and (iii) continuous-time, continuous-value processes generated by thermally driven chaotic flows. The method robustly estimates causal structure in the presence of varying external and measurement noise levels and for very high-dimensional data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hybrid programming-model strategies for GPU offloading of electronic structure calculation kernels

To address the challenge of performance portability and facilitate the implementation of electronic structure solvers, we developed the basic matrix library (BML) and Parallel, Rapid O(N), and Graph-based Recursive Electronic Structure Solver (PROGRESS) library. The BML implements linear algebra operations necessary for electronic structure kernels using a unified user interface for various matrix formats (dense and sparse) and architectures (CPUs and GPUs). Focusing on density functional theory and tight-binding models, PROGRESS implements several solvers for computing the single-particle density matrix and relies on BML. In this paper, we describe the general strategies used for these implementations on various computer architectures, using OpenMP target functionalities on GPUs, in conjunction with third-party libraries to handle performance critical numerical kernels. In this study, we demonstrate the portability of this approach and its performance in benchmark problems.

36 MATERIALS SCIENCE↗

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗