Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

SW4 Curvilinear Kernels

Five computationally expensive stencil evaluation routines from SW4(https://github.com/geodynamics/sw4 GPL license) have been extracted and packaged with a driver to create a mini-app for evaluating compiler and GPU performance. The kernels can executed on AMD and Nvidia GPUs with and without RAJA. The kernel driver generates synthetic inputs, checks for correctness and measures kernel run times.

Pankajakshan, Ramesh↗

ExaSGD: 2022 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-407 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY22 the main objective of the Kernel Thrust was (i) provide sparse optimization solver that runs efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations, (ii) strengthen the reliability and increase the performance of the mixed-dense sparse (MDS) solver of HiOp for deployment on the FY22 target architectures, Summit and Crusher, and (iii) increase performance by improving the mathematical algorithm and refining the parallel MPI-based implementation of the coarse-grain parallel solver HiOp-PriDec for capabilities deployment on the FY22 target architectures, Summit and Crusher. This document presents the developments and contributions done by the Kernels Thrust Team in FY22 toward completion of the above-mentioned objectives. These contributions progressed along four main development (sub)thrusts: (1) Design and implementation of a sparse optimization solver for use on hardware accelerators; (2) Improvement of the mathematical algorithm and of the parallel implementation of HiOp-PriDec to ensure readiness and efficient coarse-grain parallelism for FY23 target exascale machine; and (3) Support Software and Application Development Thrusts of the exaSGD project in their deployment of the project’s software stack on AMD- and NVIDIA-based architectures. The development of the sparse optimization solver (thrust 1 above) was new in FY22 and resulted in a new sparse solver in HiOp (available as of version 0.6). The second development thrust was a continuation of the efforts from FY21 and improved the mathematical algorithm and the communication strategy of the HiOp-PriDec solver. The last developement thrust is a large collaborative effort. Namely, the project’s teams from multiple labs (LLNL, PNNL, ORNL, and NREL) performed large-scale demonstration of the ExaSGD software stack, namely the optimization solvers of HiOp interfaced with the modeling front-end ExaGO and the stochastic sampler PowerScenarios. These demonstration efforts solved large-scale instances of the SC-ACOPF challenge problem of medium network sizes (10, 000-bus system) and large number of contingencies on Summit (NVIDIA accelerators) and Crusher (AMD accelerators) systems at ORNL.

97 MATHEMATICS AND COMPUTING↗

The Collins-Soper Kernel from Lattice QCD

I will present the first complete determination of the quark Collins-Soper kernel, which relates TMDs at different rapidity scales, using lattice QCD and including systematic control of quark mass, operator mixing, and discretization effects. Next-to-next-to-leading logarithmic matching is used to match lattice-calculable distributions to the corresponding TMDs. The continuum-extrapolated lattice QCD results are consistent with several recent phenomenological parametrizations of the Collins-Soper kernel and are precise enough to disfavor other parametrizations. I will also discuss a first exploration of the gluon Collins-Soper kernel.

Wagman, Michael [Fermilab]↗

Cholesky-based experimental design for Gaussian process and kernel-based emulation and calibration.

Gaussian processes and other kernel-based methods are used extensively to construct approximations of multivariate data sets. The accuracy of these approximations is dependent on the data used. This paper presents a computationally efficient algorithm to greedily select training samples that minimize the weighted L p error of kernel-based approximations for a given number of data. The method successively generates nested samples, with the goal of minimizing the error in high probability regions of densities specified by users. The algorithm presented is extremely simple and can be implemented using existing pivoted Cholesky factorization methods. Training samples are generated in batches which allows training data to be evaluated (labeled) in parallel. For smooth kernels, the algorithm performs comparably with the greedy integrated variance design but has significantly lower complexity. Numerical experiments demonstrate the efficacy of the approach for bounded, unbounded, multi-modal and non-tensor product densities. We also show how to use the proposed algorithm to efficiently generate surrogates for inferring unknown model parameters from data using Bayesian inference.

97 MATHEMATICS AND COMPUTING↗

Evaluating HPC Kernels for Processing in Memory

Memory subsystems contribute significantly to the performance and energy efficiency of high-performance computing (HPC) applications. Traditional memory technologies with conventional organization (e.g., DRAM) are struggling to keep up with the increasing memory requirements of modern applications. Techniques such as multilayer cache hierarchy and out-of-order execution are still falling short of mitigating the penalty incurred by memory accesses. Processing-in-memory (PIM), which involves moving memory-intensive kernels to memory for execution instead of bringing the data to the processing unit, is emerging as a promising technique. PIM has recently received traction among computer architecture researchers, and the increasing research activity surrounding this technique indicates its potential to alleviate main memory performance bottlenecks. In this paper, we characterize and identify memory-intensive HPC kernels, perform a first-order evaluation of the PIM technique for selected HPC kernels, quantify performance deviation, and analyze the key factors that affect PIM efficiency.

Asifuzzaman, Kazi↗

Kernel-based global sensitivity analysis obtained from a single data set

Results from global sensitivity analysis (GSA) often guide the understanding of complicated input–output systems. Kernel-based GSA methods have recently been proposed for their capability of treating a broad scope of complex systems. In this paper, we develop a new set of kernel GSA tools when only a single set of input–output data is available. Three key advances are made: (1) A new numerical estimator is proposed that demonstrates an empirical improvement over previous procedures. (2) A computational method for generating inner statistical functions from a single data set is presented. (3) A theoretical extension is made to define conditional sensitivity indices, which reveal the degree that the inputs carry shared information about the output when inherent input–input correlations are present. Utilizing these conditional sensitivity indices, a decomposition is derived for the output uncertainty based on what is called the optimal learning sequence of the input variables, which remains consistent when correlations exist between the input variables. Further, while these advances cover a range of GSA subjects, a common single data set numerical solution is provided by a technique known as the conditional mean embedding of distributions. The new methodology is implemented on benchmark systems to demonstrate the provided insights.

42 ENGINEERING↗

Digital Modeling on Large Kernel Metamaterial Neural Network

Deep neural networks (DNNs) utilized recently are physically deployed with computational units (e.g., CPUs and GPUs). Such a design might lead to a heavy computational burden, significant latency, and intensive power consumption, which are critical limitations in applications such as Internet of Things (IoT), edge computing, and usage of drones. Recent advances in optical computational units (e.g., metamaterial) have shed light on energy-free and light-speed neural networks. However, the digital design of the metamaterial neural network (MNN) is fundamentally limited by its physical limitations, such as precision, noise, and bandwidth during fabrication. Moreover, the unique advantages of MNN’s (e.g., light-speed computation) are not fully explored via standard 3×3 convolution kernels. In this paper, we propose a novel large kernel metamaterial neural network (LMNN) that maximizes the digital capacity of the state-of-the-art (SOTA) MNN with model re-parametrization and network compression, while also considering the optical limitation explicitly. The new digital learning scheme can maximize the learning capacity of MNN while modeling the physical restrictions of meta-optics. With the proposed LMNN, the computation cost of the convolutional front-end can be offloaded to fabricated optical hardware. The experimental results on two publicly available datasets demonstrate that the optimized hybrid design improved classification accuracy while reducing computational latency. In conclusion, the development of the proposed LMNN is a promising step towards the ultimate goal of energy-free and light-speed AI.

97 MATHEMATICS AND COMPUTING↗

Gaussian Kernel Methods for Seismic Fragility and Risk Assessment of Mid-Rise Buildings

Seismic fragility functions can be evaluated using the cloud analysis method with linear regression which makes three fundamental assumptions about the relation between structural response and seismic intensity: log-linear median relationship, constant standard deviation, and Gaussian distributed errors. While cloud analysis with linear regression is a popular method, the degree to which these individual and compounded assumptions affect the fragility and the risk of mid-rise buildings needs to be systematically studied. This paper conducts such a study considering three building archetypes that make up a bulk of the building stock: RC moment frame, steel moment frame, and wood shear wall. Gaussian kernel methods are employed to capture the data-driven variations in the median structural response and standard deviation and the distributions of residuals with the intensity level. With reference to the Gaussian kernels approach, it is found that while the linear regression assumptions may not affect the fragility functions of lower damage states, this conclusion does not hold for the higher damage states (such as the Complete state). In addition, the effects of linear regression assumptions on the seismic risk are evaluated. For predicting the demand hazard, it is found that the linear regression assumptions can impact the computed risk for larger structural response values. However, for predicting the loss hazard with downtime as the decision variable, linear regression can be considered adequate for all practical purposes.

58 GEOSCIENCES↗

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING↗

Multi-Kernel Adaptive Support Vector Machine for Scalable Predictive Maintenance

Application of data-driven solutions across an industry is challenging, since the data are often stored locally, and increasing privacy and security concerns restrict access to the data. In addition, it is highly unlikely that all potential data patterns are captured in a single data source. Because it is highly unlikely that all potential data patterns are captured in a single data source, machine learning (ML) models developed from a single source cannot be robust enough. An alternative is to train the ML model at each source and develop a distributed knowledge discovery and aggregation approach to build global knowledge. In this paper, we develop and demonstrate a distributed ML model, federated transfer learning (FTL), using a multi-kernel-based adaptive support vector machine (MK-A-SVM). For federated learning (FL), the multi-kernel (MK) approach enables feature-specific model aggregation under data heterogeneity; whereas for transfer learning (TL) the adaptive model enables utilization of an aggregated model from a different task. The proposed approach is validated using nuclear power plant (NPP) vertical motor-driven pump data to predict the health condition of vertical motor-driven pumps as an anomaly detection. The efficiency of the proposed approach is also quantified and compared with neural network.

42 ENGINEERING↗

Neural Network-Enhanced Reproducing Kernel Particle Method for Image-Based Multiphysics Damage Modeling of Energy Storage Materials

Energy storage materials undergo significant stresses during charge/discharge cycling, which makes understanding their reliability and durability fundamental in predicting performance and service life. Strong electrochemical-mechanical coupling and highly anisotropic material properties contribute to the formation and propagation of micro-cracking, largely along material interfaces and grain boundaries. With microstructural images supplied by the National Renewable Energy Laboratory (NREL), image-based modeling techniques are used to represent the complex material microstructures that dictate the coupled physics of these systems. Traditional electrochemical-mechanical models rely on mesh-based finite element methods, which can lead to difficulties in capturing crack propagation due to mesh dependency. Additionally, commonly used damage models, such as the continuous damage model and the cohesive zone model, often have steep tradeoffs between discontinuous field accuracy and computational expense. In this work, a neural network-enhanced reproducing kernel particle method (NN-RKPM) [1] is leveraged to accurately capture damage and crack propagation throughout the material by learning the location, orientation, and sharpness of discontinuity while allowing for a coarser nodal distribution than that necessary for capturing sharp solution transitions using traditional mesh-based methods. NN-RKPM is used to inform how crack opening and closure in turn affect the coupled chemical equations and material microstructure. Reference: [1] Baek, J., Chen, J. S., Susuki, K., "Neural Network enhanced Reproducing Kernel Particle Method for Modeling Localizations," International Journal for Numerical Methods in Engineering, Vol. 123, pp 4422-4454, https://doi.org/10.1002/nme.7040, 2022.

damage modeling↗

A note on thermal history kernel for unsteady heat transfer of a spherical particle

When a particle is subjected to an unsteady ambient flow, in terms of either time-dependent relative velocity or time-dependent temperature difference, the net heat transfer from the particle cannot be calculated based on the quasi-steady heat transfer correlation alone. Due to unsteady evolution of the thermal boundary layer, there is also a history contribution to heat transfer. The history contribution to heat transfer is expressed as a convolution integral of past evolution of temperature difference between the particle and the surrounding. While Basset history force and its finite Reynolds number extension have been well studied, similar understanding of unsteady heat transfer and thermal history kernel is lacking. Here, we use existing particle-resolved simulation results to develop a finite Peclet number thermal history kernel, which when used with the convolution integral is demonstrated to accurately predict unsteady heat transfer over a range of Peclet numbers and particle-to-fluid heat capacity ratio.

42 ENGINEERING↗

Method for measurement of TRISO kernel and layer volumes by X-ray computed tomography

Layer dimensions are key parameters for as-fabricated tristructural-isotropic (TRISO) particle fuel as well as for post-irradiation examination of particle performance. Layer thicknesses are typically measured by optical microscopy of the particle cross section near mid-plane, while layer volumes are estimated with serial sectioning and microscopy. This method for measuring layer volumes is limited due to the resolution limit imposed by slice thickness and the effect of the polishing process on delicate irradiated particle microstructures. In this study, image processing software has been developed to segregate the three-dimensional (3D) TRISO particle images provided by x-ray computed tomography (XCT) into high-resolution volumetric data for the kernel, all particle layers, and any internal voids introduced during irradiation. These data can be used to analyze key in-reactor behaviors of TRISO particles such as kernel swelling or buffer shrinkage, which are important inputs for fuel performance modeling.

36 MATERIALS SCIENCE↗

Postirradiation examination from separate effects irradiation testing of uranium nitride kernels and coated particles

An overview of postirradiation examination results for uranium nitride kernels and uranium nitride coated particles irradiated in the High Flux Isotope Reactor are presented. This is the first postirradiation examination of the MiniFuel irradiation vehicle that was recently developed to rapidly accumulate burnup during separate effects irradiation testing. In general, the burnup and fuel temperatures measured postirradiation were consistent with the design calculations. The burnup measured by mass spectrometry ranged from 5.9 to 10 MWd/kgU and was achieved after only 68 effective full-power days of irradiation. The dilatometric evaluation of passive silicon carbide thermometry indicated that the fuel was irradiated at temperatures ranging from 410 to 460 °C. Because the irradiation temperatures and burnup were low, the UN kernels showed minimal fission gas release that was within the range of the expected recoil (athermal) release. While it is possible to measure fuel swelling using x-ray computed tomography, the observed swelling was too small to quantify in this case. Extensive microstructural characterization of the irradiated fuel was performed in this study, and no significant irradiation induced changes were observed.

36 MATERIALS SCIENCE↗

Raman spectroscopy of uranium nitride kernels

Uranium nitride is an advanced fuel candidate for a wide variety of advanced nuclear reactors. This work summarizes the first characterization of UN kernels by Raman spectroscopy. First-principles density functional theory calculations were performed to predict the Raman spectra of uranium sesquinitride (U 2 N 3 ), uranium dinitride (UN 2 ), uranium mononitride (UN), uranium monocarbide (UC), as well as U-N-C (UN 1-x C x ) and a U-N-C-O mixture. Further, a core–shell structure was identified by scanning electron microscopy and Raman spectroscopy imaging. A signal at ~500 cm -1 was identified on the periphery of the core-shell structure, possibly corresponding to U 2 N 3 and/or UN 2 . This signal broadens and shifts to 470 cm -1 because of the formation of UNC, UNCO or U 2 N 3+x structures. The culmination of this work demonstrates the feasibility of using Raman spectroscopy to identify variations in composition and phases in UN kernels.

36 MATERIALS SCIENCE↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗