Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Kernel learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

On the Feasibility of Using Reduced-Precision Tensor Core Operations for Graph Analytics

Today’s data-driven analytics and machine learning workload have been largely driven by the General-PurposeGraphics Processing Units (GPGPUs). To accelerate dense matrix multiplications on the GPUs, Tensor Core Units (TCUs) have been introduced in recent years. In this paper, we study linear-algebra-based and vertex-centric algorithms for various graph kernels on the GPUs with an objective of applying this new hardware feature to graph applications. We identify the potential stages in these graph kernels that can be executed on the Tensor Core Units. In particular, we leverage the reformulation of the reduction and scan operations in terms of matrix multiplication [1]on the TCUs. We demonstrate that executing these operations on the TCUs, available inside different graph kernels, can assist in establishing an end-to-end pipeline on the GPGPUs without depending on hand-tuned external libraries and still can deliver comparable performance for various graph analytics.

Graph algorithms, GPU computing↗

Kernel-based global sensitivity analysis obtained from a single data set

Results from global sensitivity analysis (GSA) often guide the understanding of complicated input–output systems. Kernel-based GSA methods have recently been proposed for their capability of treating a broad scope of complex systems. In this paper, we develop a new set of kernel GSA tools when only a single set of input–output data is available. Three key advances are made: (1) A new numerical estimator is proposed that demonstrates an empirical improvement over previous procedures. (2) A computational method for generating inner statistical functions from a single data set is presented. (3) A theoretical extension is made to define conditional sensitivity indices, which reveal the degree that the inputs carry shared information about the output when inherent input–input correlations are present. Utilizing these conditional sensitivity indices, a decomposition is derived for the output uncertainty based on what is called the optimal learning sequence of the input variables, which remains consistent when correlations exist between the input variables. Further, while these advances cover a range of GSA subjects, a common single data set numerical solution is provided by a technique known as the conditional mean embedding of distributions. The new methodology is implemented on benchmark systems to demonstrate the provided insights.

42 ENGINEERING↗

Comparative evaluation of deep learning workloads for leadership-class systems

Deep learning (DL) workloads and their performance at scale are becoming important factors to consider as we design, develop and deploy next-generation high-performance computing systems. Since DL applications rely heavily on DL frameworks and underlying compute (CPU/GPU) stacks, it is essential to gain a holistic understanding from compute kernels, models, and frameworks of popular DL stacks, and to assess their impact on science-driven, mission-critical applications. At Oak Ridge Leadership Computing Facility (OLCF), we employ a set of micro and macro DL benchmarks established through the Collaboration of Oak Ridge, Argonne, and Livermore (CORAL) to evaluate the AI readiness of our next-generation supercomputers. In this paper, we present our early observations and performance benchmark comparisons between the Nvidia V100 based Summit system with its CUDA stack and an AMD MI100 based testbed system with its ROCm stack. We take a layered perspective on DL benchmarking and point to opportunities for future optimizations in the technologies that we consider.

Yin, Junqi↗

Learning the structure of wind: A data-driven nonlocal turbulence model for the atmospheric boundary layer

In this work, we develop a novel data-driven approach to modeling the atmospheric boundary layer. This approach leads to a nonlocal, anisotropic synthetic turbulence model which we refer to as the deep rapid distortion (DRD) model. Our approach relies on an operator regression problem that characterizes the best fitting candidate in a general family of nonlocal covariance kernels parameterized in part by a neural network. This family of covariance kernels is expressed in Fourier space and is obtained from approximate solutions to the Navier–Stokes equations at very high Reynolds numbers. Each member of the family incorporates important physical properties such as mass conservation and a realistic energy cascade. The DRD model can be calibrated with noisy data from field experiments. After calibration, the model can be used to generate synthetic turbulent velocity fields. To this end, we provide a new numerical method based on domain decomposition which delivers scalable, memory-efficient turbulence generation with the DRD model as well as others. We demonstrate the robustness of our approach with both filtered and noisy data coming from the 1968 Air Force Cambridge Research Laboratory Kansas experiments. Using these data, we witness exceptional accuracy with the DRD model, especially when compared to the International Electrotechnical Commission standard.

17 WIND ENERGY↗

An end-to-end deep learning method for solving nonlocal Allen–Cahn and Cahn–Hilliard phase-field models

Here, we propose an efficient end-to-end deep learning method for solving nonlocal Allen–Cahn (AC) and Cahn–Hilliard (CH) phase-field models. One motivation for this effort emanates from the fact that discretized partial differential equation-based AC or CH phase-field models result in diffuse interfaces between phases, with the only recourse for remediation is to severely refine the spatial grids in the vicinity of the true moving sharp interface whose width is determined by a grid-independent parameter that is substantially larger than the local grid size. In this work, we introduce non-mass conserving nonlocal AC or CH phase-field models with regular, logarithmic, or obstacle double-well potentials. Because of non-locality, some of these models feature totally sharp interfaces separating phases. The discretization of such models can lead to a transition between phases whose width is only a single grid cell wide. Another motivation is to use deep learning approaches to ameliorate the otherwise high cost of solving discretized nonlocal phase-field models. To this end, loss functions of the customized neural networks are defined using the residual of the fully discrete approximations of the AC or CH models, which results from applying a Fourier collocation method and a temporal semi-implicit approximation. To address the long-range interactions in the models, we tailor the architecture of the neural network by incorporating a nonlocal kernel as an input channel to the neural network model. We then provide the results of extensive computational experiments to illustrate the accuracy, predictive capabilities, and cost reductions of the proposed method.

42 ENGINEERING↗

Liquid–Vapor Phase Equilibrium in Molten Aluminum Chloride (AlCl 3 ) Enabled by Machine Learning Interatomic Potentials

Molten salts are promising candidates in numerous clean energy applications, where knowledge of thermophysical properties and vapor pressure across their operating temperature ranges is critical for safe operations. Due to challenges in evaluating these properties using experimental methods, fast and scalable molecular simulations are essential to complement the experimental data. In this study, we developed machine learning interatomic potentials (MLIP) to study the AlCl 3 molten salt across varied thermodynamic conditions (T = 473–613 K and P = 2.7–23.4 bar), which allowed us to predict temperature-surface tension correlations and liquid–vapor phase diagram from direct simulations of two-phase coexistence in this molten salt. Two MLIP architectures, a Kernel-based potential and neural network interatomic potential (NNIP), were considered to benchmark their performance for AlCl 3 molten salt using experimental structure and density values. The NNIP potential employed in two-phase equilibrium simulations yields the critical temperature and critical density of AlCl 3 that are within 10 K (∼3%) and 0.03 g/cm 3 (∼7%) of the reported experimental values. An accurate correlation between temperature and viscosities is obtained as well. In doing so, we report that the inclusion of low-density configurations in their training is critical to more accurately represent the AlCl 3 system across a wide phase-space. The MLIP trained using PBE-D3 functional in the ab initio molecular dynamics (AIMD) simulations (120 atoms) also showed close agreement with experimentally determined molten salt structure comprising Al 2 Cl 6 dimers, as validated using Raman spectra and neutron structure factor. Furthermore, the PBE-D3 as well as its trained MLIP showed better liquid density and temperature correlation for AlCl 3 system when compared to several other density functionals explored in this work. Overall, the demonstrated approach to predict temperature correlations for liquid and vapor densities in this study can be employed to screen nuclear reactors-relevant compositions, helping to mitigate safety concerns.

Ab initio molecular dynamics↗

Active learning of reactive Bayesian force fields applied to heterogeneous catalysis dynamics of H/Pt

Abstract Atomistic modeling of chemically reactive systems has so far relied on either expensive ab initio methods or bond-order force fields requiring arduous parametrization. Here, we describe a Bayesian active learning framework for autonomous “on-the-fly” training of fast and accurate reactive many-body force fields during molecular dynamics simulations. At each time-step, predictive uncertainties of a sparse Gaussian process are evaluated to automatically determine whether additional ab initio training data are needed. We introduce a general method for mapping trained kernel models onto equivalent polynomial models whose prediction cost is much lower and independent of the training set size. As a demonstration, we perform direct two-phase simulations of heterogeneous H 2 turnover on the Pt(111) catalyst surface at chemical accuracy. The model trains itself in three days and performs at twice the speed of a ReaxFF model, while maintaining much higher fidelity to DFT and excellent agreement with experiment.

42 ENGINEERING↗

Comparative Study on the Machine Learning-Based Prediction of Adsorption Energies for Ring and Chain Species on Metal Catalyst Surfaces

Computation of adsorption and transition state energies for a large number of surface intermediates for numerous active site models pose significant computational overhead in computational screening of catalysts. Machine learning (ML) techniques can be used to predict part of these energies. To predict the energies, ML models need to be fed appropriate metal and species descriptors. For complex surface chemistries, the structures of the intermediate species can vary greatly. In this paper, working with the hydrodeoxygenation of succinic acid on six different metal surfaces, we have studied the effect of linear and non-linear ML models used along with pen-and-paper based species descriptors and two categories of metal descriptors on two different categories of intermediate species: chain and ring. More specifically, our computations include the prediction of chain species when trained on only chain species and also when trained on both chain and ring species. Similar computations were performed for predictions of ring species. In each case, results of linear ML models were compared with kernel based non-linear models. Our results indicate that ring species data does not improve the prediction of chain species. Similarly, chain species data does not improve the prediction of ring species. The use of non-linear ML models, however, did help to minimize the prediction errors compared to the linear models. Furthermore, the study also shows that electronic or adsorption energy based metal descriptors along with bond count based species fingerprints can achieve a mean absolute error (MAE) of less than 0.2 eV for complex chain molecules when used with an appropriate machine learning model.

Adsorption↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fiber optic computing using distributed feedback

Abstract The widespread adoption of machine learning and other matrix intensive computing algorithms has renewed interest in analog optical computing, which has the potential to perform large-scale matrix multiplications with superior energy scaling and lower latency than digital electronics. However, most optical techniques rely on spatial multiplexing, requiring a large number of modulators and detectors, and are typically restricted to performing a single kernel convolution operation per layer. Here, we introduce a fiber-optic computing architecture based on temporal multiplexing and distributed feedback that performs multiple convolutions on the input data in a single layer. Using Rayleigh backscattering in standard single mode fiber, we show that this technique can efficiently apply a series of random nonlinear projections to the input data, facilitating a variety of computing tasks. The approach enables efficient energy scaling with orders of magnitude lower power consumption than GPUs, while maintaining low latency and high data-throughput.

97 MATHEMATICS AND COMPUTING↗

Classification Using Support Vector Machines with Uncertainty Quantification

Binary classification using machine learning is needed to address engineering problems such as identifying passing/failing parts based on measured features from aging hardware. In these classifications, providing the uncertainty of each prediction is essential to support engineering decision making. One popular classifier is the support vector machine (SVM). There are many variations, with the simplest being a linear division between two classes with a hyperplane. Kernel methods can be implement

Taylor, Sofia Nitsche↗

Physics-Informed Machine Learning Models for Predicting the Progress of Reactive-Mixing

This paper presents a physics-informed machine learning (ML) framework to construct reduced-order models (ROMs) for reactive-transport quantities of interest (QoIs) based on high-fidelity numerical simu-lations. QoIs include species decay, product yield, and degree of mixing. The ROMs for QoIs are applied to quantify and understand how the chemical species evolve over time. First, high-resolution datasets for constructing ROMs are generated by solving anisotropic reaction-di?usion equations using a non-negative finite element formulation for di?erent input parameters. The reactive-mixing model input parameters are: time-scale associated with flipping of velocity, spatial-scale controlling small/large vortex structures of velocity, perturbation parameter of the vortex-based velocity, anisotropic dispersion strength/contrast, and molecular diffusion. Second, random forests, F-test, and mutual information criterion are used to evaluate the importance of model inputs/features with respect to QoIs. We observed that anisotropic dispersion strength/contrast is the most important feature and time-scale associated with flipping of velocity is the least important feature. Third, Support Vector Machines (SVM) and Support Vector Regression (SVR) are used to construct ROMs based on the model inputs. The constructed SVR-ROMs are then used to predict scaling of QoIs. We also present estimates and inequalities on the QoIs, which inform that the species decay, mix, and produce in an exponential fashion. These inequalities also inform that a radial basis function is the most suitable kernel for the SVM/SVR models for QoIs. It is observed that R2-score for SVR-ROMs on unseen data is greater than 0.9, implying that the SVR-ROMs are able to predict the reaction-diffusion system state reasonably well. Finally, in terms of the computational cost, the proposed SVM-ROMs are O(107) times faster than running a high-fidelity finite element simulation for evaluating QoIs. This makes the proposed ML-based ROMs attractive for reactive-transport sensing and real-time monitoring applications as they are significantly faster yet reasonably accurate.

Mudunuru, Maruti K.↗

Machine learning force field model for kinetic Monte Carlo simulations of itinerant Ising magnets

Here, we present a scalable machine learning (ML) framework for large-scale kinetic Monte Carlo (kMC) simulations of itinerant electron Ising systems. As the effective interactions between Ising spins in such itinerant magnets are mediated by conducting electrons, the calculation of energy change due to a local spin update requires solving an electronic structure problem. Such repeated electronic structure calculations could be overwhelmingly prohibitive for large systems. Assuming the locality principle, a convolutional neural network (CNN) model is developed to directly predict the effective local field and the corresponding energy change associated with a given spin update based on Ising configuration in a finite neighborhood. As the kernel size of the CNN is fixed at a constant, the model can be directly scalable to kMC simulations of large lattices. Our approach is reminiscent of the ML force field models widely used in first-principles molecular dynamics simulations. Applying our ML framework to a square-lattice double-exchange Ising model, we uncover unusual coarsening of ferromagnetic domains at low temperatures. Our work highlights the potential of ML methods for large-scale modeling of similar itinerant systems with discrete dynamical variables.

machine learning↗

Target Detection via Cognitive Radars Using Change-Point Detection, Learning, and Adaptation

Many radar detection algorithms that assume a stationary environment (clutter) have been proposed and analyzed over the years. However, in practice, changes in the nonstationary environment can perturb the parameters of the clutter distribution, or even alter the clutter distribution family, which can greatly deteriorate the target detection capability. To avoid such potential performance degradation, cognitive radar systems are envisioned which are required to rapidly realize the nonstationarity, accurately learn the new characteristics of the environments, and adaptively update the detector. In this paper, aiming to develop a fully cognitive radar for target detection in nonstationary environments, we propose a unifying framework that integrates (i) change-point detection of clutter distributions by using a data-driven cumulative sum (CUSUM) algorithm and its extended version, (ii) learning/identification of clutter distribution by applying sparse theory and kernel density estimation methods, and (iii) adaptive target detection by automatically modifying the likelihood-ratio test and corresponding detection threshold. Further, with extensive numerical examples, we demonstrate the achieved improvements in detection performance due to the proposed framework in comparison to a nonadaptive case, an adaptive matched filter (AMF) method, and the clairvoyant case. Herein, we also use Wilcoxon rank-sum tests to evaluate the statistical significance of the performance improvements

42 ENGINEERING↗

Chemical mixture exposure patterns and obesity among U.S. adults in NHANES 2005–2012

The effect of chemical exposure on obesity has raised great concerns. Real-world chemical exposure always imposes mixture impacts, however their exposure patterns and the corresponding associations with obesity have not been fully evaluated. To discover obesity-related mixed chemical exposure patterns in the general U.S. population. Sparse Decompositional Regression (SDR), a model adapted from sparse representation learning technique, was developed to identify exposure patterns of chemical mixtures with exclusion (non-targeted model) and inclusion (targeted model) of health outcomes. We assessed the relationships between the identified chemical mixture patterns and obesity-related indexes. We also conducted a comprehensive evaluation of this SDR model by comparing to the existing models, including generalized linear regression model (GLM), principal component analysis (PCA), and Bayesian kernel machine regression (BKMR). Eight core exposure patterns were identified using the non-targeted SDR model. Patterns of high levels of MEP, high levels of naphthalene metabolites (ΣOH-Nap), and a pattern of high exposure levels of MCOP, MCNP, and MCPP were positively associated with obesity. Patterns of high levels of BP3, and a pattern of higher mixed levels of MPB, PPB, and MEP were found to have negative associations. Associations were strengthened using the targeted SDR model. In the single chemical analysis by GLM, BP3, MBP, PPB, MCOP, and MCNP showed significant associations with obesity or body indexes. The SDR model exceeded the performance of PCA in pattern identification. Both SDR and BKMR identified a positive contribution of ΣOH-Nap and MCOP, as well as a negative contribution of BP3 and PPB to obesity. Our study identified five core exposure patterns of chemical mixtures significantly associated with obesity using the newly developed SDR model. The SDR model could open a new avenue for assessing health effects of environmental mixture contaminants.

54 ENVIRONMENTAL SCIENCES↗

ECP Report: Update on Proxy Applications and Vendor Interactions

The ExaLearn miniGAN team (Ellis and Rajamanickam) have released miniGAN, a generative adversarial network(GAN) proxy application, through the ECP proxy application suite. miniGAN is the first machine learning proxy application in the suite (note: the ECP CANDLE project did previously release some benchmarks) and models the performance for training generator and discriminator networks. The GAN's generator and discriminator generate plausible 2D/3D maps and identify fake maps, respectively. miniGAN aims to be a proxy application for related applications in cosmology (CosmoFlow, ExaGAN) and wind energy (ExaWind). miniGAN has been developed so that optimized mathematical kernels (e.g., kernels provided by Kokkos Kernels) can be plugged into to the proxy application to explore potential performance improvements. miniGAN has been released as open source software and is available through the ECP proxy application website (https://proxyapps.exascaleproject.ordecp-proxy-appssuite/) and on GitHub (https://github.com/SandiaMLMiniApps/miniGAN). As part of this release, a generator is provided to generate a data set (series of images) that are inputs to the proxy application.

97 MATHEMATICS AND COMPUTING↗

Machine learning Frenkel Hamiltonian parameters to accelerate simulations of exciton dynamics

In this manuscript, we develop multiple machine learning (ML) models to accelerate a scheme for parameterizing site-based models of exciton dynamics from all-atom configurations of condensed phase sexithiophene systems. This scheme encodes the details of a system’s specific molecular morphology in the correlated distributions of model parameters through the analysis of many single-molecule excited-state electronic-structure calculations. These calculations yield excitation energies for each molecule in the system and the network of pair-wise intermolecular electronic couplings. In this work, we demonstrate that the excitation energies can be accurately predicted using a kernel ridge regression (KRR) model with Coulomb matrix featurization. We present two ML models for predicting intermolecular couplings. The first one utilizes a deep neural network and bi-molecular featurization to predict the coupling directly, which we find to perform poorly. The second one utilizes a KRR model to predict unimolecular transition densities, which can subsequently be analyzed to compute the coupling. We find that the latter approach performs excellently, indicating that an effective, generalizable strategy for predicting simple bimolecular properties is through the indirect application of ML to predict higher-order unimolecular properties. Such an approach necessitates a much smaller feature space and can incorporate the insight of well-established molecular physics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Co-optimized machine-learned manifold models for large eddy simulation of turbulent combustion

Many modeling approaches in large eddy simulation (LES) of turbulent combustion employ a projection of the thermochemical state onto a low-dimensional manifold within state space to reduce the number of transported variables and hence computational cost. Flamelet-generated manifolds (FGM) is an example of a well-established, physics-based approach, but increasingly, principal component analysis (PCA) is being used as a data-driven method for generating manifold models. For both approaches, the nonlinear relationship between the location on the predefined manifold and the outputs of interest, such as reaction rates, can be tabulated or encoded in a neural network. This work proposes a new approach for manifold modeling that extends these existing approaches. A modified neural network structure simultaneously encodes the definition of the manifold variables, the nonlinear mapping, and the subfilter closure for LES. This allows all three of these aspects of the model to be co-optimized, generating a model from any source of combustion thermochemical state data. The manifold parameterizing variables are constrained to be linear combinations of species, as in FGM and PCA-based models, to aid in interpretability and implementation. For LES, subfilter variances of the manifold variables are also included as inputs. Two types of a priori analysis are performed to evaluate the new approach. In the first, the model is trained on data from one-dimensional premixed flames. In this case, the approach recovers the behavior of flamelet-based manifold approaches, and in fact slightly improves performance by identifying an optimized progress variable. The approach is also applied to data from direct numerical simulations of spherical ignition kernels in isotropic turbulence. For any specified manifold dimensionality, the new approach provides substantially lower prediction errors than a PCA-based model developed from the same data set. Additionally, the LES formulation of the new approach can provide accurate predictions for filtered reaction rates across a variety of filter widths.

97 MATHEMATICS AND COMPUTING↗