Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

General-Purpose Bayesian Tensor Learning With Automatic Rank Determination and Uncertainty Quantification

A major challenge in many machine learning tasks is that the model expressive power depends on model size. Low-rank tensor methods are an efficient tool for handling the curse of dimensionality in many large-scale machine learning models. The major challenges in training a tensor learning model include how to process the high-volume data, how to determine the tensor rank automatically, and how to estimate the uncertainty of the results. While existing tensor learning focuses on a specific task, this paper proposes a generic Bayesian framework that can be employed to solve a broad class of tensor learning problems such as tensor completion, tensor regression, and tensorized neural networks. We develop a low-rank tensor prior for automatic rank determination in nonlinear problems. Our method is implemented with both stochastic gradient Hamiltonian Monte Carlo (SGHMC) and Stein Variational Gradient Descent (SVGD). We compare the automatic rank determination and uncertainty quantification of these two solvers. We demonstrate that our proposed method can determine the tensor rank automatically and can quantify the uncertainty of the obtained results. We validate our framework on tensor completion tasks and tensorized neural network training tasks.

Bayesian inference↗

MIONet: Learning Multiple-Input Operators via Tensor Product

As an emerging paradigm in scientific machine learning, neural operators aim to learn operators, via neural networks, that map between infinite-dimensional function spaces. Several neural operators have been recently developed. However, all the existing neural operators are only designed to learn operators defined on a single Banach space; i.e., the input of the operator is a single function. Here, for the first time, we study the operator regression via neural networks for multiple-input operators defined on the product of Banach spaces. We first prove a universal approximation theorem of continuous multiple-input operators. We also provide a detailed theoretical analysis including the approximation error, which provides guidance for the design of the network architecture. Based on our theory and a low-rank approximation, we propose a novel neural operator, MIONet, to learn multiple-input operators. MIONet consists of several branch nets for encoding the input functions and a trunk net for encoding the domain of the output function. Here, we demonstrate that MIONet can learn solution operators involving systems governed by ordinary and partial differential equations. In our computational examples, we also show that we can endow MIONet with prior knowledge of the underlying system, such as linearity and periodicity, to further improve accuracy.

97 MATHEMATICS AND COMPUTING↗

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

A Robust Event Diagnostics Platform: Integrating Tensor Analytics and Machine Learning into Real-time Grid Monitoring

The objective of this project is to develop a robust event diagnostics (RED) platform by integrating state-of-the-art tensor analytics and machine learning into real-time grid monitoring. The proposed platform can effectively analyze and discover the information hiding within the provided PMU data for effective real-time grid monitoring. The proposed RED platform provides a set of robust diagnostics tools for grid operation and management, including 1) data quality assessment, 2) data completion, 3) event detection, and 4) robust event classification. All the functionalities of the RED platform can help the operator to make informed decisions and respond in a timely manner. The developed RED platform will serve as an innovative advisory tool to reliably identify key events and discover new insights about the events and grid characteristics in the PMU data, and contribute to the efficient, safe, reliable operation and design of the nation’s electric system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Elliptically-Contoured Tensor-variate Distributions with Application to Image Learning

Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate for data arising from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) TV distributions and derive its characterizations, moments, marginal, and conditional distributions. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, for which we extend Tyler’s robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop TV classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and TV analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age, and ethnic origin than the usual TVN-based TANOVA in the celebrated labeled faces of the wild dataset.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Full NMR Chemical Shift Tensors of Silicon Oxides with Equivariant Graph Neural Networks

The nuclear magnetic resonance (NMR) chemical shift tensor is a highly sensitive probe of the electronic structure of an atom and furthermore its local structure. Recently, machine learning has been applied to NMR in the prediction of isotropic chemical shifts from a structure. Current machine learning models, however, often ignore the full chemical shift tensor for the easier-to-predict isotropic chemical shift, effectively ignoring a multitude of structural information available in the NMR chemical shift tensor. Here we use an equivariant graph neural network (GNN) to predict full 29 Si chemical shift tensors in silicate materials. The equivariant GNN model predicts full tensors to a mean absolute error of 1.05 ppm and is able to accurately determine the magnitude, anisotropy, and tensor orientation in a diverse set of silicon oxide local structures. When compared with other models, the equivariant GNN model outperforms the state-of-the-art machine learning models by 53%. The equivariant GNN model also outperforms historic analytical models by 57% for isotropic chemical shift and 91% for anisotropy. The software is available as a simple-to-use open-source repository, allowing similar models to be created and trained with ease.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY↗

Reverse-mode differentiation in arbitrary tensor network format: with application to supervised learning.

This paper describes an efficient reverse-mode differentiation algorithm for contraction operations of tensor networks that may have arbitrary and unconventional network topologies. The approach leverages the tensor contraction tree of Evenbly and Pfeifer (2014), which provides an instruction set for the contraction sequence of a network. We show that this tree can be efficiently leveraged for differentiation of a full tensor network contraction using a recursive scheme that exploits (1) the bilinear property of contraction and (2) the property that trees have single path from root to leaves. While differentiation of tensor-tensor contraction is already possible in most automatic differentiation packages, we show that exploiting these two additional properties in the specific context of contraction sequences can improve efficiency. Following a description of the algorithm and computational complexity analysis, we investigate its utility for gradient-based supervised learning for low-rank function recovery and for fitting real-world unstructured datasets. We demonstrate improved performance over alternating least-squares optimization approaches and the capability to handle heterogeneous and arbitrary tensor network formats. When compared to alternating minimization algorithms, we find that the gradient-based approach requires a smaller oversampling ratio (number of samples compared to number model parameters) for recovery. This increased efficiency extends to fitting unstructured data of varying dimensionality and when employing a variety of tensor network formats. Here, we show improved learning using the hierarchical Tucker method over the tensor-train in high-dimensional settings on a number of benchmark problems.

97 MATHEMATICS AND COMPUTING↗

Bayesian Tensor Decompositions for Scalable Supervised Learning of Scientific Data (Final Report)

In this document we highlight the detailed accomplishments and progress that we have made in this period. This progress seeks to address the three main objectives to provide new algorithms for quantifying uncertainty in low-multilinear-rank models and to leverage them for data analysis. These include: (1) develop probabilistic models for low-multilinear-rank functions; (2) develop a suite of Bayesian learning approaches to learn the probabilistic models from data; (3) apply the techniques on challenging problems arising in DOE-relevant applications.

97 MATHEMATICS AND COMPUTING↗

Atomic-scale origin of the low grain-boundary resistance in perovskite solid electrolyte Li 0.375 Sr 0.4375 Ta 0.75 Zr 0.25 O 3

Oxide solid electrolytes (OSEs) have the potential to achieve improved safety and energy density for lithium-ion batteries, but their high grain-boundary (GB) resistance generally is a bottleneck. In the well-studied perovskite oxide solid electrolyte, Li 3x La 2/3-x TiO 3 (LLTO), the ionic conductivity of grain boundaries is about three orders of magnitude lower than that of the bulk. In contrast, the related Li 0.375 Sr 0.4375 Ta 0.75 Zr 0.25 O 3 (LSTZ0.75) perovskite exhibits low grain boundary resistance for reasons yet unknown. Here, we use aberration-corrected scanning transmission electron microscopy and spectroscopy, along with an active learning moment tensor potential, to reveal the atomic scale structure and composition of LSTZ0.75 grain boundaries. Vibrational electron energy loss spectroscopy is applied for the first time to reveal atomically resolved vibrations at grain boundaries of LSTZ0.75 and to characterize the otherwise unmeasurable Li distribution therein. We find that Li depletion, which is a major reason for the low grain boundary ionic conductivity of LLTO, is absent for the grain boundaries of LSTZ0.75. Instead, the low grain boundary resistivity of LSTZ0.75 is attributed to the formation of a nanoscale defective cubic perovskite interfacial structure that contained abundant vacancies. Our study provides new insights into the atomic scale mechanisms of low grain boundary resistivity.

25 ENERGY STORAGE↗

A General Spatiotemporal Imputation Framework for Missing Sensor Data

Many applications from precision agriculture, environmental monitoring and transportation networks rely on data collected across space and time over a large geographic area. Missing data poses a significant challenge for any data-driven inference and control tasks. Data imputation or the estimation of missing data can help fill these gaps by utilizing inherent spatial relationships and temporal patterns. A variety of spatiotemporal imputation models have been developed to address missing data in spatiotemporal datasets. However, these classical methods rely on the assumption that the underlying data follows a smooth trend and fail to provide accurate estimates when there is a large number of missing points in the data. Even though there are machine learning driven tensor completion approaches such as convolutional neural network based tensor completion (CoSTCo) that capture the non-linear relationships in the dataset, the transductive nature makes the algorithm less scalable. Thus, existing approaches for estimating the missing information do not effectively capture all dimensions of the spatiotemporal data structure, resulting in erroneous predictions and poor performance. The main contributions of this paper are: (1) We propose a novel inductive framework (G-LSTM) for missing data imputation that integrates a graph neural network with LSTMs to effectively capture both spatial and temporal dependencies. (2) Experimental results on a traffic dataset demonstrate that the proposed GNN integrated with an LSTM framework achieves improved imputation and maintains steady performance even when there are extreme missing conditions in comparison with the state-of-the-art imputation framework (i.e, CoSTCo). (3) The simulation results on a traffic network show up to 69% reduction in mean absolute error and 61% reduction in root mean square error when compared to CoSTCo.

Tharzeen, Aabila↗

Deep compressed seismic learning for fast location and moment tensor inferences with natural and induced seismicity

Fast detection and characterization of seismic sources is crucial for decision-making and warning systems that monitor natural and induced seismicity. However, besides the laying out of ever denser monitoring networks of seismic instruments, the incorporation of new sensor technologies such as Distributed Acoustic Sensing (DAS) further challenges our processing capabilities to deliver short turnaround answers from seismic monitoring. In response, this work describes a methodology for the learning of the seismological parameters: location and moment tensor from compressed seismic records. In this method, data dimensionality is reduced by applying a general encoding protocol derived from the principles of compressive sensing. The data in compressed form is then fed directly to a convolutional neural network that outputs fast predictions of the seismic source parameters. Thus, the proposed methodology can not only expedite data transmission from the field to the processing center, but also remove the decompression overhead that would be required for the application of traditional processing methods. An autoencoder is also explored as an equivalent alternative to perform the same job. We observe that the CS-based compression requires only a fraction of the computing power, time, data and expertise required to design and train an autoencoder to perform the same task. Implementation of the CS-method with a continuous flow of data together with generalization of the principles to other applications such as classification are also discussed.

54 ENVIRONMENTAL SCIENCES↗

Machine learning of 27Al NMR electric field gradient tensors for crystalline structures from DFT

NMR crystallography has emerged as a promising technique for the determination and refinement of atomic coordinates in crystal structures. The crystal structure of compounds containing quadrupolar nuclei, such as 27Al, can be improved by directly comparing solid-state NMR measurements to DFT computations of the electric field gradient (EFG) tensor. The non-negligible computational cost of these first-principles calculations limits the applicability of this method to all but the most well-defined structures. We developed a fast, low-cost machine learning model to predict EFG parameters based on local structural motifs and elemental parameters. We computed 8081 EFG tensors from 1681 27Al crystalline solids using DFT and benchmarked them against 105 experimentally measured 27Al sites. Surprisingly, simple local geometric features dominate the predictive performance of the resulting random-forest model, yielding an R2 value of 0.98 and an RMSE of 0.61 MHz for CQ, the quadrupolar coupling constant. This model accuracy should enable pre-refining future structural assignments before finally validating with first-principles calculations. Such a catalogue of 27Al NMR tensors can serve as a tool for researchers assigning complex NMR spectra influenced by the nuclear electric quadrupole interaction.

Sun, He↗

Physics-informed machine learning of the Lagrangian dynamics of velocity gradient tensor

Reduced models describing the Lagrangian dynamics of the velocity gradient tensor (VGT) in homogeneous isotropic turbulence (HIT) are developed under the physics-informed machine learning (PIML) framework. We consider the VGT at both Kolmogorov scale and coarse-grained scale within the inertial range of HIT. Building reduced models requires resolving the pressure Hessian and subfilter contributions, which is accomplished by constructing them using the integrity bases and invariants of the VGT. The developed models can be expressed using the extended tensor basis neural network (TBNN) introduced by Ling et al. [J. Fluid Mech. 807, 155 (2016)]. Physical constraints, such as Galilean invariance, rotational invariance, and incompressibility condition, are thus embedded in the models explicitly. Our PIML models are trained on the Lagrangian data from a high-Reynolds number direct numerical simulation (DNS). To validate the results, we perform a comprehensive out-of-sample test. We observe that the PIML model provides an improved representation for the magnitude and orientation of the small-scale pressure Hessian contributions. Statistics of the flow, as indicated by the joint PDF of second and third invariants of the VGT, show good agreement with the “ground-truth” DNS data. A number of other important features describing the structure of HIT are reproduced by the model successfully. We have also identified challenges in modeling inertial range dynamics, which indicates that a richer modeling strategy is required. This helps us identify important directions for future research, in particular towards including inertial range geometry into the TBNN.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗