Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Geo Thermal Cloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The project is motivated by the challenges, risks, and costs associated with geothermal exploration and production. Many processes and parameters impacting geothermal conditions are poorly understood. Diverse datasets are available to help characterize subsurface geothermal conditions (public and proprietary; satellite, airborne surveys, vegetation/water sampling, geological, geophysical, etc.). Yet, it is not clear how to properly leverage these datasets for geothermal exploration due to an incomplete understanding of how physical processes impacting subsurface geothermal conditions are represented in these observations. Recent advancements in machine learning (ML) provide great promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. Our goals and work under Phases 1 and 2 (as proposed) of this project address all these needs.

15 GEOTHERMAL ENERGY↗

A High Performance Sparse Tensor Algebra Compiler in MLIR

Sparse tensor algebra is widely used in many applications, including scientific computing, machine learning, and data analytics. The performance of sparse tensor algebra kernels strongly depends on the intrinsic characteristics of the input tensors, hence many storage formats are designed for tensors to achieve optimal performance for particular applications/architectures, which makes it challenging to implement and optimize every tensor operation of interest on a given architecture. We propose a tensor algebra domain-specific language (DSL) and compiler framework to automatically generate kernels for mixed sparse-dense tensor algebra operations. The proposed DSL provides high-level programming abstractions that resemble the familiar Einstein notation to represent tensor algebra operations. The compiler introduces a new Sparse Tensor Algebra dialect built on top of LLVM's extensible MLIR compiler infrastructure for efficient code generation while covering a wide range of tensor storage formats. Our compiler also leverages input-dependent code optimization to enhance data locality for better performance. Our results show that the performance of automatically generated kernels outperforms the state-of-the-art sparse tensor algebra compiler, with up to 20.92x, 6.39x, and 13.9x performance improvement over state-of-the-art tensor algebra compilers, for parallel SpMV, SpMM, and TTM, respectively.

Tian, Ruiqin↗

Leveraging Prior Concept Learning Improves Generalization From Few Examples in Computational Models of Human Object Recognition

Humans quickly and accurately learn new visual concepts from sparse data, sometimes just a single example. The impressive performance of artificial neural networks which hierarchically pool afferents across scales and positions suggests that the hierarchical organization of the human visual system is critical to its accuracy. These approaches, however, require magnitudes of order more examples than human learners. We used a benchmark deep learning model to show that the hierarchy can also be leveraged to vastly improve the speed of learning. We specifically show how previously learned but broadly tuned conceptual representations can be used to learn visual concepts from as few as two positive examples; reusing visual representations from earlier in the visual hierarchy, as in prior approaches, requires significantly more examples to perform comparably. These results suggest techniques for learning even more efficiently and provide a biologically plausible way to learn new visual concepts from few examples.

Rule, Joshua S.↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Sparse Symmetric Format for Tucker Decomposition

Tensor-based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making them an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTM c ) in general sparse tensors, the $\underline{S}$ parse $\underline{S}$ ymmetric $\underline{T}$ ensor $\underline{T}$ imes $\underline{S}$ ame $\underline{M}$ atrix $\underline{c}$ hain (S 3 TTM c ) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. We present the novel Compressed Sparse Symmetric (CSS) format for sparse symmetric tensors, along with an efficient parallel algorithm for the S 3 TTM c operation. We theoretically establish that S 3 TTM c on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations, and visualize the variation of the performance gap over the parameter space. We demonstrate experimental findings that confirm these results and achieve up to 2.72× speedup on synthetic and real datasets. The scaling of the algorithm on different test architectures is also showcased to highlight the effect of machine characteristics on algorithm performance.

42 ENGINEERING↗

Efficient Parallel Sparse Symmetric Tucker Decomposition for High-Order Tensors

Tensor based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making it an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTMc) in general sparse tensors, the Sparse Symmetric Tensor Times Same Matrix chain (S3TTMc) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. In this work, we present a novel compressed storage format CSS for sparse symmetric tensors, along with an efficient parallel algorithm for the S3TTMc operation. We theoretically establish that S3TTMc on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations. We demonstrate experimental findings that confirm these results and achieve up to 2.9× speedup on synthetic and real datasets.

Shivakumar, Shruti↗

Exploration of Domain Aware Machine Learning for Grid Analytics: Transfer-Learnt Energy Models to Assist Buildings Control with Sparse Field Data

Buildings are a primary consumer of energy in the United States and are also increasingly being perceived as providers of grid services such as load shifting, shedding and modulation. High fidelity models of building energy consumption are needed to set appropriate baselines for measurement and verification (M&V) of controllers designed for energy efficient operation of buildings and to enable buildings to provide grid services via. participation in demand response programs. State-of-the-art building energy modeling techniques either rely on Physics based models, or extensive instrumentation of the building envelope to gather “big” data to train machine learning based models such as deep neural networks. While Physics based models are often limited by their accuracy, it is not always feasible to gather a significant amount of field data required to train machine learning based models with sufficient accuracy. In this paper, we explore the use of transfer learning-based strategies to address unsatisfactory accuracy of models for estimating building energy consumption when available field data for training is sparse or of unacceptable quality. In particular, we transfer knowledge in the form of data and parameters, from Physics based simulation frameworks to the field to improve the model accuracy, thus resulting in a Physics-informed Machine Learning framework. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that the proposed transfer learning based models provide comparative (and in some cases better) accuracy than state-of-the-art machine learning and deep learning solutions, with just one month of field data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Accelerating matrix-centric graph processing on GPUs through bit-level optimizations

Even though it is well known that binary values are common in graph applications (e.g., adjacency matrix), how to leverage the phenomenon for efficiency has not yet been adequately explored. This paper presents a systematic study on how to unlock the potential of the bit-level optimizations of graph computations that involve binary values. It proposes a two-level representation named Bit-Block Compressed Sparse Row (B2SR) and presents a series of optimizations to the graph operations on B2SR by the intrinsics of modern GPUs. It additionally introduces Deep Reinforcement Learning (DRL) as an efficient way to best configure the bit-level optimizations on the fly. Additionally, the DQN-based adaptive tile size selector with dedicated model training can reach 68% prediction accuracy. Evaluations on NVIDIA Pascal and Volta GPUs show that the optimizations bring up to 40× and 6555× for essential GraphBLAS kernels SpMV and SpGEMM, respectively, making GraphBLAS-based BFS accelerate up to 433×, SSSP, PR, and CC up to 35×, and TC up to 52×.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic locked mode predictor in the presence of a resistive wall and finite island saturation in tokamaks

We present a framework for estimating the probability of locking to an error field in a rotating tokamak plasma. This leverages machine learning methods trained on data from a mode-locking model, including an error field, resistive magnetohydrodynamics modeling of the plasma, a resistive wall, and an external vacuum region, leading to a fifth-order ordinary differential equation (ODE) system. It is an extension of the model without a resistive wall introduced by Akçay et al. [Phys. Plasmas 28, 082106 (2021)]. Tearing mode saturation by a finite island width is also modeled. We vary three pairs of control parameters in our studies: the momentum source plus either the error field, the tearing stability index, or the island saturation term. The order parameters are the time-asymptotic values of the five ODE variables. Normalization of them reduces the system to 2D and facilitates the classification into locked (L) or unlocked (U) states, as illustrated by Akçay et al., [Phys. Plasmas 28, 082106 (2021)]. This classification splits the control space into three regions: L̂, with only L states; Û, with only U states; and a hysteresis (hysteretic) region Ĥ, with both L and U states. In regions L̂ and Û, the cubic equation of torque balance yields one real root. Region Ĥ has three roots, allowing bifurcations between the L and U states. The classification of the ODE solutions into L/U is used to estimate the locking probability, conditional on the pair of the control parameters, using a neural network. We also explore estimating the locking probability for a sparse dataset, using a transfer learning method based on a dense model dataset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Collaborative Research: Enabling multi-scale studies of magnetic reconnection with interpretable data-driven models

The development of accurate reduced descriptions and improved closures for magnetic reconnection is an important and a long‐standing challenge in plasma physics. The four‐fluid approach, and associated closures, that were investigated have the potential to improve the accuracy of plasma fluid models, capturing physical effects which would otherwise require a kinetic description. If successful, this approach could have an important impact for the modeling of laboratory and space plasmas. The major goals of this project were to develop new machine learning (ML) tools based on sparse and symbolic regression techniques, and to extract interpretable and generalizable reduced models (e.g., in the form of partial differential equations - PDEs) from data generated by first principles plasma simulations. Preserving interpretability of such data‐driven models is key to addressing the long‐standing theoretical and numerical challenges. Prior proof‐of‐principle studies have demonstrated the enormous potential of this approach, by recovering the well‐established hierarchy of plasma equations (from Vlasov to MHD) from data produced by particle‐in‐cell (PIC) simulations. Our goal in this project was to extend and apply these new tools to construct better kinetic closures for magnetic reconnection; to derive better models of particle injection and acceleration by this fundamental plasma process; and to use this understanding to accelerate the development of multi‐scale plasma algorithms. While our immediate focus was on the problem of magnetic reconnection, the tools that were will developed are general and applicable to other areas of plasma physics, and more broadly to many‐body phenomena. We anticipate that the development of these multi‐scale models will have a significant impact across different areas of plasma science, from fusion to space and astrophysical plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

EXPRES. II. Searching for Planets around Active Stars: A Case Study of HD 101501

By controlling instrumental errors to below 10 cm s{sup −1}, the EXtreme PREcision Spectrograph (EXPRES) allows for a more insightful study of photospheric velocities that can mask weak Keplerian signals. Gaussian processes (GP) have become a standard tool for modeling correlated noise in radial velocity data sets. While GPs are constrained and motivated by physical properties of the star, in some cases they are still flexible enough to absorb unresolved Keplerian signals. We apply GP regression to EXPRES radial velocity measurements of the 3.5 Gyr old chromospherically active Sun-like star, HD 101501. We obtain tight constraints on the stellar rotation period and the evolution of spot distributions using 28 seasons of ground-based photometry, as well as recent Transiting Exoplanet Survey Satellite data. Light-curve inversion was carried out on both photometry data sets to reveal the spot distribution and spot evolution timescales on the star. We find that the >5 m s{sup −1} rms radial velocity variations in HD 101501 are well modeled with a GP stellar activity model without planets, yielding a residual rms scatter of 45 cm s{sup −1}. We carry out simulations, injecting and recovering signals with the GP framework, to demonstrate that high-cadence observations are required to use GPs most efficiently to detect low-mass planets around active stars like HD 101501. Sparse sampling prevents GPs from learning the correlated noise structure and can allow it to absorb prospective Keplerian signals. We quantify the moderate to high-cadence monitoring that provides the necessary information to disentangle photospheric features using GPs and to detect planets around active stars.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A dictionary learning algorithm for compression and reconstruction of streaming data in preset order

There has been an emerging interest in developing and applying dictionary learning (DL) to process massive datasets in the last decade. Many of these efforts, however, focus on employing DL to compress and extract a set of important features from data, while considering restoring the original data from this set a secondary goal. On the other hand, although several methods are able to process streaming data by updating the dictionary incrementally as new snapshots pass by, most of those algorithms are designed for the setting where the snapshots are randomly drawn from a probability distribution. In this paper, we present a new DL approach to compress and denoise massive dataset in real time, in which the data are streamed through in a preset order (instances are videos and temporal experimental data), so at any time, we can only observe a biased sample set of the whole data. Here, our approach incrementally builds up the dictionary in a relatively simple manner: if the new snapshot is adequately explained by the current dictionary, we perform a sparse coding to find its sparse representation; otherwise, we add the new snapshot to the dictionary, with a Gram-Schmidt process to maintain the orthogonality. To compress and denoise noisy datasets, we apply the denoising to the snapshot directly before sparse coding, which deviates from traditional dictionary learning approach that achieves denoising via sparse coding. Compared to full-batch matrix decomposition methods, where the whole data is kept in memory, and other mini-batch approaches, where unbiased sampling is often assumed, our approach has minimal requirement in data sampling and storage: i) each snapshot is only seen once then discarded, and ii) the snapshots are drawn in a preset order, so can be highly biased. Through experiments on climate simulations and scanning transmission electron microscopy (STEM) data, we demonstrate that the proposed approach performs competitively to those methods in data reconstruction and denoising.

97 MATHEMATICS AND COMPUTING↗

Dictionary Learning with Accumulator Neurons

The Locally Competitive Algorithm (LCA) uses local competition between non-spiking leaky integrator neurons to infer sparse representations, allowing for potentially real-time execution on massively parallel neuromorphic architectures such as Intel's Loihi processor. Here, we focus on the problem of inferring sparse representations from streaming video using dictionaries of spatiotemporal features optimized in an unsupervised manner for sparse reconstruction. Non-spiking LCA has previously been used to achieve unsupervised learning of spatiotemporal dictionaries composed of convolutional kernels from raw, unlabeled video. We demonstrate how unsupervised dictionary learning with spiking LCA (\hbox{S-LCA}) can be efficiently implemented using accumulator neurons, which combine a conventional leaky-integrate-and-fire (\hbox{LIF}) spike generator with an additional state variable that is used to minimize the difference between the integrated input and the spiking output. We demonstrate dictionary learning across a wide range of dynamical regimes, from graded to intermittent spiking, for inferring sparse representations of both static images drawn from the CIFAR database as well as video frames captured from a DVS camera. On a classification task that requires identification of the suite from a deck of cards being rapidly flipped through as viewed by a DVS camera, we find essentially no degradation in performance as the LCA model used to infer sparse spatiotemporal representations migrates from graded to spiking. We conclude that accumulator neurons are likely to provide a powerful enabling component of future neuromorphic hardware for implementing online unsupervised learning of spatiotemporal dictionaries optimized for sparse reconstruction of streaming video from event based DVS cameras.

artificial intelligence↗

Deep structural clustering for single-cell RNA-seq data jointly through autoencoder and graph neural network

Abstract Single-cell RNA sequencing (scRNA-seq) permits researchers to study the complex mechanisms of cell heterogeneity and diversity. Unsupervised clustering is of central importance for the analysis of the scRNA-seq data, as it can be used to identify putative cell types. However, due to noise impacts, high dimensionality and pervasive dropout events, clustering analysis of scRNA-seq data remains a computational challenge. Here, we propose a new deep structural clustering method for scRNA-seq data, named scDSC, which integrate the structural information into deep clustering of single cells. The proposed scDSC consists of a Zero-Inflated Negative Binomial (ZINB) model-based autoencoder, a graph neural network (GNN) module and a mutual-supervised module. To learn the data representation from the sparse and zero-inflated scRNA-seq data, we add a ZINB model to the basic autoencoder. The GNN module is introduced to capture the structural information among cells. By joining the ZINB-based autoencoder with the GNN module, the model transfers the data representation learned by autoencoder to the corresponding GNN layer. Furthermore, we adopt a mutual supervised strategy to unify these two different deep neural architectures and to guide the clustering task. Extensive experimental results on six real scRNA-seq datasets demonstrate that scDSC outperforms state-of-the-art methods in terms of clustering accuracy and scalability. Our method scDSC is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/scDSC.

Gan, Yanglan↗

Physics-informed neural networks for identification of material properties using standing waves

A metallic structure in its initial stage of failure involves plastic deformation or environmental degradation that changes the elastic modulus and density. This work presents the detection of change in wave velocity (a function of elastic modulus and density) as a system identification problem. A physics-informed neural network (PINN) is proposed to solve the system identification problem. The PINN takes the spatial coordinates of scanning locations and time as inputs and provides the displacement and wave velocity as outputs. The governing partial differential equation of standing waves in a rod is incorporated into the neural network as physics in the form of a loss function. The wave velocity vector is randomly initiated. During the training of the network, physics is used to determine and update the wave velocity target vector from the network’s displacement predictions. The measured data, comprising sparse displacement response on the rod structure, are used to train the PINN. The wave velocity at the sparse locations on the rod is learned from the predicted displacements during the training. Using the predictions of the trained network, the response of free vibration or material property variation can be reconstructed at unscanned locations on the structure to obtain high-resolution maps for full-field imaging to detect and localize the changes caused by plastic deformation. The PINN’s sparse scanning and simultaneous prediction capability during training can lead to high scanning and data-processing speeds. This capability yields a nondestructive evaluation system that can predict the presence of degraded material locations as the structural vibrations are scanned and processed in real time.

Rathod, Vivek↗

An Automated Scanning Transmission Electron Microscope Guided by Sparse Data Analytics

Abstract Artificial intelligence (AI) promises to reshape scientific inquiry and enable breakthrough discoveries in areas such as energy storage, quantum computing, and biomedicine. Scanning transmission electron microscopy (STEM), a cornerstone of the study of chemical and materials systems, stands to benefit greatly from AI-driven automation. However, present barriers to low-level instrument control, as well as generalizable and interpretable feature detection, make truly automated microscopy impractical. Here, we discuss the design of a closed-loop instrument control platform guided by emerging sparse data analytics. We hypothesize that a centralized controller, informed by machine learning combining limited a priori knowledge and task-based discrimination, could drive on-the-fly experimental decision-making. This platform may unlock practical, automated analysis of a variety of material features, enabling new high-throughput and statistical studies.

47 OTHER INSTRUMENTATION↗