Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Manifold learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Learning physics-based reduced-order models from data using nonlinear manifolds

Here we present a novel method for learning reduced-order models of dynamical systems using nonlinear manifolds. First, we learn the manifold by identifying nonlinear structure in the data through a general representation learning problem. The proposed approach is driven by embeddings of low-order polynomial form. A projection onto the nonlinear manifold reveals the algebraic structure of the reduced-space system that governs the problem of interest. The matrix operators of the reduced-order model are then inferred from the data using operator inference. Numerical experiments on a number of nonlinear problems demonstrate the generalizability of the methodology and the increase in accuracy that can be obtained over reduced-order modeling methods that employ a linear subspace approximation.

97 MATHEMATICS AND COMPUTING↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Neural embedding: learning the embedding of the manifold of physics data

In this paper, we present a method of embedding physics data manifolds with metric structure into lower dimensional spaces with simpler metrics, such as Euclidean and Hyperbolic spaces. We then demonstrate that it can be a powerful step in the data analysis pipeline for many applications. Using progressively more realistic simulated collisions at the Large Hadron Collider, we show that this embedding approach learns the underlying latent structure. With the notion of volume in Euclidean spaces, we provide for the first time a viable solution to quantifying the true search capability of model agnostic search algorithms in collider physics (i.e. anomaly detection). Finally, we discuss how the ideas presented in this paper can be employed to solve many practical challenges that require the extraction of physically meaningful representations from information in complex high dimensional datasets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data-driven models of nonautonomous systems

Nonautonomous dynamical systems are characterized by time-dependent inputs, which complicates the discovery of predictive models describing the spatiotemporal evolution of the state variables of quantities of interest from their temporal snapshots. When dynamic mode decomposition (DMD) is used to infer a linear model, this difficulty manifests itself in the need to approximate the time-dependent Koopman operators. Our approach is to approximate the original nonautonomous system with a modified system derived via a local parameterization of the time-dependent inputs. The modified system comprises a sequence of local parametric systems, which are subsequently approximated by a parametric surrogate model using the DRIPS (dimension reduction and interpolation in parameter space) framework. The offline step of DRIPS relies on DMD to build a linear surrogate model, endowed with reduced-order bases for the observables mapped from training data. The online step interpolates on suitable manifolds to construct a sequence of iterative parametric surrogate models; the target/test parameter points on these manifolds are specified by a local parameterization of the test time-dependent inputs. Here, we use numerical experimentation to demonstrate the robustness of our method and compare its performance with that of deep neural networks.

97 MATHEMATICS AND COMPUTING↗

Any Two Learning Algorithms Are (Almost) Exactly Identical

This paper shows that if one is provided with a loss function, it can be used in a natural way to specify a distance measure quantifying the similarity of any two supervised learning algorithms, even non-parametric algorithms. Intuitively, this measure gives the fraction of targets and training sets for which the expected performance of the two algorithms differs significantly. Bounds on the value of this distance are calculated for the case of binary outputs and 0-1 loss, indicating that any two learning algorithms are almost exactly identical for such scenarios. As an example, for any two algorithms A and B, even for small input spaces and training sets, for less than 2e(-50) of all targets will the difference between A's and B's generalization performance of exceed 1%. In particular, this is true if B is bagging applied to A, or boosting applied to A. These bounds can be viewed alternatively as telling us, for example, that the simple English phrase 'I expect that algorithm A will generalize from the training set with an accuracy of at least 75% on the rest of the target' conveys 20,000 bytes of information concerning the target. The paper ends by discussing some of the subtleties of extending the distance measure to give a full (non-parametric) differential geometry of the manifold of learning algorithms.

Wolpert, David H.↗

DRIPS: A framework for dimension reduction and interpolation in parameter space

Reduced-order models are often used to describe the behavior of complex systems, whose simulation with a full model is too expensive, or to extract salient features from the full model’s output. We introduce a new model-reduction framework DRIPS (dimension reduction and interpolation in parameter space) that combines the offline local model reduction with the online parameter interpolation of reduced-order bases (ROBs). The offline step of this framework relies on dynamic mode decomposition (DMD) to build a low-rank linear surrogate model, equipped with a local ROB, for quantities of interest derived from the training data generated by repeatedly solving the (nonlinear) high-fidelity model for multiple parameter points. The online step consists of the construction of a parametric reduced-order model for each target/test point in the parameter space, with the interpolation of ROBs done on a Grassman manifold and the interpolation of reduced-order operators done on a matrix manifold. The DMD component enables DRIPS to model (typically low-dimensional) quantities of interest directly, without having to access the (typically high-dimensional and possibly nonlinear) operators in a high-fidelity model that governs the dynamics of the underlying high-dimensional state variables, as required in projection-based reduced-order modeling. A series of numerical experiments suggests that DRIPS yields a model reduction, which is computationally more efficient than the commonly used projection-based proper orthogonal decomposition; it does so without requiring a prior knowledge of the governing equation for quantities of interest. Furthermore, for the nonlinear systems considered, DRIPS is more accurate than Gaussian-process interpolation (Kriging).

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Probabilistic learning and updating of a digital twin for composite material systems

This paper presents an approach for characterizing and estimating statistical dependence between a large number of observables in a composite material system. Conditional regression is carried out using the estimated joint density function, permitting a systematic exploration of interdependence between fine scale and coarse observables that can be used for both prognosis and design of complex material systems. An example demonstrates the integration of experimental data with a computational database. The statistical approach is based on the probabilistic learning on manifolds recently developed by the authors. This approach leverages intrinsic structure detected through diffusion on graphs with projected stochastic differential equations to generate samples constrained to that structure.

42 ENGINEERING↗

Grassmannian diffusion maps based surrogate modeling via geometric harmonics

Abstract A novel surrogate model based on the Grassmannian diffusion maps (GDMaps) and utilizing geometric harmonics (GH) is developed for predicting the response of complex physical phenomena. The method utilizes GDMaps to obtain a low‐dimensional representation of the underlying behavior of physical/mathematical systems with respect to uncertain input parameters. Using this representation, GH, an out‐of‐sample extension technique, is employed to create a global map from the input parameter space to a Grassmannian diffusion manifold. GH is further employed to locally map points on the diffusion manifold onto the tangent space of a Grassmann manifold. The exponential map is then used to project the points in the tangent space onto the Grassmann manifold, where reconstruction of the full solution is performed. The performance of the proposed surrogate model is verified with three examples. The first problem is a toy example used to illustrate the technique. In the second example, errors associated with the various mappings are assessed by studying response predictions of the electric potential of a dielectric cylinder in a homogeneous electric field. The last example applies the method for uncertainty prediction in the strain field evolution in a model amorphous material using the shear transformation zone theory of plasticity.

42 ENGINEERING↗

Tracking the topology of neural manifolds across populations

Neural manifolds summarize the intrinsic structure of the information encoded by a population of neurons. Advances in experimental techniques have made simultaneous recordings from multiple brain regions increasingly commonplace, raising the possibility of studying how these manifolds relate across populations. However, when the manifolds are nonlinear and possibly code for multiple unknown variables, it is challenging to extract robust and falsifiable information about their relationships. We introduce a framework, called the method of analogous cycles, for matching topological features of neural manifolds using only observed dissimilarity matrices within and between neural populations. We demonstrate via analysis of simulations and in vivo experimental data that this method can be used to correctly identify multiple shared circular coordinate systems across both stimuli and inferred neural manifolds. Conversely, the method rejects matching features that are not intrinsic to one of the systems. Further, as this method is deterministic and does not rely on dimensionality reduction or optimization methods, it is amenable to direct mathematical investigation and interpretation in terms of the underlying neural activity. We thus propose the method of analogous cycles as a suitable foundation for a theory of cross-population analysis via neural manifolds.

97 MATHEMATICS AND COMPUTING↗

Latent map Gaussian processes for mixed variable metamodeling

Gaussian processes (GPs) are ubiquitously used in sciences and engineering as metamodels. Standard GPs, however, can only handle numerical or quantitative variables. Here we introduce latent map Gaussian processes (LMGPs) that inherit the attractive properties of GPs and are also applicable to mixed data which have both quantitative and qualitative inputs. The core idea behind LMGPs is to learn a continuous, low-dimensional latent space or manifold which encodes all qualitative inputs. To learn this manifold, we first assign a unique prior vector representation to each combination of qualitative inputs. We then use a low-rank linear map to project these priors on a manifold that characterizes the posterior representations. As the posteriors are quantitative, they can be directly used in any standard correlation function such as the Gaussian or Matern. Hence, the optimal map and the corresponding manifold, along with other hyperparameters of the correlation function, can be systematically learned via maximum likelihood estimation. Through a wide range of analytic and real-world examples, we demonstrate the advantages of LMGPs over state-of-the-art methods in terms of accuracy and versatility. In particular, we show that LMGPs can handle variable-length inputs, have an explainable neural network interpretation, and provide insights into how qualitative inputs affect the response or interact with each other. We also employ LMGPs in Bayesian optimization and illustrate that they can discover optimal compound compositions more efficiently than conventional methods that convert compositions to qualitative variables via manual featurization.

42 ENGINEERING↗

Graph interpolating activation improves both natural and robust accuracies in data-efficient deep learning

Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning (DL) research. In this paper, we replace the output activation function of DNNs, typically the data-agnostic softmax function, with a graph Laplacian-based high-dimensional interpolating function which, in the continuum limit, converges to the solution of a Laplace–Beltrami equation on a high-dimensional manifold. Furthermore, we propose end-to-end training and testing algorithms for this new architecture. The proposed DNN with graph interpolating activation integrates the advantages of both deep learning and manifold learning. Compared to the conventional DNNs with the softmax function as output activation, the new framework demonstrates the following major advantages: First, it is better applicable to data-efficient learning in which we train high capacity DNNs without using a large number of training data. Second, it remarkably improves both natural accuracy on the clean images and robust accuracy on the adversarial images crafted by both white-box and black-box adversarial attacks. Third, it is a natural choice for semi-supervised learning. This paper is a significant extension of our earlier work published in NeurIPS, 2018. For reproducibility, the code is available at https://github.com/BaoWangMath/DNN-DataDependentActivation .

Mathematics↗

Reduced order modeling for flow and transport problems with Barlow Twins self-supervised learning

Abstract We propose a unified data-driven reduced order model (ROM) that bridges the performance gap between linear and nonlinear manifold approaches. Deep learning ROM (DL-ROM) using deep-convolutional autoencoders (DC–AE) has been shown to capture nonlinear solution manifolds but fails to perform adequately when linear subspace approaches such as proper orthogonal decomposition (POD) would be optimal. Besides, most DL-ROM models rely on convolutional layers, which might limit its application to only a structured mesh. The proposed framework in this study relies on the combination of an autoencoder (AE) and Barlow Twins (BT) self-supervised learning, where BT maximizes the information content of the embedding with the latent space through a joint embedding architecture. Through a series of benchmark problems of natural convection in porous media, BT–AE performs better than the previous DL-ROM framework by providing comparable results to POD-based approaches for problems where the solution lies within a linear subspace as well as DL-ROM autoencoder-based techniques where the solution lies on a nonlinear manifold; consequently, bridges the gap between linear and nonlinear reduced manifolds. We illustrate that a proficient construction of the latent space is key to achieving these results, enabling us to map these latent spaces using regression models. The proposed framework achieves a relative error of 2% on average and 12% in the worst-case scenario (i.e., the training data is small, but the parameter space is large.). We also show that our framework provides a speed-up of $$7 \times 10^{6}$$ 7 × 10 6 times, in the best case, and $$7 \times 10^{3}$$ 7 × 10 3 times on average compared to a finite element solver. Furthermore, this BT–AE framework can operate on unstructured meshes, which provides flexibility in its application to standard numerical solvers, on-site measurements, experimental data, or a combination of these sources.

97 MATHEMATICS AND COMPUTING↗

Manifold traversing as a model for learning control of autonomous robots

This paper describes a recipe for the construction of control systems that support complex machines such as multi-limbed/multi-fingered robots. The robot has to execute a task under varying environmental conditions and it has to react reasonably when previously unknown conditions are encountered. Its behavior should be learned and/or trained as opposed to being programmed. The paper describes one possible method for organizing the data that the robot has learned by various means. This framework can accept useful operator input even if it does not fully specify what to do, and can combine knowledge from autonomous, operator assisted and programmed experiences.

Szakaly, Zoltan F.↗

Operator inference for non-intrusive model reduction with quadratic manifolds

This paper proposes a novel approach for learning a data-driven quadratic manifold from high-dimensional data, then employing this quadratic manifold to derive efficient physics-based reduced-order models. The key ingredient of the approach is a polynomial mapping between high-dimensional states and a low-dimensional embedding. This mapping consists of two parts: a representation in a linear subspace (computed in this work using the proper orthogonal decomposition) and a quadratic component. The approach can be viewed as a form of data-driven closure modeling, since the quadratic component introduces directions into the approximation that lie in the orthogonal complement of the linear subspace, but without introducing any additional degrees of freedom to the low-dimensional representation. Combining the quadratic manifold approximation with the operator inference method for projection-based model reduction leads to a scalable non-intrusive approach for learning reduced-order models of dynamical systems. Further, applying the new approach to transport-dominated systems of partial differential equations illustrates the gains in efficiency that can be achieved over approximation in a linear subspace.

42 ENGINEERING↗

A Geometry-Driven Longitudal Topic Model

A simple and scalable framework for longitudinal analysis of Twitter data is developed that combines latent topic models with computational geometric methods. Dimensionality reduction tools from computational geometry are applied to learn the intrinsic manifold on which the latent, temporal topics reside. Then shortest path distances on the manifold are used to link together these topics. The proposed framework permits visualization of the low-dimensional embedding which provides clear interpretation of the complex, high-dimensional trajectories that may exist among latent topics. Practical application of the proposed framework is demonstrated through its ability to capture and effectively visualize natural progression of latent COVID-19 related topics learned from Twitter data. Interpretability of the trajectories is achieved by comparing to real-world events. In addition, the framework permits study of spatial variation in Twitter behavior for learned topics. The analysis demonstrates that the proposed framework is able to capture granular-level impact of COVID-19 on public discussions. We end by arguing that Twitter data, when analyzed within the proposed framework, can serve as a valuable supplementary data stream for COVID-related studies.

97 MATHEMATICS AND COMPUTING↗

Identifying Vehicle Signals in Continuous Seismic Data Using Unsupervised Machine-Learning Techniques

Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

GeN-ROM—An OpenFOAM®-based multiphysics reduced-order modeling framework for the analysis of Molten Salt Reactors

This work presents a projection-based multiphysics Model Order Reduction (MOR) framework for the analysis of nuclear systems and its application to parametric simulations of Molten Salt Reactors (MSR). The framework, named GeN-ROM, is developed using OpenFOAM® and employs a Proper Orthogonal Decomposition aided Reduced-Basis technique (POD-RB). It can be used to reduce steady-state and transient multiphysics problems involving parametric fluid dynamics, heat exchange, and neutronics phenomena. For the treatment of structural elements in the hydraulic systems, a porous medium approach has been adopted. The reduction process is data-driven and snapshot information is extracted via POD to learn the solution manifold and to build global spatial basis functions. At the data collection phase, GeN-ROM makes use of the solvers available in GeN-Foam, a similarly OpenFOAM®-based multiphysics framework developed for the analysis of nuclear reactors. The global bases are used both to approximate the solution fields and to project the full-order equations onto lower-dimensional subspaces, thus considerably reducing the number of unknowns in a numerical system. This reduction leads to significant computational speedups, which is ideal for multi-query applications such as uncertainty quantification or design optimization. The developed tool has been tested using a 2D multiphysics model of the Molten Salt Fast Reactor (MSFR) with steady-state and transient scenarios, with speedups on the order of 10 – 10 5 .

22 GENERAL STUDIES OF NUCLEAR REACTORS↗