Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “latent space learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Latent Representation Learning for Structural Characterization of Catalysts

Supervised machine learning-enabled mapping of the X-ray absorption near edge structure (XANES) spectra to local structural descriptors offers new methods for understanding the structure and function of working nanocatalysts. We briefly summarize a status of XANES analysis approaches by supervised machine learning methods. We present an example of an autoencoder-based, unsupervised machine learning approach for latent representation learning of XANES spectra. This new approach produces a lower-dimensional latent representation, which retains a spectrum–structure relationship that can be eventually mapped to physicochemical properties. Furthermore, the latent space of the autoencoder also provides a pathway to interpret the information content “hidden” in the X-ray absorption coefficient. Our approach (that we named latent space analysis of spectra, or LSAS) is demonstrated for the supported Pd nanoparticle catalyst studied during the formation of Pd hydride. By employing the low-dimensional representation of Pd K-edge XANES, the LSAS method was able to isolate the key factors responsible for the observed spectral changes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unsupervised machine learning discovery of structural units and transformation pathways from imaging data

We show that unsupervised machine learning can be used to learn chemical transformation pathways from observational Scanning Transmission Electron Microscopy (STEM) data. To enable this analysis, we assumed the existence of atoms, a discreteness of atomic classes, and the presence of an explicit relationship between the observed STEM contrast and the presence of atomic units. With only these postulates, we developed a machine learning method leveraging a rotationally invariant variational autoencoder (VAE) that can identify the existing molecular fragments observed within a material. The approach encodes the information contained in STEM image sequences using a small number of latent variables, allowing the exploration of chemical transformation pathways by tracing the evolution of atoms in the latent space of the system. The results suggest that atomically resolved STEM data can be used to derive fundamental physical and chemical mechanisms involved, by providing encodings of the observed structures that act as bottom-up equivalents of structural order parameters. The approach also demonstrates the potential of variational (i.e., Bayesian) methods in the physical sciences and will stimulate the development of more sophisticated ways to encode physical constraints in the encoder–decoder architectures and generative physical laws and causal relationships in the latent space of VAEs.

97 MATHEMATICS AND COMPUTING↗

Bayesian sparse learning with preconditioned stochastic gradient MCMC and its applications

Deep neural networks have been successfully employed in an extensive variety of research areas, including solving partial differential equations. Despite its significant success, there are some challenges in effectively training DNN, such as avoiding overfitting in over-parameterized DNNs and accelerating the optimization in DNNs with pathological curvature. Here, we propose a Bayesian type sparse deep learning algorithm. The algorithm utilizes a set of spike-and-slab priors for the parameters in the deep neural network. The hierarchical Bayesian mixture will be trained using an adaptive empirical method. That is, one will alternatively sample from the posterior using preconditioned stochastic gradient Langevin Dynamics (PSGLD), and optimize the latent variables via stochastic approximation. The sparsity of the network is achieved while optimizing the hyperparameters with adaptive searching and penalizing. A popular SG-MCMC approach is Stochastic gradient Langevin dynamics (SGLD). However, considering the complex geometry in the model parameter space in nonconvex learning, updating parameters using a universal step size in each component as in SGLD may cause slow mixing. To address this issue, we apply a computationally manageable preconditioner in the updating rule, which provides a step-size parameter to adapt to local geometric properties. Moreover, by smoothly optimizing the hyperparameter in the preconditioning matrix, our proposed algorithm ensures a decreasing bias, which is introduced by ignoring the correction term in the preconditioned SGLD. According to the existing theoretical framework, we show that the proposed algorithm can asymptotically converge to the correct distribution with a controllable bias under mild conditions. Numerical tests are performed on both synthetic regression problems and learning solutions of elliptic PDE, which demonstrate the accuracy and efficiency of the present work.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Materials representation and transfer learning for multi-property prediction

The adoption of machine learning in materials science has rapidly transformed materials property prediction. Hurdles limiting full capitalization of recent advancements in machine learning include the limited development of methods to learn the underlying interactions of multiple elements as well as the relationships among multiple properties to facilitate property prediction in new composition spaces. To address these issues, we introduce the Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP) framework that seamlessly integrates: (i) prediction using only a material's composition, (ii) learning and exploitation of correlations among target properties in multi-target regression, and (iii) leveraging training data from tangential domains via generative transfer learning. The model is demonstrated for prediction of spectral optical absorption of complex metal oxides spanning 69 three-cation metal oxide composition spaces. H-CLMP accurately predicts non-linear composition-property relationships in composition spaces for which no training data are available, which broadens the purview of machine learning to the discovery of materials with exceptional properties. This achievement results from the principled integration of latent embedding learning, property correlation learning, generative transfer learning, and attention models. The best performance is obtained using H-CLMP with transfer learning [H-CLMP(T)] wherein a generative adversarial network is trained on computational density of states data and deployed in the target domain to augment prediction of optical absorption from composition. H-CLMP(T) aggregates multiple knowledge sources with a framework that is well suited for multi-target regression across the physical sciences.

36 MATERIALS SCIENCE↗

Graph learning for particle accelerator operations

Particle accelerators play a crucial role in scientific research, enabling the study of fundamental physics and materials science, as well as having important medical applications. This study proposes a novel graph learning approach to classify operational beamline configurations as good or bad. By considering the relationships among beamline elements, we transform data from components into a heterogeneous graph. We propose to learn from historical, unlabeled data via our self-supervised training strategy along with fine-tuning on a smaller, labeled dataset. Additionally, we extract a low-dimensional representation from each configuration that can be visualized in two dimensions. Leveraging our ability for classification, we map out regions of the low-dimensional latent space characterized by good and bad configurations, which in turn can provide valuable feedback to operators. This research demonstrates a paradigm shift in how complex, many-dimensional data from beamlines can be analyzed and leveraged for accelerator operations.

43 PARTICLE ACCELERATORS↗

Bayesian reduced-order deep learning surrogate model for dynamic systems described by partial differential equations

We propose a reduced-order deep-learning surrogate model for dynamic systems described by time-dependent partial differential equations. This method employs space–time Karhunen–Loève expansions (KLEs) of the state variables and space-dependent KLEs of space-varying parameters to identify the reduced (latent) dimensions. Subsequently, a deep neural network (DNN) is used to map the parameter latent space to the state variable latent space. An approximate Bayesian method is developed for uncertainty quantification (UQ) in the proposed KL-DNN surrogate model. The KL-DNN method is tested for the linear advection–diffusion and nonlinear diffusion equations, and the Bayesian approach for UQ is compared with the deep ensembling (DE) approach, commonly used for quantifying uncertainty in DNN models. It was found that the approximate Bayesian method provides a more informative distribution of the PDE solutions in terms of the coverage of the reference PDE solutions (the percentage of nodes where the reference solution is within the confidence interval predicted by the UQ methods) and log predictive probability. The DE method is found to underestimate uncertainty and introduce bias. For the nonlinear diffusion equation, we compare the KL-DNN method with the Fourier Neural Operator (FNO) method and find that KL-DNN is 10% more accurate and needs less training time than the FNO method.

97 MATHEMATICS AND COMPUTING↗

Diagnostic-free onboard battery health assessment

Diverse usage patterns induce complex and variable aging behaviors in lithiumion batteries, complicating accurate health diagnosis and prognosis. Separate diagnostic cycles are often used to untangle the battery’s current state of health from prior complex aging patterns. However, these same diagnostic cycles alter the battery’s degradation trajectory, are time-intensive, and cannot be practically performed in onboard applications. Here, in this work, we leverage portions of operational measurements in combination with an interpretable machine learning model to enable rapid, onboard battery health diagnostics and prognostics without offline diagnostic testing and the requirement of historical data. We integrate mechanistic constraints within an encoder-decoder architecture to extract electrode states in a physically interpretable latent space and enable improved reconstruction of the degradation path. The health diagnosis model framework can be flexibly applied across diverse application interests with slight fine-tuning.

battery aging reconstruction↗

Predicting Future Laboratory Fault Friction Through Deep Learning Transformer Models

Machine learning models using seismic emissions as input can predict instantaneous fault characteristics such as displacement and friction in laboratory experiments, and slow slip in Earth. Here, we address whether the seismic/acoustic emission (AE) from laboratory experiments contains information about future frictional behavior. The approach uses a convolutional encoder-decoder containing a transformer model in the latent space, similar to models used for natural language processing. We test the model limits using progressively larger AE input time windows and progressively larger output friction time windows. The results demonstrate that very near-term friction predictions are indeed contained in the AE signal, and predictions are progressively worse farther into the future. The future predictions by the model of impending failure in the near-term are remarkably robust. This first effort predicting future fault frictional behavior with machine learning will aid in guiding efforts for applications in Earth.

58 GEOSCIENCES↗

Non‐Linear Dimensionality Reduction With a Variational Encoder Decoder to Understand Convective Processes in Climate Models

Abstract Deep learning can accurately represent sub‐grid‐scale convective processes in climate models, learning from high resolution simulations. However, deep learning methods usually lack interpretability due to large internal dimensionality, resulting in reduced trustworthiness in these methods. Here, we use Variational Encoder Decoder structures (VED), a non‐linear dimensionality reduction technique, to learn and understand convective processes in an aquaplanet superparameterized climate model simulation, where deep convective processes are simulated explicitly. We show that similar to previous deep learning studies based on feed‐forward neural nets, the VED is capable of learning and accurately reproducing convective processes. In contrast to past work, we show this can be achieved by compressing the original information into only five latent nodes. As a result, the VED can be used to understand convective processes and delineate modes of convection through the exploration of its latent dimensions. A close investigation of the latent space enables the identification of different convective regimes: (a) stable conditions are clearly distinguished from deep convection with low outgoing longwave radiation and strong precipitation; (b) high optically thin cirrus‐like clouds are separated from low optically thick cumulus clouds; and (c) shallow convective processes are associated with large‐scale moisture content and surface diabatic heating. Our results demonstrate that VEDs can accurately represent convective processes in climate models, while enabling interpretability and better understanding of sub‐grid‐scale physical processes, paving the way to increasingly interpretable machine learning parameterizations with promising generative properties.

54 ENVIRONMENTAL SCIENCES↗

A system identification approach for non-intrusive reduced order modeling of radiation-induced photocurrents

In this study, development of compact photocurrent models is currently dominated by analytical techniques that rely on physical assumptions to render the governing equations solvable in a closed form. Violation of these assumptions can reduce the accuracy of the models and/or limit their scope. In this paper we show that system identification of nonlinear state-space systems can serve as an alternative numerical basis for non-intrusive reduced order modeling of photocurrent effects. To that end we develop a compact gray box photocurrent model (GBPM) by using a state-space representation with a low-dimensional latent state equation that mimics a mathematical model for the response of an idealized class of devices to ionizing radiation. In so doing we obtain a model that learns the dynamics of a quantity of interest directly from its measurements without requiring snapshots of the internal device state or its discretized model, and can be inferred from very small data sets. To demonstrate the approach we train the GBPM using a small experimental data set for a Z5236 Zener diode and a small synthetic data set obtained by simulating a synthetic pn-junction device. We then compare the GBPMs with black box models trained on the same data and show that performance of the latter is limited by the size of the data set, while the former are able to achieve excellent performance in both the reproductive and the predictive regimes.

97 MATHEMATICS AND COMPUTING↗

Self-Supervised Anomaly Detection via Neural Autoregressive Flows with Active Learning

Many self-supervised methods have been proposed with the target of image anomaly detection. These methods often rely on the paradigm of data augmentation with predefined transformations such as flipping, cropping, and rotations. However, it is not straightforward to apply these techniques for non-image data, such as time series or tabular data, while the performance of the existing deep approaches has been under our expectation on tasks beyond images. In this work, we propose a novel active learning (AL) scheme that relied on neural autoregressive flows (NAF) for self-supervised anomaly detection, specifically on small-scale data. Unlike other generative models such as GANs or VAEs, flow-based models allow to explicitly learn the probability density and thus can assign accurate likelihoods to normal data which makes it usable to detect anomalies. The proposed NAF-AL method is achieved by efficiently generating random samples from latent space and transforming them into feature space along with likelihoods via invertible mapping. The samples with lower likelihoods are selected and further checked by outlier detection using Mahalanobis distance. The augmented samples incorporating with normal samples are used for training a better detector so as to approach decision boundaries. Compared with random transformations, NAF-AL can be interpreted as a likelihood-oriented data augmentation that is more efficient and robust. Extensive experiments show that our approach outperforms existing baselines on multiple time series and tabular datasets, and a real-world application in advanced manufacturing, with significant improvement on anomaly detection accuracy and robustness over the state-of-the-art.

Zhang, Jiaxin↗

Preliminary results in using Deep Learning to emulate BLOB, a nuclear interaction model

Purpose: A reliable model to simulate nuclear interactions is fundamental for Ion-therapy. We already showed how BLOB (“Boltzmann-Langevin One Body”), a model developed to simulate heavy ion interactions up to few hundreds of MeV/u, could simulate also 12 C reactions in the same energy domain. However, its computation time is too long for any medical application. For this reason we present the possibility of emulating it with a Deep Learning algorithm. Methods: The BLOB final state is a Probability Density Function (PDF) of finding a nucleon in a position of the phase space. We discretised this PDF and trained a Variational Auto-Encoder (VAE) to reproduce such a discrete PDF. As a proof of concept, we developed and trained a VAE to emulate BLOB in simulating the interactions of 12 C with 12 C at 62 MeV/u. To have more control on the generation, we forced the VAE latent space to be organised with respect to the impact parameter (b) training a classifier of b jointly with the VAE. Results: In this work, the distributions obtained from the VAE are similar to the input ones and the computation time needed to use the VAE as a generator is negligible. Conclusions: We show that it is possible to use a Deep Learning approach to emulate a model developed to simulate nuclear reactions in the energy range of interest for Ion-therapy. We foresee the implementation of the generation part in C++ and to interface it with the most used Monte Carlo toolkit: Geant4.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Divertor Plasma Detachment Control Neural Network

DivControlNN is a state-of-the-art software tool that leverages advanced machine learning techniques to predict and control divertor plasma behavior in fusion reactors. Plasma, a highly energetic and electrically charged gas, requires meticulous management to protect reactor components and maintain optimal energy production. Conventional simulation methods, although extremely detailed, typically demand extensive computational time-making them unsuitable for real-time control scenarios. DivControlNN addresses this challenge by learning from tens of thousands of high-fidelity simulations, thereby creating a rapid surrogate model that can deliver near-instantaneous predictions. At the core of its functionality is a sophisticated technique known as latent space mapping, which condenses complex, high-dimensional plasma data into a compact, lower-dimensional representation. This streamlined representation enables the system to quickly forecast essential plasma properties and determine the precise conditions required for effective detachment. Detachment is a crucial process in which the plasma is cooled before reaching the divertor plates, thereby reducing heat loads and mitigating material erosion. In recent experiments conducted on the KSTAR tokamak in South Korea, DivControlNN successfully guided the detachment process without any fine-tuning-even when applied to a new tungsten divertor configuration. By achieving a computational speed-up of over one hundred million times compared to traditional simulation methods while maintaining low prediction errors, DivControlNN stands to significantly enhance real-time control and diagnostic capabilities in future fusion reactors. This breakthrough paves the way for safer, more reliable reactor operation and represents a major advancement toward realizing fusion energy as a practical, sustainable, and clean power source.

Xu, Xueqiao [Lawrence Livermore National Laborator↗

Reification of latent microstructures: On supervised unsupervised and semi-supervised deep learning applications for microstructures in materials informatics

Machine learning (ML), including deep learning (DL), has become increasingly popular in the last few years due to its continually outstanding performance. In this context, we apply machine learning techniques to "learn" the microstructure using both supervised and unsupervised DL techniques. In particular, we focus (1) on the localization problem bridging (micro)structure (localized) property using supervised DL and (2) on the microstructure reconstruction problem in latent space using unsupervised DL. The goal of supervised and semi-supervised DL is to replace crystal plasticity finite element model (CPFEM) that maps from (micro)structure (localized) property, and implicitly the (micro)structure (homogenized) property relationships, while the goal of unsupervised DL is (1) to represent high-dimensional microstructure images in a non-linear low-dimensional manifold, and (2) to discover a way to interpolate microstructures via latent space associating with latent microstructure variables. At the heart of this report is the applications of several common DL architectures, including convolutional neural networks (CNN), autoencoder (AE), and generative adversarial network (GAN), to multiple microstructure datasets, and the quest of neural architecture search for optimal DL architectures.

36 MATERIALS SCIENCE↗

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

Learning to simulate high energy particle collisions from unlabeled data

In many scientific fields which rely on statistical inference, simulations are often used to map from theoretical models to experimental data, allowing scientists to test model predictions against experimental results. Experimental data is often reconstructed from indirect measurements causing the aggregate transformation from theoretical models to experimental data to be poorly-described analytically. Instead, numerical simulations are used at great computational cost. We introduce Optimal-Transport-based Unfolding and Simulation (OTUS), a fast simulator based on unsupervised machine-learning that is capable of predicting experimental data from theoretical models. Without the aid of current simulation information, OTUS trains a probabilistic autoencoder to transform directly between theoretical models and experimental data. Identifying the probabilistic autoencoder’s latent space with the space of theoretical models causes the decoder network to become a fast, predictive simulator with the potential to replace current, computationally-costly simulators. Here, we provide proof-of-principle results on two particle physics examples, Z-boson and top-quark decays, but stress that OTUS can be widely applied to other fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Separating Physically Distinct Mechanisms in Complex Infrared Plasmonic Nanostructures via Machine Learning Enhanced Electron Energy Loss Spectroscopy

Electron energy loss spectroscopy (EELS) enables direct exploration of plasmonic phenomena at the nanometer level. To isolate individual plasmon modes, linear unmixing methods can be used to separate different physical mechanisms, but in larger and more complex systems the interpretability of the components becomes uncertain. Here, infrared plasmonic resonances in self-assembled heterogeneous monolayer films of doped-semiconductor nanoparticles are examined beyond linear unmixing techniques, and both supervised and unsupervised machine-learning-based analyses of hyperspectral EELS datasets are demonstrated. Additionally, in the supervised approach, a human operator labels a small number of pixels in the hyperspectral dataset corresponding to features of interest which are then propagated across the entire dataset. In the unsupervised approach, non-linear autoencoders are used to create a highly-reduced latent-space representation of the dataset, within which insight into the relevant physics can be gleaned from straightforward distance metrics that do not depend on operator input and bias. The advantage of these approaches is that the labeling separates physical mechanisms without altering the data, enabling robust analyses of the influence of heterogeneities in mesoscale complex systems.

36 MATERIALS SCIENCE↗

Simulating Atmospheric Processes in Earth System Models and Quantifying Uncertainties With Deep Learning Multi‐Member and Stochastic Parameterizations

Abstract Deep learning is a powerful tool to represent subgrid processes in climate models, but many application cases have so far used idealized settings and deterministic approaches. Here, we develop stochastic parameterizations with calibrated uncertainty quantification to learn subgrid convective and turbulent processes and surface radiative fluxes of a superparameterization embedded in an Earth System Model (ESM). We explore three methods to construct stochastic parameterizations: (a) a single Deep Neural Network (DNN) with Monte Carlo Dropout; (b) a multi‐member parameterization; and (c) a Variational Encoder Decoder with latent space perturbation. We show that the multi‐member parameterization improves the representation of convective processes, especially in the planetary boundary layer, compared to individual DNNs. The respective uncertainty quantification illustrates that methods (b) and (c) are advantageous compared to a dropout‐based DNN parameterization regarding the spread of convective processes. Hybrid simulations with our best‐performing multi‐member parameterizations remained challenging and crash within the first days. Therefore, we develop a pragmatic partial coupling strategy relying on the superparameterization for condensate emulation. Partial coupling reduces the computational efficiency of hybrid Earth‐like simulations but enables model stability over 5 months with our multi‐member parameterizations. However, our hybrid simulations exhibit biases in thermodynamic fields and differences in precipitation patterns. Despite this, the multi‐member parameterizations enable improvements in reproducing tropical extreme precipitation compared to a traditional convection parameterization. Despite these challenges, our results indicate the potential of a new generation of multi‐member machine learning parameterizations leveraging uncertainty quantification to improve the representation of stochasticity of subgrid effects.

Behrens, Gunnar [Deutsches Zentrum für Luft‐ und R↗