Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “autoencoders”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Combining variational autoencoders and physical bias for improved microscopy data analysis *

Electron and scanning probe microscopy produce vast amounts of data in the form of images or hyperspectral data, such as electron energy loss spectroscopy or 4D scanning transmission electron microscope, that contain information on a wide range of structural, physical, and chemical properties of materials. To extract valuable insights from these data, it is crucial to identify physically separate regions in the data, such as phases, ferroic variants, and boundaries between them. In order to derive an easily interpretable feature analysis, combining with well-defined boundaries in a principled and unsupervised manner, here we present a physics augmented machine learning method which combines the capability of variational autoencoders to disentangle factors of variability within the data and the physics driven loss function that seeks to minimize the total length of the discontinuities in images corresponding to latent representations. Our method is applied to various materials, including NiO-LSMO, BiFeO 3 , and graphene. The results demonstrate the effectiveness of our approach in extracting meaningful information from large volumes of imaging data. The customized codes of the required functions and classes to develop phyVAE is available at https://github.com/arpanbiswas52/phy-VAE.

97 MATHEMATICS AND COMPUTING↗

Finding simplicity: unsupervised discovery of features, patterns, and order parameters via shift-invariant variational autoencoders *

Abstract Recent advances in scanning tunneling and transmission electron microscopies (STM and STEM) have allowed routine generation of large volumes of imaging data containing information on the structure and functionality of materials. The experimental data sets contain signatures of long-range phenomena such as physical order parameter fields, polarization, and strain gradients in STEM, or standing electronic waves and carrier-mediated exchange interactions in STM, all superimposed onto scanning system distortions and gradual changes of contrast due to drift and/or mis-tilt effects. Correspondingly, while the human eye can readily identify certain patterns in the images such as lattice periodicities, repeating structural elements, or microstructures, their automatic extraction and classification are highly non-trivial and universal pathways to accomplish such analyses are absent. We pose that the most distinctive elements of the patterns observed in STM and (S)TEM images are similarity and (almost-) periodicity, behaviors stemming directly from the parsimony of elementary atomic structures, superimposed on the gradual changes reflective of order parameter distributions. However, the discovery of these elements via global Fourier methods is non-trivial due to variability and lack of ideal discrete translation symmetry. To address this problem, we explore the shift-invariant variational autoencoders (shift-VAEs) that allow disentangling characteristic repeating features in the images, their variations, and shifts that inevitably occur when randomly sampling the image space. Shift-VAEs balance the uncertainty in the position of the object of interest with the uncertainty in shape reconstruction. This approach is illustrated for model 1D data, and further extended to synthetic and experimental STM and STEM 2D data. We further introduce an approach for training shift-VAEs that allows finding the latent variables that comport to known physical behavior. In this specific case, the condition is that the latent variable maps should be smooth on the length scale of the atomic lattice (as expected for physical order parameters), but other conditions can be imposed. The opportunities and limitations of the shift VAE analysis for pattern discovery are elucidated.

97 MATHEMATICS AND COMPUTING↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Inverse Analysis with Variational Autoencoders: A Comparison of Shallow and Deep Networks

Inverse problems are applied to determine unknown properties by matching observational data with a physical model that often takes many parameters as input. To overcome the underconstrained nature of inverse problems and achieve good performance, an approach is presented involving regularization with a technique known as a variational autoencoder (VAE), which is trained to map a high-dimensional parameter space with a complex structure to a low-dimensional latent space with a simple structure. We apply this approach to unconditioned realizations of the parameters (heterogeneous hydraulic fields) for a hydrogeological inverse problem. Two types of hydraulic conductivity fields are used to evaluate the characterization for the different levels of heterogeneity complexity of the physical inputs. This approach keeps the computational cost of generating the training data low. The reason is unconditioned realizations neither rely on the observational data used to perform the inverse analysis nor require any groundwater flow model forward runs. In addition, this approach applies regularization on a low-dimensional latent space from the VAE and increases optimization efficiency through automatic differentiation. Furthermore, two different neural network (NN) structures are tested for their utility in using the VAE for inverse analysis. The performance of a deep, convolutional neural network strongly depends on the dimensionality of the latent space, which requires tuning. In contrast, a shallow, dense neural network provides consistently accurate characterization without tuning. Furthermore, our approach evaluates the advantages of the shallow, dense neural network over the deep, convolutional one and enables future application to a wide range of inverse problems.

97 MATHEMATICS AND COMPUTING↗

Autoencoder Neural Network for chemically reacting systems

Incorporating detailed chemical kinetic models is critical for accurate simulations of reacting flows. However, detailed models involve a large number of thermochemical (TC) state variables. Solving the governing equations to evolve these TC variables becomes impractical for real-world applications. In this work, we propose an autoencoder (AE) neural network (NN)-based reduced model to accelerate such simulations. The AE NN is first trained to find a low-dimensional latent representation of the TC states. Then, the evolving state of a chemical system can be tracked by solving the equations of the latent variables instead of the original TC equations. We demonstrate the reduced model in a syngas CO/H 2 combustion system, using training data collected from canonical perfectly stirred reactors (PSRs). It is found that the AE model can reduce the dimension of the combustion system from 12 to 2 while maintaining low reconstruction error and excellent elemental mass conservation for the test dataset. In the a posteriori test, the combustion states obtained from solving the two latent equations are compared to those from solving the 12 equations of the full model. The AE reduced method is found to be able to capture the diverse combustion states on the top two branches of the S-curve well including the extinction turning point, but with higher prediction errors for states near the ignition turning point.

97 MATHEMATICS AND COMPUTING↗

Variational Autoencoders for Learning Nonlinear Dynamics of Physical Systems

We develop data-driven methods for incorporating physical information for priors to learn parsimonious representations of nonlinear systems arising from parameterized PDEs and mechanics. Our approach is based on Variational Autoencoders (VAEs) for learning nonlinear state space models from observations. We develop ways to incorporate geometric and topological priors through general manifold latent space representations. We investigate the performance of our methods for learning low dimensional representations for the nonlinear Burgers equation and constrained mechanical systems.

97 MATHEMATICS AND COMPUTING↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Autoencoders for semivisible jet detection

The production of dark matter particles from confining dark sectors may lead to many novel experimental signatures. Depending on the details of the theory, dark quark production in proton-proton collisions could result in semivisible jets of particles: collimated sprays of dark hadrons of which only some are detectable by particle collider experiments. The experimental signature is characterised by the presence of reconstructed missing momentum collinear with the visible components of the jets. This complex topology is sensitive to detector inefficiencies and mis-reconstruction that generate artificial missing momentum. With this work, we propose a signal-agnostic strategy to reject ordinary jets and identify semivisible jets via anomaly detection techniques. A deep neural autoencoder network with jet substructure variables as input proves highly useful for analyzing anomalous jets. The study focuses on the semivisible jet signature; however, the technique can apply to any new physics model that predicts signatures with anomalous jets from non-SM particles.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Investigation of process history and underlying phenomena associated with the synthesis of plutonium oxides using Vector Quantizing Variational Autoencoder

Accurate, high throughput, and unbiased analysis of plutonium oxide particles is needed for analysis of the phenomenology associated with process parameters in their synthesis. Compared to qualitative and taxonomic descriptors, quantitative descriptors of particle morphology through scanning electron microscopy (SEM) have shown success in analyzing process parameters of uranium oxides. Among other candidates, a neural network called a Vector Quantizing Variational Autoencoder (VQ-VAE) has shown the ability to quantitatively describe particle morphology to attain >85% accuracy in identifying uranium oxide processing routes. We utilize a VQ-VAE to quantitatively describe plutonium dioxide (PuO 2 ) particles created in a designed experiment and investigate their phenomenology and prediction of their process parameters. PuO 2 was calcined from Pu(III) oxalates that were precipitated under varying synthetic conditions that related to concentrations, temperature, addition and digestion times, precipitant feed, and strike order; the surface morphology of the resulting PuO 2 powders were analyzed by SEM. A pipeline was developed to extract and quantify useful image representations for individual particles with the VQ-VAE, then further reduce the dimensionality of the feature space using a bottlenecking neural network fit to perform multiple classification tasks simultaneously. The reduced feature space could predict process parameters with greater than 80% accuracies for some parameters with a single particle. They also showed utility for grouping particles with similar surface morphology characteristics together. Both the clustering and classification results reveal valuable information regarding which chemical process parameters chiefly influence the PuO 2 particle morphologies: strike order and oxalic acid feedstock. Doing the same analysis with multiple particles was shown to improve the classification accuracy on each process parameter over the use of a single particle, with statistically significant results generally seen with as few as four particles in a sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dynamics-Based Peptide–MHC Binding Optimization by a Convolutional Variational Autoencoder: A Use-Case Model for CASTELO

An unsolved challenge in the development of antigen-specific immunotherapies is determining the optimal antigens to target. Comprehension of antigen–major histocompatibility complex (MHC) binding is paramount toward achieving this goal. Here, we apply CASTELO, a combined machine learning-molecular dynamics (ML-MD) approach, to identify per-residue antigen binding contributions and then design novel antigens of increased MHC-II binding affinity for a type 1 diabetes-implicated system. We build upon a small-molecule lead optimization algorithm by training a convolutional variational autoencoder (CVAE) on MD trajectories of 48 different systems across four antigens and four HLA serotypes. We develop several new machine learning metrics including a structure-based anchor residue classification model as well as cluster comparison scores. ML-MD predictions agree well with experimental binding results and free energy perturbation-predicted binding affinities. Moreover, ML-MD metrics are independent of traditional MD stability metrics such as contact area and root-mean-square fluctuations (RMSF), which do not reflect binding affinity data. Finally, our work supports the role of structure-based deep learning techniques in antigen-specific immunotherapy design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Investigating Carboxysome Morphology Dynamics with a Rotationally Invariant Variational Autoencoder

Carboxysomes are a class of bacterial microcompartments that form proteinaceous organelles within the cytoplasm of cyanobacteria and play a central role in photosynthetic metabolism by defining a cellular microenvironment permissive to CO2 fixation. Critical aspects of the assembly of the carboxysomes remain relatively unknown, especially with regard to the dynamics of this microcompartment. Progress in understanding carboxysome dynamics is impeded in part because analysis of the subtle changes in carboxysome morphology with microscopy remains a low-throughput and subjective process. Here we use deep learning techniques, specifically a Rotationally Invariant Variational Autoencoder (rVAE), to analyze fluorescence microscopy images of cyanobacteria bearing a carboxysome reporter and quantitatively evaluate how carboxysome shell remodelling impacts subtle trends in the morphology of the microcompartment over time. Toward this goal, we use a recently developed tool to control endogenous protein levels, including carboxysomal components, in the model cyanobacterium Synechococcous elongatus PCC 7942. By utilization of this system, proteins that compose the carboxysome can be tuned in real time as a method to examine carboxysome dynamics. We find that rVAEs are able to assist in the quantitative evaluation of changes in carboxysome numbers, shape, and size over time. Further, we propose that rVAEs may be a useful tool to accelerate the analysis of carboxysome assembly and dynamics in response to genetic or environmental perturbation and may be more generally useful to probe regulatory processes involving a broader array of bacterial microcompartments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Design of diverse, functional mitochondrial targeting sequences across eukaryotic organisms using variational autoencoder

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

59 BASIC BIOLOGICAL SCIENCES↗

Physics and chemistry from parsimonious representations: image analysis via invariant variational autoencoders

Electron, optical, and scanning probe microscopy methods are generating ever increasing volume of image data containing information on atomic and mesoscale structures and functionalities. This necessitates the development of the machine learning methods for discovery of physical and chemical phenomena from the data, such as manifestations of symmetry breaking phenomena in electron and scanning tunneling microscopy images, or variability of the nanoparticles. Variational autoencoders (VAEs) are emerging as a powerful paradigm for the unsupervised data analysis, allowing to disentangle the factors of variability and discover optimal parsimonious representation. Here, we summarize recent developments in VAEs, covering the basic principles and intuition behind the VAEs. The invariant VAEs are introduced as an approach to accommodate scale and translation invariances present in imaging data and separate known factors of variations from the ones to be discovered. We further describe the opportunities enabled by the control over VAE architecture, including conditional, semi-supervised, and joint VAEs. Several case studies of VAE applications for toy models and experimental datasets in Scanning Transmission Electron Microscopy are discussed, emphasizing the deep connection between VAE and basic physical principles. Python codes and datasets discussed in this article are available at https://github.com/saimani5/VAE-tutorials and can be used by researchers as an application guide when applying these to their own datasets.

36 MATERIALS SCIENCE↗

Predicting wind-driven spatial deposition through simulated color images using deep autoencoders

Abstract For centuries, scientists have observed nature to understand the laws that govern the physical world. The traditional process of turning observations into physical understanding is slow. Imperfect models are constructed and tested to explain relationships in data. Powerful new algorithms can enable computers to learn physics by observing images and videos. Inspired by this idea, instead of training machine learning models using physical quantities, we used images, that is, pixel information. For this work, and as a proof of concept, the physics of interest are wind-driven spatial patterns. These phenomena include features in Aeolian dunes and volcanic ash deposition, wildfire smoke, and air pollution plumes. We use computer model simulations of spatial deposition patterns to approximate images from a hypothetical imaging device whose outputs are red, green, and blue (RGB) color images with channel values ranging from 0 to 255. In this paper, we explore deep convolutional neural network-based autoencoders to exploit relationships in wind-driven spatial patterns, which commonly occur in geosciences, and reduce their dimensionality. Reducing the data dimension size with an encoder enables training deep, fully connected neural network models linking geographic and meteorological scalar input quantities to the encoded space. Once this is achieved, full spatial patterns are reconstructed using the decoder. We demonstrate this approach on images of spatial deposition from a pollution source, where the encoder compresses the dimensionality to 0.02% of the original size, and the full predictive model performance on test data achieves a normalized root mean squared error of 8%, a figure of merit in space of 94% and a precision-recall area under the curve of 0.93.

54 ENVIRONMENTAL SCIENCES↗

VpROM: a novel variational autoencoder-boosted reduced order model for the treatment of parametric dependencies in nonlinear systems

Reduced Order Models (ROMs) are of considerable importance in many areas of engineering in which computational time presents difficulties. Established approaches employ projection-based reduction, such as Proper Orthogonal Decomposition. The limitation of the linear nature of such operators is typically tackled via a library of local reduction subspaces, which requires the assembly of numerous local ROMs to address parametric dependencies. Our work attempts to define a more generalisable mapping between parametric inputs and reduced bases for the purpose of generative modeling. We propose the use of Variational Autoencoders (VAEs) in place of the typically utilised clustering or interpolation operations, for inferring the fundamental vectors, termed as modes, which approximate the manifold of the model response for any and each parametric input state. The derived ROM still relies on projection bases, built on the basis of full-order model simulations, thus retaining the imprinted physical connotation. However, it additionally exploits a matrix of coefficients that relates each local sample response and dynamics to the global phenomena across the parametric input domain. The VAE scheme is utilised for approximating these coefficients for any input state. This coupling leads to a high-precision low-order representation, which is particularly suited for problems where model dependencies or excitation traits cause the dynamic behavior to span multiple response regimes. Moreover, the probabilistic treatment of the VAE representation allows for uncertainty quantification on the reduction bases, which may then be propagated to the ROM response. The performance of the proposed approach is validated on an open-source simulation benchmark featuring hysteresis and multi-parametric dependencies, and on a large-scale wind turbine tower characterised by nonlinear material behavior and model uncertainty.

Conditional VAEs↗

Autonomous design of new chemical reactions using a variational autoencoder

Artificial intelligence based chemistry models are a promising method of exploring chemical reaction design spaces. However, training datasets based on experimental synthesis are typically reported only for the optimal synthesis reactions. This leads to an inherited bias in the model predictions. Therefore, robust datasets that span the entirety of the solution space are necessary to remove inherited bias and permit complete training of the space. In this study, an artificial intelligence model based on a Variational AutoEncoder (VAE) has been developed and investigated to synthetically generate continuous datasets. The approach involves sampling the latent space to generate new chemical reactions. This developed technique is demonstrated by generating over 7,000,000 new reactions from a training dataset containing only 7,000 reactions. The generated reactions include molecular species that are larger and more diverse than the training set.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗