Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “autoencoders”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hierarchical-embedding autoencoder with a predictor as efficient architecture for learning time-evolution in multi-scale turbulent flows

We introduce a scale-aware, data-driven deep learning modeling framework for accurately predicting the time evolution of multi-scale turbulent plasma and liquid flows. The approach is motivated by the idea of scale separation. Structures of vastly different length scales emerge in these systems, and interactions between these structures occur only locally. To exploit this structure, the flow state is transformed by a hierarchical, fully convolutional autoencoder, not into a single embedding layer as in conventional convolutional surrogate models, but into a series of embedding layers. A stepwise training strategy ensures that fine-scale features are encoded on a high-resolution grid, while larger structures are represented on progressively coarser layers. The time evolution predictor advances all embedding layers in sync, capturing local interactions between features at the same scale as well as between all scales. This approach enables efficient modeling of multi-scale systems since negligible interactions between distant, small-scale structures do not need to be directly modeled. Our hierarchical-embedding autoencoder with a predictor framework is evaluated on canonical examples of multi-scale turbulence: two-dimensional Kolmogorov flow and Hasegawa–Wakatani plasma turbulence. In both cases, the proposed framework significantly improves predictive accuracy relative to conventional convolutional network architectures. A significant improvement in prediction accuracy was observed for crucial statistical characteristics of the Hasegawa–Wakatani plasma as well as for individual trajectories of the Kolmogorov flow turbulence. Importantly, the model's rollout for the Hasegawa–Wakatani problem demonstrates a four-order-of-magnitude speedup compared to traditional numerical solvers.

Khrabry, Alexander I. [Princeton Univ., NJ (United↗

A Conditional Autoencoder for Galaxy Photometric Parameter Estimation

Astronomical photometric surveys routinely image billions of galaxies, and traditionally infer the parameters of a parametric model for each galaxy. This approach has served us well, but the computational expense of deriving a full posterior probability distribution function is a challenge for increasingly ambitious surveys. In this paper, we use deep learning methods to characterize galaxy images, training a conditional autoencoder on mock data. The autoencoder can reconstruct and denoise galaxy images via a latent space engineered to include semantically meaningful parameters, such as brightness, location, size, and shape. Our model recovers galaxy fluxes and shapes on mock data with a lower variance than the Hyper Suprime-Cam photometry pipeline, and returns reasonable answers even for inputs outside the range of its training data. When applied to data in the training range, the regression errors on all extracted parameters are nearly unbiased with a variance near the Cramr-Rao bound.

79 ASTRONOMY AND ASTROPHYSICS↗

Enhancing Qubit Readout with Autoencoders

In addition to the need for stable and precisely controllable qubits, quantum computers take advantage of good readout schemes. Superconducting qubit states can be inferred from the readout signal transmitted through a dispersively coupled resonator. Here, this work proposes a readout classification method for superconducting qubits based on a neural network pretrained with an autoencoder approach. A neural network is pretrained with qubit readout signals as autoencoders in order to extract relevant features from the data set. Afterward, the pretrained-network inner-layer values are used to perform a classification of the inputs in a supervised manner. We demonstrate that this method can enhance classification performance, particularly for short- and long-time measurements where more traditional methods present inferior performance.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Clustering Algorithm for AM Parts using GSH and EDT with Autoencoder

SAND2025-10103O The Clustering Algorithm for AM Parts Using GSH and (EDT With Autoencoder is a software tool. It uses a clustering algorithm for additive manufacturing (AM) parts using generalized spherical harmonics (GSH) and Euclidean distance transform (EDT) with an autoencoder to quantify material microstructure. The tool offers improved sensitivity to microstructural changes compared to traditional approaches. The tool integrates multiple microstructural properties, such as grain morphology, crystallographic orientation, and material phase information, to provide a comprehensive analysis of material microstructures. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodgers, Theron [Sandia National Lab. (SNL-CA), Li↗

Multi-kernel Edge Attention Graph Autoencoder

MEAGraph (Multi-kernel Edge Attention Graph Autoencoder) is a graph-based autoencoder model designed for unsupervised data mining for datasets used in machine learning potentials. It provides accurate clustering for atomic environment identification, unsupervised and unlabeled data pruning for dataset construction.

Sun, Hong↗

An explainable variational autoencoder model for three-dimensional acoustic emission source localization in hollow cylindrical structures

We introduce an explainable variational autoencoder for three-dimensional (3D) localization of acoustic emission sources in hollow cylindrical structures, with an unsupervised approach. This research capitalizes on multi-arrival waveforms generated by helical path propagation in cylindrical geometries to enable efficient two-receiver localization. By integrating the modal characteristics of Lamb modes under multi-path conditions, we demonstrate that two sets of time-of-arrival differences and peak amplitudes extracted from one receiver can serve as effective localization features. This initial approach identifies four potential source locations, highlighting the feasibility of two-receiver source localization using traditional feature extraction methods. However, direct extraction can be challenging when mode overlaps occur, complicating the localization process. To address this, our work proposes a novel waveform-based method. This method leverages the consistent dispersion characteristics within isotropic materials, where each unique combination of mode arrival times and peak amplitudes constructs a distinct waveform. This distinctiveness overcomes the ambiguities associated with mode overlaps, significantly enhancing the method’s precision and robustness. Our approach adopts a data-driven strategy for waveform-based localization using variational autoencoder (VAE). VAE discerns waveform patterns for localization, while also addressing data uncertainties. The VAE’s encoder and decoder networks capture the localization process and the source’s influence on waveform generation, respectively, guiding latent variables to segregate waveforms by source in the latent space. The design of the learning process focuses on specific localization characteristics to enhance result explainability. Localization predictions are generated by projecting test waveforms, not included in the training set, onto a trained latent space. The prediction is determined using a nearest-neighbor approach based on the closest latent representation of a source. Validation with pencil-lead-break tests on a metallic pipe confirmed our method’s effectiveness, achieving an averaged 3D localization accuracy of 0.84.

Lee, Guan-Wei↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Improving Variational Autoencoders for New Physics Detection at the LHC With Normalizing Flows

We investigate how to improve new physics detection strategies exploiting variational autoencoders and normalizing flows for anomaly detection at the Large Hadron Collider. As a working example, we consider the DarkMachines challenge dataset. We show how different design choices (e.g., event representations, anomaly score definitions, network architectures) affect the result on specific benchmark new physics models. Once a baseline is established, we discuss how to improve the anomaly detection accuracy by exploiting normalizing flow layers in the latent space of the variational autoencoder.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Probabilistic Autoencoder for Type Ia Supernova Spectral Time Series

We construct a physically parameterized probabilistic autoencoder (PAE) to learn the intrinsic diversity of Type Ia supernovae (SNe Ia) from a sparse set of spectral time series. The PAE is a two-stage generative model, composed of an autoencoder that is interpreted probabilistically after training using a normalizing flow. We demonstrate that the PAE learns a low-dimensional latent space that captures the nonlinear range of features that exists within the population and can accurately model the spectral evolution of SNe Ia across the full range of wavelength and observation times directly from the data. By introducing a correlation penalty term and multistage training setup alongside our physically parameterized network, we show that intrinsic and extrinsic modes of variability can be separated during training, removing the need for the additional models to perform magnitude standardization. We then use our PAE in a number of downstream tasks on SNe Ia for increasingly precise cosmological analyses, including the automatic detection of SN outliers, the generation of samples consistent with the data distribution, and solving the inverse problem in the presence of noisy and incomplete data to constrain cosmological distance measurements. We find that the optimal number of intrinsic model parameters appears to be three, in line with previous studies, and show that we can standardize our test sample of SNe Ia with an rms of 0.091 ± 0.010 mag, which corresponds to 0.074 ± 0.010 mag if peculiar velocity contributions are removed.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluating the Efficacy of Conditional Variational Autoencoders in Generating Synthetic Single Nuclei RNA-Seq Data for Space Biology Research

Astronauts are subject to unique stressors during spaceflight, leading to changes in their cellular function. However, neither astronauts nor model organisms respond the same to spaceflight, and research implicates a contribution of omics components in differential responses. Understanding how gene expression affects astronaut health is critical for the success of long-term space missions, prompting interest in developing personalized predictive models leveraging artificial intelligence (AI) and machine learning (ML) techniques. Developing such models requires extensive data, which is challenging to obtain and share. This study explores the use of conditional variational autoencoders (CVAEs) to synthetically generate single-nuclei RNA-seq (snRNA-seq) data. CVAEs build on standard variational autoencoders (VAEs) by conditioning data generation on covariates like sample identity and mission parameters, enhancing the relevance of generated data for specific contexts. For our work, we built two CVAEs with varying degrees of sparsity to optimize both interpretability and generative power. We train and validate models on existing snRNA-seq data collected from the brain tissue of mice subjected to spaceflight conditions and their ground control counterparts. We evaluate model performance using statistical tests and visualizations to compare synthetic data to real data. We aim to demonstrate that these prototype CVAE architectures could be used in future space biology work and that this is a method worth further exploring.

Sarah Golts↗

Autoencoders on FPGAs for real-time, unsupervised new physics detection at 40 MHz at the Large Hadron Collider

In this paper, we show how to adapt and deploy anomaly detection algorithms based on deep autoencoders, for the unsupervised detection of new physics signatures in the extremely challenging environment of a real-time event selection system at the Large Hadron Collider (LHC). We demonstrate that new physics signatures can be enhanced by three orders of magnitude, while staying within the strict latency and resource constraints of a typical LHC event filtering system. This would allow for collecting datasets potentially enriched with high-purity contributions from new physics processes. Through per-layer, highly parallel implementations of network layers, support for autoencoder-specific losses on FPGAs and latent space based inference, we demonstrate that anomaly detection can be performed in as little as $80\,$ns using less than 3% of the logic resources in the Xilinx Virtex VU9P FPGA. Opening the way to real-life applications of this idea during the next data-taking campaign of the LHC.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Missing Photovoltaic Data Imputation

The integration of the global Photovoltaic (PV) market with real time data-loggers has enabled large scale PV data analytical pipelines for power forecasting and long-term reliability assessment of PV fleets. Nevertheless, the performance of PV data analysis heavily depends on the quality of PV timeseries data. This paper proposes a novel Spatio-Temporal Denoising Graph Autoencoder (STD-GAE) framework to impute missing PV Power Data. STDGAE exploits temporal correlation, spatial coherence, and value dependencies from domain knowledge to recover missing data. It is empowered by two modules. (1) To cope with sparse yet various scenarios of missing data, STD-GAE incorporates a domain-knowledge aware data augmentation module that creates plausible variations of missing data patterns. This generalizes STD-GAE to robust imputation over different seasons and environment. (2) STD-GAE nontrivially integrates spatiotemporal graph convolution layers (to recover local missing data by observed “neighboring” PV plants) and denoising autoencoder (to recover corrupted data from augmented counterpart) to improve the accuracy of imputation accuracy at PV fleet level. We have evaluated our proposed model on two realworld PV datasets. Experimental results show that STD-GAE can achieve a gain of 43.14% in imputation accuracy and remains less sensitive to missing rate, different seasons, and missing scenarios, compared with state-of-the-art data imputation methods such as MIDA and LRTC-TNN.

Fan, Yangxin↗

Physics-Driven Convolutional Autoencoder Approach for CFD Data Compressions: Preprint

With the growing size and complexity of turbulent flow models, data compression approaches are of the utmost importance to analyze, visualize, or restart the simulations. Recently, in-situ autoencoder-based compression approaches have been proposed and shown to be effective at producing reduced representations of turbulent flow data. However, these approaches focus solely on training the model using point-wise sample reconstruction losses that do not take advantage of the physical properties of turbulent flows. In this paper, we show that training autoencoders with additional physics-informed regularizations, e.g., enforcing incompressibility and preserving enstrophy, improves the compression model in three ways: (i) the compressed data better conform to known physics for homogeneous isotropic turbulence without negatively impacting point-wise reconstruction quality, (ii) inspection of the gradients of the trained model uncovers changes to the learned compression mapping that can facilitate the use of explainability techniques, and (iii) as a performance byproduct, training losses are shown to converge up to 12x faster than the baseline model.

auto-encoders↗

Non-intrusive reduced order modeling of natural convection in porous media using convolutional autoencoders: Comparison with linear subspace techniques

Natural convection in porous media is a highly nonlinear multiphysical problem relevant to many engineering applications (e.g., the process of CO 2 sequestration). Here, we extend and present a non-intrusive reduced order model of natural convection in porous media employing deep convolutional autoencoders for the compression and reconstruction and either radial basis function (RBF) interpolation or artificial neural networks (ANNs) for mapping parameters of partial differential equations (PDEs) on the corresponding nonlinear manifolds. To benchmark our approach, we also describe linear compression and reconstruction processes relying on proper orthogonal decomposition (POD) and ANNs. Further, we present comprehensive comparisons among different models through three benchmark problems. The reduced order models, linear and nonlinear approaches, are much faster than the finite element model, obtaining a maximum speed-up of 7 × 10 6 because our framework is not bound by the Courant–Friedrichs–Lewy condition; hence, it could deliver quantities of interest at any given time contrary to the finite element model. Our model’s accuracy still lies within a relative error of 7% in the worst-case scenario. We illustrate that, in specific settings, the nonlinear approach outperforms its linear counterpart and vice versa. We hypothesize that a visual comparison between principal component analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE) could indicate which method will perform better prior to employing any specific compression strategy.

97 MATHEMATICS AND COMPUTING↗

Learning and Predicting Photonic Responses of Plasmonic Nanoparticle Assemblies via Dual Variational Autoencoders

In this work, the application of machine learning is demonstrated for rapid and accurate extraction of plasmonic particles cluster geometries from hyperspectral image data via a dual variational autoencoder (dual-VAE). In this approach, the information is shared between the latent spaces of two VAEs acting on the particle shape data and spectral data, respectively, but enforcing a common encoding on the shape-spectra pairs. It is shown that this approach can establish the relationship between the geometric characteristics of nanoparticles and their far-field photonic responses, demonstrating that hyperspectral darkfield microscopy can be used to accurately predict the geometry (number of particles, arrangement) of a multiparticle assemblies below the diffraction limit in an automated fashion with high fidelity (for monomers (0.96), dimers (0.86), and trimers (0.58). This approach of building structure-property relationships via shared encoding is universal and should have applications to a broader range of materials science and physics problems in imaging of both molecular and nanomaterial systems.

variational autoencoder↗

Towards inverse microstructure-centered materials design using generative phase-field modeling and deep variational autoencoders

The field of Integrated Computational Materials Engineering (ICME) combines a broad range of methods to study materials’ responses over a spectrum of length scales. A relatively unexplored aspect of microstructure-sensitive materials design is uncertainty propagation and quantification (UP/UQ) of materials’ microstructure, as well as establishing process-structure–property (PSP) relationships for inverse material design. In this study, an efficient UP technique built on the idea of changing probability measures and a deep generative unsupervised representative machine learning method for microstructure-based design of thermal conductivity of materials is proposed. Probability measures are used to represent microstructure space, and Wasserstein metrics are used to test the efficiency of the UP method. By using deep Variational AutoEncoder (VAE), we identify the correlations between the material/process parameters and the thermal conductivity of heterogeneous dual-phase microstructures. Through high-throughput screening, UP, and the deep-generative VAE method, PSP relationships that are too complex can be revealed by exploiting the materials’ design space with an emphasis on microstructures. As a last point, we demonstrate generative machine learning serves as a useful tool for inverse microstructure-centered materials design, and we demonstrate this by examining the inverse design of thermal conductivity in nano-structured materials. Here, the results reveal the effects of morphology, volume fraction, characteristic length scale, and the individual thermal diffusivity of phases on the thermal conductivity of dual-phase alloys. Our findings emphasize the advantages of high-throughput phase-field modeling and generative deep learning for linking PSP and inverse microstructure-centered materials design.

36 MATERIALS SCIENCE↗

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

$\text{GPLaSDI}$: Gaussian Process-based interpretable Latent Space Dynamics Identification through deep autoencoder

Numerically solving partial differential equations (PDEs) can be challenging and computationally expensive. This has led to the development of reduced-order models (ROMs) that are accurate but faster than full order models (FOMs). Recently, machine learning advances have enabled the creation of non-linear projection methods, such as Latent Space Dynamics Identification (LaSDI). LaSDI maps full-order PDE solutions to a latent space using autoencoders and learns the system of ODEs governing the latent space dynamics. By interpolating and solving the ODE system in the reduced latent space, fast and accurate ROM predictions can be made by feeding the predicted latent space dynamics into the decoder. In this paper, we introduce GPLaSDI, a novel LaSDI-based framework that relies on Gaussian process (GP) for latent space ODE interpolations. Using GPs offers two significant advantages. First, it enables the quantification of uncertainty over the ROM predictions. Second, leveraging this prediction uncertainty allows for efficient adaptive training through a greedy selection of additional training data points. This approach does not require prior knowledge of the underlying PDEs. Consequently, GPLaSDI is inherently non-intrusive and can be applied to problems without a known PDE or its residual. Here we demonstrate the effectiveness of our approach on the Burgers equation, Vlasov equation for plasma physics, and a rising thermal bubble problem. Our proposed method achieves between 200 and 100,000 times speed-up, with up to 7% relative error.

97 MATHEMATICS AND COMPUTING↗