Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “autoencoders”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Efficient graph representation framework for chemical molecule similarity tasks

Graph data has emerged in numerous scientific domains and machine learning techniques have been widely used for analysis and learning of diverse data for prediction and decision. Machine learning techniques can readily address complex problems by leveraging their structural information. But graphs cannot be directly used for existing machine learning algorithms unless encoded as vectors. The problem of efficient representation of graphs is a substantial challenge in graph machine learning. In this paper, we propose a novel two-stage framework for the representation of chemical molecule graphs based on the strengths of Graph Isomorphism Networks (GINs) and Siamese autoencoders. In the first stage, the GIN model is constructed and trained using the structural information of chemical molecule graphs. Node attributes, edge attributes, and edge indices are used as input data, while graph attributes are used as labels. The GIN model effectively captures the structural characteristics of graphs and can accurately predict graph attributes, i.e., molecular properties. It also generates Graph Embeddings, represented as vectors that encode the structural information of graphs. In the second stage, Graph Embedding vectors are further optimized for downstream similarity tasks while preserving the graph structural information. The Siamese autoencoder is constructed and trained, which reduces the dimensionality of the Graph Embedding vectors, while maximizing the preservation of structural information in the original high-dimensional vectors. The resulting low-dimensional Graph Embeddings can be effectively utilized for tasks such as approximate nearest neighbor search. The experimental results demonstrate the effectiveness of our proposed framework in accurately predicting graph similarity.

Ma, Jiaji↗

Characterizing Sub-Cohorts via Data Normalization and Representation Learning

The process of identifying a cohort of interest is a very challenging task. It requires manually inspecting many patient records of complex structure that might include medical coding errors and missing data. This paper presents a computational pipeline for refining the process of cohort selection based on medical concepts recorded in the electronic health records (EHRs). The pipeline extracts EHR data for a given cohort and normalizes this data using standard vocabularies. Then a stacked denoising autoencoder is used to embed the normalized patient vectors in a low dimensional space, where the patients are subsequently clustered into sub-cohorts. The goal is to represent the cohort in a standard format and abstract variants of sub-populations. As a use-case, we applied the pipeline to 1.8 million Veterans diagnosed with major depressive disorder (MDD), and identified four meaningful sub-cohorts using the features learned by the autoencoder. Then, each sub-cohort was explored using a set of keywords for interpretation.

Rush III, Everett↗

Machine-Learning Architecture for Ultrasonic Thermometry

Temperature distribution in solids can be inverted from the speed of sound (SOS) measurements, as has been shown feasible by timing the propagation of the excitation pulse and the train of echoes in ultrasonically segmented waveguides (WGs) and metal components. However, complicated geometries and closely-space echogenic features (EFs) create complex waveforms, from which the segmental time of flights (TOFs) are impossible to estimate using traditional methods. This work describes a machine learning architecture shown to extract temperature information from complex ultrasonic waveforms without explicit measurements of segmental TOFs. We accomplish this by using an autoencoder neural network (NN) to map ultrasonic waveforms into a low-dimensional latent space. A second NN then maps the latent space into unknown temperature distribution along the WG. The proposed architecture was tested in simulations and experimentally. The autoencoder accurately reconstructs the waveforms from their latent representation, several orders of magnitude lower in dimensionality. The obtained latent space was successfully mapped into the temperature of the propagation path.

John, Mason↗

Detecting False Data Injection Attacks in Smart Grids: A Semi-Supervised Deep Learning Approach

The dependence on advanced information and communication technology increases the vulnerability in smart grids under cyber-attacks. Recent research on unobservable false data injection attacks (FDIAs) reveals the high risk of secure system operation, since these attacks can bypass current bad data detection mechanisms. To mitigate this risk, this paper proposes a data-driven learning-based algorithm for detecting unobservable FDIAs in distribution systems. We use autoencoders for efficient dimension reduction and feature extraction of measurement datasets. Further, we integrate the autoencoders into an advanced generative adversarial network (GAN) framework, which successfully detects anomalies under FDIAs by capturing the unconformity between abnormal and secure measurements. Also, considering that the datasets collected from practical power systems are partially labeled due to expensive labeling costs and missing labels, the proposed method only requires a few labeled measurement data in addition to unlabeled data for training. Numerical simulations in three-phase unbalanced IEEE 13-bus and 123-bus distribution systems validate the detection accuracy and efficiency of this method.

97 MATHEMATICS AND COMPUTING↗

Domain-decomposition nonlinear manifold reduced order model

This software combines nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD) techniques. NM-ROMs, which utilize a shallow, sparse autoencoder trained with full order model (FOM) snapshot data, approximate the FOM state on a nonlinear manifold. These models offer advantages over linear-subspace ROMs (LS-ROMs) particularly in scenarios with slowly decaying Kolmogorov n-width. However, the training of NM-ROMs involves a number of parameters that scale with the size of the FOM, and storing high-dimensional FOM snapshots can significantly increase the cost of ROM training for extreme-scale problems. To mitigate these costs, the software employs DD to partition the FOM into smaller subdomains, computes NM-ROMs for each, and then integrates these to form a global NM-ROM. This strategy offers multiple benefits: it enables parallel training of subdomain NM-ROMs, reduces the number of parameters needed, decreases the dimensional requirements of subdomain FOM training data, and allows for customization to the unique characteristics of each FOM subdomain. The use of a shallow, sparse autoencoder architecture in each subdomain NM-ROM facilitates the application of hyper-reduction (HR), simplifying the nonlinear complexities and enhancing computational speed. This software marks the inaugural application of NM-ROM combined with HR to a DD problem. It features an algebraic DD reformulation of the FOM, training of NM-ROMs with HR for each subdomain, and employs a sequential quadratic programming (SQP) solver for the evaluation of the coupled global NMROM. The effectiveness of the DD NM-ROM with HR is numerically demonstrated on the 2D steady-state Burgers' equation, showing an order of magnitude improvement in accuracy over the DD LS-ROM with HR.

Diaz, AlejandroN↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Learning genetic perturbation effects with variational causal inference

Advances in sequencing technologies have enhanced the understanding of gene regulation in cells. In particular, Perturb-seq has enabled high-resolution profiling of the transcriptomic response to genetic perturbations at the single-cell level. This understanding has implications in functional genomics and potentially for identifying therapeutic targets. Various computational models have been developed to predict perturbational effects. While deep learning models excel at interpolating observed perturbational data, they tend to overfit in the lack of enough data and may not generalize well to unseen perturbations. In contrast, mechanistic models, such as linear causal models based on gene regulatory networks, hold greater potential for extrapolation, as they encapsulate regulatory information that can predict responses to unseen perturbations. However, their application has been limited to small studies due to overly simplistic assumptions, making them less effective in handling noisy, large-scale single-cell data. We propose a hybrid approach that combines a mechanistic causal model with variational deep learning, termed Single Cell Causal Variational Autoencoder (SCCVAE). The mechanistic model employs a learned regulatory network to represent perturbational changes as shift interventions that propagate through the learned network. SCCVAE integrates this mechanistic causal model into a variational autoencoder, generating rich, comprehensive transcriptomic responses. Our results indicate that SCCVAE exhibits superior performance over current state-of-the-art baselines for extrapolating to predict unseen perturbational responses. Additionally, for the observed perturbations, the latent space learned by SCCVAE allows for the identification of functional perturbation modules and simulation of single-gene knockdown experiments of varying penetrance, presenting a robust tool for interpreting and interpolating perturbational responses at the single-cell level.

59 BASIC BIOLOGICAL SCIENCES↗

SRF Cavity Instability Detection with Machine Learning at CEBAF

During the operation of CEBAF, one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. Identifying an unstable cavity out of the hundreds of cavities installed at CEBAF is difficult and time-consuming. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detect fast transient instabilities. A fast data acquisition system for the legacy SRF cavities is being developed which samples and reports at 5 kHz to allow for detection of transients. A prototype chassis has been installed and tested in CEBAF. An autoencoder based machine learning model is being developed to identify anomalous SRF cavity behavior. The model is presently being trained on the slow (1 Hz) data that is currently available, and a separate model will be developed and trained using the fast (5 kHz) DAQ data once it becomes available. This paper will discuss the present status of the new fast data acquisition system and results of testing the prototype chassis. This paper will also detail the initial performance metrics of the autoencoder model.

Carpenter, A.↗

An Introduction to Word Embeddings and Language Models

Language models have advanced at a phenomenal pace over the past decade. This document provides a short introduction to terminology, word embeddings (aka low-dimensional representations), and popular large-scale language models (LMs). Word embeddings are used to represent words as numerical vectors and are context-independent, meaning a word can only have a single representation (e.g., club can only be club sandwich, not golf club ). Language models can determine the probability of a given sequence of words occurring in a sentence and can provide context to distinguish between words and phrases that sound similar. LMs are context-dependent (e.g., club can be club sandwich or golf club ) and largely fall in two main classes – autoregressive and autoencoding models. Autoregressive models are pretrained on the classic language modeling task: guess the next token having read all the previous ones. Those models can be fine-tuned and achieve great results on many tasks, the most natural application is text generation. A typical example of such models is GPT, but others include GPT-2, GPT-3, CTLR, TRANSFORMER-XL, REFORMER, XLNET. Autoencoding models are pretrained by corrupting the input tokens in some way and trying to reconstruct the original sentence. They can be fine-tuned and achieve great results on many tasks such as text generation, but their most natural application is sentence classification or token classification. A typical example of such models is BERT, but others include ROBERTA, ALBERT, XML, XML-ROBERTA, FLAUBERT AND LONGFORMER.

97 MATHEMATICS AND COMPUTING↗

All Optical Neural Networks for Low Power Edge Computing

We developed a simplistic physics-based model of an all-optical neural network that mimics the encoder part of an autoencoder neural network for image compression. Our approach relies on the generation of a MATLAB-based model for both data compression and decompression and utilizes MATLAB's built-in autoencoder networks in combination with simple propagation of optical fields between layers constituting phase elements via Fourier transform. We optimize the phase elements using the particle swarm optimization technique and using our model, we demonstrate a compression ratio of 25% for 2828-pixel input images containing numeric digits from 0 to 9.

97 MATHEMATICS AND COMPUTING↗

Fission with Exotic Nuclei (Abbreviated Report)

Nuclear fission is a key mechanism involved in the synthesis of heavy elements in the Cosmos and is the primary explanation for the stability of superheavy elements. Nevertheless, our knowledge of fission remains extremely fragmented. Most experiments have been conducted only on a tiny number of stable actinide nuclei and are often incomplete, leading to gaps in our basic understanding of the process. For many radioactive isotopes, basic fission data such as the charge or mass distribution of the fragments is unknown. These gaps cannot always be filled by simulation alone. Common fission models contain too many free parameters and lack predictive power. In contrast, the fundamental theory of fission under development at LLNL is much more predictive, but its current computational cost is too high to be used extensively for data evaluations. A unique window of opportunity to resolve these limitations has recently opened: the U.S. nuclear science community is ramping up major experimental programs at the Facility for Rare Isotope Beams (FRIB, the DOE flagship facility in low-energy nuclear science), and techniques from machine learning have shown great potential to simplify the use of a fundamental, quantum-mechanical theory of fission. This project has two components. On the experimental side, we acquired and deployed at the HIGS facility a new dual Frisch-Grid ionization chamber to measure correlated fragment-mass, kinetic energy, and angular distributions of fission fragments from induced fission. This new device was used to perform measurements of charge, mass and total kinetic energy of fission fragments in the photofission of 238 U and eight gamma-ray beam energies between 6.2 and 13 MeV, which allowed extracting high-precision independent yields for this reaction. The device was also used to perform measurements of the same quantities in the neutron-induced fission of 234 U with monoenergetic beams of energy between 5 and 8 MeV. In parallel, we collaborated with a team at Commissariat à l’énergie atomique et aux énergies alternatives (CEA) to perform a series of measurements of fission yields in inverse kinematics for the two isotopes of 236 U and 240 Pu. The experiment took place at the Grand Accélérateur National d’Ions Lourds in France in June 2023. The deployment of the VAMOS spectrometer with a new array called PISTA allowed determining the excitation energy of the fissioning system within 1 Mega-electronvolts. The second component of the project involved using deep neural networks to build fast and reliable emulators of our current fission models. In an invited paper published in Frontier in Physics, we showed that autoencoders could successfully compress nuclear wavefunctions in nuclear density functional theory. We achieved a dimensionality reduction of the order of two orders of magnitude while keeping the error in the total energy to less than 0.01%. In a second paper submitted to Physical Review Letters in June 2023 with our collaborators at CEA, we showed that variational autoencoders can learn the collective degrees of freedom driving the fission process.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SRF Cavity Instability Detection with Machine Learning at CEBAF

During the operation of CEBAF, one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. Identifying an unstable cavity out of the hundreds of cavities installed at CEBAF is difficult and time-consuming. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detect fast transient instabilities. A fast data acquisition system for the legacy SRF cavities is being developed which samples and reports at 5 kHz to allow for detection of transients. A prototype chassis has been installed and tested in CEBAF. An autoencoder based machine learning model is being developed to identify anomalous SRF cavity behavior. The model is presently being trained on the slow (1 Hz) data that is currently available, and a separate model will be developed and trained using the fast (5 kHz) DAQ data once it becomes available. This paper will discuss the present status of the new fast data acquisition system and results of testing the prototype chassis. This paper will also detail the initial performance metrics of the autoencoder model.

Carpenter, A.↗

Low Energy LArTPC Signal Detection Using Anomaly Detection

Extracting low-energy signals from LArTPC detectors is useful, for example, for detecting supernova events or calibrating the energy scale with argon-39. However, it is difficult to efficiently extract the signals because of noise. We propose using a 1DCNN to select wire traces that have a signal. This efficiently suppresses the background while still being efficient for the signal. This is then followed by a 1D autoencoder to denoise the wire traces. At that point the signal waveform can be cleanly extracted. In order to make this processing efficient, we implement the two networks on an FPGA. In particular we use hls4ml to produce HLS from the Keras models for both the 1DCNN and the autoencoder. We deploy them on an AMD/Xilinx Alveo U55C using the Vitis software platform.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Efficiency of ML Anomaly Detection Triggers for Emerging Jets

Novel machine learning-based anomaly detection Level 1 (L1) triggers are currently under development at CMS, namely AXOL1TL and CICADA. The former employs a variational autoencoder, while the latter utilizes a convolutional autoencoder. These triggers aim to balance rate reduction with model independence, enabling the selection of potentially significant events that might be overlooked by traditional triggers relying on basic kinematic variable selections. Consequently, they have the potential to enhance signals indicative of physics beyond the Standard Model, such as those associated with emerging jets. Such signals are predicted by models featuring a composite dark sector where long-lived particles decay into Standard Model jets with displaced tracks and numerous vertices. This study evaluates the efficiency of these anomaly detection triggers in selecting events with emerging jets produced via the s-channel production of two dark quarks.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Real-Time Anomaly Detection for Beyond Standard Model Searches in ProtoDUNE Horizontal Drift

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events---making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving $31.9 \pm 0.2$\% ($26.6 \pm 0.2$\%) $\nu$ efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, $17.5 \pm 0.3$\% ($18.3 \pm 0.3$\%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, Cameron C. [Cincinnati U., RWC]↗