Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Uncovering electronic and geometric descriptors of chemical activity for metal alloys and oxides using unsupervised machine learning

Here, we show that unsupervised machine learning (ML) using principal component analysis (PCA) provides a straightforward pathway for developing accurate and interpretable electronic-structure descriptors of the chemical and catalytic properties of materials. We demonstrate the approach by finding chemisorption descriptors for metal alloys and surface oxygens on metals and metal oxides. In both cases, the principal component (PC) descriptors yield ML models that predict the material’s chemical properties with competitive accuracy compared to ML models built using established descriptors. Importantly, interpreting the electronic-structure patterns captured by each PC descriptor via signal reconstruction suggests potential design motifs for future electronic-structure descriptor design and allows us to identify links between a material’s geometric and catalytic properties. Ultimately, we show that the unsupervised ML approach provides a route to find electronic-structure descriptors of the catalytic properties of materials that readily connect to geometric structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.↗

Monitoring Fracture Hydromechanical Evolution in the Lab and Field Using Unsupervised Metric Learning

Fractures evolve in time through thermal‐hydraulic‐mechanical‐chemical (THMC) processes that alter their long‐range hydraulic transport properties and modify subsurface behavior and activities. The location of subsurface fractures makes it necessary to use remote sensing techniques such as passive or active seismic monitoring for fracture characterization. In this paper, we develop a machine learning approach to monitor the evolution of fracture properties using passive seismic sources in a laboratory setting and using active seismic monitoring from the Sanford Underground Research Facility in Lead, South Dakota, at a depth of 1.25 km in amphibolite rock during stimulation of natural fractures as well as during induced fracturing. The unsupervised metric learning technique applies tandem neural networks (twin (Siamese) or triplet) with contrastive loss and adaptive margins to track slowly varying systems for which class or similarity labels are not available. The approach adopts locality‐sensitive hashing to divide time‐ordered contiguous data into an arbitrary number of pseudo‐classes. Contrastive‐loss training with many hash bins generates an evolving latent‐space trajectory. This approach enables unsupervised metric learning for seismic data stacks under the condition of contiguous state sampling and slowly varying fracture properties. The displacement discontinuity theory provides a mechanistic foundation for the fracture‐dependent trajectories that are related to relaxation of fractures with time‐dependent specific stiffness responding to changes in stress or fluid saturation.

02 PETROLEUM↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

Detecting anomalous SRF cavity behavior with unsupervised learning

We present an unsupervised learning framework for detecting anomalous superconducting radio-frequency (SRF) cavity behavior at the Continuous Electron Beam Accelerator Facility (CEBAF), emphasizing its initial performance and effectiveness. Key to the system’s success was the development of data acquisition systems (DAQs) that capture fast-sampled, information-rich signals, essential for detecting transient effects. The approach involves creating daily cavity-specific models using principal component analysis to handle variations in rf signal behavior and mitigate performance degradation from data drift. This unsupervised method eliminates the need for expensive labeling by continuously updating models with recent data. Deployed and operational for 3 months before a scheduled shutdown, the system successfully identified several issues with DAQ signals, confirming its effectiveness. Despite access to only a fraction of CEBAF’s SRF cavity signals, the framework efficiently detected several instances requiring intervention, demonstrating a significant improvement over traditional, labor-intensive methods of manual plot inspection. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

Supervised and unsupervised machine learning of structural phases of polymers adsorbed to nanowires

Here, we identify configurational phases and structural transitions in a polymer nanotube composite by means of machine learning. We employ various unsupervised dimensionality reduction methods, conventional neural networks, as well as the confusion method, an unsupervised neural-network-based approach. We find neural networks are able to reliably recognize all configurational phases that have been found previously in experiment and simulation. Furthermore, we locate the boundaries between configurational phases in a way that removes human intuition or bias. This could be done before only by relying on preconceived, ad hoc order parameters.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

Exploring unsupervised top tagging using Bayesian inference

Recognizing hadronically decaying top-quark jets in a sample of jets, or even its total fraction in the sample, is an important step in many LHC searches for Standard Model and Beyond Standard Model physics as well. Although there exists outstanding top-tagger algorithms, their construction and their expected performance rely on Montecarlo simulations, which may induce potential biases. For these reasons we develop two simple unsupervised top-tagger algorithms based on performing Bayesian inference on a mixture model. In one of them we use as the observed variable a new geometrically-based observable \tilde{A}_{3} A ̃ 3 , and in the other we consider the more traditional \tau_{3}/\tau_{2} τ 3 / τ 2 N N -subjettiness ratio, which yields a better performance. As expected, we find that the unsupervised tagger performance is below existing supervised taggers, reaching expected Area Under Curve AUC \sim 0.80-0.81 ∼ 0.80 − 0.81 and accuracies of about 69% - − 75% in a full range of sample purity. However, these performances are more robust to possible biases in the Montecarlo that their supervised counterparts. Our findings are a step towards exploring and considering simpler and unbiased taggers.

Alvarez, Ezequiel↗

Development of Gamma Background Radiation Digital Twin with Machine Learning Algorithms: Application of Unsupervised Machine Learning to Detection of Anomalies and Nuisances in Gamma Background Radiation Environmental Screening Data

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for development of a digital twin of gamma radiation background, and for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. In one segment of work, we developed a gamma background estimation model using a Longshort term memory (LSTM) network for one-step CPS time series prediction. The LSTM model was validated with two data sets of measurements from two independent NaI detectors positioned on a mobile platform. The data sets contained background radiation only and no orphan isotope sources. The LSTM model was constructed and tested using data from one of the detectors. Performance of the LSTM model was validate through one-step prediction of CPS time series of another NaI detector without re-training. This approach allows to create a digital twin for nuclear background estimation. Using LSTM, it could be possible to detect a source through subtraction of the estimated counts from the measured background. In another segment of work, we investigated detection of gamma emitting sources in the presence of complex background using unsupervised machine learning. Spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

54 ENVIRONMENTAL SCIENCES↗

Exploring variability in seasonal average and extreme precipitation using unsupervised machine learning.

Focal Area(s): We will use unsupervised machine learning methods to identify and quantify the influence of large scale natural modes of climate variability to gain insight into the observed and simulated seasonal average and extreme precipitation changes. Science Challenge: A recent paper, led by co-PI Mark Risser, finds that although much of the variability in seasonal average and extreme precipitation over CONUS is unforced, the effect of large-scale modes of circulation variability (such as ENSO, AMO, PNA, etc.) can be detected and attributed. However, it is unclear whether or not unsupervised learning methods can (a) replicate this finding or (b) yield insight into possible nonlinear behavior that was not captured in the initial statistical analysis. Further work would entail extending this framework to other global land areas.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised probabilistic models for sequential Electronic Health Records

We develop an unsupervised probabilistic model for heterogeneous Electronic Health Record (EHR) data. Utilizing a mixture model formulation, our approach directly models sequences of arbitrary length, such as medications and laboratory results. This allows for subgrouping and incorporation of the dynamics underlying heterogeneous data types. The model consists of a layered set of latent variables that encode underlying structure in the data. These variables represent subject subgroups at the top layer, and unobserved states for sequences in the second layer. We train this model on episodic data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The resulting properties of the trained model generate novel insight from these complex and multifaceted data. In addition, we show how the model can be used to analyze sequences that contribute to assessment of mortality likelihood.

59 BASIC BIOLOGICAL SCIENCES↗

An unsupervised machine-learning checkpoint-restart algorithm using Gaussian mixtures for particle-in-cell simulations

We propose an unsupervised machine-learning checkpoint-restart (CR) algorithm for particle-in-cell (PIC) algorithms using Gaussian mixtures (GM). The algorithm compresses the particle population per spatial cell by constructing a velocity distribution function using GM. Particles are reconstructed at restart time by local resampling of the Gaussians. To guarantee fidelity of the CR process, we ensure the exact preservation of invariants such as charge, momentum, and energy for both compression and reconstruction stages, everywhere on the mesh. We also ensure the preservation of Gauss' law after particle reconstruction by exactly matching the density profile at restart time. As a result, the GM CR algorithm is shown to provide a clean, conservative restart capability while potentially affording orders of magnitude savings in input/output requirements. Here, we demonstrate the algorithm using a recently developed exactly energy- and charge-conserving PIC algorithm using both electrostatic and electromagnetic tests. The tests demonstrate not only a high-fidelity CR capability, but also its potential for enhancing the fidelity of the PIC solution for a given particle resolution.

97 MATHEMATICS AND COMPUTING↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

Uranium Oxide Synthetic Pathway Discernment through Unsupervised Morphological Analysis

We present a novel unsupervised machine learning method for quantitative representation of scanning electron micrographs and its applications and performance for nuclear forensic analysis of uranium ore concentrates. The method uses a vector quantizing variational autoencoder followed by a histogram operation to encode a micrograph into a single dimensional representation, called the latent vector. The method requires no extant labeling of the data and can be applied over large datasets of micrographs with minimal human interaction. The representations generated are broadly descriptive of each micrograph and the microstructure of the material imaged. In the case of uranium ore concentrate analysis, the representations were amenable to processing reagent and ore concentrate species classification with accuracy of 81:8%, which is competitive with state-of-the-art supervised networks. The representations were also used to classify previously unseen processing routes, were able to classify imaging parameters such as magnification (to 76:0% accuracy), were able to classify fine grained process parameters such as calcining temperature (to 74:4% accuracy), and their informatic properties indicate that they are generally descriptive of the image represented. This method can be applied across microstructure analysis fields to perform quantitative analysis without the need for labor intensive and possibly biased human analysis.

Scanning Electron Microscopy, Vector Quantizing Va↗