Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Neural network reconstruction of the dense matter equation of state from neutron star observables

The Equation of State (EoS) of strongly interacting cold and hot ultra-dense QCD matter remains a major challenge in the field of nuclear astrophysics. With the advancements in measurements of neutron star masses, radii, and tidal deformabilities, from electromagnetic and gravitational wave observations, neutron stars play an important role in constraining the ultra-dense QCD matter EoS. Here, in this work, we present a novel method that exploits deep learning techniques to reconstruct the neutron star EoS from mass-radius (M-R) observations. We employ neural networks (NNs) to represent the EoS in a model-independent way, within the range of ~1-7 times the nuclear saturation density. The unsupervised Automatic Differentiation (AD) framework is implemented to optimize the EoS, so as to yield through TOV equations, an M-R curve that best fits the observations. We demonstrate that this method works by rebuilding the EoS on mock data, i.e., mass-radius pairs derived from a randomly generated polytropic EoS. The reconstructed EoS fits the mock data with reasonable accuracy, using just 11 mock M-R pairs observations, close to the current number of actual observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Single Channel Infrasound Detection Using Machine Learning

Infrasound, low frequency sound less than 20 Hz, is generated by both natural and anthropogenic sources. Infrasound sensors measure pressure fluctuations only in the vertical plane and are single channel. However, the most robust infrasound signal detection methods rely on stations with multiple sensors (arrays), despite the fact that these are sparse. Automated methods developed for seismic data, such as short-term average to long-term average ratio (STA/LTA), often have a high false alarm rate when applied to infrasound data. Leveraging single channel infrasound stations has the potential to decrease signal detection limits, though this cannot be done without a reliable detection method. Therefore, this report presents initial results using (1) a convolutional neural network (CNN) to detect infrasound signals and (2) unsupervised learning to gain insight into source type.

47 OTHER INSTRUMENTATION↗

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Physics-constrained superresolution diffusion for six-dimensional phase space diagnostics

Adaptive physics-constrained superresolution diffusion is developed for noninvasive virtual diagnostics of the six-dimensional (6D) phase space density of charged particle beams. An adaptive variational autoencoder embeds initial beam condition images and scalar measurements to a low-dimensional latent space from which a 32 6 pixel 6D tensor representation of the beam's 6D phase space density is generated. Projecting from a 6D tensor generates physically consistent two-dimensional projections. Physics-guided superresolution diffusion transforms low-resolution images of the 6D density to high resolution 256 × 256 pixel images. Unsupervised adaptive latent space tuning enables tracking of time-varying beams without knowledge of time-varying initial conditions. The method is demonstrated with experimental data and multiparticle simulations at the HiRES UED. The general approach is applicable to a wide range of complex dynamic systems evolving in high-dimensional phase space. The method is shown to be robust to distribution shift without retraining. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Classifying biophysical subpopulations of insulin secretory granules using quantitative whole-cell structure analysis

Pancreatic beta cells contain insulin secretory granules (ISGs), organelles where proinsulin is converted into insulin. As ISGs mature, they undergo extensive biophysical remodeling, producing a spectrum of subpopulations with heterogeneous molecular and spatial characteristics. However, systematic methods to define ISG subpopulations remain underdeveloped. To address this gap in knowledge, we employed soft X-ray tomography (SXT), which can quantitatively measure the biochemical density of ISGs within whole beta cells. Using unsupervised clustering, we classified subpopulations based on molecular density, size, and spatial positioning. Across different insulin secretory stimuli, we observed shifts toward mature and releasable subtypes, demonstrating that exogenous signals can dynamically remodel ISG subpopulation distributions. We extended this methodology to primary beta cells characterized using volume electron microscopy (vEM). Integrating subpopulations from SXT and vEM uncovered insights inaccessible by a single method in isolation. This strategy establishes a framework for defining therapeutic approaches aimed at enriching physiologically beneficial ISG subpopulations.

dense-core granules↗

Cy-Phy ADS: Cyber Physical Anomaly Detection Framework for EV Charging Systems

Today’s large-scale Electric Vehicle (EV) infrastructures are heavily dependent on information communication technologies to maintain their operation and to support communication within sub-system components as well as the outside world. These technologies are vulnerable to various cyber and physical threats. Timely identification and mitigation of these threats are critical for improving human safety, avoiding economic losses, and preventing catastrophic system failures. By addressing this, our work presents a ResNet Autoencoder (AE) based Cyber-Physical Anomaly Detection System (Cy-Phy ADS) for detecting anomalies in EV Controller Area Network (CAN) protocol communication. It consists of four main components: Cyber-Physical Feature Extractor, ResNet AE-based Anomaly Detection Framework, Cyber-Physical Health Metric (CPHM), and Visualization Dashboard. The presented framework was trained and tested using CAN data collected from the EV charging system testbed at the Idaho National Laboratory. The presented Cy-Phy ADS compared against six widely used unsupervised anomaly detection algorithms: One Class Support Vector Machine (OCSVM), Variational Autoencoder (VAE), LSTM Autoencoder (LSTM AE), Isolation Forest (IForest), Principle Component Analysis (PCA) and Local Outlier Factor (LOF). Here the presented approach showed the highest accuracy among the compared methods. Further, the proposed approach showed comparable performance in terms of precision, F1, and False positive rate. It also showed the lowest training and inference time compared to the neural network-based baseline algorithms compared against with. Additionally, the Cy-Phy ADS has advantages such as unsupervised training, the ability to provide a holistic metric for system health characterization, and non-linear feature extraction.

99 GENERAL AND MISCELLANEOUS↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Earthquake Phase Association Using a Bayesian Gaussian Mixture Model

Earthquake phase association algorithms aggregate picked seismic phases from a network of seismometers into individual seismic events and play an important role in earthquake monitoring and research. Dense seismic networks and improved phase picking methods produce massive seismic phase datasets, particularly for earthquake swarms and aftershocks occurring closely in time and space, making phase association a challenging problem. Here, we present a new association method, the Gaussian Mixture Model Association (GaMMA), that combines the Gaussian mixture model with earthquake location, origin time, and magnitude estimation. We treat earthquake phase association as an unsupervised clustering problem in a probabilistic framework, where each earthquake corresponds to a cluster of P and S phases with a hyperbolic moveout of arrival times and a decay of amplitude with distance. We use the multivariate Gaussian distribution to model the collection of phase picks of an event; and the mean of the multivariate Gaussian distribution is given by the predicted arrival time and amplitude from the causative event. We carry out the pick assignment to each earthquake and determine earthquake source parameters (i.e., earthquake location, origin time, and magnitude) under the maximum likelihood criterion using the Expectation-Maximization algorithm. The GaMMA method does not require typical association steps of other algorithms, such as grid-search or supervised training. The results for both synthetic tests and for the 2019 Ridgecrest earthquake sequence show that GaMMA effectively associates phases from a temporally and spatially dense earthquake sequence while producing useful estimates of earthquake location and magnitude.

58 GEOSCIENCES↗

Coordinate-Based Seismic Interpolation in Irregular Land Survey: A Deep Internal Learning Approach

Physical and budget constraints often result in irregular sampling, which complicates accurate subsurface imaging. Preprocessing approaches, such as missing trace or shot interpolation, are typically employed to enhance seismic data in such cases. Recently, deep learning has been used to address the trace interpolation problem at the expense of large amounts of training data to adequately represent typical seismic events. Nonetheless, most research in this area has focused on trace reconstruction, with little attention having been devoted to shot interpolation. Furthermore, existing methods assume regularly spaced receivers/sources failing in approximating seismic data from real (irregular) surveys. This work presents a novel shot gather interpolation approach which uses a continuous coordinate-based representation of the acquired seismic wavefield parameterized by a neural network. The proposed unsupervised approach, which we call coordinate-based seismic interpolation (CoBSI), enables the prediction of specific seismic characteristics in irregular land surveys without using external data during neural network training. Importantly, experimental results on real and synthetic 3-D data validate the ability of the proposed method to estimate continuous smooth seismic events in the time-space and frequency-wavenumber domains, improving sparsity or low-rank-based interpolation methods.

58 GEOSCIENCES↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Protein Conformational States—A First Principles Bayesian Method

Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naïve Bayes classifier from the machine learning community for use on atom-to-atom pairwise contacts. The result is an unsupervised learning algorithm that samples a ‘distribution’ over potential classification schemes. We apply the classifier to a series of test structures and one real protein, showing that it identifies the conformational transition with >95% accuracy in most cases. A nontrivial feature of our adaptation is a new connection to information entropy that allows us to vary the level of structural detail without spoiling the categorization. This is confirmed by comparing results as the number of atoms and time-samples are varied over 1.5 orders of magnitude. Further, the method’s derivation from Bayesian analysis on the set of inter-atomic contacts makes it easy to understand and extend to more complex cases.

97 MATHEMATICS AND COMPUTING↗

STSR-INR: Spatiotemporal super-resolution for multivariate time-varying volumetric data via implicit neural representation

Implicit neural representation (INR) has surfaced as a promising direction for solving different scientific visualization tasks due to its continuous representation and flexible input and output settings. We present STSR-INR, an INR solution for generating simultaneous spatiotemporal super-resolution for multivariate time-varying volumetric data. Inheriting the benefits of the INR-based approach, STSR-INR supports unsupervised learning and permits data upscaling with arbitrary spatial and temporal scale factors. Unlike existing GAN- or INR-based super-resolution methods, STSR-INR focuses on tackling variables or ensembles and enabling joint training across datasets of various spatiotemporal resolutions. Here we achieve this capability via a variable embedding scheme that learns latent vectors for different variables. In conjunction with a modulated structure in the network design, we employ a variational auto-decoder to optimize the learnable latent vectors to enable latent-space interpolation. To combat the slow training of INR, we leverage a multi-head strategy to improve training and inference speed with significant speedup. We demonstrate the effectiveness of STSR-INR with multiple scalar field datasets and compare it with conventional tricubic+linear interpolation and state-of-the-art deep-learning-based solutions (STNet and CoordNet).

97 MATHEMATICS AND COMPUTING↗

Event-Based Analysis of Solar Power Distribution Feeder Using Micro-PMU Measurements

Solar distribution feeders are commonly used in solar farms that are integrated into distribution substations. In this paper, we focus on a real-world solar distribution feeder and conduct an event-based analysis by using micro-PMU measurements. The solar distribution feeder of interest is a behind-the-meter solar farm with a generation capacity of over 4 MW that has about 200 low-voltage distributed photovoltaic (PV) inverters. The event-based analysis in this study seeks to address the following practical matters. First, we conduct event detection by using an unsupervised machine learning approach. For each event, we determine the event’s source region by an impedancebased analysis, coupled with a descriptive analytic method. We segregate the events that are caused by the solar farm, i.e., locallyinduced events, versus the events that are initiated in the grid, i.e., grid-induced events, which caused a response by the solar farm. Second, for the locally-induced events, we examine the impact of solar production level and other significant parameters to make statistical conclusions. Third, for the grid-induced events, we characterize the response of the solar farm; and make comparisons with the response of an auxiliary neighboring feeder to the same events. Fourth, we scrutinize multiple specific events; such as by revealing the dynamics to the control system of the solar distribution feeder. The results and discoveries in this study are informative to utilities and solar power industry.

14 SOLAR ENERGY↗

SRF cavity instability detection with machine learning at CEBAF

During the operation of the Continuous Electron Beam Accelerator Facility (CEBAF), one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detectfast transient instabilities during these trip events. These challenges make the identification of an unstable cavity out of the hundreds installed at CEBAF a difficult and time-consuming task. To tackle these issues, a fast data acquisition system (DAQ) for the legacy SRF cavities has been developed, which records the sample at 5 kHz. An unsupervised learning framework has been developed to identify anomalous SRF cavity behavior. We will discuss the present status of the DAQ system and our framework, along with recent successes in detecting anomalous cavity behavior. Overall, our method offers a practical solution for identifying unstable SRF cavities, contributing to increased beam availability and machine reliability.

Accelerator Physics↗