Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Comparative Analysis of TRGBs (CATs) from Unsupervised, Multi-halo-field Measurements: Contrast is Key

The tip of the red giant branch (TRGB) is an apparent discontinuity of the luminosity function (LF) due to the end of the red giant evolutionary phase and is used to measure distances in the local universe. In practice, tip localization via edge detection response (EDR) relies on several methods applied on a case-by-case basis. It is hard to evaluate how individual choices affect a distance estimation using only a single host field while also avoiding confirmation bias. To devise a standardized approach, we compare unsupervised, algorithmic analyses of the TRGB in multiple halo fields per galaxy. We first optimize methods for the lowest field-to-field dispersion, including spatial filtering, smoothing, and weighting of LF, color band selection, and tip selection based on the number of likely RGB stars and the ratio of stars below versus above the tip (R). We find R, which we call the tip contrast, to be the most important indicator of the quality of EDR measurements; higher R selection can decrease field-to-field dispersion. Further, since R is found to correlate with the age or metallicity of the stellar population based on theoretical modeling, it might result in a displacement of the detected tip magnitude. We find a tip-contrast relation with a slope of -0.023 ± 0.0046 mag/ratio, an ~5σ result that can be used to correct these variations in the detections. When using TRGB to establish a distance ladder, consistent TRGB standardization using tip-contrast relation across rungs is vital to make robust cosmological measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

AutoPhaseNN: unsupervised physics-aware deep learning of 3D nanoscale Bragg coherent diffraction imaging

Abstract The problem of phase retrieval underlies various imaging methods from astronomy to nanoscale imaging. Traditional phase retrieval methods are iterative and are therefore computationally expensive. Deep learning (DL) models have been developed to either provide learned priors or completely replace phase retrieval. However, such models require vast amounts of labeled data, which can only be obtained through simulation or performing computationally prohibitive phase retrieval on experimental datasets. Using 3D X-ray Bragg coherent diffraction imaging (BCDI) as a representative technique, we demonstrate AutoPhaseNN, a DL-based approach which learns to solve the phase problem without labeled data. By incorporating the imaging physics into the DL model during training, AutoPhaseNN learns to invert 3D BCDI data in a single shot without ever being shown real space images. Once trained, AutoPhaseNN can be effectively used in the 3D BCDI data inversion about 100× faster than iterative phase retrieval methods while providing comparable image quality.

36 MATERIALS SCIENCE↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗

Topological and magnetic properties of the interacting Bernevig-Hughes-Zhang model

We investigate the effects of electronic correlations on the Bernevig-Hughes-Zhang model using the real-space density matrix renormalization group (DMRG) algorithm. We introduce a method to probe topological phase transitions in systems with strong correlations using DMRG, substantiated by an unsupervised machine learning methodology that analyzes the orbital structure of the real-space edges. Including the full multi-orbital Hubbard interaction term, we construct a phase diagram as a function of a gap parameter (m) and the Hubbard interaction strength (U) via exact DMRG simulations on N×4 cylinders. Our analysis confirms that the topological phase persists in the presence of interactions, consistent with previous studies, but it also reveals an intriguing phase transition from a paramagnetic to a stripey antiferromagnetic topological insulator. The combination of the magnetic structure factor, strength of magnetic moments, and the orbitally resolved density, provides real-space information on both topology and magnetism in a strongly correlated system.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Sparse and Low Rank Penalized Signal Decomposition Model with Constraints: Anomaly Detection in PV Systems

Recently, robust PCA has seen its wide application in various industries for its ability to perform the task of anomaly detection. The essence of robust PCA approach is to break down the signal into a low rank component and sparse component. In many applications, a simple breakdown of the signal without accounting for the signs of low rank components and sparse components would violate the physical constraints of the decomposed signal. In addition, often times, the signals in the real world collected for a long duration has smooth changes within a day and between days. As an example, the power signals collected in a photovoltaic (PV) system are cyclostationary, exhibiting these characteristics. Neglecting the smoothness of signals would result in miss detection of anomalous signals which are smooth within a day but non-smooth between days and vice versa. In this paper, we developed a signal decomposition approach for the purpose of anomaly detection based on the idea of low rank and sparse decomposition taking into consideration the signs of the decomposed low rank and sparse components and the within-day and between-day smooth changes in the original signals. The proposed unsupervised approach for fault detection eliminates the need for faulty samples required by other machine learning methods. It does not require the full I-V characteristics to work. Furthermore, there is no need for complex modelling of PV systems as in the case of power loss analysis. Using Monte Carlo simulations, we demonstrate the ability of our proposed approach for detecting anomalies of different duration and severity in PV systems.

14 SOLAR ENERGY↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Nondestructive Damage Detection of Concrete With Alkali-Silica Reactions Using Coda Wave and Anomaly Detection

An anomaly detection model for early damage detection for concrete structures undergoing alkali-silica reaction (ASR) is presented. It is difficult to detect ASR initiation and early damage without a reference expansion measurement. Coda waves, or the multiply scattered portion of ultrasonic waves, have been found to be indicative of small changes in complex material such as concrete. The relationship between concrete damage and relative velocity change and decorrelation of coda waves has been studied, but a generalized model which detects when damage occurs in a concrete structure is still lacking. The presented method uses features extracted from coda waves to detect early damage in concrete structures. The model uses unsupervised learning and only requires data from undamaged structures for training. During the training process, the reconstruction error of the training data is minimized. When the data collected from damaged concrete structures is used as an input of the model, it returns high reconstruction errors that indicate the occurrence of damage in the structures. The performance of the model is validated using experimental studies and has been shown to generalize across two different ASR specimens.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unsupervised learning from three-component accelerometer data to monitor the spatiotemporal evolution of meso-scale hydraulic fractures

Enhanced geothermal systems can provide a substantial share of the global energy demand. There exist several hurdles in the engineering implementations of such geothermal systems. One such hurdle is the accurate monitoring of the fracture networks created in subsurface through hydraulic stimulation of these systems. Micro seismicity associated with the stimulation is the primary means to locate the event hypocenters for estimating the stimulated rock volume. Existing methods for location the hypocenters are restricted to only the highest amplitude impulsive signals that are simultaneously detected on several sensors. Consequently, a large portion (usually ~99%) of the measurements are left unused. In this paper, an unsupervised manifold-approximation followed by clustering of 3-component accelerometer data is used to analyze the seismicity recorded on a monitoring well. With this method, a larger portion of the measured signal is used for the monitoring of the hydraulic fracture network. We analyze the EGS Collab experiment 1 microseismic data, recorded at the Sanford Underground Research Facility, South Dakota. Using the data from a single three-component accelerometer, the polarization features viz. Azimuth, incidence, rectilinearity, and planarity are used as inputs for the unsupervised manifold approximation followed by clustering. Our study shows that density-based clusters in the projected 3D space correspond to distinct types of hydraulically fractured zones around the injection point. Finally, we show that the temporal evolution of these clusters can be used to track fracture creation and propagation.

58 GEOSCIENCES↗

Data-Driven Smoothers for Extreme-Scale Computing

Patch-based relaxation refers to a family of methods for solving linear systems which partitions the matrix into smaller pieces often corresponding to groups of adjacent degrees of freedom residing within patches of the computational domain. The two most common families of patch-based methods are block-Jacobi and Schwarz methods, where the former typically corresponds to non-overlapping domains and the later implies some overlap. We focus on cases where each patch consists of the degrees of freedom on a finite element method mesh cell. Patch methods often capture complex local physics much more effectively than simpler point-smoothers such as Jacobi; however, forming, inverting, and applying each patch can be prohibitively expensive in terms of both storage and computation time. To this end, we propose several approaches for performing analysis on these patches and constructing a reduced representation. The compression techniques rely on either matrix norm comparisons or unsupervised learning via a clustering approach. We illustrate how it is frequently possible to retain/factor less than 5% of all patches and still develop a method that converges only a little slower than when all patches are stored/factored.

97 MATHEMATICS AND COMPUTING↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗