Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Beyond the hubble sequence – exploring galaxy morphology with unsupervised machine learning

We explore unsupervised machine learning for galaxy morphology analyses using a combination of feature extraction with a vector-quantized variational autoencoder (VQ-VAE) and hierarchical clustering (HC). We propose a new methodology that includes: (1) consideration of the clustering performance simultaneously when learning features from images; (2) allowing for various distance thresholds within the HC algorithm; (3) using the galaxy orientation to determine the number of clusters. This set-up provides 27 clusters created with this unsupervised learning that we show are well separated based on galaxy shape and structure (e.g. Sérsic index, concentration, asymmetry, Gini coefficient). These resulting clusters also correlate well with physical properties such as the colour–magnitude diagram, and span the range of scaling relations such as mass versus size amongst the different machine-defined clusters. When we merge these multiple clusters into two large preliminary clusters to provide a binary classification, an accuracy of $\sim 87{{\ \rm per\ cent}}$ is reached using an imbalanced data set, matching real galaxy distributions, which includes 22.7 per cent early-type galaxies and 77.3 per cent late-type galaxies. Comparing the given clusters with classic Hubble types (ellipticals, lenticulars, early spirals, late spirals, and irregulars), we show that there is an intrinsic vagueness in visual classification systems, in particular galaxies with transitional features such as lenticulars and early spirals. Based on this, the main result in this work is not how well our unsupervised method matches visual classifications and physical properties, but that the method provides an independent classification that may be more physically meaningful than any visually based ones.

79 ASTRONOMY AND ASTROPHYSICS↗

Estimating Permafrost Distribution Using Co‐Located Temperature and Electrical Resistivity Measurements

Abstract Assessing the lateral and vertical extent of permafrost is critical to understanding the fate of Arctic ecosystems under climate change. Yet, direct measurements of permafrost distribution and temperature are often limited to a small number of borehole locations. Here, we assess the use of co‐located shallow temperature and electrical resistivity tomography (ERT) measurements to estimate at high‐resolution the distribution of permafrost in three watersheds underlain by discontinuous permafrost. Synthetic modeling shows that co‐located temperature and ERT measurements allow for supervised classification schemes that provide 60% higher accuracy compared to unsupervised methods. Linking resistivity and size of the identified permafrost bodies to surface observations, we show that tall vegetation (>0.5 m) and gentle slopes (<15°) are related to warmer and smaller permafrost bodies, and a more frequent occurrence of taliks.

54 ENVIRONMENTAL SCIENCES↗

Low Activity Tritium Detection in CCDs Using Deep Learning Techniques

Here, this study explores the use of charge-coupled devices (CCDs) for detecting low-energy beta particles from tritium decay - a critical signal for nuclear safety, nuclear nonproliferation, and environmental monitoring. We employ a dual approach utilizing both measured CCD data and detailed Geant4 simulations. Our analysis compares classical techniques with advanced deep learning methods, including convolutional neural networks (CNNs), autoencoders trained exclusively on tritium data, and preliminary studies on boosted decision trees (BDTs). The CNN, trained on mixed signal/background datasets, demonstrates superior classification performance, while the autoencoder shows the potential of unsupervised, background-agnostic strategies when background characteristics are poorly defined. These results highlight the excellent sensitivity achievable thanks to the background rejection made possible by information-rich CCD data, paving the way for improved portable tritium monitoring.

Autoencoder↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Behavioral Segmentation and Clustering of Geospatial Trajectories

The rapid growth of global positioning system (GPS) devices has led to a corresponding increase in the size of GPS datasets. While these large GPS datasets contain a wealth of information about the behaviors of the moving objects in them, manual classification and anomaly detection are prohibitively time consuming. We utilize unsupervised machine learning techniques to first identify the behaviors for individual moving objects and then cluster those objects by their behavioral sequences. In this way, trajectories behaving unusually as well as common patterns of behavior are both detectable in large datasets without requiring an a priori definition of "unusual" or "common."

97 MATHEMATICS AND COMPUTING↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

A robust approach to 3D neuron shape representation for quantification and classification

We consider the problem of finding an accurate representation of neuron shapes, extracting sub-cellular features, and classifying neurons based on neuron shapes. In neuroscience research, the skeleton representation is often used as a compact and abstract representation of neuron shapes. However, existing methods are limited to getting and analyzing “curve” skeletons which can only be applied for tubular shapes. This paper presents a 3D neuron morphology analysis method for more general and complex neuron shapes. First, we introduce the concept of skeleton mesh to represent general neuron shapes and propose a novel method for computing mesh representations from 3D surface point clouds. A skeleton graph is then obtained from skeleton mesh and is used to extract sub-cellular features. Finally, an unsupervised learning method is used to embed the skeleton graph for neuron classification. Extensive experiment results are provided and demonstrate the robustness of our method to analyze neuron morphology.

3D neuron morphology↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

A network approach for multiscale catchment classification using traits

Abstract. The classification of river catchments into groups with similar biophysical characteristics is useful to understand and predict their hydrological behavior. The increasing availability of remote sensing and other large-scale geospatial datasets has enabled the use of advanced data-driven approaches to classify catchments using traits such as topography, geology, climate, land cover, land use, and human influence. Unsupervised clustering algorithms based on the Euclidean distance are commonly used for trait-based classification but are not suitable for highly dimensional data. In this study we present a new network-based method for multi-scale catchment classification, which can be applied to large datasets and used to determine the traits associated with different catchment groups. In this framework, two networks are analyzed in parallel: the first being where the nodes are traits and the second being where the nodes are catchments. In both cases, edges represent pairwise similarity, and a network cluster detection algorithm is used for the classification. The trait network is used to investigate redundancy in the trait data and to condense this information into a small number of interpretable categories. The catchments network is used to classify the catchments into clusters and to identify representative catchments for the different groups using the degree centrality metric. We apply this method to classify 9067 river catchments across the contiguous United States at both regional and continental scales using 274 non-categorical traits. At the continental scale, we identify 25 interpretable trait categories and 34 catchment clusters of sizes greater than 50. We find that catchments with similar trait categories are typically located in the same region, with different spatial patterns emerging among clusters dominated by natural and anthropogenic traits. We also find that the catchment clusters exhibit distinct hydrological behavior based on an analysis of streamflow indices. This network approach provides several advantages over traditional means of classification, including better separation of clusters, the use of alternate similarity metrics that are more suitable for highly dimensional data, and reducing redundancy in the trait information. The paired catchment–trait networks enable analysis of hydrological behavior using the dominant trait categories for each catchment cluster. The approach can be used at multiple spatial scales since the network topologies adjust automatically to reflect the trait patterns at the scale of investigation. Finally, the representative catchments identified as hub nodes in the network can be used to guide transferable observational and modeling strategies. The method is broadly applicable beyond hydrology for classification of other complex systems that utilize different types of trait datasets.

54 ENVIRONMENTAL SCIENCES↗

Intelligent Energy Optimizer for Residential Buildings

Demand-side management in the buildings is essential for meeting grid flexibility needs in a highly renewable energy scenario. Appliance load monitoring helps decision making for demand-side management by providing the information on operation status/power consumption from different appliances in the buildings. Nonintrusive load monitoring (NILM) is an attractive option for appliance load monitoring using because it has lower cost for sensors and helps mitigate privacy concerns. In this study, the team used an event detection technique followed by two different methods for event classification. The results from k-means clustering showed that the events from a single appliance are often distributed in multiple clusters. Thus, the unsupervised method of NILM using k-means clustering used in this study was not very suitable for load disaggregation. The results from NILM showed that the F1 score for event classification was 0.77 for a heat pump water heater and very low for other appliances using the rule-based classification.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Comparing weak- and unsupervised methods for resonant anomaly detection

Abstract Anomaly detection techniques are growing in importance at the Large Hadron Collider (LHC), motivated by the increasing need to search for new physics in a model-agnostic way. In this work, we provide a detailed comparative study between a well-studied unsupervised method called the autoencoder (AE) and a weakly-supervised approach based on the Classification Without Labels (CWoLa) technique. We examine the ability of the two methods to identify a new physics signal at different cross sections in a fully hadronic resonance search. By construction, the AE classification performance is independent of the amount of injected signal. In contrast, the CWoLa performance improves with increasing signal abundance. When integrating these approaches with a complete background estimate, we find that the two methods have complementary sensitivity. In particular, CWoLa is effective at finding diverse and moderately rare signals while the AE can provide sensitivity to very rare signals, but only with certain topologies. We therefore demonstrate that both techniques are complementary and can be used together for anomaly detection at the LHC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data driven discovery and quantification of hyperspectral leaf reflectance phenotypes across a maize diversity panel

Abstract Estimates of plant traits derived from hyperspectral reflectance data have the potential to efficiently substitute for traits, which are time or labor intensive to manually score. Typical workflows for estimating plant traits from hyperspectral reflectance data employ supervised classification models that can require substantial ground truth datasets for training. We explore the potential of an unsupervised approach, autoencoders, to extract meaningful traits from plant hyperspectral reflectance data using measurements of the reflectance of 2151 individual wavelengths of light from the leaves of maize ( Zea mays ) plants harvested from 1658 field plots in a replicated field trial. A subset of autoencoder‐derived variables exhibited significant repeatability, indicating that a substantial proportion of the total variance in these variables was explained by difference between maize genotypes, while other autoencoder variables appear to capture variation resulting from changes in leaf reflectance between different batches of data collection. Several of the repeatable latent variables were significantly correlated with other traits scored from the same maize field experiment, including one autoencoder‐derived latent variable (LV8) that predicted plant chlorophyll content modestly better than a supervised model trained on the same data. In at least one case, genome‐wide association study hits for variation in autoencoder‐derived variables were proximal to genes with known or plausible links to leaf phenotypes expected to alter hyperspectral reflectance. In aggregate, these results suggest that an unsupervised, autoencoder‐based approach can identify meaningful and genetically controlled variation in high‐dimensional, high‐throughput phenotyping data and link identified variables back to known plant traits of interest.

Tross, Michael C.↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗