Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Unsupervised machine learning for unbiased chemical classification in X-ray absorption spectroscopy and X-ray emission spectroscopy

Here we report a comprehensive computational study of unsupervised machine learning for extraction of chemically relevant information in X-ray absorption near edge structure (XANES) and in valence-to-core X-ray emission spectra (VtC-XES) for classification of a broad ensemble of sulphorganic molecules. By progressively decreasing the constraining assumptions of the unsupervised machine learning algorithm, moving from principal component analysis (PCA) to a variational autoencoder (VAE) to t-distributed stochastic neighbour embedding (t-SNE), we find improved sensitivity to steadily more refined chemical information. Surprisingly, when embedding the ensemble of spectra in merely two dimensions, t-SNE distinguishes not just oxidation state and general sulphur bonding environment but also the aromaticity of the bonding radical group with 87% accuracy as well as identifying even finer details in electronic structure within aromatic or aliphatic sub-classes. We find that the chemical information in XANES and VtC-XES is very similar in character and content, although they unexpectedly have different sensitivity within a given molecular class. We also discuss likely benefits from further effort with unsupervised machine learning and from the interplay between supervised and unsupervised machine learning for X-ray spectroscopies. Our overall results, i.e., the ability to reliably classify without user bias and to discover unexpected chemical signatures for XANES and VtC-XES, likely generalize to other systems as well as to other one-dimensional chemical spectroscopies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unified architecture for data-driven metadata tagging of building automation systems

This article presents a Unified Architecture (UA) for automated point tagging of Building Automation System (BAS) data, based on a combination of data-driven approaches. Advanced energy analytics applications—including fault detection and diagnostics and supervisory control—have emerged as a significant opportunity for improving the performance of our built environment. Effective application of these analytics depends on harnessing structured data from the various building control and monitoring systems, but typical BAS implementations do not employ any standardized metadata schema. While standards such as Project Haystack and Brick Schema have been developed to address this issue, the process of structuring the data, i.e., tagging the points to apply a standard metadata schema, has, to date, been a manual process. This process is typically costly, labor-intensive, and error-prone. In this work we address this gap by proposing a UA that automates the process of point tagging by leveraging the data accessible through connection to the BAS, including time-series data and the raw point names. The UA intertwines supervised classification and unsupervised clustering techniques from machine learning and leverages both their deterministic and probabilistic outputs to inform the point tagging process. Furthermore, we extend the UA to embed additional input and output data-processing modules that are designed to address the challenges associated with the real-time deployment of this automation solution. We test the UA on two datasets for real-life buildings: (i) commercial retail buildings and (ii) office buildings from the National Renewable Energy Laboratory (NREL) campus. We report the proposed methodology correctly applied 85–90% and 70–75% of the tags in each of these test scenarios, respectively for two significantly different building types used for testing UA's fully-functional prototype. The proposed UA, therefore, offers promising approach for automatically tagging BAS data as it reaches close to 90% accuracy. Further building upon this framework to algorithmically identify the equipment type and their relationships is an apt future research direction to pursue.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Robust Group Subspace Recovery: A New Approach for Multi-Modality Data Fusion

Robust Subspace Recovery (RoSuRe) algorithm was recently introduced as a principled and numerically efficient algorithm that unfolds underlying Unions of Subspaces (UoS) structure, present in the data. The union of Subspaces (UoS) is capable of identifying more complex trends in data sets than simple linear models. In this work, we build on and extend RoSuRe to prospect the structure of different data modalities individually. We propose a novel multi-modal data fusion approach based on group sparsity which we refer to as Robust Group Subspace Recovery (RoGSuRe). Relying on a bi-sparsity pursuit paradigm and non-smooth optimization techniques, the introduced framework learns a new joint representation of the time series from different data modalities, respecting an underlying UoS model. We subsequently integrate the obtained structures to form a unified subspace structure. The proposed approach exploits the structural dependencies between the different modalities data to cluster the associated target objects. The resulting fusion of the unlabeled sensors’ data from experiments on audio and magnetic data has shown that our method is competitive with other state of the art subspace clustering methods. The resulting UoS structure is employed to classify newly observed data points, highlighting the abstraction capacity of the proposed method.

47 OTHER INSTRUMENTATION↗

International Symposium on Remote Sensing of Environment, 9th, University of Michigan, Ann Arbor, Mich., April 15-19, 1974, Proceedings. Volumes 1, 2 & 3

The present work gathers together numerous papers describing the use of remote sensing technology for mapping, monitoring, and management of earth resources and man's environment. Studies using various types of sensing equipment are described, including multispectral scanners, radar imagery, spectrometers, lidar, and aerial photography, and both manual and computer-aided data processing techniques are described. Some of the topics covered include: estimation of population density in Tokyo districts from ERTS-1 data, a clustering algorithm for unsupervised crop classification, passive microwave sensing of moist soils, interactive computer processing for land use planning, the use of remote sensing to delineate floodplains, moisture detection from Skylab, scanning thermal plumes, electrically scanning microwave radiometers, oil slick detection by X-band synthetic aperture radar, and the use of space photos for search of oil and gas fields. Individual items are announced in this issue.

Source record↗

Cooperative processes in image segmentation

Research into the role of cooperative, or relaxation, processes in image segmentation is surveyed. Cooperative processes can be employed at several levels of the segmentation process as a preprocessing enhancement step, during supervised or unsupervised pixel classification and, finally, for the interpretation of image segments based on segment properties and relations.

Davis, L. S.↗

Classification Of Terrain In Polarimetric SAR Images

Two algorithms processing polarimetric synthetic-aperture-radar data found effective in assigning various parts of SAR images to classes representing different types of terrain. Partially automate interpretation of SAR imagery, reducing amount of photointerpretation needed and putting whole interpretation process on more quantitative and systematic basis. First algorithm implements Bayesian classification scheme "supervised" by use of training data. Second algorithm implements classification procedure unsupervised.

Van Zyl, Jakob J.↗

Characterizing Detailed Grain Shape and Size Distribution Properties of Lunar Regolith

Introduction: As the nation prepares to return to the Moon, there is an increasing need for testing tools, instruments, and equipment in simulated environments on Earth to ensure successful operations during lunar missions. Regolith will affect all aspects of future lunar missions, from plume interactions during landing to space suit and tool design [1]. Because of this, it is important to understand the grain shape and size properties of lunar regolith and how those influence regolith behavior in order to prepare for these missions. This knowledge is also vital to create more accurate lunar regolith simulants for testing equipment in a lunar environment. While particle size analyses have been performed on most Apollo soils using simple sieving, shape has only been crudely addressed [2]. This work analyzes 4 lunar regolith samples to provide a better understanding of these size and shape parameters and will provide a more accurate baseline of data to create high-fidelity lunar regolith simulants. Methods: New technologies exist today that are capable of measuring size and shape simultaneously for hundreds of thousands of particles in a single measurement. We conducted a rigorous analysis of the particle size distribution (PSD) as well as the size-dependent 2D and 3D shape parameters of lunar regolith samples of different compositions and maturity levels. This analysis was done using a Microtrac SYNC which provides a unique combination of tri-laser diffraction and Dynamic Image Analysis (DIA). Sample Selection. Four regolith samples were selected for analysis based on maturity level and lunar terrain type: • 10084 – Mature high-Ti mare regolith • 15601 – Immature low-Ti mare regolith • 64501 – Mature highland regolith • 67461 – Immature highland regolith Sample Analysis. After receiving the samples, each sample was imaged with an optical microscope (Figure 1). To obtain 2D and 3D particle size and shape, 0.1 g of each sample was then added to the SYNC for analysis, which outputs more than 30 size and shape parameters for each individual grain as well as the complete size distribution from 0.01-2000 μm by blending laser diffraction and DIA together. For the 0.1 g sample masses requested, ~100,000 grains per sample were captured by DIA. Results: From the PSD analysis, it can be seen that the average particle size of 10084 is ~24.5 µm which is much smaller than the average particle size of the Apollo sample collection (~72 µm) [3], while the average particle size of samples 15601, 64501, and 67461 are larger than the Apollo sample average at ~106 µm, ~103 µm, and ~118.5 µm respectively. The size and shape measurements for the samples output ~30 parameters, four of which were focused on for this study: sphericity, aspect ratio, roundness, and concavity (Figure 3). However, after analyzing the plots it was found that only sphericity and aspect ratio showed differences between the samples. These two parameters are measured on a scale from 0 to 1, with 1 being a perfect sphere with equal dimensions. The results show that the sphericity values are slightly high-er for the mature samples 10084 and 64501 (~0.96) than they are for samples 15601 and 67461 (~0.95) (Figure 4). The aspect ratio values are slightly lower for samples 64501 (~0.7) and 67461 (~0.74) than they are for samples 10084 (~0.76) and 15601 (~0.8) (Figure 5). Discussion: Mature regoliths are those that have been exposed to micrometeorite and solar wind bombardment for long periods of time, breaking up particles and causing them to become more rounded [4]. Therefore, the smaller particle sizes and higher sphericity values for samples 10084 and 64501 are expected due to their higher maturity level compared to 15601 and 67461. However, the aspect ratio values are not dependent on maturity and are instead dependent on terrain type. The lower aspect ratio values for the high-land samples is potentially due to the higher plagioclase content, which occurs in elongated particles and does not break down as easily as the pyroxenes or olivines that are present in the mare. Conclusion and Future Work: The results of this work provide a baseline of high-quality data that will contribute to the creation of high-fidelity lunar simulants and will greatly benefit NASA’s efforts of establishing a human presence on the Moon. Future work includes performing unsupervised image classification on the ~105 particle images per sample in order to identify different classes of grains. These grain classes can then be linked to detailed shape properties, and the relative abundance of each class in the samples can be compared. Acknowledgments: We would like to thank the Extraterrestrial Materials Analysis Group (ExMAG) and the Astromaterials Allocation Review Board (AARB) for allocating lunar regolith samples 10084,9010, 15601,365, 64501,249, and 67461,171 for this work. References: [1] Taylor, L. A. et al. (2005) AIAA #2510. [2] Katagiri et al. (2015) ASCE. [3] Carrier III, W. D. (2005) Lunar Geotech Institute Tech Report. [4] McKay, D. S. et al. (1991) The Lunar Sourcebook, Chapter 7.

S R Deitrick↗

Lake classification in Vermont

In order to comply with the Federal Clean Water Act and, in so doing, develop a procedure to periodically update the classification, the State of Vermont evaluated the ability of LANDSAT to detect general water quality and specific water quality parameters in Vermont lakes. Unsupervised and supervised classifications as well as regression analyses were used to examine LANDSAT data from Lake Champlain and from four small nearby lakes. Unsupervised and supervised classifications were found to be of somewhat limited value. Regression analyses revealed a good correlation between depth-integrated total phosphorus concentrations and LANDSAT band 4 data (r2= 0.92) and between Secchi disk transparencies and LANDSAT band 4 data (r2 - 0.85). No correlation was found between depth-integrated chlorophyll-a samples and LANDSAT data. Vermont is expanding this LANDSAT evaluation to include the remaining lakes in the state greater than twenty acres and steps are being taken to incorporate LANDSAT into the state's ongoing water quality monitoring programs.

Garrison, V.↗

Solar forecasting using machine learned cloudiness classification

Methods and systems for predicting irradiance include learning a classification model using unsupervised learning based on historical irradiance data. The classification model is updated using supervised learning based on an association between known cloudiness states and historical weather data. A cloudiness state is predicted based on forecasted weather data. An irradiance is predicted using a regression model associated with the cloudiness state.

Hamann, Hendrik F.↗

Universal and interpretable classification of atomistic structural transitions via unsupervised graph learning

Materials processing often occurs under extreme dynamic conditions leading to a multitude of unique structural environments. These structural environments generally occur at high temperatures and/or high pressures, often under non-equilibrium conditions, which results in drastic changes in the material's structure over time. Computational techniques, such as molecular dynamics simulations, can probe the atomic regime under these extreme conditions. However, characterizing the resulting diverse atomistic structures as a material undergoes extreme changes in its structure has proved challenging due to the inherently non-linear relationship between structures as large-scale changes occur. Here, we introduce SODAS++, a universal graph neural network framework, that can accurately and intuitively quantify the atomistic structural evolution corresponding to the transition between any two arbitrary phases. We showcase SODAS++ for both solid–solid and solid–liquid transitions for systems of increasing geometric and chemical complexity, such as colloidal systems, elemental Al, rutile and amorphous TiO 2 , and the non-stoichiometric ternary alloy Ag 26 Au 5 Cu 19 . Finally, we show that SODAS++ can accurately quantify all transitions in a physically interpretable manner, showcasing the power of unsupervised graph neural network encodings for capturing the complex and non-linear pathway, a material's structure takes as it evolves.

36 MATERIALS SCIENCE↗

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.↗

Identifying Different Classes of Seismic Noise Signals Using Unsupervised Learning

Abstract Proper classification of nontectonic seismic signals is critical for detecting microearthquakes and developing an improved understanding of ongoing weak ground motions. We use unsupervised machine learning to label five classes of nonstationary seismic noise common in continuous waveforms. Temporal and spectral features describing the data are clustered to identify separable types of emergent and impulsive waveforms. The trained clustering model is used to classify every 1 s of continuous seismic records from a dense seismic array with 10–30 m station spacing. We show that dominate noise signals can be highly localized and vary on length scales of hundreds of meters. The methodology demonstrates the complexity of weak ground motions and improves the standard of analyzing seismic waveforms with a low signal‐to‐noise ratio. Application of this technique will improve the ability to detect genuine microseismic events in noisy environments where seismic sensors record earthquake‐like signals originating from nontectonic sources.

Johnson, Christopher W.↗

Wildlife management by habitat units: A preliminary plan of action

Procedures for yielding vegetation type maps were developed using LANDSAT data and a computer assisted classification analysis (LARSYS) to assist in managing populations of wildlife species by defined area units. Ground cover in Travis County, Texas was classified on two occasions using a modified version of the unsupervised approach to classification. The first classification produced a total of 17 classes. Examination revealed that further grouping was justified. A second analysis produced 10 classes which were displayed on printouts which were later color-coded. The final classification was 82 percent accurate. While the classification map appeared to satisfactorily depict the existing vegetation, two classes were determined to contain significant error. The major sources of error could have been eliminated by stratifying cluster sites more closely among previously mapped soil associations that are identified with particular plant associations and by precisely defining class nomenclature using established criteria early in the analysis.

Frentress, C. D.↗