Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Anomaly Detection, Localization and Classification using Drifting Synchrophasor Data Streams

With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.

adversarial deep learning↗

Comparative techniques used to evaluate Thematic Mapper data for land cover classification in Logan County, West Virginia

Several digital data processing techniques were evaluated in an effort to identify and map active/abandoned, partially reclaimed, and fully revegetated surface mine areas in the central portion of Logan County. The TM data were first subjected to various enhancement procedures, including a linear contrast stretch, principal components and canonical analysis transformations. At the same time, four general procedures were followed to produce six classifications as a means of comparing the techniques involved. Preliminary results show that various feature extraction/data reduction techniques provide classification results equal or superior to the more straightforward unsupervised clustering technique. Analyst interaction time for labelling clusters is reduced using the canonical analysis and principal components procedures, though the canonical technique has clearly produced better results to date.

Brumfield, J. O.↗

Global radar units on Venus derived from statistical analysis of Pioneer Venus Orbiter radar data

The classification of surface radar units on Venus using an unsupervised cluster analysis of Pioneer Venus radar reflectivity and root-mean-square (rms)-slope data is described. The advantages of the unsupervised analysis are discussed. F tests are utilized to evaluate the numerical significance of the clusters. The derived rms-slope data and reflectivity for 15 radar units are presented. The relations between radar data bases and elevation are studied. The lowlands, rolling plains, highlands, and mountainous surface of Venus are examined. The geology of Venus landing sites and radar properties, and the surface radar reflectivity images and earth-based images are compared. The spatial relations between classification units are calculated. It is concluded that the unsupervised analysis data correlate well with Head et al. (1985b) data and produce more detailed classification images.

Davis, P. A.↗

Classification of multispectral image data by the Binary Diamond neural network and by nonparametric, pixel-by-pixel methods

The classification of multispectral image data obtained from satellites has become an important tool for generating ground cover maps. This study deals with the application of nonparametric pixel-by-pixel classification methods in the classification of pixels, based on their multispectral data. A new neural network, the Binary Diamond, is introduced, and its performance is compared with a nearest neighbor algorithm and a back-propagation network. The Binary Diamond is a multilayer, feed-forward neural network, which learns from examples in unsupervised, 'one-shot' mode. It recruits its neurons according to the actual training set, as it learns. The comparisons of the algorithms were done by using a realistic data base, consisting of approximately 90,000 Landsat 4 Thematic Mapper pixels. The Binary Diamond and the nearest neighbor performances were close, with some advantages to the Binary Diamond. The performance of the back-propagation network lagged behind. An efficient nearest neighbor algorithm, the binned nearest neighbor, is described. Ways for improving the performances, such as merging categories, and analyzing nonboundary pixels, are addressed and evaluated.

Salu, Yehuda↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

Classification of remotely sensed data using OCR-inspired neural network techniques

Neural networks have been applied to classifications of remotely sensed data with some success. To improve the performance of this approach, an examination was made of how neural networks are applied to the optical character recognition (OCR) of handwritten digits and letters. A three-layer, feedforward network, along with techniques adopted from OCR, was used to classify Landsat-4 Thematic Mapper data. Good results were obtained. To overcome the difficulties that are characteristic of remote sensing applications and to attain significant improvements in classification accuracy, a special network architecture may be required.

Kiang, Richard K.↗

Contextual classification of multispectral image data: An unbiased estimator for the context distribution

A key input to a statistical classification algorithm, which exploits the tendency of certain ground cover classes to occur more frequently in some spatial context than in others, is a statistical characterization of the context: the context distribution. An unbiased estimator of the context distribution is discussed which, besides having the advantage of statistical unbiasedness, has the additional advantage over other estimation techniques of being amenable to an adaptive implementation in which the context distribution estimate varies according to local contextual information. Results from applying the unbiased estimator to the contextual classification of three real LANDSAT data sets are presented and contrasted with results from non-contextual classifications and from contextual classifications utilizing other context distribution estimation techniques.

Tilton, J. C.↗

An improvement in land cover classification achieved by merging microwave data with Landsat multispectral scanner data

The improvement in land cover classification achieved by merging microwave data with Landsat MSS data is examined. To produce a merged data set for analysis and comparison, a registration procedure by which a set of Seasat SAR digital data was merged with the MSS data is described. The Landsat MSS data and the merged Landsat/Seasat data sets were processed using conventional multichannel spectral pattern recognition techniques. An analysis of the classified data sets indicates that while Landsat data delineate different forest types (i.e., deciduous/coniferous) and allow some species separation, SAR data provide additional information related to plant canopy configuration and vegetation density as associated with varying water regimes, and therefore allow for further subdivision in the classification of forested wetlands of the coastal region of the southern United States.

Wu, S. T.↗

Drop Size Distribution - Based Separation of Stratiform and Convective Rain

For applications in hydrology and meteorology, it is often desirable to separate regions of stratiform and convective rain from meteorological radar observations, both from ground-based polarimetric radars and from space-based dual frequency radars. In a previous study by Bringi et al. (2009), dual frequency profiler and dual polarization radar (C-POL) observations in Darwin, Australia, had shown that stratiform and convective rain could be separated in the log10(Nw) versus Do domain, where Do is the mean volume diameter and Nw is the scaling parameter which is proportional to the ratio of water content to the mass weighted mean diameter. Note, Nw and Do are two of the main drop size distribution (DSD) parameters. In a later study, Thurai et al (2010) confirmed that both the dual-frequency profiler based stratiform-convective rain separation and the C-POL radar based separation were consistent with each other. In this paper, we test this separation method using DSD measurements from a ground based 2D video disdrometer (2DVD), along with simultaneous observations from a collocated, vertically-pointing, X-band profiling radar (XPR). The measurements were made in Huntsville, Alabama. One-minute DSDs from 2DVD are used as input to an appropriate gamma fitting procedure to determine Nw and Do. The fitted parameters - after averaging over 3-minutes - are plotted against each other and compared with a predefined separation line. An index is used to determine how far the points lie from the separation line (as described in Thurai et al. 2010). Negative index values indicate stratiform rain and positive index indicate convective rain, and, moreover, points which lie somewhat close to the separation line are considered 'mixed' or 'transition' type precipitation. The XPR observations are used to evaluate/test the 2DVD data-based classification. A 'bright-band' detection algorithm was used to classify each vertical reflectivity profile as either stratiform or convective, depending on whether or not a clearly-defined melting layer is present at an expected height, and if present, maximum reflectivity within the melting layer as well as the corresponding height are determined. We will present results of quantitative comparisons between the XPR observations-based classifications and the simultaneous 2DVD data-based classifications. Time series comparisons will be presented for thirteen events in Huntsville.

Thurai, Merhala↗

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised classification of MSS Landsat data for mapping spatially complex vegetation

The various stages in carrying out a monocluster block unsupervised classification using Landsat MSS data are described. Procedures for carrying out these various stages were found to be far from well-established for the type of terrain being investigated, which is rugged and contains many small land cover units. Two particular difficulties were encountered: first, that of precise ground location of pixels; and, secondly, that of objectively evaluating the results. Ways in which these can be surmounted are suggested.

Townshend, J. R. G.↗

Neural network approaches versus statistical methods in classification of multisource remote sensing data

Neural network learning procedures and statistical classificaiton methods are applied and compared empirically in classification of multisource remote sensing and geographic data. Statistical multisource classification by means of a method based on Bayesian classification theory is also investigated and modified. The modifications permit control of the influence of the data sources involved in the classification process. Reliability measures are introduced to rank the quality of the data sources. The data sources are then weighted according to these rankings in the statistical multisource classification. Four data sources are used in experiments: Landsat MSS data and three forms of topographic data (elevation, slope, and aspect). Experimental results show that two different approaches have unique advantages and disadvantages in this classification application.

Benediktsson, Jon A.↗

Data resolution versus forestry classification and modeling

This paper examines the effects on timber stand computer classification accuracies caused by changes in the resolution of remotely sensed multispectral data. This investigation is valuable, especially for determining optimal sensor and platform designs. Theoretical justification and experimental verification support the finding that classification accuracies for low resolution data could be better than the accuracies for data with higher resolution. The increase in accuracy is constructed as due to the reduction of scene inhomogeneity at lower resolution. The computer classification scheme was a maximum likelihood classifier.

Kan, E. P.↗

Multiple Spectral-Spatial Classification Approach for Hyperspectral Data

A .new multiple classifier approach for spectral-spatial classification of hyperspectral images is proposed. Several classifiers are used independently to classify an image. For every pixel, if all the classifiers have assigned this pixel to the same class, the pixel is kept as a marker, i.e., a seed of the spatial region, with the corresponding class label. We propose to use spectral-spatial classifiers at the preliminary step of the marker selection procedure, each of them combining the results of a pixel-wise classification and a segmentation map. Different segmentation methods based on dissimilar principles lead to different classification results. Furthermore, a minimum spanning forest is built, where each tree is rooted on a classification -driven marker and forms a region in the spectral -spatial classification: map. Experimental results are presented for two hyperspectral airborne images. The proposed method significantly improves classification accuracies, when compared to previously proposed classification techniques.

Tarabalka, Yuliya↗

Classification of corn and soybeans using multitemporal Thematic Mapper data

The multitemporal classification approach based on the greenness profile derived from Landsat Multispectral Scanner (MSS) spectral bands has proved successful in effectively separating and identifying corn, soybean, and other ground cover classes. Features derived from these profiles have been shown to carry virtually all the information contained in the original data and, in addition, have been shown to be stable over a large geographic area of the United States. The objective of this investigation was to determine if the same features derived from multitemporal Thematic Mapper (TM) data would also prove effective in separating these two crop types, and, in fact, if algorithms developed for MSS could be directly applied to TM. It is shown that this is indeed the case. In addition, because of greater spatial and spectral resolution, the accuracy of TM classifications is better than in MSS.

Badhwar, G. D.↗