Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A comparison of unsupervised classification procedures on LANDSAT MSS data for an area of complex surface conditions in Basilicata, Southern Italy

Two unsupervised classification procedures were applied to ratioed and unratioed LANDSAT multispectral scanner data of an area of spatially complex vegetation and terrain. An objective accuracy assessment was undertaken on each classification and comparison was made of the classification accuracies. The two unsupervised procedures use the same clustering algorithm. By on procedure the entire area is clustered and by the other a representative sample of the area is clustered and the resulting statistics are extrapolated to the remaining area using a maximum likelihood classifier. Explanation is given of the major steps in the classification procedures including image preprocessing; classification; interpretation of cluster classes; and accuracy assessment. Of the four classifications undertaken, the monocluster block approach on the unratioed data gave the highest accuracy of 80% for five coarse cover classes. This accuracy was increased to 84% by applying a 3 x 3 contextual filter to the classified image. A detailed description and partial explanation is provided for the major misclassification. The classification of the unratioed data produced higher percentage accuracies than for the ratioed data and the monocluster block approach gave higher accuracies than clustering the entire area. The moncluster block approach was additionally the most economical in terms of computing time.

Justice, C.↗

Detection of aspen-conifer forest mixes from LANDSAT digital data

Aspen, conifer and mixed aspen/conifer forests were mapped for a 15-quadrangle study area in the Utah-Idaho Bear River Range using LANDSAT multispectral scanner data. Digital classification and statistical analysis of LANDSAT data allowed the identification of six groups of signatures which reflect different types of aspen/conifer forest mixing. Photo interpretations of the print symbols suggest that such classes are indicative of mid to late seral aspen forests. Digital print map overlays and acreage calculations were prepared for the study area quadrangles. Further field verification is needed to acquire additional information about the nature of the forests. Single date LANDSAT analysis should be a cost effective means to index aspen forests which are at least in the mid seral phase of conifer invasion. Since aspen canopies tend to obscure understory conifers for early seral forests, a second date analysis, using data taken when aspens are leafless, could provide information about early seral aspen forests.

Jaynes, R. A.↗

Detection of aspen/conifer forest mixes from multitemporal Landsat digital data

Aspen, conifer and mixed aspen/conifer forests were mapped for a 15-quadrangle study area in the Utah-Idaho Bear River Range using Landsat multispectral scanner data. Digital classification and statistical analysis of Landsat data allowed the identification of six groups of signatures which reflect different types of aspen/conifer forest mixing. Photo interpretations of the print symbols suggest that such classes are indicative of mid to late seral aspen forests. Digital print map overlayes and acreage calculations were prepared for the study area quadrangles. Further field verification is needed to acquire additional information about the nature of the forests. Single data Landsat analysis should be a cost effective means to index aspen forests which are at least in the mid seral phase of conifer invasion. Since aspen canopies tend to obscure understory conifers for early seral forests, a second data analysis, using data taken when aspens are leafless, could provide information about early seral aspen forests.

Merola, J. A.↗

Spacecube: A Family of Reconfigurable Hybrid On-Board Science Data Processors

SpaceCube is a family of Field Programmable Gate Array (FPGA) based on-board science data processing systems developed at the NASA Goddard Space Flight Center (GSFC). The goal of the SpaceCube program is to provide 10x to 100x improvements in on-board computing power while lowering relative power consumption and cost. SpaceCube is based on the Xilinx Virtex family of FPGAs, which include processor, FPGA logic and digital signal processing (DSP) resources. These processing elements are leveraged to produce a hybrid science data processing platform that accelerates the execution of algorithms by distributing computational functions to the most suitable elements. This approach enables the implementation of complex on-board functions that were previously limited to ground based systems, such as on-board product generation, data reduction, calibration, classification, eventfeature detection, data mining and real-time autonomous operations. The system is fully reconfigurable in flight, including data parameters, software and FPGA logic, through either ground commanding or autonomously in response to detected eventsfeatures in the instrument data stream.

reconfigurable computing↗

Comparative Analysis of ML Techniques for Data-Driven Anomaly Detection, Classification and Localization in Distribution System

High penetration of Distributed Energy Resources (DERs), fundamental load behavior changes, controllable loads, and significant increase in Electrical Vehicles (EVs) lead to complex dynamic behavior of the electric distribution system. Increasing number of components also means more measurements, more data and more data anomalies. Detecting, classifying and localizing these anomalies are important for situational awareness, and at the same time, very challenging given increasing complexity of the system. Highly accurate and high-resolution analytical techniques are needed to support anomaly detection, classification and localization (AD-C-L) for monitoring, root cause analysis and decision making. This paper provides comprehensive review and analysis of the existing spatio-temporal AD-C-L techniques within the distribution system. Challenges for specific problems in AD-C- L have been also discussed in this paper. Existing AD-C- L techniques have been categorized and synthesized for specific merits and limitations of multiple Machine Learning (ML) methodologies using common developed metrics of performance. The comparative analysis is summarized and presented with the open research challenges and path forward for future research needs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Training data selection for event classification in a highly variable environment

A problem of interest for nuclear nonproliferation is monitoring activities at nuclear facilities, where proliferation events may only take place a few times and often under variable conditions. Machine learning has revolutionized data analytics by enabling the use of measurable signatures to generate predictive models of facility operations. However, traditional methods for training these models require large, reliable data sets with labeled observations, a challenge for nonproliferation. Highly variable conditions further complicate this as events from training data may have occurred in conditions quite different from the event of interest. Our hypothesis is that when events occur in a highly variable environment, careful training data selection for each test event could outperform the standard approach of using all available training data. We developed a method to optimize training data selection for the given test event and applied it to predicting the power level of the High Flux Isotope Reactor (HFIR) at Oak Ridge National Laboratory. In this study, the reactor startup exhibits variability between occurrences due to natural variability in environmental conditions and operational procedures. Using a combination of analysis techniques, a similitude assessment was performed on data collected from HFIR to isolate clusters that were optimal for training a predictive model. Concepts such as dynamic time warping and Jaccard similarity were used in conjunction with clustering analysis. In order to validate this approach, the model was trained on every combination of unique training events and the predictive performance was compared to the performance using a subset of the training data selected by isolated clusters found through the similitude assessment.

Iyer, A↗

A Deterministic Self-Organizing Map Approach and its Application on Satellite Data based Cloud Type Classification

A self-organizing map (SOM) is a type of competitive artificial neural network, which projects the high dimensional input space of the training samples into a low dimensional space with the topology relations preserved. This makes SOMs supportive of organizing and visualizing complex data sets and have been pervasively used among numerous disciplines with different applications. Notwithstanding its wide applications, the self-organizing map is perplexed by its inherent randomness, which produces dissimilar SOM patterns even when being trained on identical training samples with the same parameters every time, and thus causes usability concerns for other domain practitioners and precludes more potential users from exploring SOM based applications in a broader spectrum. Motivated by this practical concern, we propose a deterministic approach as a supplement to the standard self-organizing map. In accordance with the theoretical design, the experimental results with satellite cloud data demonstrate the effective and efficient organization as well as simplification capabilities of the proposed approach.

Initialization method↗

Spectral analysis for automated exploration and sample acquisition

Future space exploration missions will rely heavily on the use of complex instrument data for determining the geologic, chemical, and elemental character of planetary surfaces. One important instrument is the imaging spectrometer, which collects complete images in multiple discrete wavelengths in the visible and infrared regions of the spectrum. Extensive computational effort is required to extract information from such high-dimensional data. A hierarchical classification scheme allows multispectral data to be analyzed for purposes of mineral classification while limiting the overall computational requirements. The hierarchical classifier exploits the tunability of a new type of imaging spectrometer which is based on an acousto-optic tunable filter. This spectrometer collects a complete image in each wavelength passband without spatial scanning. It may be programmed to scan through a range of wavelengths or to collect only specific bands for data analysis. Spectral classification activities employ artificial neural networks, trained to recognize a number of mineral classes. Analysis of the trained networks has proven useful in determining which subsets of spectral bands should be employed at each step of the hierarchical classifier. The network classifiers are capable of recognizing all mineral types which were included in the training set. In addition, the major components of many mineral mixtures can also be recognized. This capability may prove useful for a system designed to evaluate data in a strange environment where details of the mineral composition are not known in advance.

Eberlein, Susan↗

Application of Landsat data to wetland study and land use classification in West Tennessee

Landsat data were employed in determining land use of a 32,300-hectare watershed area within the Obion-Forked Deer River Basin in northwest Tennessee. Black and white transparency chips for all four wavelength bands were interpreted by use of a video-input analog/digital automatic analysis and classification facility; densitometric methods showed that wetlands, urban areas, agricultural lands and forests could be discriminated by analysis of band 6 or 7 together with band 4 or 5. Comparison with high- and low-altitude photography indicated that the Landsat data could provide sufficiently accurate resource information and determine drainage trends.

Jones, N. L.↗

A Framework for Land Cover Classification Using Discrete Return LiDAR Data: Adopting Pseudo-Waveform and Hierarchical Segmentation

Acquiring current, accurate land-use information is critical for monitoring and understanding the impact of anthropogenic activities on natural environments.Remote sensing technologies are of increasing importance because of their capability to acquire information for large areas in a timely manner, enabling decision makers to be more effective in complex environments. Although optical imagery has demonstrated to be successful for land cover classification, active sensors, such as light detection and ranging (LiDAR), have distinct capabilities that can be exploited to improve classification results. However, utilization of LiDAR data for land cover classification has not been fully exploited. Moreover, spatial-spectral classification has recently gained significant attention since classification accuracy can be improved by extracting additional information from the neighboring pixels. Although spatial information has been widely used for spectral data, less attention has been given to LiDARdata. In this work, a new framework for land cover classification using discrete return LiDAR data is proposed. Pseudo-waveforms are generated from the LiDAR data and processed by hierarchical segmentation. Spatial featuresare extracted in a region-based way using a new unsupervised strategy for multiple pruning of the segmentation hierarchy. The proposed framework is validated experimentally on a real dataset acquired in an urban area. Better classification results are exhibited by the proposed framework compared to the cases in which basic LiDAR products such as digital surface model and intensity image are used. Moreover, the proposed region-based feature extraction strategy results in improved classification accuracies in comparison with a more traditional window-based approach.

Light Detection & Ranging (LIDAR)↗

Engagement Assessment Using EEG Signals

In this paper, we present methods to analyze and improve an EEG-based engagement assessment approach, consisting of data preprocessing, feature extraction and engagement state classification. During data preprocessing, spikes, baseline drift and saturation caused by recording devices in EEG signals are identified and eliminated, and a wavelet based method is utilized to remove ocular and muscular artifacts in the EEG recordings. In feature extraction, power spectrum densities with 1 Hz bin are calculated as features, and these features are analyzed using the Fisher score and the one way ANOVA method. In the classification step, a committee classifier is trained based on the extracted features to assess engagement status. Finally, experiment results showed that there exist significant differences in the extracted features among different subjects, and we have implemented a feature normalization procedure to mitigate the differences and significantly improved the engagement assessment performance.

Li, Feng↗

Unsupervised classification of earth resources data.

A new clustering technique is presented. It consists of two parts: (a) a sequential statistical clustering which is essentially a sequential variance analysis and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by existing supervised maximum liklihood classification technique.

Su, M. Y.↗

Automated Classification of Transient Contamination in Stationary Acoustic Data

An automated procedure for the classification of transient contamination of stationary acoustic data is proposed and analyzed. The procedure requires the assumption that the stationary acoustic data of interest can be modeled as a band-limited, Gaussian random process. It also requires that the transient contamination be of higher variance than the acoustic data of interest. When these assumptions are satisfied, it is a blind separation procedure, aside from the initial input specifying how to subdivide the time series of interest. No a priori threshold criterion is required. Simulation results show that for a sufficient number of blocks, the method performs well, as long as the occasional false positive or false negative is acceptable. The effectiveness of the procedure is demonstrated with an application to experimental wind tunnel acoustic test data which are contaminated by hydrodynamic gusts.

binary classification↗

Identification and area estimation of agricultural crops by computer classification of Landsat MSS data

Landsat Multispectral Scanner (MSS) data covering a three-county area in northern Illinois were classified using computer-aided techniques as corn, soybeans, or 'other.' Recognition of test fields was 80% accurate. County estimates of the area of corn and soybeans agreed closely with those made by the USDA. Results of the use of a priori information in classification, techniques to produce unbiased area estimates, and the use of temporal and spatial features for classification are discussed. The extendability, variability, and size of training sets, wavelength band selection, and spectral characteristics of crops were also investigated.

Bauer, M. E.↗

Experimental study of digital image processing techniques for LANDSAT data

The author has identified the following significant results. Results are reported for: (1) subscene registration, (2) full scene rectification and registration, (3) resampling techniques, (4) and ground control point (GCP) extraction. Subscenes (354 pixels x 234 lines) were registered to approximately 1/4 pixel accuracy and evaluated by change detection imagery for three cases: (1) bulk data registration, (2) precision correction of a reference subscene using GCP data, and (3) independently precision processed subscenes. Full scene rectification and registration results were evaluated by using a correlation technique to measure registration errors of 0.3 pixel rms thoughout the full scene. Resampling evaluations of nearest neighbor and TRW cubic convolution processed data included change detection imagery and feature classification. Resampled data were also evaluated for an MSS scene containing specular solar reflections.

Rifman, S. S.↗

Design of a graphical user interface for few-shot machine learning classification of electron microscopy data

The recent growth in data generation by modern electron microscopes requires rapid, scalable, and flexible approaches to image segmentation and analysis. Few-shot machine learning, which can richly classify images from a handful of user-provided examples, is a promising route to high-throughput analysis. However, current command-line implementations of such approaches can be slow and unintuitive to use, lacking the real-time feedback necessary to perform effective classification. Here we report on the development of a Python-based graphical user interface that enables end users to easily conduct and visualize the output of few-shot learning models. This interface is portable and can be hosted locally or on the web, providing the opportunity to reproducibly conduct, share, and crowd-source few-shot analyses.

97 MATHEMATICS AND COMPUTING↗

Evaluation of several schemes for classification of remotely sensed data: Their parameters and performance

The author has identified the following significant results. Data sets for corn, soybeans, winter wheat, and spring wheat were used to evaluate the following schemes for crop identification: (1) per point Gaussian maximum classifier; (2) per point sum of normal densities classifiers; (3) per point linear classifier; (4) per point Gaussian maximum likelihood decision tree classifiers; and (5) texture sensitive per field Gaussian maximum likelihood classifier. Test site location and classifier both had significant effects on classification accuracy of small grains; classifiers did not differ significantly in overall accuracy, with the majority of the difference among classifiers being attributed to training method rather than to the classification algorithm applied. The complexity of use and computer costs for the classifiers varied significantly. A linear classification rule which assigns each pixel to the class whose mean is closest in Euclidean distance was the easiest for the analyst and cost the least per classification.

Scholz, D.↗