Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Geophysical phenomena classification by artificial neural networks

Space science information systems involve accessing vast data bases. There is a need for an automatic process by which properties of the whole data set can be assimilated and presented to the user. Where data are in the form of spectrograms, phenomena can be detected by pattern recognition techniques. Presented are the first results obtained by applying unsupervised Artificial Neural Networks (ANN's) to the classification of magnetospheric wave spectra. The networks used here were a simple unsupervised Hamming network run on a PC and a more sophisticated CALM network run on a Sparc workstation. The ANN's were compared in their geophysical data recognition performance. CALM networks offer such qualities as fast learning, superiority in generalizing, the ability to continuously adapt to changes in the pattern set, and the possibility to modularize the network to allow the inter-relation between phenomena and data sets. This work is the first step toward an information system interface being developed at Sussex, the Whole Information System Expert (WISE). Phenomena in the data are automatically identified and provided to the user in the form of a data occurrence morphology, the Whole Information System Data Occurrence Morphology (WISDOM), along with relationships to other parameters and phenomena.

Gough, M. P.↗

Global radar units on Venus derived from statistical analysis of Pioneer Venus Orbiter radar data

The classification of surface radar units on Venus using an unsupervised cluster analysis of Pioneer Venus radar reflectivity and root-mean-square (rms)-slope data is described. The advantages of the unsupervised analysis are discussed. F tests are utilized to evaluate the numerical significance of the clusters. The derived rms-slope data and reflectivity for 15 radar units are presented. The relations between radar data bases and elevation are studied. The lowlands, rolling plains, highlands, and mountainous surface of Venus are examined. The geology of Venus landing sites and radar properties, and the surface radar reflectivity images and earth-based images are compared. The spatial relations between classification units are calculated. It is concluded that the unsupervised analysis data correlate well with Head et al. (1985b) data and produce more detailed classification images.

Davis, P. A.↗

Automated Classification of Transient Contamination in Stationary Acoustic Data

An automated procedure for the classification of transient contamination of stationary acoustic data is proposed and analyzed. The procedure requires the assumption that the stationary acoustic data of interest can be modeled as a band-limited, Gaussian random process. It also requires that the transient contamination be of higher variance than the acoustic data of interest. When these assumptions are satisfied, it is a blind separation procedure, aside from the initial input specifying how to subdivide the time series of interest. No a priori threshold criterion is required. Simulation results show that for a sufficient number of blocks, the method performs well, as long as the occasional false positive or false negative is acceptable. The effectiveness of the procedure is demonstrated with an application to experimental wind tunnel acoustic test data which are contaminated by hydrodynamic gusts.

binary classification↗

Automated Grain Yield Behavior Classification

A method for classifying grain stress evolution behaviors using unsupervised learning techniques is presented. The method is applied to analyze grain stress histories measured in-situ using high-energy X-ray diffraction microscopy (HEDM) from the aluminum-lithium alloy Al-Li 2099 at the elastic-plastic transition (yield). The unsupervised learning process automatically classified the grain stress histories into four groups: major softening, no work-hardening or softening, moderate work-hardening, and major work-hardening. The orientation and spatial dependence of these four groups are discussed. In addition, the generality of the classification process to other samples is explored.

Pagan, Darren C↗

Characterizing Interference in Radio Astronomy Observations through Active and Unsupervised Learning

In the process of observing signals from astronomical sources, radio astronomers must mitigate the effects of manmade radio sources such as cell phones, satellites, aircraft, and observatory equipment. Radio frequency interference (RFI) often occurs as short bursts (< 1 ms) across a broad range of frequencies, and can be confused with signals from sources of interest such as pulsars. With ever-increasing volumes of data being produced by observatories, automated strategies are required to detect, classify, and characterize these short "transient" RFI events. We investigate an active learning approach in which an astronomer labels events that are most confusing to a classifier, minimizing the human effort required for classification. We also explore the use of unsupervised clustering techniques, which automatically group events into classes without user input. We apply these techniques to data from the Parkes Multibeam Pulsar Survey to characterize several million detected RFI events from over a thousand hours of observation.

Doran, G.↗

Characterizing Interference in Radio Astronomy Observations through Active and Unsupervised Learning

In the process of observing signals from astronomical sources, radio astronomers must mitigate the effects of man-made radio sources such as cell phones, satellites, aircraft, and observatory equipment. Radio frequency interference (RFI) often occurs as short bursts (< 1 ms) across a broad range of frequencies, and can be confused with signals from sources of interest such as pulsars. With ever-increasing volumes of data being produced by observatories, automated strategies are required to detect, classify, and characterize these short “transient” RFI events. We investigate an active learning approach in which an astronomer labels events that are most confusing to a classifier, minimizing the human effort required for classification. We also explore the use of unsupervised clustering techniques, which automatically group events into classes without user input. We apply these techniques to data from the Parkes Multibeam Pulsar Survey to characterize several million detected RFI events from over a thousand hours of observation

Doran, G.↗

Adaptive fuzzy leader clustering of complex data sets in pattern recognition

A modular, unsupervised neural network architecture for clustering and classification of complex data sets is presented. The adaptive fuzzy leader clustering (AFLC) architecture is a hybrid neural-fuzzy system that learns on-line in a stable and efficient manner. The initial classification is performed in two stages: a simple competitive stage and a distance metric comparison stage. The cluster prototypes are then incrementally updated by relocating the centroid positions from fuzzy C-means system equations for the centroids and the membership values. The AFLC algorithm is applied to the Anderson Iris data and laser-luminescent fingerprint image data. It is concluded that the AFLC algorithm successfully classifies features extracted from real data, discrete or continuous.

Newton, Scott C.↗

Landsat Thematic Mapper digital information content for agricultural environments

Landsat Thematic Mapper (TM) data collected for Imperial Valley, California in December, 1982 were digitally examined to assess their utility to distinguish among agricultural and other land-covers. Statistics for thirty-seven training sites representing a variety of crops plus urban, water and desert land-covers were obtained and analyzed using transformed divergence (TD) calculations. TD values were employed to assess intraclass variability and the best bands for classification. Four subscenes were selected for clustering or unsupervised signature extraction. These areas were agriculture, urban, desert and water land-covers. The number of clusters for these subscenes were examined and the best TM bands for interclass separability were identified. The results of the clustering and training site analyses for interclass separability were compared. The TM data were useful for the digital delimitation of most crops and other cover types in this analysis. Four bands of data are adequate for classification with the best results obtained by the selection of one band from each of the available portions of the electromagnetic spectrum. Different band combinations are best for various land-cover intraclass separability.

Haack, Barry↗

Assessment of Computer-based Geologic Mapping of Rock Units in the LANDSAT-4 Scene of Northern Death Valley, California

Geologists obtain low accuracy levels when maps derived from LANDSAT MSS data are compared with those made by conventional methods. Procedures developed for the IDIMS computer system and used to classify a subset of a TM image of the Death Valley, California - Nevada border are described. Despite the superior resolution, broader spectral coverage, and greater sensitivity inherent to the TM, the actual recorded measured accuracy was in the same narrow range (30 to 60%) recorded for MSS data from earlier LANDSATs. The supervised classification approach appears to be superior to the unsupervised approach when applied to vegetation-sparse surfaces composed of spectrally contrasting rock/soil units distributed in relatively flat to low relief terrain. As spatial resolution improves and optimal spectral bands for identifying rock materials are specified, use of classified multispectral remote sensing data from air and space when coupled with supporting field calibration and checks should become the dominant way in which geologic mapping is carried out in future decades.

Short, N. M.↗

Developing Land Use Land Cover Maps for the Lower Mekong Basin to Aid Hydrologic Modeling and Basin Planning

This paper discusses research methodology to develop Land Use Land Cover (LULC) mapsfor the Lower Mekong Basin (LMB) for basin planning, using both MODIS and Landsat satellitedata. The 2010 MODIS MOD09 and MYD09 8-day reflectance data was processed into monthlyNDVI maps with the Time Series Product Tool software package and then used to classify regionallycommon forest and agricultural LULC types. Dry season circa 2010 Landsat top of atmosphere reflectance mosaics were classified to map locally common LULC types. Unsupervised ISODATAclustering was used to derive most LULC classifications. MODIS and Landsat classifications werecombined with GIS methods to derive final 250-m LULC maps for Sub-basins (SBs) 1–8 of the LMB.The SB 7 LULC map with 14 classes was assessed for accuracy. This assessment compared randomlocations for sampled types on the SB 7 LULC map to geospatial reference data such as Landsat RGBs,MODIS NDVI phenologic profiles, high resolution satellite data, and Mekong River Commissiondata (e.g., crop calendars). The SB 7 LULC map showed an overall agreement to reference data of~81%. By grouping three deciduous forest classes into one, the overall agreement improved to ~87%.The project enabled updated regional LULC maps that included more detailed agriculture LULCtypes. LULC maps were supplied to project partners to improve use of Soil andWater AssessmentTool for modeling hydrology and water use, plus enhance LMB water and disaster managementin a region vulnerable to flooding, droughts, and anthropogenic change as part of basin planningand assessment.

land use land cover mapping; SWAT hydrologic model↗

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity↗

Comparative techniques used to evaluate Thematic Mapper data for land cover classification in Logan County, West Virginia

Several digital data processing techniques were evaluated in an effort to identify and map active/abandoned, partially reclaimed, and fully revegetated surface mine areas in the central portion of Logan County. The TM data were first subjected to various enhancement procedures, including a linear contrast stretch, principal components and canonical analysis transformations. At the same time, four general procedures were followed to produce six classifications as a means of comparing the techniques involved. Preliminary results show that various feature extraction/data reduction techniques provide classification results equal or superior to the more straightforward unsupervised clustering technique. Analyst interaction time for labelling clusters is reduced using the canonical analysis and principal components procedures, though the canonical technique has clearly produced better results to date.

Brumfield, J. O.↗

Development and application of operational techniques for the inventory and monitoring of resources and uses for the Texas coastal zone

The author has identified the following significant results. Image interpretation mapping techniques were successfully applied to test site 5, an area with a semi-arid climate. The land cover/land use classification required further modification. A new program, HGROUP, added to the ADP classification schedule provides a convenient method for examining the spectral similarity between classes. This capability greatly simplifies the task of combining 25-30 unsupervised subclasses into about 15 major classes that approximately correspond to the land use/land cover classification scheme.

Harwood, P.↗

Comparison of MSS and TM Data for Landcover Classification in the Chesapeake Bay Area: a Preliminary Report

An area bordering the Eastern Shore of the Chesapeake Bay was selected for study and classified using unsupervised techniques applied to LANDSAT-2 MSS data and several band combinations of LANDSAT-4 TM data. The accuracies of these Level I land cover classifications were verified using the Taylor's Island USGS 7.5 minute topographic map which was photointerpreted, digitized and rasterized. The the Taylor's Island map, comparing the MSS and TM three band (2 3 4) classifications, the increased resolution of TM produced a small improvement in overall accuracy of 1% correct due primarily to a small improvement, and 1% and 3%, in areas such as water and woodland. This was expected as the MSS data typically produce high accuracies for categories which cover large contiguous areas. However, in the categories covering smaller areas within the map there was generally an improvement of at least 10%. Classification of the important residential category improved 12%, and wetlands were mapped with 11% greater accuracy.

Mulligan, P. J.↗

Remote sensing of submerged aquatic vegetation in lower Chesapeake Bay - A comparison of Landsat MSS to TM imagery

Landsat MSS and TM imagery, obtained simultaneously over Guinea Marsh, VA, as analyzed and compares for its ability to detect submerged aquatic vegetation (SAV). An unsupervised clustering algorithm was applied to each image, where the input classification parameters are defined as functions of apparent sensor noise. Class confidence and accuracy were computed for all water areas by comparing the classified images, pixel-by-pixel, to rasterized SAV distributions derived from color aerial photography. To illustrate the effect of water depth on classification error, areas of depth greater than 1.9 m were masked, and class confidence and accuracy recalculated. A single-scattering radiative-transfer model is used to illustrate how percent canopy cover and water depth affect the volume reflectance from a water column containing SAV. For a submerged canopy that is morphologically and optically similar to Zostera marina inhabiting Lower Chesapeake Bay, dense canopies may be isolated by masking optically deep water. For less dense canopies, the effect of increasing water depth is to increase the apparent percent crown cover, which may result in classification error.

Ackleson, S. G.↗

Bayesian Fusion of Color and Texture Segmentations

In many applications one would like to use information from both color and texture features in order to segment an image. We propose a novel technique to combine "soft" segmentations computed for two or more features independently. Our algorithm merges models according to a mean entropy criterion, and allows to choose the appropriate number of classes for the final grouping. This technique also allows to improve the quality of supervised classification based on one feature (e.g. color) by merging information from unsupervised segmentation based on another feature (e.g., texture.)

Manduchi, Roberto↗

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing↗

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing↗