Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Application of SEASAT-1 Synthetic Aperture Radar (SAR) data to enhance and detect geological lineaments and to assist LANDSAT landcover classification mapping

Digital SEASAT-1 synthetic aperture radar (SAR) data were used to enhance linear features to extract geologically significant lineaments in the Appalachian region. Comparison of Lineaments thus mapped with an existing lineament map based on LANDSAT MSS images shows that appropriately processed SEASAT-1 SAR data can significantly improve the detection of lineaments. Merge MSS and SAR data sets were more useful fo lineament detection and landcover classification than LANDSAT or SEASAT data alone. About 20 percent of the lineaments plotted from the SEASAT SAR image did not appear on the LANDSAT image. About 6 percent of minor lineaments or parts of lineaments present in the LANDSAT map were missing from the SEASAT map. Improvement in the landcover classification (acreage and spatial estimation accuracy) was attained by using MSS-SAR merged data. The aerial estimation of residential/built-up and forest categories was improved. Accuracy in estimating the agricultural and water categories was slightly reduced.

Sekhon, R.↗

Delineation of soil temperature regimes from HCMM data

Evaluation of LANDSAT and Heat Capacity Mapping Mission (HCMM) data as input into National Cooperative Soil Survey is discussed. Signature classification techniques were applied to 13 May 76 LANDSAT data. LANDSAT data was overlaid with HCMM data, revealing registration problems caused by a shortage of control points in LANDSAT data, and the WARP program developed to improve registration accuracy. Initial images for control point selection were produced using digital terrain elevation data. Statistical procedures for evaluating data classification and to describe spatial distribution of surface temperature and its correlation with soil surface conditions were investigated.

Day, R. L.↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗

Interpretation and automatic image enhancement facility

The author has identified the following significant results. Efforts to provide data processing support for ERTS-1 investigators in Kansas are summarized. Programs have been developed for data retrieval, feature extraction, and classification of digital MSS data. The IDECS/PDP-15 facility at the University of Kansas Remote Sensing Laboratory has been used for quick look analysis of ERTS-1 imagery. Programs have been developed for studying fresh water bodies in ERTS-1 imagery over Kansas on the IDECS.

Haralick, R. M.↗

An automated approach to the design of decision tree classifiers

The classification of large dimensional data sets arising from the merging of remote sensing data with more traditional forms of ancillary data is considered. Decision tree classification, a popular approach to the problem, is characterized by the property that samples are subjected to a sequence of decision rules before they are assigned to a unique class. An automated technique for effective decision tree design which relies only on apriori statistics is presented. This procedure utilizes a set of two dimensional canonical transforms and Bayes table look-up decision rules. An optimal design at each node is derived based on the associated decision table. A procedure for computing the global probability of correct classfication is also provided. An example is given in which class statistics obtained from an actual LANDSAT scene are used as input to the program. The resulting decision tree design has an associated probability of correct classification of .76 compared to the theoretically optimum .79 probability of correct classification associated with a full dimensional Bayes classifier. Recommendations for future research are included.

Argentiero, P.↗

One of These Things IS Like the Other: Pursuing a New Taxonomy of Industry for Improved Energy System Modeling

Industrial processes drive the exchange of materials, energy, and currency throughout the economy. These processes are powered by electricity and direct combustion, with variation in their operation even within the same industry. This heterogeneity makes it difficult for large models, including the National Energy Modeling System (US), to project their energy use while remaining tractable. Decarbonization and ensuing changes to the energy system require changes to industrial processes while offering opportunities for process innovation, but the extent and nature of changes are difficult to model with current classification schemes and corresponding data. The North American Industrial Classification (NAICS) is an economic taxonomy of industries, but its categories are less meaningful from an energy and material flow perspective. For example, a facility that makes steel from iron ore in a blast furnace/basic oxygen furnace is categorized under the same NAICS code as a facility that makes steel from scrap in an electric arc furnace despite the scale, use of recycled scrap versus iron ore, and energy use differences in the two facility types. Exploratory analysis is performed on a large dataset used for plant-level energy assessment in order to detect clusters that can aid in better modeling of industry for energy analysis in an evolving system with breakthrough technologies.

28 EE - Advanced Manufacturing Office (EE-5A)↗

Adaptive fuzzy leader clustering of complex data sets in pattern recognition

A modular, unsupervised neural network architecture for clustering and classification of complex data sets is presented. The adaptive fuzzy leader clustering (AFLC) architecture is a hybrid neural-fuzzy system that learns on-line in a stable and efficient manner. The initial classification is performed in two stages: a simple competitive stage and a distance metric comparison stage. The cluster prototypes are then incrementally updated by relocating the centroid positions from fuzzy C-means system equations for the centroids and the membership values. The AFLC algorithm is applied to the Anderson Iris data and laser-luminescent fingerprint image data. It is concluded that the AFLC algorithm successfully classifies features extracted from real data, discrete or continuous.

Newton, Scott C.↗

Necessity to adapt land use and land cover classification systems to readily accept radar data

A hierarchial, four level, standardized system for classifying land use/land cover primarily from remote-sensor data (USGS system) is described. The USGS system was developed for nonmicrowave imaging sensors such as camera systems and line scanners. The USGS system is not compatible with the land use/land cover classifications at different levels that can be made from radar imagery, and particularly from synthetic-aperture radar (SAR) imagery. The use of radar imagery for classifying land use/land cover at different levels is discussed, and a possible revision of the USGS system to more readily accept land use/land cover classifications from radar imagery is proposed.

Drake, B.↗

Feature Identification and Location Experiment

The Feature Identification and Location Experiment (FILE), which was flown on the second Space Shuttle flight to test a technique for real-time, autonomous classification of water, vegetation and bare land as well as clouds, snow and ice, senses earth radiation in spectral bands centered at 0.65 and 0.85 microns. The radiance ratio classification algorithm has successfully made automatic data selection decisions. A classification image obtained on the mission is providing data needed to evaluate the FILE algorithm and overall system performance.

Sivertson, W. E., Jr.↗

Classification and Fusion of Two Disparate Data Streams and Nuclear Dissolutions Application

We consider two streams of data or measurements with disparate qualities and time resolutions that need to be classified. The first stream consists of higher quality data at a coarser time resolution, and the other consists of lower quality data at a finer time resolution. We present a fuser-switch method that fuses the set of classifiers of each stream separately and switches between them. We show that this method provides classification decisions at a finer time resolution with superior detection and false alarm probabilities compared to individual classifiers, under the statistical independence and time resolution ratio conditions. When classifiers are trained using machine learning methods, we show that this superior performance is guaranteed with a confidence probability specified by the classifiers' generalization equations. We use these results to provide analytical foundations for previous practical results that achieved significant performance improvements in classifying Pu/Np target dissolution events at a radiochemical processing facility.

Rao, Nageswara↗

Identification of sea ice types in spaceborne synthetic aperture radar data

This study presents an approach for identification of sea ice types in spaceborne SAR image data. The unsupervised classification approach involves cluster analysis for segmentation of the image data followed by cluster labeling based on previously defined look-up tables containing the expected backscatter signatures of different ice types measured by a land-based scatterometer. Extensive scatterometer observations and experience accumulated in field campaigns during the last 10 yr were used to construct these look-up tables. The classification approach, its expected performance, the dependence of this performance on radar system performance, and expected ice scattering characteristics are discussed. Results using both aircraft and simulated ERS-1 SAR data are presented and compared to limited field ice property measurements and coincident passive microwave imagery. The importance of an integrated postlaunch program for the validation and improvement of this approach is discussed.

Kwok, Ronald↗

Statistical classification techniques for engineering and climatic data samples

Fisher's sample linear discriminant function is modified through an appropriate alteration of the common sample variance-covariance matrix. The alteration consists of adding nonnegative values to the eigenvalues of the sample variance covariance matrix. The desired results of this modification is to increase the number of correct classifications by the new linear discriminant function over Fisher's function. This study is limited to the two-group discriminant problem.

Temple, E. C.↗

Computer program documentation for the patch subsampling processor

The programs presented are intended to provide a way to extract a sample from a full-frame scene and summarize it in a useful way. The sample in each case was chosen to fill a 512-by-512 pixel (sample-by-line) image since this is the largest image that can be displayed on the Integrated Multivariant Data Analysis and Classification System. This sample size provides one megabyte of data for manipulation and storage and contains about 3% of the full-frame data. A patch image processor computes means for 256 32-by-32 pixel squares which constitute the 512-by-512 pixel image. Thus, 256 measurements are available for 8 vegetation indexes over a 100-mile square.

Nieves, M. J.↗

Mapping permafrost in the boreal forest with Thematic Mapper satellite data

A geographic data base incorporating Landsat TM data was used to develop and evaluate logistic discriminant functions for predicting the distribution of permafrost in a boreal forest watershed. The data base included both satellite-derived information and ancillary map data. Five permafrost classifications were developed from a stratified random sample of the data base and evaluated by comparison with a photo-interpreted permafrost map using contingency table analysis and soil temperatures recorded at sites within the watershed. A classification using a TM thermal band and a TM-derived vegetation map as independent variables yielded the highest mapping accuracy for all permafrost categories.

Morrissey, L. A.↗

Multisensor data analysis and its application to monitoring of cropland, forest, strip mines and cultural targets

Seasat L-band and aircraft X-band dual polarized synthetic aperture (SAR) data of the Western Kentucky Coal Region were examined, preprocessed and combined with Landsat Multispectral Scanner (MSS) data to form a seven-band multisensor data set. Analysis of classified data sets show that the three-band SAR data contain moderate discrimination accuracy for the strip mine class but low accuracy for the residential class. The four-band MSS data contain low classification accuracy for the strip mine and residential classes. The integrated five-band SAR/MSS data show that significant improvement in classification accuracy is obtained for both strip mine and residential classes.

Wu, S. T.↗

An unsupervised classification technique for multispectral remote sensing data.

Description of a two-part clustering technique consisting of (a) a sequential statistical clustering, which is essentially a sequential variance analysis, and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum-likelihood classification techniques.

Su, M. Y.↗

Crop classification using airborne radar and LANDSAT data

Airborne radar data acquired with a 13.3 GHz scatterometer over a test-site near Colby, Kansas were used to investigate the statistical properties of the scattering coefficient of three types of vegetation cover and of bare soil. A statistical model for radar data was developed that incorporates signal-fading and natural within-field variabilities. Estimates of the within-field and between-field coefficients of variation were obtained for each cover-type and compared with similar quantities derived from LANDSAT images of the same fields. The classification accuracy provided by LANDSAT alone, radar alone, and both sensors combined was investigated. The results indicate that the addition of radar to LANDSAT improves the classification accuracy by about 10; percentage-points when the classification is performed on a pixel basis and by about 15 points when performed on a field-average basis.

Ulaby, F. T.↗