Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Application of LANDSAT data to wetland study and land use classification in west Tennessee

The Obion-Forked Deer River Basin in northwest Tennessee is confronted with several acute land use problems which result in excessive erosion, sedimentation, pollution, and hydrologic runoff. LANDSAT data was applied to determine land use of selected watershed areas within the basin, with special emphasis on determining wetland boundaries. Densitometric analysis was performed to allow numerical classification of objects observed in the imagery on the basis of measurements of optical densities. Multispectral analysis of the LANDSAT imagery provided the capability of altering the color of the image presentation in order to enhance desired relationships. Manual mapping and classification techniques were performed in order to indicate a level of accuracy of the LANDSAT data as compared with high and low altitude photography for land use classification.

Jones, N. L.↗

Tests of a Semi-Analytical Case 1 and Gelbstoff Case 2 SeaWiFS Algorithm with a Global Data Set

A semi-analytical algorithm was tested with a total of 733 points of either unpackaged or packaged-pigment data, with corresponding algorithm parameters for each data type. The 'unpackaged' type consisted of data sets that were generally consistent with the Case 1 CZCS algorithm and other well calibrated data sets. The 'packaged' type consisted of data sets apparently containing somewhat more packaged pigments, requiring modification of the absorption parameters of the model consistent with the CalCOFI study area. This resulted in two equally divided data sets. A more thorough scrutiny of these and other data sets using a semianalytical model requires improved knowledge of the phytoplankton and gelbstoff of the specific environment studied. Since the semi-analytical algorithm is dependent upon 4 spectral channels including the 412 nm channel, while most other algorithms are not, a means of testing data sets for consistency was sought. A numerical filter was developed to classify data sets into the above classes. The filter uses reflectance ratios, which can be determined from space. The sensitivity of such numerical filters to measurement resulting from atmospheric correction and sensor noise errors requires further study. The semi-analytical algorithm performed superbly on each of the data sets after classification, resulting in RMS1 errors of 0.107 and 0.121, respectively, for the unpackaged and packaged data-set classes, with little bias and slopes near 1.0. In combination, the RMS1 performance was 0.114. While these numbers appear rather sterling, one must bear in mind what mis-classification does to the results. Using an average or compromise parameterization on the modified global data set yielded an RMS1 error of 0.171, while using the unpackaged parameterization on the global evaluation data set yielded an RMS1 error of 0.284. So, without classification, the algorithm performs better globally using the average parameters than it does using the unpackaged parameters. Finally, the effects of even more extreme pigment packaging must be examined in order to improve algorithm performance at high latitudes. Note, however, that the North Sea and Mississippi River plume studies contributed data to the packaged and unpackaged classess, respectively, with little effect on algorithm performance. This suggests that gelbstoff-rich Case 2 waters do not seriously degrade performance of the semi-analytical algorithm.

Carder, Kendall L.↗

A review and analysis of neural networks for classification of remotely sensed multispectral imagery

A literature survey and analysis of the use of neural networks for the classification of remotely sensed multispectral imagery is presented. As part of a brief mathematical review, the backpropagation algorithm, which is the most common method of training multi-layer networks, is discussed with an emphasis on its application to pattern recognition. The analysis is divided into five aspects of neural network classification: (1) input data preprocessing, structure, and encoding; (2) output encoding and extraction of classes; (3) network architecture, (4) training algorithms; and (5) comparisons to conventional classifiers. The advantages of the neural network method over traditional classifiers are its non-parametric nature, arbitrary decision boundary capabilities, easy adaptation to different types of data and input structures, fuzzy output values that can enhance classification, and good generalization for use with multiple images. The disadvantages of the method are slow training time, inconsistent results due to random initial weights, and the requirement of obscure initialization values (e.g., learning rate and hidden layer size). Possible techniques for ameliorating these problems are discussed. It is concluded that, although the neural network method has several unique capabilities, it will become a useful tool in remote sensing only if it is made faster, more predictable, and easier to use.

Paola, Justin D.↗

Identification of phenological stages and vegetative types for land use classification

The author has identified the following significant results. Classification of digital data for mapping Alaskan vegetation has been compared to ground truth data and found to have accuracies as high as 90%. These classifications are broad scale types as are currently being used on the Major Ecosystems of Alaska map prepared by the Joint Federal-State Land Use Planning Commission for Alaska. Cost estimates for several options using the ERTS-1 digital data to map the Alaskan land mass at the 1:250,000 scale ranged between $2.17 to $1.49 per square mile.

Mckendrick, J. D.↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

Computer-aided classification for remote sensing in agriculture and forestry in Northern Italy

A set of results concerning the processing and analysis of data from LANDSAT satellite and airborne scanner is presented. The possibility of performing inventories of irrigated crops-rice, planted groves-poplars, and natural forests in the mountians-beeches and chestnuts, is investigated in the Po valley and in an alphine site of Northern Italy. Accuracies around 95% or better, 70% and 60% respectively are achieved by using LANDSAT data and supervised classification. Discrimination of rice varieties is proved with 8 channels data from airborne scanner, processed after correction of the atmospheric effect due to the scanning angle, with and without linear feature selection of the data. The accuracies achieved range from 65% to more than 80%. The best results are obtained with the maximum likelihood classifier for normal parameters but rather close results are derived by using a modified version of the weighted euclidian distance between points, with consequent decrease in computing time around a factor 3.

Dejace, J.↗

Wheat classification exercise, using 11 June 1973, ERTS MSS data for Fayette County, Illinois (for CITARS task)

The prime emphasis was on classification of pixels in field centers, away from boundary effects. Results were encouraging in both training and test field centers for wheat and other major types of vegetation present. However, the location of fields was found to be a serious problem and it was even more difficult to select field-center pixels for fields of sizes less than 20 acres (or even larger, depending upon field shape) for use in the field-center analysis. The majority of fields in the segment are less than 20 acres in size. ERTS-1 data were received on 12 September 1973. Ground truth information and aerial photography were received on 9 and 15 September. The data were analyzed and processed digitally using the ERIM multispectral software system.

Malila, W. A.↗

Evaluation of change detection techniques for monitoring coastal zone environments

The author has identified the following significant results. Four change detection techniques were designed and implemented for evaluation: (1) post classification comparison change detection, (2) delta data change detection, (3) spectral/temporal change classification, and (4) layered spectral/temporal change classification. The post classification comparison technique reliably identified areas of change and was used as the standard for qualitatively evaluating the other three techniques. The layered spectral/temporal change classification and the delta data change detection results generally agreed with the post classification comparison technique results; however, many small areas of change were not identified. Major discrepancies existed between the post classification comparison and spectral/temporal change detection results.

Weismiller, R. A.↗

Artificial neural network classification using a minimal training set - Comparison to conventional supervised classification

Recent research has shown an artificial neural network (ANN) to be capable of pattern recognition and the classification of image data. This paper examines the potential for the application of neural network computing to satellite image processing. A second objective is to provide a preliminary comparison and ANN classification. An artificial neural network can be trained to do land-cover classification of satellite imagery using selected sites representative of each class in a manner similar to conventional supervised classification. One of the major problems associated with recognition and classifications of pattern from remotely sensed data is the time and cost of developing a set of training sites. This reseach compares the use of an ANN back propagation classification procedure with a conventional supervised maximum likelihood classification procedure using a minimal training set. When using a minimal training set, the neural network is able to provide a land-cover classification superior to the classification derived from the conventional classification procedure. This research is the foundation for developing application parameters for further prototyping of software and hardware implementations for artificial neural networks in satellite image and geographic information processing.

Hepner, George F.↗

Per-point and per-field contextual classification of multipolarization and multiple incidence angle aircraft L-band radar data

Multipolarized aircraft L-band radar data are classified using two different image classification algorithms: (1) a per-point classifier, and (2) a contextual, or per-field, classifier. Due to the distinct variations in radar backscatter as a function of incidence angle, the data are stratified into three incidence-angle groupings, and training and test data are defined for each stratum. A low-pass digital mean filter with varied window size (i.e., 3x3, 5x5, and 7x7 pixels) is applied to the data prior to the classification. A predominately forested area in northern Florida was the study site. The results obtained by using these image classifiers are then presented and discussed.

Hoffer, Roger M.↗

The role of spatial, spectral and radiometric resolution on information content

The results of a factorial experiment to evaluate the effects of spatial, spectral, and radiometric resolution on training-data spectral separability and classification accuracy are reported. Aircraft scanner data from five flightlines at 19.8 km over California including croplands, rangeland, forest, water, and urban areas were systematically degraded over a range approximately from Landsat MSS to Thematic Mapper specifications. Reference data were collected on the ground and from aerial photography. The degradations, training-site delineation, data-analysis procedures, and accuracy-assessment techniques are described; the results are presented in tables and graphs and discussed. It is found that while accuracy was increased by higher spectral resolution in 70 percent of the cases and uniformly by increased radiometric resolution, it was decreased by higher spatial resolution. This phenomenon is attributed to classification methods.

Buis, J. S.↗

Sea Ice Identification using Dual-Polarized Ku-Band Scatterometer Data

This paper describes a classification algorithm using dual-polarized scatterometer measurements to identify the edge of the sea ice cover. The distinct polarization scattering of sea ice and open water are discussed and illustrated with the dual-polarized radar measurements from the Seasat-A scatterometer (SASS).

Sea↗

Design of partially supervised classifiers for multispectral image data

A partially supervised classification problem is addressed, especially when the class definition and corresponding training samples are provided a priori only for just one particular class. In practical applications of pattern classification techniques, a frequently observed characteristic is the heavy, often nearly impossible requirements on representative prior statistical class characteristics of all classes in a given data set. Considering the effort in both time and man-power required to have a well-defined, exhaustive list of classes with a corresponding representative set of training samples, this 'partially' supervised capability would be very desirable, assuming adequate classifier performance can be obtained. Two different classification algorithms are developed to achieve simplicity in classifier design by reducing the requirement of prior statistical information without sacrificing significant classifying capability. The first one is based on optimal significance testing, where the optimal acceptance probability is estimated directly from the data set. In the second approach, the partially supervised classification is considered as a problem of unsupervised clustering with initially one known cluster or class. A weighted unsupervised clustering procedure is developed to automatically define other classes and estimate their class statistics. The operational simplicity thus realized should make these partially supervised classification schemes very viable tools in pattern classification.

Jeon, Byeungwoo↗

Automated classification of big X-ray diffraction data using deep learning models

Abstract In current in situ X-ray diffraction (XRD) techniques, data generation surpasses human analytical capabilities, potentially leading to the loss of insights. Automated techniques require human intervention, and lack the performance and adaptability required for material exploration. Given the critical need for high-throughput automated XRD pattern analysis, we present a generalized deep learning model to classify a diverse set of materials’ crystal systems and space groups. In our approach, we generate training data with a holistic representation of patterns that emerge from varying experimental conditions and crystal properties. We also employ an expedited learning technique to refine our model’s expertise to experimental conditions. In addition, we optimize model architecture to elicit classification based on Bragg’s Law and use evaluation data to interpret our model’s decision-making. We evaluate our models using experimental data, materials unseen in training, and altered cubic crystals, where we observe state-of-the-art performance and even greater advances in space group classification.

Chemistry↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Multimodal Few-Shot Segmentation of Electron Micrographs

Scanning transmission electron microscopy (STEM) is one of the most used methods of analyzing the chemistry and composition of materials. By analyzing microstructures, these microscopes can help scientists better understand the molecular underpinnings of microelectronics, batteries, and more. However, STEM data can be difficult to interpret, so recent developments have been made in applications of machine learning to analyze these images. The PNNL-developed pyCHIP Classifier has achieved results in segmenting STEM these images via few-shot learning, a method which requires little data and human input, perfect for quickly analysis. In my internship I (Eli Meyers) investigated a multimodal improvement of this classifier by incorporating energy dispersive x-ray spectroscopy (EDS) data into the classification process for a more accurate segmentation. Furthermore, I encoded the spectral data by training a mass spectrometry encoder on the EDS data to extract a more meaningful representation of the data.

36 MATERIALS SCIENCE↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗