Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Misclassification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An Expression for the Transformed Covariance Matrix of Multivariate Normal Populations

The problem is considered of classifying into one of m distinct n-variate classes an arbitrary n-channel multispectral measurement vector x. The classification procedure used is the maximum likelihood procedure. Information loss in compressing the n-channel data to k channels is taken to be the difference in the average interclass divergences (or probability of misclassification) in n-space and in k-space. Data compression is accomplished by kxn linear transformation i.e., multiplication of the spectral n-vector by a kxn matrix of rank k.

Decell, H. P., Jr.

An Iterative Approach to the Feature Selection Problem

The problem dealt with concerns feature selection or reducing the dimension of the data to be processed from n to k. By reducing the dimension of the data from n to k, classification time is generally reduced. Yet the dimension reduction should not be so great that classification accuracy is impaired. Thus, the general problem is considered of classifying an n-dimensional observation vector x into one of m-distinct classes where each class is normally distributed with mean and covariance. It is shown that the probability of misclassification is minimized if a maximum likelihood classification procedure is used to classify the data. The dimension of each observation vector to be processed is conveniently reduced by performing the transformation y = Bx, where B is a K by n matrix of rank k. Thus, the n-dimensional classification problem transforms into a k-dimensional classification problem.

Decell, H. P., Jr.

A study of techniques for processing multispectral scanner data

A linear decision rule to reduce the time required for processing multispectral scanner data is developed. Test results are presented which justify the use of the new rule for digital processing whenever both accuracy and processing time are important. A method of evaluating the performance of the rule is also developed and applied to the problem of choosing a subset of channels. A technique used to find linear combinations of channels is described. The ability to extend signatures throughout a small area of approximately fifty square miles is tested. After preprocessing, signatures derived from the first of seven overlapping data sets are applied to all data sets. The test results show that the average probability of misclassification tends to increase with an increase in the number of data sets over which the signatures are extended.

Crane, R. B.

An iterative approach to the feature selection problem

The B-average divergence for m-distinct classes, resulting from the linear transformation y = Bx, is proposed as a feature selection criterion, where B is a k by n matrix of rank k not greater than n. It is shown that if the B-average divergence resulting from B is large enough, then the probability of misclassification, considered as a function f the class of all k by n matrices, is essentially minimized by B. A computer program, utilizing a gradient procedure, is developed to numerically maximize the B-average divergence and results are presented for the Cl flight line. For this example, corresponding to 9-distinct classes, most of the discriminatory information is found to lie in a 3-dimensional subspace, defined by an appropriately chosen 3 by 12 matrix B.

Decell, H. P., Jr.

A Kalman filter approach to adaptive estimation of multispectral signatures

The signatures of remote sensing data from agricultural crops exhibit significant non-stationarity, so that the performance of fixed parameter classifiers degenerates with time and distance from the initial training data. A class of adaptive decision-directed classifiers are being developed, based on Kalman filter theory. Limited results to date on two data sets indicate approximately a 25 to 40% reduction in rates of misclassification.

Crane, R. B.

Land use mapping in Erie County, Pennsylvania: A pilot study

The author has identified the following significant results. A pilot study was conducted to determine the feasibility of mapping land use in the Great Lakes Basin area utilizing ERTS-1 data. Small streams were clearly defined by the presence of trees along their length in predominantly agricultural country. Field patterns were easily differentiated from forested areas; dairy and beef farms were differentiated from other farmlands, but no attempt was made to identify crops. Large railroad lines and major highway systems were identified. The city of Erie and several smaller towns were identified, as well as residential areas between these towns, and docks along the shoreline in Erie. Marshes, forests, and beaches within Presque Isle State Park were correctly identified, using the DCLUS program. Bay water was differentiated from lake water, with a small amount of misclassification.

Mcmurtry, G. J.

A method for estimating proportions

A proportion estimation procedure is presented which requires only on set of ground truth data for determining the error matrix. The error matrix is then used to determine an unbiased estimate. The error matrix is shown to be directly related to the probability of misclassifications, and is more diagonally dominant with the increase in the number of passes used.

Guseman, L. F., Jr.

A sequential nonparametric pattern classification algorithm based on the Wald SPRT

A sequential nonparametric pattern classification procedure is presented. The method presented is an estimated version of the Wald sequential probability ratio test (SPRT). This method utilizes density function estimates, and the density estimate used is discussed, including a proof of convergence in probability of the estimate to the true density function. The classification procedure proposed makes use of the theory of order statistics, and estimates of the probabilities of misclassification are given. The procedure was tested on discriminating between two classes of Gaussian samples and on discriminating between two kinds of electroencephalogram (EEG) responses.

Poage, J. L.

Mapping of the wildland fuel characteristics of the Santa Monica mountains of Southern California

LANDSAT digital data was successfully used to map and evaluate the wildland fuels of the Santa Monica Mountains in Southern California. A mixed classification scheme was used where training areas of known vegetation types were entered and the maximum likelihood classifier run, followed by an evaluation of the results and an unsupervised retraining of the classifier using an image of the probability of misclassification. Estimation of maturity class and crown closure percents of the major cover types were assigned to each computer class by associating the photointerpretation of 159 large scale photo samples with the resultant computer classes using analysis of variance and analysis of categorized data. The result of the computer classification and statistical analysis were then transformed from the LANDSAT Coordinate California State Plane Coordinate system for use in a digital format in the FIRESCOPE data retrieval and fire modeling system.

Nichols, J. D.

Acreage estimation, feature selection, and signature extension dependent upon the maximum likelihood decision rule

A maximum likelihood estimation technique is used for the analysis of agricultural remote sensor data. The m-class probability of misclassification is estimated using unlabeled test samples and labeled training samples. A bound on the variance of a proposed unbiased estimator of the m-class probability of error is derived. The particular case in which each class density is assumed to be a mixture of multivariate normal densities is considered. The extension of spectral signatures in space and time is discussed.

Quirein, J. A.

A general non-parametric classifier applied to discriminating surface water from terrain shadows

A general non-parametric classifier is described in the context of discriminating surface water from terrain shadows. In addition to using non-parametric statistics, this classifier permits the use of a cost matrix to assign different penalties to various types of misclassifications. The approach also differs from conventional classifiers in that it applies the maximum-likelihood criterion to overall class probabilities as opposed to the standard practice of choosing the most likely individual subclass. The classifier performance is evaluated using two different effectiveness measures for a specific set of ERTS data.

Eppler, W. G.

A procedure used for a ground truth study of a land use map of North Alabama generated from LANDSAT data

A land use map of a five county area in North Alabama was generated from LANDSAT data using a supervised classification algorithm. There was good overall agreement between the land use designated and known conditions, but there were also obvious discrepancies. In ground checking the map, two types of errors were encountered - shift and misclassification - and a method was developed to eliminate or greatly reduce the errors. Randomly selected study areas containing 2,525 pixels were analyzed. Overall, 76.3 percent of the pixels were correctly classified. A contingency coefficient of correlation was calculated to be 0.7 which is significant at the alpha = 0.01 level. The land use maps generated by computers from LANDSAT data are useful for overall land use by regional agencies. However, care must be used when making detailed analysis of small areas. The procedure used for conducting the ground truth study together with data from representative study areas is presented.

Downs, S. W., Jr.

Number of signatures necessary for accurate classification

This paper presents a procedure for determining the number of signatures to use in classifying multispectral scanner data. A large initial set of signatures is obtained by clustering the training points within each category (such as 'wheat' or 'other') to be recognized. These clusters are then combined into broader signatures by a program that considers each pair of signatures within a category, combines the best pair in the light of certain criteria, saves the combined signature and repeats the procedure until there is one signature for each category. The result is a collection of sets of signatures, one set for each number between the number of initial clusters and the number of categories. With the aid of statistics such as an estimate of the probability of misclassification between categories, the user can choose the smallest set satisfying his requirements for classification accuracy.

Richardson, W.

Tabular data base construction and analysis from thematic classified Landsat imagery of Portland, Oregon

A systematic verification of Landsat data classifications of the Portland, Oregon metropolitan area has been undertaken on the basis of census tract data. The degree of systematic misclassification due to the Bayesian classifier used to process the Landsat data was noted for the various suburban, industrialized and central business districts of the metropolitan area. The Landsat determinations of residential land use were employed to estimate the number of automobile trips generated in the region and to model air pollution hazards.

Bryant, N. A.

ISODATA: Thresholds for splitting clusters

The author has identified the following significant results. The parameter AD (average distance) as used in the ISODATA program was critically examined. Thresholds of AD to decide on the splitting of clusters were obtained. For the univariate case, 0.84 was established as a sound choice, after examining several simple, as well as composite, distributions and also after investigating the probability of misclassification when points have to be reassigned to the newly identified clusters. For the multivariate case, the empirical threshold (N-0.16)/square root of N was extrapolated. A final criticism on AD was that AD would lose its effectiveness as a discriminative measure for the present purpose when N was large.

Kan, E. P. F.

Speech as a pilot input medium

The speech recognition system under development is a trainable pattern classifier based on a maximum-likelihood technique. An adjustable uncertainty threshold allows the rejection of borderline cases for which the probability of misclassification is high. The syntax of the command language spoken may be used as an aid to recognition, and the system adapts to changes in pronunciation if feedback from the user is available. Words must be separated by .25 second gaps. The system runs in real time on a mini-computer (PDP 11/10) and was tested on 120,000 speech samples from 10- and 100-word vocabularies. The results of these tests were 99.9% correct recognition for a vocabulary consisting of the ten digits, and 99.6% recognition for a 100-word vocabulary of flight commands, with a 5% rejection rate in each case. With no rejection, the recognition accuracies for the same vocabularies were 99.5% and 98.6% respectively.

Plummer, R. P.