Development of mathematical techniques for the analysis of remote sensing data
There are no author-identified significant results in this report.
Engineering topics
Publications and source records attributed to Decell, H. P., Jr..
There are no author-identified significant results in this report.
Several ways in which feature selection techniques were used in LACIE are discussed. In all cases, the methods require some a priori information and assumptions; in most, the classification procedure (Bayes optimal) was chosen in advance. The transformations used for dimensionality reduction are linear, that is, the variables in feature space are always linear combinations of the original measurements. Several numerically tractable criteria developed for LACIE, which provide information about the probability of misclassification, are discussed. Recent results on linear feature selection techniques are included. Their use in LACIE is discussed. Related open questions are mentioned.
There are no author-identified significant results in this report.
There are no author-identified significant results in this report.
An explicit expression for a compression matrix T of smallest possible left dimension K consistent with preserving the n variate normal Bayes assignment of X to a given one of a finite number of populations and the K variate Bayes assignment of TX to that population was developed. The Bayes population assignment of X and TX were shown to be equivalent for a compression matrix T explicitly calculated as a function of the means and covariances of the given populations.
A surjective bounded linear operator T from a Banach space X to a Banach space Y must be a sufficient statistic for a dominated family of probability measures defined on the Borel sets of X. These results were applied, so that they characterize linear sufficient statistics for families of the exponential type, including as special cases the Wishart and multivariate normal distributions. The latter result was used to establish precisely which procedures for sampling from a normal population had the property that the sample mean was a sufficient statistic.
Several candidate signature extensions algorithms were evaluated. Four simulated data sets and seven consecutive day data sets were used. Analysis and evaluation of the UHMLE test and recommendations on changes in the UHMLE algorithm motivated by the test are given for each data set. The criterion for evaluation of each algorithm was overall classification accuracy.
A necessary and sufficient condition is developed such that there exists a continous linear sufficient statistic T for a dominated collection of totally finite measures defined on the Borel field generated by the open sets of a Banach space X. In particular, corollary necessary and sufficient conditions are given so that there exists a rank K linear sufficient statistic T for any finite collection of probability measures having n-variate normal densities. In this case a simple calculation, involving only the population means and covariances, determines the smallest integer K for which there exists a rank K linear sufficient statistic T (as well as an associated statistic T itself).
A procedure for calculating a kxn rank k matrix B for data compression using the Bhattacharyya bound on the probability of error and an iterative construction using Householder transformation was developed. Two sets of remotely sensed agricultural data are used to demonstrate the application of the procedure. The results of the applications gave some indication of the extent to which the Bhattacharyya bound on the probability of error is affected by such transformations for multivariate normal populations.
Classifying large quantities of multidimensional remotely sensed agricultural data requires efficient and effective classification techniques and the construction of certain transformations of a dimension reducing, information preserving nature. The construction of transformations that minimally degrade information (i.e., class separability) is described. Linear dimension reducing transformations for multivariate normal populations are presented. Information content is measured by divergence.
Results that suggest the possibility of using a sequential monotone process for solving the feature selection problem using Householder transformations are applied to the divergence separability criterion and an expression for the gradient of the divergence with respect to the generator of a single Householder transformation will be developed. This expression for the gradient is used in any number of differential correction schemes (iterators) that attempt to extremize the divergence. Data sets provided by the Earth Observations Division-JSC are used to demonstrate selecting the Householder transformations that generate the kxn matrix defining the best (in the sense of extremizing the divergence) k linear combinations of features. The tests allow initial comparisons to be made with results. In particular, this technique does not appear to require initial guesses for the iterator to be generated without replacement, exhaustive search, or other similar schemes.
An outline for an Image 100 procedures manual for Earth Resources Program image analysis was developed which sets forth guidelines that provide a basis for the preparation and updating of an Image 100 Procedures Manual. The scope of the outline was limited to definition of general features of a procedures manual together with special features of an interactive system. Computer programs were identified which should be implemented as part of an applications oriented library for the system.
Several theorems related to the Householder transformation and separability criteria are proven. Orthogonal transformations, topology, divergence, mathematical matrices, and group theory are discussed.
The digital calculations that drive the classification portion of the Hybrid Pattern Recognition System (HYMPS) are considered. A revision of these calculations is suggested that involves three items: matrix inversion, det calculations and singularity, and covariance factorization. It is shown that it is more economical to first factor the covariance matrix, and by so doing, delete the matrix inversion routine. In addition, the necessary det calculations can be more easily realized by use of simple theoretical facts about the factorization.
Classical iterative methods in nonlinear regression are reviewed and improved upon. This is accomplished by discussion of the geometrical and theoretical motivation for introducing modifications using generalized matrix inversion. Examples having inherent pitfalls are presented and compared in terms of results obtained using classical and modified techniques. The modification is shown to be useful alone or in conjunction with other modifications appearing in the literature.
The problem dealt with concerns feature selection or reducing the dimension of the data to be processed from n to k. By reducing the dimension of the data from n to k, classification time is generally reduced. Yet the dimension reduction should not be so great that classification accuracy is impaired. Thus, the general problem is considered of classifying an n-dimensional observation vector x into one of m-distinct classes where each class is normally distributed with mean and covariance. It is shown that the probability of misclassification is minimized if a maximum likelihood classification procedure is used to classify the data. The dimension of each observation vector to be processed is conveniently reduced by performing the transformation y = Bx, where B is a K by n matrix of rank k. Thus, the n-dimensional classification problem transforms into a k-dimensional classification problem.
The B-average divergence for m-distinct classes, resulting from the linear transformation y = Bx, is proposed as a feature selection criterion, where B is a k by n matrix of rank k not greater than n. It is shown that if the B-average divergence resulting from B is large enough, then the probability of misclassification, considered as a function f the class of all k by n matrices, is essentially minimized by B. A computer program, utilizing a gradient procedure, is developed to numerically maximize the B-average divergence and results are presented for the Cl flight line. For this example, corresponding to 9-distinct classes, most of the discriminatory information is found to lie in a 3-dimensional subspace, defined by an appropriately chosen 3 by 12 matrix B.
A technique is developed for selecting from n-channel multispectral data some k combinations of the n-channels upon which to base a given classification technique so that some measure of the loss of the ability to distinguish between classes, using the compressed k-dimensional data, is minimized. Information loss in compressing the n-channel data to k channels is taken to be the difference in the average interclass divergences (or probability of misclassification) in n-space and in k-space.