Engineering PapersSearch

SEARCH · Engineering Papers

Results for “multiclass”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multiclass Reduced-Set Support Vector Machines

There are well-established methods for reducing the number of support vectors in a trained binary support vector machine, often with minimal impact on accuracy. We show how reduced-set methods can be applied to multiclass SVMs made up of several binary SVMs, with significantly better results than reducing each binary SVM independently. Our approach is based on Burges' approach that constructs each reduced-set vector as the pre-image of a vector in kernel space, but we extend this by recomputing the SVM weights and bias optimally using the original SVM objective function. This leads to greater accuracy for a binary reduced-set SVM, and also allows vectors to be 'shared' between multiple binary SVMs for greater multiclass accuracy with fewer reduced-set vectors. We also propose computing pre-images using differential evolution, which we have found to be more robust than gradient descent alone. We show experimental results on a variety of problems and find that this new approach is consistently better than previous multiclass reduced-set methods, sometimes with a dramatic difference.

reduced set methods

A parametric multiclass Bayes error estimator for the multispectral scanner spatial model performance evaluation

The author has identified the following significant results. The probability of correct classification of various populations in data was defined as the primary performance index. The multispectral data being of multiclass nature as well, required a Bayes error estimation procedure that was dependent on a set of class statistics alone. The classification error was expressed in terms of an N dimensional integral, where N was the dimensionality of the feature space. The multispectral scanner spatial model was represented by a linear shift, invariant multiple, port system where the N spectral bands comprised the input processes. The scanner characteristic function, the relationship governing the transformation of the input spatial, and hence, spectral correlation matrices through the systems, was developed.

Mobasseri, B. G.

Multiclass Flight Anomaly Detection Using Sensor Fusion Based on Dempster-Shafer Theory

As aviation systems in commercial operations continue to grow in complexity, the anomalies exhibited by these systems become more elaborate and difficult to detect. To address the challenge of detecting these complex anomalies, deep learning models have been used extensively in aviation anomaly detection studies, at the expense of end-user interpretability. Aiming to maintain the same level of interpretability as traditional threshold-exceedance methods, we continue our development of prediction models using ordinal patterns and their distributions throughout the flight. Specifically, this study extends our work into multiclass anomaly detection using sensor fusion based on Dempster-Shafer theory (DST), a second-order probability theory used to combine information from different sources of evidence. Our approach uses DST to reduce the uncertainty in the class predictions of an ensemble of classifiers. These classifiers rely on the similarity between flight data and class templates to make a prediction of the state of the aircraft. Our approach aims to take advantage of simple models trained on interpretable features (ordinal patterns) to correctly predict an anomaly and identify the flight dynamics linked to the anomaly. Our results show an improvement when using DST-based sensor fusion over simple majority voting. Additionally, our results provide insight into aircraft states linked to rare high-risk anomalies.

Risk detection

Multiclass Flight Anomaly Detection Using Sensor Fusion Based on Dempster-Shafer Theory

As aviation systems in commercial operations continue to grow in complexity, the anomalies exhibited by these systems become more elaborate and difficult to detect. To address the challenge of detecting these complex anomalies, deep learning models have been used extensively in aviation anomaly detection studies, at the expense of end-user interpretability. Aiming to maintain the same level of interpretability as traditional threshold-exceedance methods, we continue our development of prediction models using ordinal patterns and their distributions throughout the flight. Specifically, this study extends our work into multiclass anomaly detection using sensor fusion based on Dempster-Shafer theory (DST), a second-order probability theory used to combine information from different sources of evidence. Our approach uses DST toreduce the uncertainty in the class predictions of an ensemble of classifiers. These classifiers rely on the similarity between flight data and class templates to make a prediction of the state of the aircraft. Our approach aims to take advantage of simple models trained on interpretable features (ordinal patterns) to correctly predict an anomaly and identify the flight dynamics linked to the anomaly. Our results show an improvement when using DST-based sensor fusion over simple majority voting. Additionally, our results provide insight into aircraft states linked to rare high-risk anomalies.

Risk detection

Multiclass Bayes error estimation by a feature space sampling technique

A general Gaussian M-class N-feature classification problem is defined. An algorithm is developed that requires the class statistics as its only input and computes the minimum probability of error through use of a combined analytical and numerical integration over a sequence simplifying transformations of the feature space. The results are compared with those obtained by conventional techniques applied to a 2-class 4-feature discrimination problem with results previously reported and 4-class 4-feature multispectral scanner Landsat data classified by training and testing of the available data.

Mobasseri, B. G.

Multiclass Continuous Correspondence Learning

We extend the Structural Correspondence Learning (SCL) domain adaptation algorithm of Blitzer er al. to the realm of continuous signals. Given a set of labeled examples belonging to a 'source' domain, we select a set of unlabeled examples in a related 'target' domain that play similar roles in both domains. Using these 'pivot samples, we map both domains into a common feature space, allowing us to adapt a classifier trained on source examples to classify target examples. We show that when between-class distances are relatively preserved across domains, we can automatically select target pivots to bring the domains into correspondence.

correspondence learning

Investigation of spatial misregistration effects in multispectral scanner data

The author has identified the following significant results. A model for estimating the expected proportion of multiclass pixels in a scene was generalized and extended to include misregistration effects. Another substantial effort was the development of a simulation model to generate signatures to represent the distributions of signals from misregistered multiclass pixels, based on single class signatures. Spatial misregistration causes an increase in the proportion of multiclass pixels in a scene and a decorrelation between signals in misregistered data channels. The multiclass pixel proportion estimation model indicated that this proportion is strongly dependent on the pixel perimeter and on the ratio of the total perimeter of the fields in the scene to the area of the scene. Test results indicated that expected values computed with this model were similar to empirical measurements made of this proportion in four LACIE data segments.

Nalepka, R. F.

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity

Two effective feature selection criteria for multispectral remote sensing

Distance measures which are useful for feature selection are considered, giving attention to the divergence distance measure and the Jeffreys-Matusita (JM) distance. Experimental studies show that the JM-distance yields more reliable results than other distance measures. A number of questions which are not solved by the experiments are discussed. The investigation provides an explanation for previous observations that the JM-distance and a saturating transform of divergence are highly useful for feature selection in the multiclass case.

Swain, P. H.

Data processing large quantities of multispectral information

Method is combination of digital and optical techniques. Multispectral data is coded into binary matrix format and then encoded onto photographic film. Film is holographically correlated with spectral signature to generate single-class classification map. Number of maps are optically superimposed to produce full-color, multiclass classification map.

Haskell, R. E.

The decision tree approach to classification

A class of multistage decision tree classifiers is proposed and studied relative to the classification of multispectral remotely sensed data. The decision tree classifiers are shown to have the potential for improving both the classification accuracy and the computation efficiency. Dimensionality in pattern recognition is discussed and two theorems on the lower bound of logic computation for multiclass classification are derived. The automatic or optimization approach is emphasized. Experimental results on real data are reported, which clearly demonstrate the usefulness of decision tree classifiers.

Wu, C.

Development and application of operational techniques for the inventory and monitoring of resources and uses for the Texas coastal zone

The author has identified the following significant results. The most significant ADP result was the modification of the DAM package to produce classified printouts, scaled and registered to U.S.G.S., 71/2 minute topographic maps from LARSYS-type classification files. With this modification, all the powerful scaling and registration capabilities of DAM become available for multiclass classification files. The most significant results with respect to image interpretation were the application of mapping techniques to a new, more complex area, and the refinement of an image interpretation procedure which should yield the best results.

Jones, R.

A new computer approach to mixed feature classification for forestry application

A computer approach for mapping mixed forest features (i.e., types, classes) from computer classification maps is discussed. Mixed features such as mixed softwood/hardwood stands are treated as admixtures of softwood and hardwood areas. Large-area mixed features are identified and small-area features neglected when the nominal size of a mixed feature can be specified. The computer program merges small isolated areas into surrounding areas by the iterative manipulation of the postprocessing algorithm that eliminates small connected sets. For a forestry application, computer-classified LANDSAT multispectral scanner data of the Sam Houston National Forest were used to demonstrate the proposed approach. The technique was successful in cleaning the salt-and-pepper appearance of multiclass classification maps and in mapping admixtures of softwood areas and hardwood areas. However, the computer-mapped mixed areas matched very poorly with the ground truth because of inadequate resolution and inappropriate definition of mixed features.

Kan, E. P.

A counter-example in linear feature selection theory

The paper shows that it is possible to construct two k x n matrices, both of which maximize divergence in the transformed space of the linear feature selection problem in multiclass pattern recognition, and which are not row equivalent. Thus, even under extremely strong conditions, it is not possible to assume that all matrix solutions which maximize transformed divergence are row equivalent.

Brown, D. R.

Wheat signature modeling and analysis for improved training statistics

The author has identified the following significant results. The spectral, spatial, and temporal characteristics of wheat and other signatures in LANDSAT multispectral scanner data were examined through empirical analysis and simulation. Irrigation patterns varied widely within Kansas; 88 percent of wheat acreage in Finney was irrigated and 24 percent in Morton, as opposed to less than 3 percent for western 2/3's of the State. The irrigation practice was definitely correlated with the observed spectral response; wheat variety differences produced observable spectral differences due to leaf coloration and different dates of maturation. Between-field differences were generally greater than within-field differences, and boundary pixels produced spectral features distinct from those within field centers. Multiclass boundary pixels contributed much of the observed bias in proportion estimates. The variability between signatures obtained by different draws of training data decreased as the sample size became larger; also, the resulting signatures became more robust and the particular decision threshold value became less important.

Nalepka, R. F.