Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Safe and Robust Binary Classification and Fault Detection Using Reinforcement Learning

In this paper, we propose a learning-based method utilizing the Soft Actor-Critic (SAC) algorithm to train a binary Support Vector Machine (SVM) classifier. This classifier is designed to identify valid input spaces in high-dimensional, highly constrained systems while minimizing the total runtime of offline simulations. The simulations adapt their runtime based on the likelihood that a given training input will be informative to the classifier. Furthermore, we introduce a method for using the trained SAC model to predict whether a desired system input is likely to violate constraints, along with a technique to adjust the input as necessary. Additionally, we explore the potential of this model to detect faults or adversarial attacks within the system. The effectiveness of our approach is demonstrated through various simulations of challenging classification problems and a constrained quadrotor model.

Netter, Josh [Georgia Institute of Technology, Atl↗

Training a Quantum Annealing Based Restricted Boltzmann Machine on Cybersecurity Data

A restricted Boltzmann machine (RBM) is a generative model that could be used in effectively balancing a cybersecurity dataset because the synthetic data a RBM generates follows the probability distribution of the training data. RBM training can be performed using contrastive divergence (CD) and quantum annealing (QA). QA-based RBM training is fundamentally different from CD and requires samples from a quantum computer. We present a real-world application that uses a quantum computer. Specifically, we train a RBM using QA for cybersecurity applications. The D-Wave 2000Q has been used to implement QA. RBMs are trained on the ISCX data, which is a benchmark dataset for cybersecurity. For comparison, RBMs are also trained using CD. CD is a commonly used method for RBM training. Our analysis of the ISCX data shows that the dataset is imbalanced. We present two different schemes to balance the training dataset before feeding it to a classifier. The first scheme is based on the undersampling of benign instances. The imbalanced training dataset is divided into five sub-datasets that are trained separately. A majority voting is then performed to get the result. Our results show the majority vote increases the classification accuracy up from 90.24% to 95.68%, in the case of CD. For the case of QA, the classification accuracy increases from 74.14% to 80.04%. In the second scheme, a RBM is used to generate synthetic data to balance the training dataset. We show that both QA and CD-trained RBM can be used to generate useful synthetic data. Balanced training data is used to evaluate several classifiers. Among the classifiers investigated, K-Nearest Neighbor (KNN) and Neural Network (NN) perform better than other classifiers. They both show an accuracy of 93%. Our results show a proof-of-concept that a QA-based RBM can be trained on a 64-bit binary dataset. The illustrative example suggests the possibility to migrate many practical classification problems to QA-based techniques. Further, we show that synthetic data generated from a RBM can be used to balance the original dataset.

97 MATHEMATICS AND COMPUTING↗

GEDet: Detecting Erroneous Nodes with A Few Examples

Detecting nodes with erroneous values in real-world graphs re- mains challenging due to the lack of examples and various error scenarios. We demonstrate GEDet, an error detection engine that can detect erroneous nodes in graphs with a few examples. The GEDet framework tackles error detection as a few-shot node classification problem. We invite the attendees to experience the following unique features. (1) Few-shot detection. Users only need to provide a few examples of erroneous nodes to perform error detection with GEDet. GEDet achieves desirable accuracy with (a) a graph augmentation module, which automatically generates synthetic examples to learn the classifier, and (b) an adversarial detection module, which improves classifiers to better distinguish erroneous nodes from both cleaned nodes and synthetic examples. We show that GEDet significantly improves the state-of-the-art error detection methods. (2) Diverse error scenarios. GEDet profiles data errors with a built-in library of transformation functions from correct values to errors. Users can also easily “plug in” new error types or examples. (3) User-centric detection. GEDet supports (a) an active learning mode to engage users to verify detected results, and adapts the error detection process accordingly; and (b) visual interfaces to interpret and track detected errors.

Guan, Sheng↗

Improving ADAM through an implicit-explicit (IMEX) time-stepping approach

The ADAM optimizer, often used in machine learning for neural network training, corresponds to an underlying ordinary differential equation (ODE) in the limit of very small learning rates. Here, this work shows that the classical ADAM algorithm is a first-order implicit-explicit (IMEX) Euler discretization of the underlying ODE. Employing the time discretization point of view, we propose new extensions of the ADAM scheme obtained by using higher-order IMEX methods to solve the ODE. Based on this approach, we derive a new optimization algorithm for neural network training that performs better than classical ADAM on several regression and classification problems.

97 MATHEMATICS AND COMPUTING↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

Scaling laws in jet classification

We demonstrate the emergence of scaling laws in the benchmark top versus QCD jet classification problem in collider physics. Six distinct physically-motivated classifiers exhibit power-law scaling of the binary cross-entropy test loss as a function of training set size, with distinct power law indices. This result highlights the importance of comparing classifiers as a function of dataset size rather than for a fixed training set, as the optimal classifier may change considerably as the dataset is scaled up. We speculate on the interpretation of our results in terms of previous models of scaling laws observed in natural language and image datasets.

Batson, Joshua↗

On minimizing the probability of misclassification for linear feature selection

The use of techniques for feature selection permits treatment of classification problems in spaces of reduced dimensions. A method is considered of linear feature selection for n-dimensional observation vectors which belong to one of two populations, where each population is described by a known multivariate normal density function. More specifically, the problem of finding a 1xn transformation matrix B for which the probability of misclassification with respect to the one-dimensional transformed density functions was minimized was considered. Theoretical results are presented which give rise to a numerically tractable expression for the variation in the probability of misclassification with respect to B. Using this expression a computational procedure is discussed for obtaining a B which minimizes the probability of misclassification. Preliminary numerical results are discussed.

Guseman, L. F., Jr.↗

The preprocessing of multispectral data. II

It is pointed out that a correction of atmospheric effects is an important requirement for a full utilization of the possibilities provided by preprocessing techniques. The most significant characteristics of original and preprocessed data are considered, taking into account the solution of classification problems by means of the preprocessing procedure. Improvements obtainable with different preprocessing techniques are illustrated with the aid of examples involving Landsat data regarding an area in Colorado.

Quiel, F.↗

Satellite onboard feature classification using CCD's

In this paper discrete time analog signal processing applications of CCD's are discussed. The types of computations which can be performed in the CCD technology are discussed and compared with respect to their accuracy and required support hardware. The classification problem encountered in the reduction of multispectral data into features of terrestrial interest are decomposed into a CCD technology based hardware solution. The resulting system is readily implemented with existing and anticipated CCD devices. This system will be projected to spacecraft application.

Benz, H. F.↗

Structured estimation - Sample size reduction for adaptive pattern classification

The Gaussian two-category classification problem with known category mean value vectors and identical but unknown category covariance matrices is considered. The weight vector depends on the unknown common covariance matrix, so the procedure is to estimate the covariance matrix in order to obtain an estimate of the optimum weight vector. The measure of performance for the adapted classifier is the output signal-to-interference noise ratio (SIR). A simple approximation for the expected SIR is gained by using the general sample covariance matrix estimator; this performance is both signal and true covariance matrix independent. An approximation is also found for the expected SIR obtained by using a Toeplitz form covariance matrix estimator; this performance is found to be dependent on both the signal and the true covariance matrix.

Morgera, S.↗

Binary classification of real sequences by discrete-time systems

This paper considers a novel approach to coding or classifying sequences of real numbers through the use of (generally nonlinear) finite-dimensional discrete-time systems. This approach involves a finite-dimensional discrete-time system (which we call a real acceptor) in cascade with a threshold type device (which we call a discriminator). The proposed classification scheme and the exact nature of the classification problem are described, along with two examples illustrating its applicability. Suggested approaches for further research are given.

Kaliski, M. E.↗

Multiclass Bayes error estimation by a feature space sampling technique

A general Gaussian M-class N-feature classification problem is defined. An algorithm is developed that requires the class statistics as its only input and computes the minimum probability of error through use of a combined analytical and numerical integration over a sequence simplifying transformations of the feature space. The results are compared with those obtained by conventional techniques applied to a 2-class 4-feature discrimination problem with results previously reported and 4-class 4-feature multispectral scanner Landsat data classified by training and testing of the available data.

Mobasseri, B. G.↗

Radiometric calibration of Seasat SAR imagery

A goal of future synthetic aperture radar (SAR) system design is to yield the radiometrically accurate imagery necessary for use in classification problems. To accomplish this, a reliable calibration technique must be developed. Many such techniques have been proposed (e.g., Stiles and Holtzman, 1982); however, until now, none have been tested on spaceborne SAR digital data. In this paper, a method of significantly increasing the reproducibility of several independent Seasat SAR data sets of a test site using an automated correction procedure is described.

Croft, C.↗

The evaluative imaging of mental models - Visual representations of complexity

The paper deals with some design issues involved in building a system that could visually represent the semantic structures of training materials and their underlying mental models. In particular, hypermedia-based semantic networks that instantiate classification problem solving strategies are thought to be a useful formalism for such representations; the complexity of these web structures can be best managed through visual depictions. It is also noted that a useful approach to implement in these hypermedia models would be some metrics of conceptual distance.

Dede, Christopher↗

A Bayesian classification of the IRAS LRS Atlas

The availability of a reclassification of the IRAS LRS Atlas of spectra using a new Bayesian classification procedure (AutoClass) is announced. The classes of objects which result from the application of the AutoClass algorithm include many of the previously known LRS classes. New classes which have interesting astronomical and astrophysical interpretations were also found. Techniques, such as the AutoClass algorithm, have a bright future in the arena of astronomical classification problems.

Goebel, J.↗

Hierarchical Pattern Classifier

Hierarchical pattern classifier reduces number of comparisons between input and memory vectors without reducing detail of final classification by dividing classification process into coarse-to-fine hierarchy that comprises first "grouping" step and second classification step. Three-layer neural network reduces computation further by reducing number of vector dimensions in processing. Concept applicable to pattern-classification problems with need to reduce amount of computation necessary to classify, identify, or match patterns to desired degree of resolution.

Yates, Gigi L.↗

Feature detection in satellite images using neural network technology

A feasibility study of automated classification of satellite images is described. Satellite images were characterized by the textures they contain. In particular, the detection of cloud textures was investigated. The method of second-order gray level statistics, using co-occurrence matrices, was applied to extract feature vectors from image segments. Neural network technology was employed to classify these feature vectors. The cascade-correlation architecture was successfully used as a classifier. The use of a Kohonen network was also investigated but this architecture could not reliably classify the feature vectors due to the complicated structure of the classification problem. The best results were obtained when data from different spectral bands were fused.

Augusteijn, Marijke F.↗

Experiment and simulation for CSI: What are the missing links?

Viewgraphs on experiment and simulation for control structure interaction (CSI) are presented. Topics covered include: control structure interaction; typical control/structure interaction system; CSI problem classification; actuator/sensor models; modeling uncertainty; noise models; real-time computations; and discrete versus continuous.

Belvin, W. Keith↗