Engineering PapersSearch

SEARCH · Engineering Papers

Results for “binary classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Three-Dimensional Receiver Operator Characteristic Surface Diagnostic Metric

Receiver Operator Characteristic (ROC) curves are commonly applied as metrics for quantifying the performance of binary fault detection systems. An ROC curve provides a visual representation of a detection system s True Positive Rate versus False Positive Rate sensitivity as the detection threshold is varied. The area under the curve provides a measure of fault detection performance independent of the applied detection threshold. While the standard ROC curve is well suited for quantifying binary fault detection performance, it is not suitable for quantifying the classification performance of multi-fault classification problems. Furthermore, it does not provide a measure of diagnostic latency. To address these shortcomings, a novel three-dimensional receiver operator characteristic (3D ROC) surface metric has been developed. This is done by generating and applying two separate curves: the standard ROC curve reflecting fault detection performance, and a second curve reflecting fault classification performance. A third dimension, diagnostic latency, is added giving rise to 3D ROC surfaces. Applying numerical integration techniques, the volumes under and between the surfaces are calculated to produce metrics of the diagnostic system s detection and classification performance. This paper will describe the 3D ROC surface metric in detail, and present an example of its application for quantifying the performance of aircraft engine gas path diagnostic methods. Metric limitations and potential enhancements are also discussed

Simon, Donald L.

Fourier domain target transformation analysis in the thermal infrared

Remote sensing uses of principal component analysis (PCA) of multispectral images include band selection and optimal color selection for display of information content. PCA has also been used for quantitative determination of mineral types and abundances given end member spectra. The preliminary results of the investigation of target transformation PCA (TTPCA) in the fourier domain to both identify end member spectra in an unknown spectrum, and to then calculate the relative concentrations of these selected end members are presented. Identification of endmember spectra in an unknown sample has previously been performed through bandmatching, expert systems, and binary classifiers. Both bandmatching and expert system techniques require the analyst to select bands or combinations of bands unique to each endmember. Thermal infrared mineral spectra have broad spectral features which vary subtly with composition. This makes identification of unique features difficult. Alternatively, whole spectra can be used in the classification process, in which case there is not need for an expert to identify unique spectra. Use of binary classifiers on whole spectra to identify endmember components has met with some success. These techniques can be used, along with a least squares fit approach on the endmembers identified, to derive compositional information. An alternative to the approach outlined above usese target transformation in conjunction with PCA to both identify and quantify the composition of unknown spectra. Preprocessing of the library and unknown spectra into the fourier domain, and using only a specific number of the components, allows for significant data volume reduction while maintaining a linear relationship in a Beer's Law sense. The approach taken here is to iteratively calculate concentrations, reducing the number of endmember components until only non-negative concentrations remain.

Anderson, D. L.

Nearly sonic and transsonic convective motions in the solar atmosphere related to the solar wind origin

MHD equations are considered for the solar atmosphere. 15 different simplest MHD regimes are indicated for the momentum transport equation depending on the mutual binary interplay between 6 possible and locally dominant terms: non-stationarity and inhomogeneity of the flow, gas pressure, magnetic tensions, viscous and gravity forces. These regimes are delimited by five physically independent dimensionless parameters, for example, Strouhal, sonic Mach, alfvenic Mach, Reynolds and Froude numbers or their combinations. Another partially overlapping classification of the simplest regimes may be introduced based on the energy conservation equation. There are also 15 independent binary combinations between nonstationary and inhomogeneous convective terms, dissipative energy sinks and sources (viscous, heat-conductive, Joule and radiative ones) in the energy equation. More complicated regimes are considered with multiple dominated terms. All these MHD regimes play their important role somewhere in the solar atmosphere complicated by the tensor transport coefficients in the magnetically dominated regions of the upper atmosphere. Nearly sonic and transsonic nonstationary convective motions with ascending and descending flows are observed in the solar chromosphere. the transition region and the lower pans of the solar corona together with related horizontal velocity components. This convection represents a kind of the 'cocoonery' manufacturing nonstationary vortices generated here and partially connected to the photosphere and to the solar wind. The solar wind originates from this powerful transsonic muddle in the solar atmosphere as a tiny fraction of the streamlines which are temporarily getting detached from the 'cocoons; and going to the infinity. The topologically complicated instantaneous 'runaway surface' around the Sun, i.e., the surface which separates outgoing to the infinity streams from other finite flows in the solar atmosphere was not described in the literature and needs additional investigation. We conclude that a simple one-connected smooth and quasistationary 'critical surface' which is often supposed to be placed somewhere in the solar corona (say, at 4 solar radii) in reality does not exist. Observational and theoretical arguments favor instead of this a highly structured, disordered, patched and time-dependent sonic transition surface in the solar atmosphere permanently fluctuating at positional dependent heights starting sometimes from photospheric (and maybe subphotospheric) levels and more frequently from the chromosphenc network up to several solar radii in the solar corona.

Veselovsky, I. S.

Identifying and Mapping Agricultural Areas Using Synthetic Aperture Radar Time Series

This study uses Synthetic Aperture Radar (SAR) data to identify active agricultural extent in Mato Grosso, Brazil. Coefficient of variation (CoVar) calculations were applied to year-long time series stacks to identify areas of crop growth. A CoVar threshold value was determined using crop reference layers, then applied to create a crop/non-crop binary map. In the Mato Grosso region, CoVar-based processing resulted in high performance crop classification and allowed for the monitoring of dynamic cropland expansion between 2016 and 2021.

Jordan Bell

Multistage classification of multispectral Earth observational data: The design approach

An algorithm is proposed which predicts the optimal features at every node in a binary tree procedure. The algorithm estimates the probability of error by approximating the area under the likelihood ratio function for two classes and taking into account the number of training samples used in estimating each of these two classes. Some results on feature selection techniques, particularly in the presence of a very limited set of training samples, are presented. Results comparing probabilities of error predicted by the proposed algorithm as a function of dimensionality as compared to experimental observations are shown for aircraft and LANDSAT data. Results are obtained for both real and simulated data. Finally, two binary tree examples which use the algorithm are presented to illustrate the usefulness of the procedure.

Bauer, M. E.

Classifying Unidentified X-Ray Sources in the Chandra Source Catalog Using A Multiwavelength Machine-Learning Approach

The rapid increase in serendipitous X-ray source detections requires the development of novel approaches to efficiently explore the nature of X-ray sources. If even a fraction of these sources could be reliably classified, it would enable population studies for various astrophysical source types on a much larger scale than currently possible. Classification of large numbers of sources from multiple classes characterized by multiple properties (features) must be done automatically and supervised machine learning (ML) seems to provide the only feasible approach. We perform classification of Chandra Source Catalog version 2.0 (CSCv2) sources to explore the potential of the ML approach and identify various biases, limitations, and bottlenecks that present themselves in these kinds of studies. We establish the framework and present a flexible and expandable Python pipeline, which can be used and improved by others. We also release the training data set of 2941 X-ray sources with confidently established classes. In addition to providing probabilistic classifications of 66,369 CSCv2 sources (21% of the entire CSCv2 catalog), we perform several narrower-focused case studies (high-mass X-ray binary candidates and X-ray sources within the extent of the H.E.S.S. TeV sources) to demonstrate some possible applications of our ML approach. We also discuss future possible modifications of the presented pipeline, which are expected to lead to substantial improvements in classification confidences.

Hui Yang

Southern RS CVn systems - Candidate list

A list of 43 candidate RS CVn binary systems in the far southern hemisphere of the sky (south of -40 deg declination) is presented. The candidate systems were selected from the first two volumes of the Michigan Spectral Catalog (1975, 1978), which provides MK classifications for southern HD stars and identifies any unusual characteristics noted for individual stellar spectra. The selection criteria used were: (1) the occurrence of Ca II H and K emission; (2) known or suspected binary nature; (3) regular light variations of zero to one magnitude; and (4) spectral type between F0 and K2 and luminosity less than bright giant (II).

Weiler, E. J.

Scaling laws in jet classification

We demonstrate the emergence of scaling laws in the benchmark top versus QCD jet classification problem in collider physics. Six distinct physically-motivated classifiers exhibit power-law scaling of the binary cross-entropy test loss as a function of training set size, with distinct power law indices. This result highlights the importance of comparing classifiers as a function of dataset size rather than for a fixed training set, as the optimal classifier may change considerably as the dataset is scaled up. We speculate on the interpretation of our results in terms of previous models of scaling laws observed in natural language and image datasets.

Batson, Joshua

The optical counterparts of compact galactic X-ray sources

Ninety-six optical identifications of X-ray sources presumed to be compact objects are presented and discussed. X-ray and optical classifications of the sources are considered, and the several types of systems are discussed. These include neutron star binaries with massive and low-mass stellar components, neutron stars in supernova remnants, cataclysmic variables, and 'isolated' hot white dwarfs. The information that can be derived from optical observations of neutron-star systems with low-mass optical companions is emphasized.

Bradt, H. V. D.

Automated Classification of ROSAT Sources Using Heterogeneous Multiwavelength Source Catalogs

We describe an on-line system for automated classification of X-ray sources, ClassX, and present preliminary results of classification of the three major catalogs of ROSAT sources, RASS BSC, RASS FSC, and WGACAT, into six class categories: stars, white dwarfs, X-ray binaries, galaxies, AGNs, and clusters of galaxies. ClassX is based on a machine learning technology. It represents a system of classifiers, each classifier consisting of a considerable number of oblique decision trees. These trees are built as the classifier is 'trained' to recognize various classes of objects using a training sample of sources of known object types. Each source is characterized by a preselected set of parameters, or attributes; the same set is then used as the classifier conducts classification of sources of unknown identity. The ClassX pipeline features an automatic search for X-ray source counterparts among heterogeneous data sets in on-line data archives using Virtual Observatory protocols; it retrieves from those archives all the attributes required by the selected classifier and inputs them to the classifier. The user input to ClassX is typically a file with target coordinates, optionally complemented with target IDs. The output contains the class name, attributes, and class probabilities for all classified targets. We discuss ways to characterize and assess the classifier quality and performance and present the respective validation procedures. Based on both internal and external validation, we conclude that the ClassX classifiers yield reasonable and reliable classifications for ROSAT sources and have the potential to broaden class representation significantly for rare object types.

McGlynn, Thomas

Classifying multispectral data by neural networks

Several energy functions for synthesizing neural networks are tested on 2-D synthetic data and on Landsat-4 Thematic Mapper data. These new energy functions, designed specifically for minimizing misclassification error, in some cases yield significant improvements in classification accuracy over the standard least mean squares energy function. In addition to operating on networks with one output unit per class, a new energy function is tested for binary encoded outputs, which result in smaller network sizes. The Thematic Mapper data (four bands were used) is classified on a single pixel basis, to provide a starting benchmark against which further improvements will be measured. Improvements are underway to make use of both subpixel and superpixel (i.e. contextual or neighborhood) information in tile processing. For single pixel classification, the best neural network result is 78.7 percent, compared with 71.7 percent for a classical nearest neighbor classifier. The 78.7 percent result also improves on several earlier neural network results on this data.

Telfer, Brian A.

The ultraviolet spectra of four binaries observed with the S59 spectrometer

Ultraviolet spectra of Omicron And, Alpha CrB, Eta Ori A, and Alpha Vir, which were obtained with the S59 spectrometer at a resolution of 1.7 A in three 100-A-wide regions centered at 2110, 2454, and 2825 A, have been studied for the presence or absence of effects due to their binary nature. As may have been anticipated from their orbital and other characteristics, no indication of strong binary interactions were seen in these observations. However, there are certain spectral peculiarities suggesting the possibility of modifications of spectral classifications for some of these objects. A rather unusual spectral behavior in Alpha Vir is also noted. In addition, based primarily on a review of available literature, attention is drawn to a remarkable property of the third component in Eta Ori A.

Herczeg, T. J.

Stereo Vision Based Terrain Mapping for Off-Road Autonomous Navigation

Successful off-road autonomous navigation by an unmanned ground vehicle (UGV) requires reliable perception and representation of natural terrain. While perception algorithms are used to detect driving hazards, terrain mapping algorithms are used to represent the detected hazards in a world model a UGV can use to plan safe paths. There are two primary ways to detect driving hazards with perception sensors mounted to a UGV: binary obstacle detection and traversability cost analysis. Binary obstacle detectors label terrain as either traversable or non-traversable, whereas, traversability cost analysis assigns a cost to driving over a discrete patch of terrain. In uncluttered environments where the non-obstacle terrain is equally traversable, binary obstacle detection is sufficient. However, in cluttered environments, some form of traversability cost analysis is necessary. The Jet Propulsion Laboratory (JPL) has explored both approaches using stereo vision systems. A set of binary detectors has been implemented that detect positive obstacles, negative obstacles, tree trunks, tree lines, excessive slope, low overhangs, and water bodies. A compact terrain map is built from each frame of stereo images. The mapping algorithm labels cells that contain obstacles as no-go regions, and encodes terrain elevation, terrain classification, terrain roughness, traversability cost, and a confidence value. The single frame maps are merged into a world map where temporal filtering is applied. In previous papers, we have described our perception algorithms that perform binary obstacle detection. In this paper, we summarize the terrain mapping capabilities that JPL has implemented during several UGV programs over the last decade and discuss some challenges to building terrain maps with stereo range data.

passive perception

VHClass

The code is used to predict the taxonomic source of an antibody heavy chain sequence. The code assigns a binary label to the input set of sequences - camelid or human. This prediction is generated using a random-forest based classification algorithm which is the backbone of the code. A complementary code splits the antibody sequence into antibody features - framework regions and CDR regions.

Davis, Anastasiia

Data processing large quantities of multispectral information

Method is combination of digital and optical techniques. Multispectral data is coded into binary matrix format and then encoded onto photographic film. Film is holographically correlated with spectral signature to generate single-class classification map. Number of maps are optically superimposed to produce full-color, multiclass classification map.

Haskell, R. E.

A VLA survey of unidentified HEAO-1 X-ray sources

Radio observations toward 47 unidentified sources from the HEAO-1 all-sky X-ray survey (Wood et al., 1984), obtained at 1.418 GHz using the C configuration of the VLA on May 11-14, 1983 are reported. Of the 238 radio sources detected at flux density 1-3 mJy or greater, 26 are found to be within 5 arcsec of optical objects, which are thereby considered to be candidate X-ray sources. Tentative classifications of the candidates include five RS CVn systems, three AGNs, three galaxy or cluster sources, and two X-ray binaries.

Schmelz, J. T.

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS

Quantitative spectral types for 19 Algol secondaries

Time-resolved spectra of 19 short-period Algol-type binary star systems obtained during total eclipse are used to derive the temperature spectral class of the mass-losing secondary component. The spectral classifications employed a quantitative comparison of the strengths of absorption features in stars of known spectral class with those of the program stars. The luminosity spectral class can not be determined from these data, so both main-sequence and giant stars were used for the comparison. Our spectral types are compared with published types and found to be generally in good agreement, unless the published types are derived from the light curves. The photometrically determined types are systematically later than our directly determined types. This effect is shown also to exist in catalogs of Algol parameters.

Yoon, Tae S.