Engineering PapersSearch

SEARCH · Engineering Papers

Results for “binary classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fast Solution in Sparse LDA for Binary Classification

An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.

Moghaddam, Baback

Binary classification of real sequences by discrete-time systems

This paper considers a novel approach to coding or classifying sequences of real numbers through the use of (generally nonlinear) finite-dimensional discrete-time systems. This approach involves a finite-dimensional discrete-time system (which we call a real acceptor) in cascade with a threshold type device (which we call a discriminator). The proposed classification scheme and the exact nature of the classification problem are described, along with two examples illustrating its applicability. Suggested approaches for further research are given.

Kaliski, M. E.

Binary image classification

Motivated by the LANDSAT problem of estimating the probability of crop or geological types based on multi-channel satellite imagery data, Morris and Kostal (1983), Hill, Hinkley, Kostal, and Morris (1984), and Morris, Hinkley, and Johnston (1985) developed an empirical Bayes approach to this problem. Here, researchers return to those developments, making certain improvements and extensions, but restricting attention to the binary case of only two attributes.

Morris, Carl N.

Testing of the Support Vector Machine for Binary-Class Classification

The Support Vector Machine is a powerful algorithm, useful in classifying data in to species. The Support Vector Machines implemented in this research were used as classifiers for the final stage in a Multistage Autonomous Target Recognition system. A single kernel SVM known as SVMlight, and a modified version known as a Support Vector Machine with K-Means Clustering were used. These SVM algorithms were tested as classifiers under varying conditions. Image noise levels varied, and the orientation of the targets changed. The classifiers were then optimized to demonstrate their maximum potential as classifiers. Results demonstrate the reliability of SMV as a method for classification. From trial to trial, SVM produces consistent results

autonomous target recognition systemr

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration

Spectral and Timing Nature of the Symbiotic X-Ray Binary 4U 1954+319: The Slowest Rotating Neutron Star in AN X-Ray Binary System

The symbiotic X-ray binary (SyXB) 4U 1954+319 is a rare system hosting a peculiar neutron star (NS) and an M-type optical companion. Its approx. 5.4 hr NS spin period is the longest among all known accretion-powered pulsars and exhibited large (is approx. 7%) fluctuations over 8 yr. A spin trend transition was detected with Swift/BAT around an X-ray brightening in 2012. The source was in quiescent and bright states before and after this outburst based on 60 ks Suzaku observations in 2011 and 2012. The observed continuum is well described by a Comptonized model with the addition of a narrow 6.4 keV Fe-K alpha line during the outburst. Spectral similarities to slowly rotating pulsars in high-mass X-ray binaries, its high pulsed fraction (approx. 60%-80%), and the location in the Corbet diagram favor high B-field (approx. greater than 10(exp12) G) over a weak field as in low-mass X-ray binaries. The observed low X-ray luminosity (10(exp33)-10(exp35) erg s(exp−1)), probable wide orbit, and a slow stellar wind of this SyXB make quasi-spherical accretion in the subsonic settling regime a plausible model. Assuming a approx. 10(exp13) G NS, this scheme can explain the approx. 5.4 hr equilibrium rotation without employing the magnetar-like field (approx. 10(exp16) G) required in the disk accretion case. The timescales of multiple irregular flares (approx. 50 s) can also be attributed to the free-fall time from the Alfv´en shell for a approx. 10(exp13) G field. A physical interpretation of SyXBs beyond the canonical binary classifications is discussed.

accretion disks aEuro" binaries: symbiotic aEuro"

The Gaseous Content of the Universe at Zeta less than 1.6

Together with graduate student Hsiao-Wen Chen, I have measured and analyzed structural and morphological parameters of 38 galaxies in eight fields for which sensitive measurements of corresponding Ly(alpha) absorption toward background QSOs are available. These measurements are based on Wide Field Planetary Camera 2 (WFPC2) observations obtained with the Hubble Space Telescope (HST) and provide a first look at how the incidence and extent of tenuous gas around galaxies depends on galaxy luminosity, size, and morphological type and on geometry of the impact. The primary result of the analysis is that the amount of gas encountered along the line of sight depends on the galaxy impact parameter and B-band luminosity but does not depend strongly on the galaxy average surface brightness, disk-to-bulge ratio, or redshift. This result confirms and improves upon an anti-correlation between Ly(alpha) absorption equivalent width and galaxy impact parameter found previously. More importantly, this result provides the first quantitative means of relating statistics of faint galaxies to statistics of Ly(alpha) absorption systems. which we plan to exploit to constrain the luminosity function of galaxies beyond the realm of current surveys. Results have been submitted for publication and will greatly improve our statistical conclusions . Together with graduate student Noriaki Yahata. I have measured and classified spectral properties of over 1000 faint galaxies and stars obtained in our low-resolution spectroscopic survey. The goal of this project is two-fold: (1) to exhaustively characterize the spectral properties of all faint galaxies that comprise our current survey, and (2) to gain experience with our measurement and classification code. which ultimately will be used on a data base of 20,000 galaxies to be obtained with the Two-Degree Field (2df) spectrograph at the Anglo-Australian Telescope (AAT). The results will ultimately be used for many goals, but so far we have concentrated on using the results to make a binary classification of the galaxies (i.e. early type versus late type) and to then exploit the density-morphology relationship to obtain a crude density indicator. The primary result of the analysis is that the incidence and extent of tenuous gas around galaxies shows no strong preference for local galaxy environment, at least over the range of densities spanned by the current observations. Along similar lines, two instances of Ly(alpha) absorption lines that arise in groups or clusters were examined. Analysis demonstrates that some can produce corresponding absorption lines and that LY(alpha) absorption lines do not avoid a high-density environment. A new measure of the galaxy-absorber cross-correlation function defines the statistical criterion by which galaxies and absorber pairs are to be matched. I have identified a damped Ly(alpha) absorption system at redshift z equals approximately 0.16, the lowest redshift confirmed to date. The most important results of the analysis are learning that the metal abundances of the absorption system are less than 10 percent of the solar metal abundance and that the absorbing gas is not rotating with the galaxy disk.

Source record

State Predictor of Classification Cognitive Engine Applied to Channel Fading

This study presents the application of machine learning (ML) to a space-to-ground communication link, showing how ML can be used to detect the presence of detrimental channel fading. Using this channel state information, the communication link can be used more efficiently by reducing the amount of lost data during fading. The motivation for this work is based on channel fading observed during on-orbit operations with NASA's Space Communication and Navigation (SCaN) testbed on the International Space Station (ISS). This paper presents the process to extract a target concept (fading and not-fading) from the raw data. The pre-processing and data exploration effort is explained in detail, with a list of assumptions made for parsing and labelling the dataset. The model selection process is explained, specifically emphasizing the benefits of using an ensemble of algorithms with majority voting for binary classification of the channel state. Experimental results are shown, highlighting how an end-to-end communication system can utilize knowledge of the channel fading status to identity fading and take appropriate action. With a laboratory testbed to emulate channel fading, the overall performance is compared to standard adaptive methods without fading knowledge, such as adaptive coding and modulation.

Fading

Automated Pneumothorax Diagnosis using Deep Neural Networks

Thoracic ultrasound can provide information leading to rapid diagnosis of pneumothorax with improved accuracy over the standard physical examination and with higher sensitivity than anteroposterior chest radiography. However, the clinical We have Furthermore, remote environments, such as the battlefield or deep-space exploration, may lack expertise for diagnosing developed an automated image interpretation pipeline for the analysis of thoracic ultrasound data and the classification of pneumothorax events to provide decision support in such situations. Our pipeline consists of image preprocessing, data augmentation, and deep learning architectures for medical diagnosis. In this work, we demonstrate that robust, accurate interpretation of chest images and video can be achieved using deep neural networks. A number of novel image processing techniques were employed to achieve this result. Affine transformations were applied for data augmentation. Hyperparameters were optimized for learning rate, dropout regularization, batch size, and epoch iteration by a sequential model-based Bayesian approach. In addition, we utilized pretrained architecturesinterpretation of a patient medical image is highly operator dependent. certain pathologies., applying transfer learning and fine-tuning techniques to fully connected layers. Our pipeline yielded binary classification validation accuracies of 98.3% for M-mode images and 99.8% with B-mode video frames.

US Army collaboration

Calibration or inverse regression: Which is appropriate for crop surveys using LANDSAT data?

Calibration and inverse regression estimators of crop proportions are investigated where the auxiliary variable is obtained from binary classification of multivariate LANDSAT data. The appropriate model relating classifier proportions and ground observed proportions for a given crop type is the calibration model. Under this model the inverse regression estimator is superior to the calibration estimator in estimating the crop acreage or proportion for a region of interest.

Chhikara, R. S.

Stochastic robustness

To carry out stochastic robustness analysis, an expected probability distribution is assigned to each uncertain parameter in the system. The Monte Carlo analysis proceeds by repeatedly assigning shaped random values to each plant parameter, evaluating the stability of performance metric, and performing the binary classification (stable/unstable, etc.). If the system is stable, the state response to a unit disturbance impulse can be propagated to establish whether the response would violate settling time envelopes and whether peak actuator use would violate predetermined maximums. The final estimates of the probability of each form of unacceptable behavior are found by dividing the number of cases in which the overall system had that form of unacceptability by the number of cases run. Stability robustness can be portrayed graphically using the stochastic root locus and by using histograms of parameter values found in the unacceptable cases.

Marrison, C.

Experiments on Supervised Learning Algorithms for Text Categorization

Modern information society is facing the challenge of handling massive volume of online documents, news, intelligence reports, and so on. How to use the information accurately and in a timely manner becomes a major concern in many areas. While the general information may also include images and voice, we focus on the categorization of text data in this paper. We provide a brief overview of the information processing flow for text categorization, and discuss two supervised learning algorithms, viz., support vector machines (SVM) and partial least squares (PLS), which have been successfully applied in other domains, e.g., fault diagnosis [9]. While SVM has been well explored for binary classification and was reported as an efficient algorithm for text categorization, PLS has not yet been applied to text categorization. Our experiments are conducted on three data sets: Reuter's- 21578 dataset about corporate mergers and data acquisitions (ACQ), WebKB and the 20-Newsgroups. Results show that the performance of PLS is comparable to SVM in text categorization. A major drawback of SVM for multi-class categorization is that it requires a voting scheme based on the results of pair-wise classification. PLS does not have this drawback and could be a better candidate for multi-class text categorization.

Namburu, Setu Madhavi

NASA Tech Briefs, May 2010

Topics covered include: Instrument for Analysis of Greenland's Glacier Mills Cryogenic Moisture Apparatus; A Transportable Gravity Gradiometer Based on Atom Interferometry; Three Methods of Detection of Hydrazines; Crossed, Small-Deflection Energy Analyzer for Wind/Temperature Spectrometer; Wavefront Correction for Large, Flexible Antenna Reflector; Novel Micro Strip-to-Waveguide Feed Employing a Double-Y Junction; Thin-Film Ferro Electric-Coupled Microstripline Phase Shifters With Reduced Device Hysteresis; Two-Stage, 90-GHz, Low-Noise Amplifier; A 311-GHz Fundamental Oscillator Using InP HBT Technology; FPGA Coprocessor Design for an Onboard Multi-Angle Spectro-Polarimetric Imager; Serrating Nozzle Surfaces for Complete Transfer of Droplets; Turbomolecular Pumps for Holding Gases in Open Containers; Triaxial Swirl Injector Element for Liquid-Fueled Engines; Integrated Budget Office Toolbox; PLOT3D Export Tool for Tecplot; Math Description Engine Software Development Kit; Astronaut Office Scheduling System Software; ISS Solar Array Management; Probabilistic Structural Analysis Program; SPOT Program; Integrated Hybrid System Architecture for Risk Analysis; System for Packaging Planetary Samples for Return to Earth; Offset Compound Gear Drive; Low-Dead-Volume Inlet for Vacuum Chamber; Simple Check Valves for Microfluidic Devices; A Capillary-Based Static Phase Separator for Highly Variable Wetting Conditions; Gimballing Spacecraft Thruster; Finned Carbon-Carbon Heat Pipe with Potassium Working Fluid; Lightweight Heat Pipes Made from Magnesium; Ceramic Rail-Race Ball Bearings; Improved OTEC System for a Submarine Robot; Reflector Surface Error Compensation in Dual-Reflector Antennas; Enriched Storable Oxidizers for Rocket Engines; Planar Submillimeter-Wave Mixer Technology with Integrated Antenna; Widely Tunable Mode-Hop-Free External-Cavity Quantum Cascade Laser; Non-Geiger-Mode Single-Photon Avalanche Detector with Low Excess Noise; Using Whispering-Gallery-Mode Resonators for Refractometry; RF Device for Acquiring Images of the Human Body; Reactive Collision Avoidance Algorithm; Fast Solution in Sparse LDA for Binary Classification; Modeling Common-Sense Decisions in Artificial Intelligence; Graph-Based Path-Planning for Titan Balloons; Nanolaminate Membranes as Cylindrical Telescope Reflectors; Air-Sea Spray Airborne Radar Profiler Characterizes Energy Fluxes in Hurricanes; Large Telescope Segmented Primary Mirror Alignment; and Simplified Night Sky Display System.

Source record

Convective Weather Forecast Accuracy Analysis at Center and Sector Levels

This paper presents a detailed convective forecast accuracy analysis at center and sector levels. The study is aimed to provide more meaningful forecast verification measures to aviation community, as well as to obtain useful information leading to the improvements in the weather translation capacity models. In general, the vast majority of forecast verification efforts over past decades have been on the calculation of traditional standard verification measure scores over forecast and observation data analyses onto grids. These verification measures based on the binary classification have been applied in quality assurance of weather forecast products at the national level for many years. Our research focuses on the forecast at the center and sector levels. We calculate the standard forecast verification measure scores for en-route air traffic centers and sectors first, followed by conducting the forecast validation analysis and related verification measures for weather intensities and locations at centers and sectors levels. An approach to improve the prediction of sector weather coverage by multiple sector forecasts is then developed. The weather severe intensity assessment was carried out by using the correlations between forecast and actual weather observation airspace coverage. The weather forecast accuracy on horizontal location was assessed by examining the forecast errors. The improvement in prediction of weather coverage was determined by the correlation between actual sector weather coverage and prediction. observed and forecasted Convective Weather Avoidance Model (CWAM) data collected from June to September in 2007. CWAM zero-minute forecast data with aircraft avoidance probability of 60% and 80% are used as the actual weather observation. All forecast measurements are based on 30-minute, 60- minute, 90-minute, and 120-minute forecasts with the same avoidance probabilities. The forecast accuracy analysis for times under one-hour showed that the errors in intensity and location for center forecast are relatively low. For example, 1-hour forecast intensity and horizontal location errors for ZDC center were about 0.12 and 0.13. However, the correlation between sector 1-hour forecast and actual weather coverage was weak, for sector ZDC32, about 32% of the total variation of observation weather intensity was unexplained by forecast; the sector horizontal location error was about 0.10. The paper also introduces an approach to estimate the sector three-dimensional actual weather coverage by using multiple sector forecasts, which turned out to produce better predictions. Using Multiple Linear Regression (MLR) model for this approach, the correlations between actual observation and the multiple sector forecast model prediction improved by several percents at 95% confidence level in comparison with single sector forecast.

Wang, Yao

Modeling Weather Impact on Airport Arrival Miles-in-Trail Restrictions

When the demand for either a region of airspace or an airport approaches or exceeds the available capacity, miles-in-trail (MIT) restrictions are the most frequently issued traffic management initiatives (TMIs) that are used to mitigate these imbalances. Miles-intrail operations require aircraft in a traffic stream to meet a specific inter-aircraft separation in exchange for maintaining a safe and orderly flow within the stream. This stream of aircraft can be departing an airport, over a common fix, through a sector, on a specific route or arriving at an airport. This study begins by providing a high-level overview of the distribution and causes of arrival MIT restrictions for the top ten airports in the United States. This is followed by an in-depth analysis of the frequency, duration and cause of MIT restrictions impacting the Hartsfield-Jackson Atlanta International Airport (ATL) from 2009 through 2011. Then, machine-learning methods for predicting (1) situations in which MIT restrictions for ATL arrivals are implemented under low demand scenarios, and (2) days in which a large number of MIT restrictions are required to properly manage and control ATL arrivals are presented. More specifically, these predictions were accomplished by using an ensemble of decision trees with Bootstrap aggregation (BDT) and supervised machine learning was used to train the BDT binary classification models. The models were subsequently validated using data cross validation methods. When predicting the occurrence of arrival MIT restrictions under low demand situations, the model was able to achieve over all accuracy rates ranging from 84% to 90%, with false alarm ratios ranging from 10% to 15%. In the second set of studies designed to predict days on which a high number of MIT restrictions were required, overall accuracy rates of 80% were achieved with false alarm ratios of 20%. Overall, the predictions proposed by the model give better MIT usage information than what has been currently provided under current day operations. Traffic flow managers can use these predictions to identify potential MIT restrictions to eliminate (e.g., those occurring during low arrival demand periods), and to determine the days in which a significant number of restrictions may be required

Operation

Long-Term Impacts of Selective Logging on Amazon Forest Dynamics from Multi-Temporal Airborne LiDAR

Forest degradation is common in tropical landscapes, but estimates of the extent and duration of degradation impacts are highly uncertain. In particular, selective logging is a form of forest degradation that alters canopy structure and function, with persistent ecological impacts following forest harvest. In this study, we employed airborne laser scanning in 2012 and 2014 to estimate three-dimensional changes in the forest canopy and understory structure and aboveground biomass following reduced-impact selective logging in a site in Eastern Amazon. Also, we developed a binary classification model to distinguish intact versus logged forests. We found that canopy gap frequency was significantly higher in logged versus intact forests even after 8 years (the time span of our study). In contrast, the understory of logged areas could not be distinguished from the understory of intact forests after 6–7 years of logging activities. Measuring new gap formation between LiDAR acquisitions in 2012 and 2014, we showed rates 2 to 7 times higher in logged areas compared to intact forests. New gaps were spatially clumped with 76 to 89% of new gaps within 5 m of prior logging damage. The biomass dynamics in areas logged between the two LiDAR acquisitions was clearly detected with an average estimated loss of -4.14 +/- 0.76 MgC/hay. In areas recovering from logging prior to the first acquisition, we estimated biomass gains close to zero. Together, our findings unravel the magnitude and duration of delayed impacts of selective logging in forest structural attributes, confirm the high potential of airborne LiDAR multitemporal data to characterize forest degradation in the tropics, and present a novel approach to forest classification using LiDAR data.

Pinage, Ekena Rangel

Quality of Candidate Flights and Submission Prediction in Collaborative Digital Departure Reroute

Collaborative Digital Departure Reroute (CDDR) enables the reroute of flights using a flight operator proposed set of alternative route options, referred to as Trajectory Option Set (TOS), in order to reduce delay on the airport's surface and in the Metroplex environment. The reroute functionality is enabled through NASA's Digital Information Platform (DIP). TOS candidate flights are defined as flights with an alternative route with delay savings greater than the flight operator defined relative trajectory cost. This paper analyzes the TOS candidate flights at Dallas/Fort Worth International Airport (KDFW) in the North Texas Metroplex to gain insight into which candidate flights are higher quality through a scoring method. This insight will inform refinements to help CDDR focus on high quality reroute opportunities. Binary classification models for predicting the flight operator's submission of candidate flights are also explored in this paper.

Machine Learning

Quality of Candidate Flights and Submission Prediction in Collaborative Digital Departure Reroute

Collaborative Digital Departure Reroute (CDDR) enables the reroute of flights using a flight operator proposed set of alternative route options, referred to as Trajectory Option Set (TOS), in order to reduce delay on the airport's surface and in the Metroplex environment. The reroute functionality is enabled through NASA's Digital Information Platform (DIP). TOS candidate flights are defined as flights with an alternative route with delay savings greater than the flight operator defined relative trajectory cost. This paper analyzes the TOS candidate flights at Dallas/Fort Worth International Airport (KDFW) in the North Texas Metroplex to gain insight into which candidate flights are higher quality through a scoring method. This insight will inform refinements to help CDDR focus on high quality reroute opportunities. Binary classification models for predicting the flight operator's submission of candidate flights are also explored in this paper.

Machine Learning