Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A comparison of Landsat point and rectangular field training sets for land-use classification

Rectangular training fields of homogeneous spectroreflectance are commonly used in supervised pattern recognition efforts. Trial image classification with manually selected training sets gives irregular and misleading results due to statistical bias. A self-verifying, grid-sampled training point approach is proposed as a more statistically valid feature extraction technique. A systematic pixel sampling network of every ninth row and ninth column efficiently replaced the full image scene with smaller statistical vectors which preserved the necessary characteristics for classification. The composite second- and third-order average classification accuracy of 50.1 percent for 331,776 pixels in the full image substantially agreed with the 51 percent value predicted by the grid-sampled, 4,100-point training set.

Tom, C. H.↗

Automated Cardiovascular Pathology Assessment using Semantic Segmentation and Ensemble Learning

Cardiac magnetic resonance imaging provides high spatial resolution, enabling improved extraction of important functional and morphological features for cardiovascular disease staging. Segmentation of ventricular cavities and myocardium in cardiac cine sequencing provides a basis to quantify cardiac measures such as ejection fraction. A method is presented that curtails the expense and observer bias of manual cardiac evaluation by combining semantic segmentation and disease classification into a fully automatic processing pipeline. The initial processing element consists of a robust dilated convolutional neural network architecture for voxel-wise segmentation of the myocardium and ventricular cavities. The resulting comprehensive volumetric feature matrix captures diagnostic clinical procedure data and is utilized by the final processing element to model a cardiac pathology classifier. Our approach evaluated anonymized cardiac images from a training data set of 100 patients (4 pathology groups, 1 healthy group, 20 patients per group) examined at the University Hospital of Dijon. The top average Dice index scores achieved were 0.940, 0.886, 0.849 for structure segmentation of the left ventricle (LV), myocardium and right ventricle (RV) respectively. A 5-ary pathology classification accuracy of 90% was recorded on an independent test set using the trained model. Performance results demonstrate potential for advanced machine learning methods to deliver accurate, efficient and reproducible cardiac pathological assessment.

Semantic Segmentation↗

Automated Semantic Segmentation for Volumetric Cardiovascular Feature Quantification and Pathology Assessment

We present a pipeline method that curtails the expense and observer bias of manual cardiac evaluation by combining semantic segmentation and disease classification as a fully automatic processing pipeline. The initial element consists of a 2D U-Net convolutional neural network architecture for voxel-wise segmentation of the myocardium and ventricular cavities. The results of the segmentation were used to compute a comprehensive volumetric feature matrix that captured diagnostic clinical procedure data and that was used to model a cardiac pathology classifier.Our approach evaluated anonymized parasternal MRI cardiac images from a database of 100 patients (4 pathology groups, 1 healthy group, 20 patients per group) examined at the University Hospital of Dijon. We achieved top average Dice index scores of 0.939, 0.849, 0.886 for structure segmentation of the left ventricle (LV), right ventricle (RV) and myocardium respectively. A 5-ary pathology classification accuracy of 90% was recorded on an independent test set using our trained model.

Lindsey, Tony↗

Automatic Detection of Large-scale Flux Ropes and Their Geoeffectiveness with a Machine-learning Approach

Detecting large-scale flux ropes (FRs) embedded in interplanetary coronal mass ejections (ICMEs) and assessing their geoeffectiveness are essential, since they can drive severe space weather. At 1 au, these FRs have an average duration of 1 day. Their most common magnetic features are large, smoothly rotating magnetic fields. Their manual detection has become a relatively common practice over decades, although visual detection can be time-consuming and subject to observer bias. Our study proposes a pipeline that utilizes two supervised binary classification machine-learning models trained with solar wind magnetic properties to automatically detect large-scale FRs and additionally determine their geoeffectiveness. The first model is used to generate a list of autodetected FRs. Using the properties of the southward magnetic field, the second model determines the geoeffectiveness of FRs. Our method identifies 88.6% and 80% of large-scale ICMEs (duration 1day) observed at 1au by the Wind and the Solar TErrestrial RElations Observatory missions, respectively. While testing with continuous solar wind data obtained from Wind, our pipeline detected 56 of the 64 large-scale ICMEs during the 2008–2014 period (recall = 0.875), but also many false positives (precision = 0.56), as we do not take into account any additional solar wind properties other than the magnetic properties. We find an accuracy of 0.88 when estimating the geoeffectiveness of the autodetected FRs using our method. Thus, in space-weather nowcasting and forecasting at L1 or any planetary missions, our pipeline can be utilized to offer a first-order detection of large-scale FRs and their geoeffectiveness.

Sanchita Pal↗

Tests of a Semi-Analytical Case 1 and Gelbstoff Case 2 SeaWiFS Algorithm with a Global Data Set

A semi-analytical algorithm was tested with a total of 733 points of either unpackaged or packaged-pigment data, with corresponding algorithm parameters for each data type. The 'unpackaged' type consisted of data sets that were generally consistent with the Case 1 CZCS algorithm and other well calibrated data sets. The 'packaged' type consisted of data sets apparently containing somewhat more packaged pigments, requiring modification of the absorption parameters of the model consistent with the CalCOFI study area. This resulted in two equally divided data sets. A more thorough scrutiny of these and other data sets using a semianalytical model requires improved knowledge of the phytoplankton and gelbstoff of the specific environment studied. Since the semi-analytical algorithm is dependent upon 4 spectral channels including the 412 nm channel, while most other algorithms are not, a means of testing data sets for consistency was sought. A numerical filter was developed to classify data sets into the above classes. The filter uses reflectance ratios, which can be determined from space. The sensitivity of such numerical filters to measurement resulting from atmospheric correction and sensor noise errors requires further study. The semi-analytical algorithm performed superbly on each of the data sets after classification, resulting in RMS1 errors of 0.107 and 0.121, respectively, for the unpackaged and packaged data-set classes, with little bias and slopes near 1.0. In combination, the RMS1 performance was 0.114. While these numbers appear rather sterling, one must bear in mind what mis-classification does to the results. Using an average or compromise parameterization on the modified global data set yielded an RMS1 error of 0.171, while using the unpackaged parameterization on the global evaluation data set yielded an RMS1 error of 0.284. So, without classification, the algorithm performs better globally using the average parameters than it does using the unpackaged parameters. Finally, the effects of even more extreme pigment packaging must be examined in order to improve algorithm performance at high latitudes. Note, however, that the North Sea and Mississippi River plume studies contributed data to the packaged and unpackaged classess, respectively, with little effect on algorithm performance. This suggests that gelbstoff-rich Case 2 waters do not seriously degrade performance of the semi-analytical algorithm.

Carder, Kendall L.↗

On the error in crop acreage estimation using satellite (LANDSAT) data

The problem of crop acreage estimation using satellite data is discussed. Bias and variance of a crop proportion estimate in an area segment obtained from the classification of its multispectral sensor data are derived as functions of the means, variances, and covariance of error rates. The linear discriminant analysis and the class proportion estimation for the two class case are extended to include a third class of measurement units, where these units are mixed on ground. Special attention is given to the investigation of mislabeling in training samples and its effect on crop proportion estimation. It is shown that the bias and variance of the estimate of a specific crop acreage proportion increase as the disparity in mislabeling rates between two classes increases. Some interaction is shown to take place, causing the bias and the variance to decrease at first and then to increase, as the mixed unit class varies in size from 0 to 50 percent of the total area segment.

Chhikara, R.↗

Distribution and evolution of asteroid rotation rates

Data on the rotational characteristics of more than 300 asteroids are currently available, and it is now clear that the distribution of the rotation rates is nonrandom. A plot of rotation rate against asteroid diameter shows large dispersion but is distinctly V-shaped. The minimum of this curve at about 120 km may separate primordial asteroids from their collision products. There is also evidence that rotation rate depends on type classification, and weak evidence that it may also depend on family membership. Recent bias-free observations suggest that the marked rise of rotation rate with decreasing diameter D for those asteroids with D less than 120 km cannot be completely accounted for by observational-selection effects. A significantly large subset of the small asteroids have exceptionally long rotation periods suggestive of either a different nature and origin or a peculiar history. Models that have been proposed to account for these results are discussed.

Dermott, S. F.↗

Classifying Unidentified X-Ray Sources in the Chandra Source Catalog Using A Multiwavelength Machine-Learning Approach

The rapid increase in serendipitous X-ray source detections requires the development of novel approaches to efficiently explore the nature of X-ray sources. If even a fraction of these sources could be reliably classified, it would enable population studies for various astrophysical source types on a much larger scale than currently possible. Classification of large numbers of sources from multiple classes characterized by multiple properties (features) must be done automatically and supervised machine learning (ML) seems to provide the only feasible approach. We perform classification of Chandra Source Catalog version 2.0 (CSCv2) sources to explore the potential of the ML approach and identify various biases, limitations, and bottlenecks that present themselves in these kinds of studies. We establish the framework and present a flexible and expandable Python pipeline, which can be used and improved by others. We also release the training data set of 2941 X-ray sources with confidently established classes. In addition to providing probabilistic classifications of 66,369 CSCv2 sources (21% of the entire CSCv2 catalog), we perform several narrower-focused case studies (high-mass X-ray binary candidates and X-ray sources within the extent of the H.E.S.S. TeV sources) to demonstrate some possible applications of our ML approach. We also discuss future possible modifications of the presented pipeline, which are expected to lead to substantial improvements in classification confidences.

Hui Yang↗

The Ocean Colour Climate Change Initiative: III. A Round-Robin Comparison on In-Water Bio-Optical Algorithms

Satellite-derived remote-sensing reflectance (Rrs) can be used for mapping biogeochemically relevant variables, such as the chlorophyll concentration and the Inherent Optical Properties (IOPs) of the water, at global scale for use in climate-change studies. Prior to generating such products, suitable algorithms have to be selected that are appropriate for the purpose. Algorithm selection needs to account for both qualitative and quantitative requirements. In this paper we develop an objective methodology designed to rank the quantitative performance of a suite of bio-optical models. The objective classification is applied using the NASA bio-Optical Marine Algorithm Dataset (NOMAD). Using in situ Rrs as input to the models, the performance of eleven semianalytical models, as well as five empirical chlorophyll algorithms and an empirical diffuse attenuation coefficient algorithm, is ranked for spectrally-resolved IOPs, chlorophyll concentration and the diffuse attenuation coefficient at 489 nm. The sensitivity of the objective classification and the uncertainty in the ranking are tested using a Monte-Carlo approach (bootstrapping). Results indicate that the performance of the semi-analytical models varies depending on the product and wavelength of interest. For chlorophyll retrieval, empirical algorithms perform better than semi-analytical models, in general. The performance of these empirical models reflects either their immunity to scale errors or instrument noise in Rrs data, or simply that the data used for model parameterisation were not independent of NOMAD. Nonetheless, uncertainty in the classification suggests that the performance of some semi-analytical algorithms at retrieving chlorophyll is comparable with the empirical algorithms. For phytoplankton absorption at 443 nm, some semi-analytical models also perform with similar accuracy to an empirical model. We discuss the potential biases, limitations and uncertainty in the approach, as well as additional qualitative considerations for algorithm selection for climate-change studies. Our classification has the potential to be routinely implemented, such that the performance of emerging algorithms can be compared with existing algorithms as they become available. In the long-term, such an approach will further aid algorithm development for ocean-colour studies.

Phytoplankton↗

Myths and legends in learning classification rules

A discussion is presented of machine learning theory on empirically learning classification rules. Six myths are proposed in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, universal learning algorithms, and interactive learning. Some of the problems raised are also addressed from a Bayesian perspective. Questions are suggested that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

LACIE/ERIPS software system summary

The Earth resources interactive processing system (ERIPS) supports LACIE by classifying LANDSAT sensed data on the basis of the statistical similarity to those portions which were identified by analysts. The development and capabilities of the ERIPS software system are described with emphasis on (1) system requirements; (2) LACIE/ERIPS hardware; (3) system functions; (4) pattern recognition concept; and (5) LACIE/ERIPS data bases. Algorithms used in LACIE/ERIPS for statistics, divergence, feature selection, classification, registration, adaptive clustering, iterative clustering, clustering report functions, Sun angle correction, mean level adjustment, and bias correction are appended.

Johnson, C. L.↗

Myths and legends in learning classification rules

This paper is a discussion of machine learning theory on empirically learning classification rules. The paper proposes six myths in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, 'universal' learning algorithms, and interactive learnings. Some of the problems raised are also addressed from a Bayesian perspective. The paper concludes by suggesting questions that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

Sampling Landsat classifications for crop area estimation

An investigation was conducted to evaluate the effect of several sampling alternatives on the accuracy of crop area estimates made from classification of Landsat Multispectral Scanner (MSS) data. The specific objective was to assess the precision and the bias associated with alternative sampling schemes involving different numbers of several sampling unit sizes. The estimates achieved using the 5 by 6 nm segments were found to have the least precision of any sampling scheme tested. The estimates become more precise as the segment size decreases and more segments are taken. The precision of the 5 by 6 nm segments was significantly less than that of the pixel samples. None of the sampling schemes was significantly biased on the average, and none of the average estimates differed significantly from the population parameter. The maximum absolute deviation, however, was directly related to sampling unit size and should be considered in selection of a sampling unit.

Hixson, M. M.↗

Automatic discovery of optimal classes

A criterion, based on Bayes' theorem, is described that defines the optimal set of classes (a classification) for a given set of examples. This criterion is transformed into an equivalent minimum message length criterion with an intuitive information interpretation. This criterion does not require that the number of classes be specified in advance, this is determined by the data. The minimum message length criterion includes the message length required to describe the classes, so there is a built in bias against adding new classes unless they lead to a reduction in the message length required to describe the data. Unfortunately, the search space of possible classifications is too large to search exhaustively, so heuristic search methods, such as simulated annealing, are applied. Tutored learning and probabilistic prediction in particular cases are an important indirect result of optimal class discovery. Extensions to the basic class induction program include the ability to combine category and real value data, hierarchical classes, independent classifications and deciding for each class which attributes are relevant.

Cheeseman, Peter↗

Results from the Crop Identification Technology Assessment for Remote Sensing (CITARS) project

The author has identified the following significant results. It was found that several factors had a significant effect on crop identification performance: (1) crop maturity and site characteristics, (2) which of several different single date automatic data processing procedures was used for local recognition, (3) nonlocal recognition, both with and without preprocessing for the extension of recognition signatures, and (4) use of multidate data. It also was found that classification accuracy for field center pixels was not a reliable indicator of proportion estimation performance for whole areas, that bias was present in proportion estimates, and that training data and procedures strongly influenced crop identification performance.

Bauer, M. E.↗

Improved Estimates of Clear Sky Longwave Flux and Application to the Tropical Greenhouse Effect

The first objective of this investigation is to eliminate the clear-sky offset introduced by the scene-identification procedures developed for the Earth Radiation Budget Experiment (ERBE). Estimates of this systematic bias range from 10 to as high as 30 W/sq m. The initial version of the ScaRaB data is being processed with the original ERBE algorithm. Since the ERBE procedure for scene identification is based upon zonal flux averages, clear scenes with longwave emission well below the zonal mean value are mistakenly classified as cloudy. The erroneous classification is more frequent in regions with deep convection and enhanced mid- and upper-tropospheric humidity. We will develop scene identification parameters with zonal and/or time dependence to reduce or eliminate the bias in the clear- sky data. The modified scene identification procedure could be used for the ScaRaB-specific version of the Earth-radiation products. The second objective is to investigate changes in the clear-sky Outgoing Longwave Radiation (OLR) associated with decadal variations in the tropical and subtropical climate. There is considerable evidence for a shift in the climate state starting in approximately 1977. The shift is accompanied by higher SSTs in the equatorial Pacific, increased tropical convection, and higher values of atmospheric humidity. Other evidence indicates that the humidity in the tropical troposphere has been steadily increasing over the last 30 years. It is not known whether the atmospheric greenhouse effect has increased during this period in response to these changes in SST and precipitable water. We will investigate the decadal-scale fluctuations in the greenhouse effect using Nimbus-7, ERBE, and ScaRaB measurements spaning 1979 to the present. The data from the different satellites will be intercalibrated by comparison with model calculations based upon ship radiosonde observations. The fluxes calculated from the radiation model will also be used for validation of the ScaRaB fluxes.

Collins, W. D.↗

Geography of the asteroid belt

The CSM classification serves as the starting point on the geography of the asteroid belt. Raw data on asteroid types are corrected for observational biases (against dark objects, for instance) to derive the distribution of types throughout the belt. Recent work on family members indicates that dynamical families have a true physical relationship, presumably indicating common origin in the breakup of a parent asteroid.

Zellner, B. H.↗

AgRISTARS: Foreign commodity production forecasting. The 1980 US corn and soybeans exploratory experiment

The U.S. corn and soybeans exploratory experiment is described which consisted of evaluations of two technology components of a production forecasting system: classification procedures (crop labeling and proportion estimation at the level of a sampling unit) and sampling and aggregation procedures. The results from the labeling evaluations indicate that the corn and soybeans labeling procedure works very well in the U.S. corn belt with full season (after tasseling) LANDSAT data. The procedure should be readily adaptable to corn and soybeans labeling required for subsequent exploratory experiments or pilot tests. The machine classification procedures evaluated in this experiment were not effective in improving the proportion estimates. The corn proportions produced by the machine procedures had a large bias when the bias correction was not performed. This bias was caused by the manner in which the machine procedures handled spectrally impure pixels. The simulation test indicated that the weighted aggregation procedure performed quite well. Although further work can be done to improve both the simulation tests and the aggregation procedure, the results of this test show that the procedure should serve as a useful baseline procedure in future exploratory experiments and pilot tests.

Malin, J. T.↗