Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Selecting class weights to minimize classification bias in acreage estimation

Preliminary results of experiments being performed to select optimal class weights for use with the maximum likelihood classifier in acreage estimation using remote sensor imagery are presented. These weights will be optimal in the sense that the bias will be minimized in the proportion estimate obtained from the classification results by sample counting. The procedure was tested using Landsat MSS data from an 8 by 9.6 km area of ground truth in Finney County, Kansas.

Belcher, W. M.↗

Evaluating algorithmic bias on biomarker classification of breast cancer pathology reports

Objectives: This work evaluated algorithmic bias in biomarkers classification using electronic pathology reports from female breast cancer cases. Bias was assessed across 5 subgroups: cancer registry, race, Hispanic ethnicity, age at diagnosis, and socioeconomic status. Materials and Methods: We utilized 594 875 electronic pathology reports from 178 121 tumors diagnosed in Kentucky, Louisiana, New Jersey, New Mexico, Seattle, and Utah to train 2 deep-learning algorithms to classify breast cancer patients using their biomarkers test results. We used balanced error rate (BER), demographic parity (DP), equalized odds (EOD), and equal opportunity (EOP) to assess bias. Results: We found differences in predictive accuracy between registries, with the highest accuracy in the registry that contributed the most data (Seattle Registry, BER ratios for all registries >1.25). BER showed no significant algorithmic bias in extracting biomarkers (estrogen receptor, progesterone receptor, human epidermal growth factor receptor 2) for race, Hispanic ethnicity, age at diagnosis, or socioeconomic subgroups (BER ratio <1.25). DP, EOD, and EOP all showed insignificant results. Discussion: We observed significant differences in BER by registry, but no significant bias using the DP, EOD, and EOP metrics for socio-demographic or racial categories. This highlights the importance of employing a diverse set of metrics for a comprehensive evaluation of model fairness. Conclusion: A thorough evaluation of algorithmic biases that may affect equality in clinical care is a critical step before deploying algorithms in the real world. We found little evidence of algorithmic bias in our biomarker classification tool. Artificial intelligence tools to expedite information extraction from clinical records could accelerate clinical trial matching and improve care.

60 APPLIED LIFE SCIENCES↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A spatiotemporally explicit and scalable indicator of intact lands across the conterminous United States, 1986–2023

Globally, ecologically intact areas are increasingly scarce. Agricultural expansion into previously uncultivated areas drives the loss of intact lands that might otherwise exhibit high levels of ecological integrity. Thus, the absence of cultivation can be an indicator of intact lands as measured from remote sensing data and thematic maps. Our objective for this study was to develop and compare tractable approaches based on remotely sensed satellite data to map spatial patterns of potentially intact lands across the conterminous U.S. (CONUS). Using annual cultivation probabilities derived from satellite observations, we classified and mapped potentially intact lands across CONUS from 1986 to 2023 at 30 m resolution. We created three maps, first by applying a constant cultivation probability threshold across CONUS, second by varying the threshold state-by-state to maximize state-level overall accuracies, and third by equalizing the state-level user's and producer's accuracies to minimize classification bias. Validation against 800,000+ independent ground samples resulted in CONUS-level overall accuracies ≥85% for the roughly 660 million ha of potentially intact land. Map accuracy varied with the proportion of potentially intact lands across regions, with the Pacific-Mountain and Great Plains regions exhibiting the highest accuracies, while Eastern CONUS exhibited a greater mix of potentially intact and non-intact lands and more moderate map accuracies. These novel maps and approaches can be adapted to different spatiotemporal extents to support conservation and production decisions ranging from species and ecosystems protection to reducing land conversion and climate mitigation.

agriculture↗

Crop identification technology assessment for remote sensing (CITARS). Volume 10: Interpretation of results

The CITARS was an experiment designed to quantitatively evaluate crop identification performance for corn and soybeans in various environments using a well-defined set of automatic data processing (ADP) techniques. Each technique was applied to data acquired to recognize and estimate proportions of corn and soybeans. The CITARS documentation summarizes, interprets, and discusses the crop identification performances obtained using (1) different ADP procedures; (2) a linear versus a quadratic classifier; (3) prior probability information derived from historic data; (4) local versus nonlocal recognition training statistics and the associated use of preprocessing; (5) multitemporal data; (6) classification bias and mixed pixels in proportion estimation; and (7) data with differnt site characteristics, including crop, soil, atmospheric effects, and stages of crop maturity.

Bizzell, R. M.↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

The Dark Energy Survey supernova program: cosmological biases from supernova photometric classification

ABSTRACT Cosmological analyses of samples of photometrically identified type Ia supernovae (SNe Ia) depend on understanding the effects of ‘contamination’ from core-collapse and peculiar SN Ia events. We employ a rigorous analysis using the photometric classifier SuperNNova on state-of-the-art simulations of SN samples to determine cosmological biases due to such ‘non-Ia’ contamination in the Dark Energy Survey (DES) 5-yr SN sample. Depending on the non-Ia SN models used in the SuperNNova training and testing samples, contamination ranges from 0.8 to 3.5 per cent, with a classification efficiency of 97.7–99.5 per cent. Using the Bayesian Estimation Applied to Multiple Species (BEAMS) framework and its extension BBC (‘BEAMS with Bias Correction’), we produce a redshift-binned Hubble diagram marginalized over contamination and corrected for selection effects, and use it to constrain the dark energy equation-of-state, w. Assuming a flat universe with Gaussian ΩM prior of 0.311 ± 0.010, we show that biases on w are <0.008 when using SuperNNova, with systematic uncertainties associated with contamination around 10 per cent of the statistical uncertainty on w for the DES-SN sample. An alternative approach of discarding contaminants using outlier rejection techniques (e.g. Chauvenet’s criterion) in place of SuperNNova leads to biases on w that are larger but still modest (0.015–0.03). Finally, we measure biases due to contamination on w0 and wa (assuming a flat universe), and find these to be <0.009 in w0 and <0.108 in wa, 5 to 10 times smaller than the statistical uncertainties for the DES-SN sample.

79 ASTRONOMY AND ASTROPHYSICS↗

A bootstrapping approach to social media quantification

Abstract This work considers the use of classifiers in a downstream aggregation task estimating class proportions, such as estimating the percentage of reviews for a movie with positive sentiment. We derive the bias and variance of the class proportion estimator when taking classification error into account to determine how to best trade off different error types when tuning a classifier for these tasks. Additionally, we propose a method for constructing confidence intervals that correctly adjusts for classification error when estimating these statistics. We conduct experiments on four document classification tasks comparing our methods to prior approaches across classifier thresholds, sample sizes, and label distributions. Prior approaches have focused on providing the most accurate point estimate while this work focuses on the creation of correct confidence intervals that appropriately account for classifier error. Compared to the prior approaches, our methods provide lower error and more accurate confidence intervals.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The intercrater plains of Mercury and the Moon: Their nature, origin and role in terrestrial planet evolution. Measurement and errors of crater statistics

Planetary imagery techniques, errors in measurement or degradation assignment, and statistical formulas are presented with respect to cratering data. Base map photograph preparation, measurement of crater diameters and sampled area, and instruments used are discussed. Possible uncertainties, such as Sun angle, scale factors, degradation classification, and biases in crater recognition are discussed. The mathematical formulas used in crater statistics are presented.

Leake, M. A.↗

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES↗

Creating a Canonical Scientific and Technical Information Classification System for NCSTRL+

The purpose of this paper is to describe the new subject classification system for the NCSTRL+ project. NCSTRL+ is a canonical digital library (DL) based on the Networked Computer Science Technical Report Library (NCSTRL). The current NCSTRL+ classification system uses the NASA Scientific and Technical (STI) subject classifications, which has a bias towards the aerospace, aeronautics, and engineering disciplines. Examination of other scientific and technical information classification systems showed similar discipline-centric weaknesses. Traditional, library-oriented classification systems represented all disciplines, but were too generalized to serve the needs of a scientific and technically oriented digital library. Lack of a suitable existing classification system led to the creation of a lightweight, balanced, general classification system that allows the mapping of more specialized classification schemes into the new framework. We have developed the following classification system to give equal weight to all STI disciplines, while being compact and lightweight.

Tiffany, Melissa E.↗

Uncertainty methodology for in-flight thrust determination

A methodology is proposed for the evaluation of uncertainty in the in-flight determination of aircraft thrust, which provides error traceability to a national standards laboratory, and is independent of the procedure used to calculate or measure thrust in flight, thereby yielding a consistent means for the evaluation of measurement capabilities. Attention is given to the factors of measurement error, precision, bias, uncertainty, error estimation and classification, error propagation, ground testing, and the related problems of model bias error, model precision error, and the uncertainty limit.

Adams, G. R.↗

An ad hoc map evaluation procedure

An ad hoc map evaluation procedure is proposed which is most suitable for evaluating low-resolution classification maps against high resolution ground truth maps, such as maps against interpreted aircraft photographs. Commonly practiced sampling and evaluation procedures are impracticable in this context because of difficulties in registration and in comparing the samples. This ad hoc procedure is designed to overcome these two major problems, and its practicability is discussed. Two widely accepted parameters are estimated by the new procedure; namely, the probability of correct classification and the proportion biases. Statistical qualifications are also provided.

Kan, E. P.↗

The Dark Energy Survey Supernova Programme: Modelling Selection Efficiency and Observed Core-collapse Supernova Contamination

The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accuratephotometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes anessential step for training and testing photometric classification algorithms, and for correcting biases introduced by selectioneffects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-seriesspectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies toconstruct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark EnergySurvey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations.We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data.We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and beforephotometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation,we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 percent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveyswithout fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmologyanalysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associatedsystematic uncertainty.

M Vincenzi↗

The Dark Energy Survey supernova programme: modelling selection efficiency and observed core-collapse supernova contamination

ABSTRACT The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accurate photometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes an essential step for training and testing photometric classification algorithms, and for correcting biases introduced by selection effects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-series spectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies to construct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark Energy Survey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations. We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data. We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and before photometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation, we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 per cent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveys without fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmology analysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associated systematic uncertainty.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of signature extension algorithms

The author has identified the following significant results. One of the major findings was that nearly all of the bias in the proportion estimates of the multisegment training and classification procedure resulted from the particular configuration of the signature set used for classification, rather than from peculiarities of the recognition sample segments. This meant that the proportion estimation bias could be accurately corrected simply by estimating the bias on the original six training segments. The bias corrected proportion estimates of the multisegment training and classification procedure were extremely accurate and had a low variance when compared to local training and classification. This finding may have important ramifications for reducing the cost and increasing the accuracy of bias correction procedures.

Nalepka, R. F.↗

ZTF SN Ia DR2 follow-up: Exploring the origin of the Type Ia supernova host galaxy step through Si II velocities

The relation between Type Ia supernovae (SNe Ia) and the stellar masses of their host galaxy is well documented. In particular, Hubble residuals display a distinct luminosity shift based on host mass. This is known as the mass step. This effect is widely used as an additional correction factor in the standardisation of SN Ia luminosities. We investigate the Hubble residuals and the mass step of normal SNe Ia in the context of Si IIλ6355 velocities based on 277 normal SNe Ia that are near their peak in the second data release (DR2) of the Zwicky Transient Facility (ZTF). We divided the sample into high-velocity (HV) and normal-velocity (NV) SNe Ia, separated at 12,000 km s −1 . This produced a sample of 70 HV and 207 NV objects. We then explored potential environment- and/or progenitor-related effects by investigating the Si IIλ6355 velocities with parameters such as the light-curve stretch x 1 , the colour c, and the host galaxy properties. Although we only find a marginal difference between the Hubble residuals of HV and NV SNe Ia, the NV mass step is 0.149 ± 0.024 mag (6.3σ). The HV mass step is smaller, 0.046 ± 0.041 mag (1.1σ), and is consistent with zero. The difference between the NV and HV mass steps is modest, at ∼2.2σ. Moreover, the clearest subtype difference appears for SNe in central regions (d DLR < 1), where NV SNe Ia show a large mass step, whereas HV SNe Ia are consistent with no step, yielding a difference of 3.1–3.6σ between NV and HV SNe Ia. We observe a host-colour step for both subtypes. NV SNe Ia show a step of 0.142 ± 0.024 mag (5.9σ), while HV SNe Ia show a step of 0.158 ± 0.042 mag (3.8σ), where the HV SNe Ia step appears to be larger, but the significance is lower because the sample size is smaller. Overall, the NV and HV colour steps are statistically consistent. HV SNe Ia also show modest (∼2.5–3σ) steps in certain subsets, such as those in outer regions (d DLR > 1), whereas NV SNe display stronger environmental trends. Our results indicate that NV SNe Ia appear to be more environmentally sensitive, particularly in central likely metal-rich and older regions, while HV SNe Ia show weaker and subset-dependent trends. This suggests that applying a universal mass-step correction might introduce biases, and that incorporating refined classifications and/or environment-dependent factors, such as the location within the host, might improve future cosmological analyses beyond the standard x 1 and c cuts.

supernovae: general↗

Some properties of the Svalgaard A/C index

Several properties of the A/C index of polar cap variations introduced by L. Svalgaard in 1972 have been found to vary with time. In the presatellite era, C days, as measured by the Ap index, are almost twice as geomagnetically active as A days, while in the modern epoch they have essentially identical activity. Prior to 1962 there were over 40% more A days than C days per year, while during the modern epoch there are essentially equal numbers of A days and C days. In view of this strong bias to assigning a 'toward' classification on geomagnetically active days it is recommended that the Svalgaard A/C classification not be used in studies of geomagnetic activity. It is also recommended that a full and thorough documentation of the index be prepared and/or that others undertake to compile such a classification separately.

Russell, C. T.↗