Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluating algorithmic bias on biomarker classification of breast cancer pathology reports

Objectives: This work evaluated algorithmic bias in biomarkers classification using electronic pathology reports from female breast cancer cases. Bias was assessed across 5 subgroups: cancer registry, race, Hispanic ethnicity, age at diagnosis, and socioeconomic status. Materials and Methods: We utilized 594 875 electronic pathology reports from 178 121 tumors diagnosed in Kentucky, Louisiana, New Jersey, New Mexico, Seattle, and Utah to train 2 deep-learning algorithms to classify breast cancer patients using their biomarkers test results. We used balanced error rate (BER), demographic parity (DP), equalized odds (EOD), and equal opportunity (EOP) to assess bias. Results: We found differences in predictive accuracy between registries, with the highest accuracy in the registry that contributed the most data (Seattle Registry, BER ratios for all registries >1.25). BER showed no significant algorithmic bias in extracting biomarkers (estrogen receptor, progesterone receptor, human epidermal growth factor receptor 2) for race, Hispanic ethnicity, age at diagnosis, or socioeconomic subgroups (BER ratio <1.25). DP, EOD, and EOP all showed insignificant results. Discussion: We observed significant differences in BER by registry, but no significant bias using the DP, EOD, and EOP metrics for socio-demographic or racial categories. This highlights the importance of employing a diverse set of metrics for a comprehensive evaluation of model fairness. Conclusion: A thorough evaluation of algorithmic biases that may affect equality in clinical care is a critical step before deploying algorithms in the real world. We found little evidence of algorithmic bias in our biomarker classification tool. Artificial intelligence tools to expedite information extraction from clinical records could accelerate clinical trial matching and improve care.

60 APPLIED LIFE SCIENCES↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A spatiotemporally explicit and scalable indicator of intact lands across the conterminous United States, 1986–2023

Globally, ecologically intact areas are increasingly scarce. Agricultural expansion into previously uncultivated areas drives the loss of intact lands that might otherwise exhibit high levels of ecological integrity. Thus, the absence of cultivation can be an indicator of intact lands as measured from remote sensing data and thematic maps. Our objective for this study was to develop and compare tractable approaches based on remotely sensed satellite data to map spatial patterns of potentially intact lands across the conterminous U.S. (CONUS). Using annual cultivation probabilities derived from satellite observations, we classified and mapped potentially intact lands across CONUS from 1986 to 2023 at 30 m resolution. We created three maps, first by applying a constant cultivation probability threshold across CONUS, second by varying the threshold state-by-state to maximize state-level overall accuracies, and third by equalizing the state-level user's and producer's accuracies to minimize classification bias. Validation against 800,000+ independent ground samples resulted in CONUS-level overall accuracies ≥85% for the roughly 660 million ha of potentially intact land. Map accuracy varied with the proportion of potentially intact lands across regions, with the Pacific-Mountain and Great Plains regions exhibiting the highest accuracies, while Eastern CONUS exhibited a greater mix of potentially intact and non-intact lands and more moderate map accuracies. These novel maps and approaches can be adapted to different spatiotemporal extents to support conservation and production decisions ranging from species and ecosystems protection to reducing land conversion and climate mitigation.

agriculture↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

The Dark Energy Survey supernova program: cosmological biases from supernova photometric classification

ABSTRACT Cosmological analyses of samples of photometrically identified type Ia supernovae (SNe Ia) depend on understanding the effects of ‘contamination’ from core-collapse and peculiar SN Ia events. We employ a rigorous analysis using the photometric classifier SuperNNova on state-of-the-art simulations of SN samples to determine cosmological biases due to such ‘non-Ia’ contamination in the Dark Energy Survey (DES) 5-yr SN sample. Depending on the non-Ia SN models used in the SuperNNova training and testing samples, contamination ranges from 0.8 to 3.5 per cent, with a classification efficiency of 97.7–99.5 per cent. Using the Bayesian Estimation Applied to Multiple Species (BEAMS) framework and its extension BBC (‘BEAMS with Bias Correction’), we produce a redshift-binned Hubble diagram marginalized over contamination and corrected for selection effects, and use it to constrain the dark energy equation-of-state, w. Assuming a flat universe with Gaussian ΩM prior of 0.311 ± 0.010, we show that biases on w are <0.008 when using SuperNNova, with systematic uncertainties associated with contamination around 10 per cent of the statistical uncertainty on w for the DES-SN sample. An alternative approach of discarding contaminants using outlier rejection techniques (e.g. Chauvenet’s criterion) in place of SuperNNova leads to biases on w that are larger but still modest (0.015–0.03). Finally, we measure biases due to contamination on w0 and wa (assuming a flat universe), and find these to be <0.009 in w0 and <0.108 in wa, 5 to 10 times smaller than the statistical uncertainties for the DES-SN sample.

79 ASTRONOMY AND ASTROPHYSICS↗

A bootstrapping approach to social media quantification

Abstract This work considers the use of classifiers in a downstream aggregation task estimating class proportions, such as estimating the percentage of reviews for a movie with positive sentiment. We derive the bias and variance of the class proportion estimator when taking classification error into account to determine how to best trade off different error types when tuning a classifier for these tasks. Additionally, we propose a method for constructing confidence intervals that correctly adjusts for classification error when estimating these statistics. We conduct experiments on four document classification tasks comparing our methods to prior approaches across classifier thresholds, sample sizes, and label distributions. Prior approaches have focused on providing the most accurate point estimate while this work focuses on the creation of correct confidence intervals that appropriately account for classifier error. Compared to the prior approaches, our methods provide lower error and more accurate confidence intervals.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES↗

The Dark Energy Survey supernova programme: modelling selection efficiency and observed core-collapse supernova contamination

ABSTRACT The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accurate photometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes an essential step for training and testing photometric classification algorithms, and for correcting biases introduced by selection effects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-series spectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies to construct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark Energy Survey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations. We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data. We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and before photometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation, we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 per cent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveys without fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmology analysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associated systematic uncertainty.

79 ASTRONOMY AND ASTROPHYSICS↗

ZTF SN Ia DR2 follow-up: Exploring the origin of the Type Ia supernova host galaxy step through Si II velocities

The relation between Type Ia supernovae (SNe Ia) and the stellar masses of their host galaxy is well documented. In particular, Hubble residuals display a distinct luminosity shift based on host mass. This is known as the mass step. This effect is widely used as an additional correction factor in the standardisation of SN Ia luminosities. We investigate the Hubble residuals and the mass step of normal SNe Ia in the context of Si IIλ6355 velocities based on 277 normal SNe Ia that are near their peak in the second data release (DR2) of the Zwicky Transient Facility (ZTF). We divided the sample into high-velocity (HV) and normal-velocity (NV) SNe Ia, separated at 12,000 km s −1 . This produced a sample of 70 HV and 207 NV objects. We then explored potential environment- and/or progenitor-related effects by investigating the Si IIλ6355 velocities with parameters such as the light-curve stretch x 1 , the colour c, and the host galaxy properties. Although we only find a marginal difference between the Hubble residuals of HV and NV SNe Ia, the NV mass step is 0.149 ± 0.024 mag (6.3σ). The HV mass step is smaller, 0.046 ± 0.041 mag (1.1σ), and is consistent with zero. The difference between the NV and HV mass steps is modest, at ∼2.2σ. Moreover, the clearest subtype difference appears for SNe in central regions (d DLR < 1), where NV SNe Ia show a large mass step, whereas HV SNe Ia are consistent with no step, yielding a difference of 3.1–3.6σ between NV and HV SNe Ia. We observe a host-colour step for both subtypes. NV SNe Ia show a step of 0.142 ± 0.024 mag (5.9σ), while HV SNe Ia show a step of 0.158 ± 0.042 mag (3.8σ), where the HV SNe Ia step appears to be larger, but the significance is lower because the sample size is smaller. Overall, the NV and HV colour steps are statistically consistent. HV SNe Ia also show modest (∼2.5–3σ) steps in certain subsets, such as those in outer regions (d DLR > 1), whereas NV SNe display stronger environmental trends. Our results indicate that NV SNe Ia appear to be more environmentally sensitive, particularly in central likely metal-rich and older regions, while HV SNe Ia show weaker and subset-dependent trends. This suggests that applying a universal mass-step correction might introduce biases, and that incorporating refined classifications and/or environment-dependent factors, such as the location within the host, might improve future cosmological analyses beyond the standard x 1 and c cuts.

supernovae: general↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Evaluation of data collection bias of third molar stages of mineralisation for age estimation in the living

Abstract Age assessment of the living is a fundamental procedure in the process of human identification, in order to guarantee fair treatment of individuals, which has ethical, civil, legal, and medical repercussions. The careful selection of the appropriate methods requires evaluation of several parameters: accuracy, precision of the method, as well as its reproducibility. The approach proposed by Mincer et al. adapted from Demirjian et al. exploring third molar mineralisation, is one of the most frequently considered for age estimation of the living. Thus, this work aims to assess potential bias in the data collection when applying the classification stages for dental mineralisation adapted by Mincer et al. A total of 102 orthopantomographs, of clinical origin, belonging to individuals aged between 12 and 25 years ($ \bar{\textit x} $ = 20.12 years, SD = 3.49 years; 65 females, 37 males, all of Portuguese nationality) were included and a retrospective analysis performed by five observers with different levels of experience (high, average, and basic). The performance and agreement between five observers were evaluated using Weighted Cohen’s Kappa and the Intraclass Correlation Coefficient. To access the influence of impaction on third molar classification, variables were tested using ordinal logistic regression Generalised Linear Model. It was observed that there were variations in the number of teeth identified among the observers, but the agreement levels ranged from moderate to substantial (0.4–0.8). Upon closer examination of the results, it was observed that although there were discernible differences between highly experienced observers and those with less experience, the gap was not as significant as initially hypothesised, and a greater disparity between the classifications of the upper (0.24–0.49) and lower third molars (>0.55) was observed. When bone superimposition is present, the classification process is not significantly influenced; however, variation in teeth angulation affects the assessment. The results suggest that with an efficient preparation, the level of experience as a factor can be overcome. Mincer and colleague's classification system can be replicated with ease and consistency, even though the classification of upper and lower third molars presents distinct challenges.

de Oliveira Santos, Inês (ORCID:0000000267324347)↗

Characteristics and selection of near-fault simulated earthquake ground motions for nonlinear analysis of buildings

Earthquake-induced ground shaking near rupturing faults is highly sensitive to the rupture characteristics, seismic wave propagation patterns and site conditions, and field recordings of near-fault shaking are relatively sparse. These challenges complicate the assessment of the seismic performance of near-fault structures. A common approach to representing near-fault ground motion in engineering analysis is to explicitly consider and select records with strong directivity pulses (pulse records). We use three-dimensional high-resolution physics-based earthquake simulations to test this approach in the context of scenario-based ground motion record selection, and to study the important characteristics of near-fault ground shaking. We highlight the deficiencies associated with classifying near-fault simulated records as “pulse” or “non-pulse,” based on the presence of a single dominating pulse in the velocity time history. We show that this approach is inadequate for characterizing near-fault shaking on soft soils which can be dominated by both forward rupture directivity and basin amplification effects. We conduct ground motion selection experiments for the analysis of near-fault structures with and without explicit classification of the pulse features in the records, and evaluate the bias in the predicted structural demands. We find that the maximum interstory drift demands on building structures imposed by unscaled site-specific simulated ground motion records selected based on relevant spectral shape features are not sensitive to the classification of records as pulse/non-pulse. Therefore, with regard to predicting the maximum interstory drifts in near-fault buildings, we do not find justification for the binary pulse classification of near-fault records.

42 ENGINEERING↗

Galaxy Spin Classification. I. Z-wise versus S-wise Spirals with the Chirality Equivariant Residual Network

The angular momentum of galaxies (galaxy spin) contains rich information about the initial condition of the universe, yet it is challenging to efficiently measure the spin direction for the tremendous amount of galaxies that are being mapped by ongoing and forthcoming cosmological surveys. We present a machine-learning-based classifier for the Z-wise versus S-wise spirals, which can help to break the degeneracy in the galaxy spin direction measurement. The proposed chirality equivariant residual network (CE-ResNet) is manifestly equivariant under a reflection of the input image, which guarantees that there is no inherent asymmetry between the Z-wise and S-wise probability estimators. We train the model with Sloan Digital Sky Survey images, with the training labels given by the Galaxy Zoo 1 project. A combination of data augmentation techniques is used during the training, making the model more robust to be applied to other surveys. We find an ~30% increase in both types of spirals when Dark Energy Spectroscopic Instrument (DESI) images are used for classification, due to the better imaging quality of DESI. We verify that the ~7σ difference between the numbers of Z-wise and S-wise spirals is due to human bias, since the discrepancy drops to <1.8σ with our CE-ResNet classification results. We discuss the potential systematics relevant to future cosmological applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Heterogeneous Graph Neural Network for identifying hadronically decayed tau leptons at the High Luminosity LHC

Here, we present a new algorithm that identifies reconstructed jets originating from hadronic decays of tau leptons against those from quarks or gluons. No tau lepton reconstruction algorithm is used. Instead, the algorithm represents jets as heterogeneous graphs with tracks and energy clusters as nodes and trains a Graph Neural Network to identify tau jets from other jets. Different attributed graph representations and different GNN architectures are explored. We propose to use differential track and energy cluster information as node features and a heterogeneous sequentially-biased encoding for the inputs to final graph-level classification.

47 OTHER INSTRUMENTATION↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Integrating NDVI-Based Within-Wetland Vegetation Classification in a Land Surface Model Improves Methane Emission Estimations

Earth system models (ESMs) are a common tool for estimating local and global greenhouse gas emissions under current and projected future conditions. Efforts are underway to expand the representation of wetlands in the Energy Exascale Earth System Model (E3SM) Land Model (ELM) by resolving the simultaneous contributions to greenhouse gas fluxes from multiple, different, sub-grid-scale patch-types, representing different eco-hydrological patches within a wetland. However, for this effort to be effective, it should be coupled with the detection and mapping of within-wetland eco-hydrological patches in real-world wetlands, providing models with corresponding information about vegetation cover. In this short communication, we describe the application of a recently developed NDVI-based method for within-wetland vegetation classification on a coastal wetland in Louisiana and the use of the resulting yearly vegetation cover as input for ELM simulations. Processed Harmonized Landsat and Sentinel-2 (HLS) datasets were used to drive the sub-grid composition of simulated wetland vegetation each year, thus tracking the spatial heterogeneity of wetlands at sufficient spatial and temporal resolutions and providing necessary input for improving the estimation of methane emissions from wetlands. Our results show that including NDVI-based classification in an ELM reduced the uncertainty in predicted methane flux by decreasing the model’s RMSE when compared to Eddy Covariance measurements, while a minimal bias was introduced due to the resampling technique involved in processing HLS data. Our study shows promising results in integrating the remote sensing-based classification of within-wetland vegetation cover into earth system models, while improving their performances toward more accurate predictions of important greenhouse gas emissions.

54 ENVIRONMENTAL SCIENCES↗