Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

Active Learning with Irrelevant Examples

Active learning algorithms attempt to accelerate the learning process by requesting labels for the most informative items first. In real-world problems, however, there may exist unlabeled items that are irrelevant to the user's classification goals. Queries about these points slow down learning because they provide no information about the problem of interest. We have observed that when irrelevant items are present, active learning can perform worse than random selection, requiring more time (queries) to achieve the same level of accuracy. Therefore, we propose a novel approach, Relevance Bias, in which the active learner combines its default selection heuristic with the output of a simultaneously trained relevance classifier to favor items that are likely to be both informative and relevant. In our experiments on a real-world problem and two benchmark datasets, the Relevance Bias approach significantly improved the learning rate of three different active learning approaches.

machine learning↗

A catalogue of cataclysmic variables from 20 yr of the Sloan Digital Sky Survey with new classifications, periods, trends, and oddities

ABSTRACT We present a catalogue of 507 cataclysmic variables (CVs) observed in SDSS I to IV including 70 new classifications collated from multiple archival data sets. This represents the largest sample of CVs with high-quality and homogeneous optical spectroscopy. We have used this sample to derive unbiased space densities and period distributions for the major sub-types of CVs. We also report on some peculiar CVs, period bouncers and also CVs exhibiting large changes in accretion rates. We report 70 new CVs, 59 new periods, 178 unpublished spectra, and 262 new or updated classifications. From the SDSS spectroscopy, we also identified 18 systems incorrectly identified as CVs in the literature. We discuss the observed properties of 13 peculiar CVS, and we identify a small set of eight CVs that defy the standard classification scheme. We use this sample to investigate the distribution of different CV sub-types, and we estimate their individual space densities, as well as that of the entire CV population. The SDSS I to IV sample includes 14 period bounce CVs or candidates. We discuss the variability of CVs across the Hertzsprung–Russell diagram, highlighting selection biases of variability-based CV detection. Finally, we searched for, and found eight tertiary companions to the SDSS CVs. We anticipate that this catalogue and the extensive material included in the Supplementary Data will be useful for a range of observational population studies of CVs.

79 ASTRONOMY AND ASTROPHYSICS↗

Unsupervised Classification of Global Radar Units on Venus

Characterization of the Venusian surface in terms of its radar properties was accomplished by application of an unsupervised, linear discriminant algorithm to two Pioneer-Venus (PV) Orbiter radar data sets: the RMS-slope (surface roughness) and reflectivity. Both databases were spatially filtered to the same effective resolution of 100 km prior to classification. A recent supervised classification study using these data was based on presupposed morphologic significance of selected data ranges. The knowledge of both Venusian geology and the geologic significance of the radar data is so limited that the data warrant a more unsupervised approach; for this study a linear discriminant classifier was chosen. This approach is purely statistical, thereby removing any observer bias. Statistical significance of the resulting clusters was evaluated by an ancillary program in which an F test utilizing the Mahalanobis' distance.

Kozak, R. C.↗

Sampling for area estimation: A comparison of full-frame sampling with the sample segment approach

The author has identified the following significant results. Full-frame classifications of wheat and non-wheat for eighty counties in Kansas were repetitively sampled to simulate alternative sampling plans. Evaluation of four sampling schemes involving different numbers of samples and different size sampling units shows that the precision of the wheat estimates increased as the segment size decreased and the number of segments was increased. Although the average bias associated with the various sampling schemes was not significantly different, the maximum absolute bias was directly related to sampling size unit.

Hixson, M.↗

Radar Retrieval Evaluation and Investigation of Dendritic Growth Layer Polarimetric Signatures in a Winter Storm

Abstract This study evaluates ice particle size distribution and aspect ratio φ Multi-Radar Multi-Sensor (MRMS) dual-polarization radar retrievals through a direct comparison with two legs of observational aircraft data obtained during a winter storm case from the Investigation of Microphysics and Precipitation for Atlantic Coast-Threatening Snowstorms (IMPACTS) campaign. In situ cloud probes, satellite, and MRMS observations illustrate that the often-observed K dp and Z DR enhancement regions in the dendritic growth layer can either indicate a local number concentration increase of dry ice particles or the presence of ice particles mixed with a significant number of supercooled liquid droplets. Relative to in situ measurements, MRMS retrievals on average underestimated mean volume diameters by 50% and overestimated number concentrations by over 100%. IWC retrievals using Z DR and K dp within the dendritic growth layer were minimally biased relative to in situ calculations where retrievals yielded −2% median relative error for the entire aircraft leg. Incorporating φ retrievals decreased both the magnitude and spread of polarimetric retrievals below the dendritic growth layer. While φ radar retrievals suggest that observed dendritic growth layer particles were nonspherical (0.1 ≤ φ ≤ 0.2), in situ projected aspect ratios, idealized numerical simulations, and habit classifications from cloud probe images suggest that the population mean φ was generally much higher. Coordinated aircraft radar reflectivity with in situ observations suggests that the MRMS systematically underestimated reflectivity and could not resolve local peaks in mean volume diameter sizes. These results highlight the need to consider particle assumptions and radar limitations when performing retrievals. significance statement Developing snow is often detectable using weather radars. Meteorologists combine these radar measurements with mathematical equations to study how snow forms in order to determine how much snow will fall. This study evaluates current methods for estimating the total number and mass, sizes, and shapes of snowflakes from radar using images of individual snowflakes taken during two aircraft legs. Radar estimates of snowflake properties were most consistent with aircraft data inside regions with prominent radar signatures. However, radar estimates of snowflake shapes were not consistent with observed shapes estimated from the snowflake images. Although additional research is needed, these results bolster understanding of snow-growth physics and uncertainties between radar measurements and snow production that can improve future snowfall forecasting.

Meteorology & Atmospheric Sciences↗

Geomorphic classification of Icelandic and Martian volcanoes: Limitations of comparative planetology research from LANDSAT and Viking orbiter images

Some limitations in using orbital images of planetary surfaces for comparative landform analyses are discussed. The principal orbital images used were LANDSAT MSS images of Earth and nominal Viking Orbiter images of Mars. Both are roughly comparable in having a pixel size which corresponds to about 100 m on the planetary surface. A volcanic landform on either planet must have a horizontal dimension of at least 200 m to be discernible on orbital images. A twofold bias is directly introduced into any comparative analysis of volcanic landforms on Mars versus those in Iceland because of this scale limitation. First, the 200-m cutoff of landforms may delete more types of volcanic landforms on Earth than on Mars or vice versa. Second, volcanic landforms in Iceland, too small to be resolved or orbital images, may be represented by larger counterparts on Mars or vice versa.

Williams, R. S., Jr.↗

Toward Complete Merger Identification at Cosmic Noon with Deep Learning

As we enter the era of large imaging surveys such as Roman, Rubin, and Euclid, a deeper understanding of potential biases and selection effects in optical astronomical catalogs created with the use of ML-based methods is paramount. This work focuses on a deeper understanding of the performance and limitations of deep learning-based classifiers as tools for galaxy merger identification. We train a ConvNeXT-Pico model on mock HST CANDELS images from the IllustrisTNG50 simulation. Our focus is on a more challenging classification of galaxy mergers and non-mergers at higher redshifts 1 < z < 1.5, including minor mergers and lower mass galaxies down to the stellar mass of 108M⊙. We demonstrate, for the first time, that a deep learning model, such as the one developed in this work, can successfully identify even minor and low mass mergers even at these redshifts. Our model achieves overall accuracy, purity, and completeness of over 73%. We show that some galaxy mergers can only be identified from certain observation angles, leading to a potential upper limit in overall accuracy. Using Grad-CAMs and UMAPs, we more deeply examine the performance and observe a visible gradient in the latent space with stellar mass and specific star formation rate, but no visible gradient with merger mass ratio or merger stage.

Schechter, Aimee L. [U. Colorado, Boulder]↗

An expert system based software sizing tool, phase 2

A software tool was developed for predicting the size of a future computer program at an early stage in its development. The system is intended to enable a user who is not expert in Software Engineering to estimate software size in lines of source code with an accuracy similar to that of an expert, based on the program's functional specifications. The project was planned as a knowledge based system with a field prototype as the goal of Phase 2 and a commercial system planned for Phase 3. The researchers used techniques from Artificial Intelligence and knowledge from human experts and existing software from NASA's COSMIC database. They devised a classification scheme for the software specifications, and a small set of generic software components that represent complexity and apply to large classes of programs. The specifications are converted to generic components by a set of rules and the generic components are input to a nonlinear sizing function which makes the final prediction. The system developed for this project predicted code sizes from the database with a bias factor of 1.06 and a fluctuation factor of 1.77, an accuracy similar to that of human experts but without their significant optimistic bias.

Friedlander, David↗

Toward Complete Merger Identification at Cosmic Noon with Deep Learning

As we enter the era of large imaging surveys such as $\textit{Roman}$, Rubin, and $\textit{Euclid}$, a deeper understanding of potential biases and selection effects in optical astronomical catalogs created with the use of ML-based methods is paramount. This work focuses on a deeper understanding of the performance and limitations of deep learning-based classifiers as tools for galaxy merger identification. We train a ResNet18 model on mock Hubble Space Telescope CANDELS images from the IllustrisTNG50 simulation. Our focus is on a more challenging classification of galaxy mergers and nonmergers at higher redshifts $1

Schechter, Aimee [Colorado U.] (ORCID:000000017120↗

Online and Offline Identification of False Data Injection Attacks in Battery Sensors Using a Single Particle Model

The cells in battery energy storage systems are monitored, protected, and controlled by battery management systems whose sensors are susceptible to cyberattacks. False data injection attacks (FDIAs) targeting batteries’ voltage sensors affect cell protection functions and the estimation of critical battery states like the state of charge (SoC). Inaccurate SoC estimation could result in battery overcharging and over discharging, which can have disastrous consequences on grid operations. This paper proposes a three-pronged online and offline method to detect, identify, and classify FDIAs corrupting the voltage sensors of a battery stack. To accurately model the dynamics of the series-connected cells a single particle model is used and to estimate the SoC, the unscented Kalman filter is employed. FDIA detection, identification, and classification was accomplished using a tuned cumulative sum (CUSUM) algorithm, which was compared with a baseline method, the chi-squared error detector. Online simulations and offline batch simulations were performed to determine the effectiveness of the proposed approach. Throughout the batch simulations, the CUSUM algorithm detected attacks, with no false positives, in 99.83% of cases, identified the corrupted sensor in 97% of cases, and determined if the attack was positively or negatively biased in 97% of cases.

25 ENERGY STORAGE↗

Prediction of Cognitive States During Flight Simulation Using Multimodal Psychophysiological Sensing

The Commercial Aviation Safety Team found the majority of recent international commercial aviation accidents attributable to loss of control inflight involved flight crew loss of airplane state awareness (ASA), and distraction was involved in all of them. Research on attention-related human performance limiting states (AHPLS) such as channelized attention, diverted attention, startle/surprise, and confirmation bias, has been recommended in a Safety Enhancement (SE) entitled "Training for Attention Management." To accomplish the detection of such cognitive and psychophysiological states, a broad suite of sensors was implemented to simultaneously measure their physiological markers during a high fidelity flight simulation human subject study. Twenty-four pilot participants were asked to wear the sensors while they performed benchmark tasks and motion-based flight scenarios designed to induce AHPLS. Pattern classification was employed to predict the occurrence of AHPLS during flight simulation also designed to induce those states. Classifier training data were collected during performance of the benchmark tasks. Multimodal classification was performed, using pre-processed electroencephalography, galvanic skin response, electrocardiogram, and respiration signals as input features. A combination of one, some or all modalities were used. Extreme gradient boosting, random forest and two support vector machine classifiers were implemented. The best accuracy for each modality-classifier combination is reported. Results using a select set of features and using the full set of available features are presented. Further, results are presented for training one classifier with the combined features and for training multiple classifiers with features from each modality separately. Using the select set of features and combined training, multistate prediction accuracy averaged 0.64 +/- 0.14 across thirteen participants and was significantly higher than that for the separate training case. These results support the goal of demonstrating simultaneous real-time classification of multiple states using multiple sensing modalities in high fidelity flight simulators. This detection is intended to support and inform training methods under development to mitigate the loss of ASA and thus reduce accidents and incidents.

Harrivel, Angela R.↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗

The DESI Early Data Release white dwarf catalogue

The Early Data Release (EDR) of the Dark Energy Spectroscopic Instrument (DESI) comprises spectroscopy obtained from 2020 December 14 to 2021 June 10. White dwarfs were targeted by DESI both as calibration sources and as science targets and were selected based on Gaia photometry and astrometry. Here, we present the DESI EDR white dwarf catalogue, which includes 2706 spectroscopically confirmed white dwarfs of which approximately 60 per cent have been spectroscopically observed for the first time, as well as 66 white dwarf binary systems. We provide spectral classifications for all white dwarfs, and discuss their distribution within the Gaia Hertzsprung–Russell diagram. We provide atmospheric parameters derived from spectroscopic and photometric fits for white dwarfs with pure hydrogen or helium photospheres, a mixture of those two, and white dwarfs displaying carbon features in their spectra. We also discuss the less abundant systems in the sample, such as those with magnetic fields, and cataclysmic variables. The DESI EDR white dwarf sample is significantly less biased than the sample observed by the Sloan Digital Sky Survey, which is skewed to bluer and therefore hotter white dwarfs, making DESI more complete and suitable for performing statistical studies of white dwarfs.

79 ASTRONOMY AND ASTROPHYSICS↗

SCA Test Report: H4RG-20828 WFIRST

This report summarizes the measured performance for the Sensor Chip Assembly (SCA), which is identified in Table 1. The SCA architecture is a substrate-removed HgCdTe detector with an area of4096x4096 pixels (with a reference pixel area of four pixels deep around all four sides, available to substitute corresponding image pixels) and a pixel pitch of 10 μm. This SCA has been tested at the Goddard Space Flight Center (GSFC/NASA) Detector Characterization Laboratory (DCL). Teledyne (the vendor) classifies its detectors into different grades based on testing performed at its facility. The classification of this SCA and the tested dates are included in Table 1 below.The tests performed on this array are derived from the document WFIRST-PROC-09220_WFIRST-SCA-ATP_-.docx. A summary of the test parameters, requirements, and test results is presented in Table 2. The details of each test are subsequently described in the report. In the Summary (Table 2), SCA results are reported at an operating temperature of 95K, and 1.0 V bias voltage. In the detailed section of each test, the results for 90 K (1.0V bias voltage) and 95 K (0.5V bias voltage) are also reported. For all cases, the frame time used is 2.764 seconds. Reference pixel correction was applied to every raw frame.

Laddawan Miko↗

SCA Test Report: H4RG-20833 WFIRST

This report summarizes the measured performance for the Sensor Chip Assembly (SCA), which is identified in Table 1. The SCA architecture is a substrate-removed HgCdTe detector with an area of 4096x4096 pixels (with a reference pixel area of four pixels deep around all four sides, available to substitute corresponding image pixels) and a pixel pitch of 10 μm. This SCA has been tested at the Goddard Space Flight Center (GSFC/NASA) Detector Characterization Laboratory (DCL). Teledyne (the vendor) classifies its detectors into different grades based on testing performed at its facility. The classification of this SCA and the tested dates are included in Table 1 below.The tests performed on this array are derived from the document WFIRST-PROC-09220_WFIRST-SCA-ATP_-.docx. A summary of the test parameters, requirements, and test results is presented in Table 2. The details of each test are subsequently described in the report. In the Summary (Table 2), SCA results are reported at an operating temperature of 95K, and 1.0 V bias voltage. In the detailed section of each test, the results for 90 K (1.0V bias voltage) and 95 K (0.5V bias voltage) are also reported. For all cases, the frame time used is 2.764 seconds. Reference pixel correction was applied to every raw frame.

Laddawan Miko↗

SCA Test Report: H4RG-21224 WFIRST

This report summarizes the measured performance for the Sensor Chip Assembly (SCA), which is identified in Table 1. The SCA architecture is a substrate-removed HgCdTe detector with an area of 4096x4096 pixels (with a reference pixel area of four pixels deep around all four sides, available to substitute corresponding image pixels) and a pixel pitch of 10 μm.This SCA has been tested at the Goddard Space Flight Center (GSFC/NASA) Detector Characterization Laboratory (DCL). Teledyne (the vendor) classifies its detectors into different grades based on testing performed at its facility. The classification of this SCA and the tested dates are included in Table 1 below.The tests performed on this array are derived from the document WFIRST-PROC-09220_WFIRST-SCA-ATP_-.docx. A summary of the test parameters, requirements, and test results is presented in Table 2. The details of each test are subsequently described in the report. In the Summary (Table 2), SCA results are reported at an operating temperature of 95K, and 1.0 V bias voltage. In the detailed section of each test, the results for 95 K and 0.5V bias voltage is also reported. The data for the result reported in this document was acquired in pixel reset mode. For all cases, the frame time used is 2.830 seconds. Reference pixel correction was applied to every raw frame.

Laddawan Miko↗

SCA Test Report H4RG-21645 Roman Space Telescope

This report summarizes the measured performance for the Sensor Chip Assembly (SCA), which is identified in Table 1. The SCA architecture is a substrate-removed HgCdTe detector with an area of 4096x4096 pixels (with a reference pixel area of four pixels deep around all four sides, available to substitute corresponding image pixels) and a pixel pitch of 10 μm. This SCA has been tested at the Goddard Space Flight Center (GSFC/NASA) Detector Characterization Laboratory (DCL). Teledyne (the vendor) classifies its detectors into different grades based on testing performed at its facility. The classification of this SCA and the tested dates are included in Table 1 below. The tests performed on this array are derived from the document WFIRST-PROC-09220_WFIRST-SCAATP_-.docx. A summary of the test parameters, requirements, and test results is presented in Table 2. The details of each test are subsequently described in the report. In the Summary (Table 2), SCA results are reported at an operating temperature of 95K, and 1.0 V bias voltage. In the detailed section of each test, the results for 95 K and 0.5V bias voltage is also reported. The data for the result reported in this document was acquired in pixel reset mode. For all cases, the frame time used is 2.830 seconds. Reference pixel correction was applied to every raw frame.

Laddawan Miko↗