Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Field validation of data-driven BSDF and peak extraction models for light-scattering fabric shades

Shading and daylighting systems affect cooling, heating, and lighting energy use by modulating solar radiation through the building façade. Characterizing shading systems holistically and accurately helps designers and engineers evaluate shading systems to achieve energy and non-energy performance goals. These complex fenestration systems can be modeled using Bidirectional Scattering Distribution Functions (BSDF), which map incident radiation to hemispherical distributions of outgoing radiation. Data-driven, tabulated BSDFs are derived from interpolated goniophotometer measured data, then sampled during the raytracing calculation. A peak extraction (PE) algorithm was developed to circumvent limits in BSDF angular resolution, where the specular peak is extracted during simulation by evaluating the BSDF in the through direction and surrounding region. The objective of this study was to validate this measurement and modeling workflow using field monitored data from a full scale testbed with eleven installed fabrics of different weaves, openness factors, and colors and assess the accuracy of the workflow under different adaptation and contrast conditions. Test conditions were limited to clear sky conditions with the sun in the field of view. Results showed that, for tensor tree datasets, vertical illuminance, solar luminance (2.5° apex), and daylight glare probability (DGP) were predicted to within a mean bias error (MBE) error of -456 lx (-12.3%), -3.46e5 (-38.4%), and -0.042 (-7.8%) when full PE occurred. With a binary classification of glare/ no glare, DGP was predicted accurately with a true positive rate of 0.98 and true negative rate of 1.0 using tensor tree data and less accurately with Klems BSDF data, particularly for cases of no glare. The workflow may be of insufficient accuracy to distinguish borderline performance between fabrics using the four-point glare scale, particularly under low adaptation, high contrast daylit conditions. Errors were due to reductions in peak shape and intensity across the BSDF interpolation and data reduction workflow. Future work is needed to better preserve measurement fidelity during interpolation and sampling, which in turn will improve PE performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Improved Estimates of Clear Sky Longwave Flux and Application to the Tropical Greenhouse Effect

The first objective of this investigation is to eliminate the clear-sky offset introduced by the scene-identification procedures developed for the Earth Radiation Budget Experiment (ERBE). Estimates of this systematic bias range from 10 to as high as 30 W/sq m. The initial version of the ScaRaB data is being processed with the original ERBE algorithm. Since the ERBE procedure for scene identification is based upon zonal flux averages, clear scenes with longwave emission well below the zonal mean value are mistakenly classified as cloudy. The erroneous classification is more frequent in regions with deep convection and enhanced mid- and upper-tropospheric humidity. We will develop scene identification parameters with zonal and/or time dependence to reduce or eliminate the bias in the clear- sky data. The modified scene identification procedure could be used for the ScaRaB-specific version of the Earth-radiation products. The second objective is to investigate changes in the clear-sky Outgoing Longwave Radiation (OLR) associated with decadal variations in the tropical and subtropical climate. There is considerable evidence for a shift in the climate state starting in approximately 1977. The shift is accompanied by higher SSTs in the equatorial Pacific, increased tropical convection, and higher values of atmospheric humidity. Other evidence indicates that the humidity in the tropical troposphere has been steadily increasing over the last 30 years. It is not known whether the atmospheric greenhouse effect has increased during this period in response to these changes in SST and precipitable water. We will investigate the decadal-scale fluctuations in the greenhouse effect using Nimbus-7, ERBE, and ScaRaB measurements spaning 1979 to the present. The data from the different satellites will be intercalibrated by comparison with model calculations based upon ship radiosonde observations. The fluxes calculated from the radiation model will also be used for validation of the ScaRaB fluxes.

Collins, W. D.↗

CSAPR2 Optimized Convective Cell Tracking Data during TRACER

One of the challenges in analyzing convective cell properties is to observe the quick evolution of individual convective cells. While the operational radar data provide a volumetric data set to analyze radar observables of convective precipitation clouds, previous studies also suggested the quick evolution of cell life cycle that might not be captured by conventional radar volume scan strategies that take ~5-7 minutes. Aiming at enhancing our understanding of the links between convective cloud kinematic and microphysical processes as well as life cycles, the Tracking Aerosol Convection Interactions ExpeRiment (TRACER; Jensen et al. 2019) was conducted at Houston, Texas, in 2022. The TRACER campaign deployed the 2nd generation C-band Scanning ARM Precipitation Radar (CSAPR2), which performed frequent updates of range height indicator (RHI) and sector plan position indicator (PPI) scans to track individual convective cells every < 2 minutes, guided by a new cell tracking framework, Multisensor Agile Adaptive Sampling (MAAS; Kollias et al. 2020). This allows for capturing fast-evolving radar observables. We provide the processed CSAPR2 cell tracking data in CfRadial format collected during the TRACER field campaign from June to September 2022. The data files include processed radar variables: noise-masked reflectivity and differential reflectivity corrected for rain attenuation and systematic biases, noise-masked dealiased radial velocity, specific differential phase, locations of target cells (latitude, longitude, radar range), and radar-echo classification. Figure 1 provides an example of a 3D image of CSAPR2 reflectivity from the lowest PPI scan and an RHI scan after data processing.

54 ENVIRONMENTAL SCIENCES↗

CSAPR2 cell-tracking data collected during TRACER

One of the challenges of analyzing convective cell properties is quick evolution of the individual convective cells. While the operational radar data provide great a data set to analyze the evolution of radar observables of convective precipitation clouds statistically, previous studies also suggested that, because of the quick evolution of cell life cycle, conventional radar volume scan strategies taking ~5-7 minutes might not capture the detailed evolution. The TRACER campaign deployed CSAPR2, which performed frequent update of RHI and sector PPI scans to track convective cells every < 2 minutes guided by a new cell-tracking framework, Multisensor Agile Adaptive Sampling (MAAS; Kollias et al. 2020). This allows for capturing fast-evolving radar observables. The submitted data files are CSAPR2 data in CfRadial format collected during the TRACER field campaign from June to September 2020. The data files include processed radar variables including: noise-masked reflectivity and differential reflectivity corrected for rain attenuation and systematic biases, noise-masked dealiased radial velocity, specific differential phase, locations of target cells (latitude, longitude, radar range), and radar-echo classification.

54 ENVIRONMENTAL SCIENCES↗

Geography of the asteroid belt

The CSM classification serves as the starting point on the geography of the asteroid belt. Raw data on asteroid types are corrected for observational biases (against dark objects, for instance) to derive the distribution of types throughout the belt. Recent work on family members indicates that dynamical families have a true physical relationship, presumably indicating common origin in the breakup of a parent asteroid.

Zellner, B. H.↗

Subfield crop yields and temporal stability in thousands of US Midwest fields

Understanding subfield crop yields and temporal stability is critical to better manage crops. Several algorithms have proposed to study within-field temporal variability but they were mostly limited to few fields. In this study, a large dataset composed of 5520 yield maps from 768 fields provided by farmers was used to investigate the influence of subfield yield distribution skewness on temporal variability. The data are used to test two intuitive algorithms for mapping stability: one based on standard deviation and the second based on pixel ranking and percentiles. The analysis of yield monitor data indicates that yield distribution is asymmetric, and it tends to be negatively skewed (p < 0.05) for all of the four crops analyzed, meaning that low yielding areas are lower in frequency but cover a larger range of low values. The mean yield difference between the pixels classified as high-and-stable and the pixels classified as low-and-stable was 1.04 Mg ha –1 for maize, 0.39 Mg ha –1 for cotton, 0.34 Mg ha –1 for soybean, and 0.59 Mg ha —1 for wheat. The yield of the unstable zones was similar to the pixels classified as low-and-stable by the standard deviation algorithm, whereas the two-way outlier algorithm did not exhibit this bias. Furthermore, the increase in the number years of yield maps available induced a modest but significant increase in the certainty of stability classifications, and the proportion of unstable pixels increased with the precipitation heterogeneity between the years comprising the yield maps.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

Lorentz group equivariant autoencoders

Abstract There has been significant work recently in developing machine learning (ML) models in high energy physics (HEP) for tasks such as classification, simulation, and anomaly detection. Often these models are adapted from those designed for datasets in computer vision or natural language processing, which lack inductive biases suited to HEP data, such as equivariance to its inherent symmetries. Such biases have been shown to make models more performant and interpretable, and reduce the amount of training data needed. To that end, we develop the Lorentz group autoencoder (LGAE), an autoencoder model equivariant with respect to the proper, orthochronous Lorentz group $$\textrm{SO}^+(3,1)$$ SO + ( 3 , 1 ) , with a latent space living in the representations of the group. We present our architecture and several experimental results on jets at the LHC and find it outperforms graph and convolutional neural network baseline models on several compression, reconstruction, and anomaly detection metrics. We also demonstrate the advantage of such an equivariant model in analyzing the latent space of the autoencoder, which can improve the explainability of potential anomalies discovered by such ML models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

AgRISTARS: Foreign commodity production forecasting. The 1980 US corn and soybeans exploratory experiment

The U.S. corn and soybeans exploratory experiment is described which consisted of evaluations of two technology components of a production forecasting system: classification procedures (crop labeling and proportion estimation at the level of a sampling unit) and sampling and aggregation procedures. The results from the labeling evaluations indicate that the corn and soybeans labeling procedure works very well in the U.S. corn belt with full season (after tasseling) LANDSAT data. The procedure should be readily adaptable to corn and soybeans labeling required for subsequent exploratory experiments or pilot tests. The machine classification procedures evaluated in this experiment were not effective in improving the proportion estimates. The corn proportions produced by the machine procedures had a large bias when the bias correction was not performed. This bias was caused by the manner in which the machine procedures handled spectrally impure pixels. The simulation test indicated that the weighted aggregation procedure performed quite well. Although further work can be done to improve both the simulation tests and the aggregation procedure, the results of this test show that the procedure should serve as a useful baseline procedure in future exploratory experiments and pilot tests.

Malin, J. T.↗

Bridging the Gap Between Astronomical Datasets: From Proof-of-Concept to AI Model Deployment with Domain Adaptation

Artificial Intelligence is transforming astrophysics, from studying stars and galaxies to analyzing cosmic large-scale structures. However, a critical challenge arises when AI models trained on simulations or past observational data are applied to new observation— leading to domain shifts, reduced robustness, and increased uncertainty of model predictions. This talk will explore these issues, highlighting examples such as galaxy morphology classification and cosmological parameter inference, where AI struggles to adapt across different datasets. We will discuss domain adaptation as a strategy to improve model generalization and mitigate biases—essential for making AI-driven discoveries reliable. Notably, these challenges extend beyond astrophysics, affecting AI applications across physics and other scientific domains. Addressing them is essential for maximizing AI’s impact in advancing scientific research.

Ćiprijanović, Aleksandra [Fermilab]↗

On the distribution of pitch angles in external galactic spirals NGC 1232 and NGC 5457

A numerical method, originally developed to analyze the morphology of global and local structure in prototype galaxies, is modified for analyzing observed disk-shape galaxies. Two digitized spiral galaxies NGC 1232 and NGC 5457 with varying degrees of contrast between arm and interarm regions are analyzed. A synergism of partitioning methods and a geometric mean least-squares regression algorithm serves to isolate local arm segments, spurs, feathers, and secondary features and to measure their pitch angles and lengths. The global arms are actually highly disjointed, with arm segments frequently revealing pitch angles between 30 and 50 deg, certainly greater than those of the parent arms. Prominent spurs tend to exhibit a much greater pitch angle. The automated mathematical algorithm is shown to have negligible numerical biasing and could be applied to any number of spiral galaxies manifesting flocculent structure, either prototype or observed, and could possibly be used as a tool for classification of multiple-armed-type galaxies.

Russell, William S.↗

Evaluating cosmological biases using photometric redshifts for Type Ia Supernova cosmology with the Dark Energy Survey Supernova Program

Cosmological analyses with Type Ia Supernovae (SNe Ia) have traditionally been reliant on spectroscopy for both classifying the type of supernova and obtaining reliable redshifts to measure the distance–redshift relation. While obtaining a host-galaxy spectroscopic redshift for most SNe is feasible for small-area transient surveys, it will be too resource intensive for upcoming large-area surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time, which will observe on the order of millions of SNe. Here, we use data from the Dark Energy Survey (DES) to address this problem with photometric redshifts (photo-z) inferred directly from the SN light curve in combination with Gaussian and full p(z) priors from host-galaxy photo-z estimates. Using the DES 5-yr photometrically classified SN sample, we consider several photo-z algorithms as host-galaxy photo-z priors, including the Self-Organizing Map redshifts (SOMPZ), Bayesian Photometric Redshifts (BPZ), and Directional-Neighbourhood Fitting (DNF) redshift estimates employed in the DES 3 × 2 point analyses. With detailed catalogue-level simulations of the DES 5-yr sample, we find that the simulated w can be recovered within ±0.02 when using SN+SOMPZ or DNF prior photo-z, smaller than the average statistical uncertainty for these samples of 0.03. With data, we obtain biases in w consistent with simulations within ~1σ for three of the five photo-z variants. We further evaluate how photo-z systematics interplay with photometric classification and find classification introduces a subdominant systematic component. This work lays the foundation for next-generation fully photometric SNe Ia cosmological analyses.

(cosmology:) dark energy↗

Gamma-Ray Burst Class Properties

Guided by the Supervised pattern recognition algorithm C4.5, we examine the three gamma-ray burst classes identified by Mukherjee et al. C4.5 provides strong statistical support for this classification. However, with C4.5 and our knowledge of the BATSE instrument, we demonstrate that Class 3 (intermediate fluence, intermediate duration, soft) does not have to be a distinct source population: statistical/systematic errors in measuring burst attributes combined with the well-known hardness/intensity correlation can cause low peak flux Class I (high fluence, long, intermediate hardness) bursts to take on Class 3 characteristics naturally. Based on our hypothesis that the third class is not a distinct one, we provide rules so that future events can be placed in either Class I or Class 2 (low fluence, short, hard). Using classified bursts from the BATSE 4B Catalog, we plot log(N>P) vs. log(P) curves and study spectral features of each class. We find that the two classes are relatively distinct on the basis of spectral parameters, alpha, Beta, and E(sub peak) alone. Although this does not indicate a better basis for classification, it does suggest that different physical conditions exist for Class I and Class 2 bursts. In the process of studying burst class characteristics, we identify a new bias that affects measurement of burst fluences and durations. Using a simple model of how burst duration can be underestimated, we generally characterize how this fluence duration bias affects BATSE measurements, and demonstrate the type of effect it can have on the BATSE fluence vs. peak flux diagram.

Hakkila, Jon↗

Gamma-Ray Burst Class Properties

Guided by the supervised pattern recognition algorithm C4.5 developed by Quinlan in 1986, we examine the three gamma-ray burst classes identified by Mukherjee et al. in 1998. C4.5 provides strong statistical support for this classification. However, with C4.5 and our knowledge of the Burst and Transient Source Experiment (BATSE) instrument, we demonstrate that class 3 (intermediate fluence, intermediate duration, soft) does not have to be a distinct source population: statistical/systematic errors in measuring burst attributes combined with the well-known hardness/intensity correlation can cause low peak flux class 1 (high fluence, long, intermediate hardness) bursts to take on class 3 characteristics naturally. Based on our hypothesis that the third class is not a distinct one, we provide rules so that future events can be placed in either class 1 or class 2 (low fluence, short, hard). We find that the two classes are relatively distinct on the basis of Band's work in 1993 on spectral parameters alpha, beta, and E (sub peak) alone. Although this does not indicate a better basis for classification, it does suggest that different physical conditions exist for class 1 and class 2 bursts. In the process of studying burst class characteristics, we identify a new bias affecting burst fluence and duration measurements. Using a simple model of how burst duration can be underestimated, we show how this fluence duration bias can affect BATSE measurements and demonstrate the type of effect it can have on the BATSE fluence versus peak flux diagram.

Hakkila, Jon↗

Enhanced read resolution in reconfigurable memristive synapses for Spiking Neural Networks

Abstract The synapse is a key element circuit in any memristor-based neuromorphic computing system. A memristor is a two-terminal analog memory device. Memristive synapses suffer from various challenges including high voltage, SET or RESET failure, and READ margin issues that can degrade the distinguishability of stored weights. Enhancing READ resolution is very important to improving the reliability of memristive synapses. Usually, the READ resolution is very small for a memristive synapse with a 4-bit data precision. This work considers a step-by-step analysis to enhance the READ current resolution or the read current difference between two resistance levels for a current-controlled memristor-based synapse. An empirical model is used to characterize the $${\hbox {HfO}}_{2}$$ HfO 2 based memristive device. $$1\textrm{st}$$ 1 st and $$2\textrm{nd}$$ 2 nd stage device of our proposed synapse design can be scaled to enhance the READ current margin up to $$\sim$$ ∼ 4.3 $$\times$$ × and $$\sim$$ ∼ 21%, respectively. Moreover, READ current resolution can be enhanced with run-time adaptation techniques such as READ voltage scaling and body biasing. The READ voltage scaling and body biasing can improve the READ current resolution by about 46% and 15%, respectively. TENNLab’s neuromorphic computing framework is leveraged to evaluate the effect of READ current resolution on classification, control, and reservoir computing applications. Higher READ current resolution shows better accuracy than lower resolution even when facing different levels of read noise.

97 MATHEMATICS AND COMPUTING↗

Measuring Cosmological Parameters with Type Ia Supernovae in redMaGiC Galaxies

Abstract Current and future cosmological analyses with Type Ia supernovae (SNe Ia) face three critical challenges: (i) measuring the redshifts from the SNe or their host galaxies; (ii) classifying the SNe without spectra; and (iii) accounting for correlations between the properties of SNe Ia and their host galaxies. We present here a novel approach that addresses each of these challenges. In the context of the Dark Energy Survey (DES), we analyze an SN Ia sample with host galaxies in the redMaGiC galaxy catalog, a selection of luminous red galaxies. redMaGiC photo- z estimates are expected to be accurate to σ Δ z /(1+ z ) ∼ 0.02. The DES-5YR photometrically classified SN Ia sample contains approximately 1600 SNe, and 125 of these SNe are in redMaGiC galaxies. We demonstrate that redMaGiC galaxies almost exclusively host SNe Ia, reducing concerns relating to classification uncertainties. With this subsample, we find similar Hubble scatter (to within ∼0.01 mag) using photometric redshifts in place of spectroscopic redshifts. With detailed simulations, we show that the bias due to using redMaGiC photo- z s on the measurement of the dark energy equation of state w is up to Δ w ∼ 0.01–0.02. With real data, we measure a difference in w when using the redMaGiC photo- z s versus the spec- z s of Δ w = 0.005. Finally, we discuss how SNe in redMaGiC galaxies appear to comprise a more standardizable population, due to a weaker relation between color and luminosity ( β ) compared to the DES-3YR population by ∼5 σ . These results establish the feasibility of performing redMaGiC SN cosmology with photometric survey data in the absence of spectroscopic data.

79 ASTRONOMY AND ASTROPHYSICS↗

Strong dependence of Type Ia supernova standardization on the local specific star formation rate

As part of an on-going effort to identify, understand and correct for astrophysics biases in the standardization of Type Ia supernovae (SN Ia) for cosmology, we have statistically classified a large sample of nearby SNe Ia into those that are located in predominantly younger or older environments. This classification is based on the specific star formation rate measured within a projected distance of 1 kpc from each SN location (LsSFR). This is an important refinement compared to using the local star formation rate directly, as it provides a normalization for relative numbers of available SN progenitors and is more robust against extinction by dust. We find that the SNe Ia in predominantly younger environments are Δ Y = 0.163 ± 0.029 mag (5.7 σ ) fainter than those in predominantly older environments after conventional light-curve standardization. This is the strongest standardized SN Ia brightness systematic connected to the host-galaxy environment measured to date. The well-established step in standardized brightnesses between SNe Ia in hosts with lower or higher total stellar masses is smaller, at Δ M = 0.119 ± 0.032 mag (4.5 σ ), for the same set of SNe Ia. When fit simultaneously, the environment-age offset remains very significant, with Δ Y = 0.129 ± 0.032 mag (4.0 σ ), while the global stellar mass step is reduced to Δ M = 0.064 ± 0.029 mag (2.2 σ ). Thus, approximately 70% of the variance from the stellar mass step is due to an underlying dependence on environment-based progenitor age. Also, we verify that using the local star formation rate alone is not as powerful as LsSFR at sorting SNe Ia into brighter and fainter subsets. Standardization that only uses the SNe Ia in younger environments reduces the total dispersion from 0.142 ± 0.008 mag to 0.120 ± 0.010 mag. Overall, we show that as environment-ages evolve with redshift, a strong bias, especially on the measurement of the derivative of the dark energy equation of state, can develop. Fortunately, data that measure and correct for this effect using our local specific star formation rate indicator, are likely to be available for many next-generation SN Ia cosmology experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

Psychophysiological Sensing and State Classification for Attention Management in Commercial Aviation

Attention-related human performance limiting states (AHPLS) can cause pilots to lose airplane state awareness (ASA), and their detection is important to improving commercial aviation safety. The Commercial Aviation Safety Team found that the majority of recent international commercial aviation accidents attributable to loss of control inflight involved flight crew loss of airplane state awareness, and that distraction of various forms was involved in all of them. Research on AHPLS, including channelized attention, diverted attention, startle / surprise, and confirmation bias, has been recommended in a Safety Enhancement (SE) entitled "Training for Attention Management." To accomplish the detection of such cognitive and psychophysiological states, a broad suite of sensors has been implemented to simultaneously measure their physiological markers during high fidelity flight simulation human subject studies. Pilot participants were asked to perform benchmark tasks and experimental flight scenarios designed to induce AHPLS. Pattern classification was employed to distinguish the AHPLS induced by the benchmark tasks. Unimodal classification using pre-processed electroencephalography (EEG) signals as input features to extreme gradient boosting, random forest and deep neural network multiclass classifiers was implemented. Multi-modal classification using galvanic skin response (GSR) in addition to the same EEG signals and using the same types of classifiers produced increased accuracy with respect to the unimodal case (90 percent vs. 86 percent), although only via the deep neural network classifier. These initial results are a first step toward the goal of demonstrating simultaneous real time classification of multiple states using multiple sensing modalities in high-fidelity flight simulators. This detection is intended to support and inform training methods under development to mitigate the loss of ASA and thus reduce accidents and incidents.

Harrivel, Angela R.↗