Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Classification bias”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

WISE-PS1-STRM: neural network source classification and photometric redshifts for WISE×PS1

ABSTRACT We cross-match between the WISE All-Sky and PS1 3π DR2 source catalogues. The resulting catalogue has 354 590 570 objects, significantly fewer than the parent PS1 catalogue, but its combination of optical and infrared colours facilitate both better source classification and photometric redshift estimation. We perform a neural network-based classification of the objects into galaxies, quasars, and stars, then run neural network-based photometric redshift estimation for the galaxies. The star sample purity and quasar sample completeness measures improve substantially, and the resulting photo-z’s are significantly more accurate in terms of statistical scatter and bias than those calculated from PS1 properties alone. The catalogue will be a basis for future large-scale structure studies, and will be made available as a high-level science product via the Mikulski Archive for Space Telescopes.

79 ASTRONOMY AND ASTROPHYSICS↗

Chemical Heterointerface Engineering on Hybrid Electrode Materials for Electrochemical Energy Storage

Abstract The chemical heterointerfaces in hybrid electrode materials play an important role in overcoming the intrinsic drawbacks of individual materials and thus expedite the in‐depth development of electrochemical energy storage. Benefiting from the three enhancement effects of accelerating charge transport, increasing the number of storage sites, and reinforcing structural stability, the chemical heterointerfaces have attracted extensive interest and the electrochemical performances of hybrid electrode materials have been significantly optimized. In this review, recent advances regarding chemical heterointerface engineering in hybrid electrode materials are systematically summarized. Especially, the intrinsic behaviors of chemical heterointerfaces on hybrid electrode materials are refined based on built‐in electric field, van der Waals interaction, lattice mismatch and connection, electron cloud bias and chemical bond, and their combination. The strategies for introducing chemical heterointerfaces are classified into in situ local transformation, in situ growth, cosynthesis, and other strategy. The recent progress about the chemical heterointerfaces engineering specially focusing on metal‐ion batteries, supercapacitors, and Li–S batteries are introduced in detail. Furthermore, the classification and characterization of chemical heterointerfaces are briefly described. Finally, the emerging challenges and perspectives about future directions of chemical heterointerface engineering are proposed.

Li, Wenbin↗

A map of roadmaps for zero and low energy and carbon buildings worldwide

Formulation of targets and establishing which factors in different contexts will achieve these targets are critical to successful decarbonization of the building sector. To contribute to this, we have performed an evidence map of roadmaps for zero and low energy and carbon buildings (ZLECB) worldwide, including a list and classification of documents in an on-line geographical map, a description of gaps, and a narrative review of the knowledge gluts. We have retrieved 1219 scientific documents from Scopus, extracted metadata from 274 documents, and identified 117 roadmaps, policies or plans from 27 countries worldwide. We find that there is a coverage bias towards more developed regions. The identified scientific studies are mostly recommendations to policy makers, different types of case studies, and demonstration projects. The geographical inequalities found in the coverage of the scientific literature are even more extreme in the coverage of the roadmaps. These underexplored world regions represent an area for further investigation and increased research/policy attention. Our review of the more substantial amount of literature and roadmaps for developed regions shows differences in target metrics and enforcement mechanisms but that all regions dedicate some efforts at national and local levels. Roadmaps generally focus more on new and public buildings than existing buildings, despite the fact that the latter are naturally larger in number and total floor area, and perform less energy efficiently. A combination of efficiency, technical upgrades, and renewable generation is generally proposed in the roadmaps, with behavioral measures only reflected in the use of information and communication technologies, and minimal focus being placed on lifecycle perspectives. We conclude that insufficient progress is being made in the implementation of ZLECB. More work is needed to couple the existing climate goals, with realistic, enforceable policies to make the carbon savings a reality for different contexts and stakeholders worldwide.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

A catalogue of cataclysmic variables from 20 yr of the Sloan Digital Sky Survey with new classifications, periods, trends, and oddities

ABSTRACT We present a catalogue of 507 cataclysmic variables (CVs) observed in SDSS I to IV including 70 new classifications collated from multiple archival data sets. This represents the largest sample of CVs with high-quality and homogeneous optical spectroscopy. We have used this sample to derive unbiased space densities and period distributions for the major sub-types of CVs. We also report on some peculiar CVs, period bouncers and also CVs exhibiting large changes in accretion rates. We report 70 new CVs, 59 new periods, 178 unpublished spectra, and 262 new or updated classifications. From the SDSS spectroscopy, we also identified 18 systems incorrectly identified as CVs in the literature. We discuss the observed properties of 13 peculiar CVS, and we identify a small set of eight CVs that defy the standard classification scheme. We use this sample to investigate the distribution of different CV sub-types, and we estimate their individual space densities, as well as that of the entire CV population. The SDSS I to IV sample includes 14 period bounce CVs or candidates. We discuss the variability of CVs across the Hertzsprung–Russell diagram, highlighting selection biases of variability-based CV detection. Finally, we searched for, and found eight tertiary companions to the SDSS CVs. We anticipate that this catalogue and the extensive material included in the Supplementary Data will be useful for a range of observational population studies of CVs.

79 ASTRONOMY AND ASTROPHYSICS↗

Radar Retrieval Evaluation and Investigation of Dendritic Growth Layer Polarimetric Signatures in a Winter Storm

Abstract This study evaluates ice particle size distribution and aspect ratio φ Multi-Radar Multi-Sensor (MRMS) dual-polarization radar retrievals through a direct comparison with two legs of observational aircraft data obtained during a winter storm case from the Investigation of Microphysics and Precipitation for Atlantic Coast-Threatening Snowstorms (IMPACTS) campaign. In situ cloud probes, satellite, and MRMS observations illustrate that the often-observed K dp and Z DR enhancement regions in the dendritic growth layer can either indicate a local number concentration increase of dry ice particles or the presence of ice particles mixed with a significant number of supercooled liquid droplets. Relative to in situ measurements, MRMS retrievals on average underestimated mean volume diameters by 50% and overestimated number concentrations by over 100%. IWC retrievals using Z DR and K dp within the dendritic growth layer were minimally biased relative to in situ calculations where retrievals yielded −2% median relative error for the entire aircraft leg. Incorporating φ retrievals decreased both the magnitude and spread of polarimetric retrievals below the dendritic growth layer. While φ radar retrievals suggest that observed dendritic growth layer particles were nonspherical (0.1 ≤ φ ≤ 0.2), in situ projected aspect ratios, idealized numerical simulations, and habit classifications from cloud probe images suggest that the population mean φ was generally much higher. Coordinated aircraft radar reflectivity with in situ observations suggests that the MRMS systematically underestimated reflectivity and could not resolve local peaks in mean volume diameter sizes. These results highlight the need to consider particle assumptions and radar limitations when performing retrievals. significance statement Developing snow is often detectable using weather radars. Meteorologists combine these radar measurements with mathematical equations to study how snow forms in order to determine how much snow will fall. This study evaluates current methods for estimating the total number and mass, sizes, and shapes of snowflakes from radar using images of individual snowflakes taken during two aircraft legs. Radar estimates of snowflake properties were most consistent with aircraft data inside regions with prominent radar signatures. However, radar estimates of snowflake shapes were not consistent with observed shapes estimated from the snowflake images. Although additional research is needed, these results bolster understanding of snow-growth physics and uncertainties between radar measurements and snow production that can improve future snowfall forecasting.

Meteorology & Atmospheric Sciences↗

Toward Complete Merger Identification at Cosmic Noon with Deep Learning

As we enter the era of large imaging surveys such as Roman, Rubin, and Euclid, a deeper understanding of potential biases and selection effects in optical astronomical catalogs created with the use of ML-based methods is paramount. This work focuses on a deeper understanding of the performance and limitations of deep learning-based classifiers as tools for galaxy merger identification. We train a ConvNeXT-Pico model on mock HST CANDELS images from the IllustrisTNG50 simulation. Our focus is on a more challenging classification of galaxy mergers and non-mergers at higher redshifts 1 < z < 1.5, including minor mergers and lower mass galaxies down to the stellar mass of 108M⊙. We demonstrate, for the first time, that a deep learning model, such as the one developed in this work, can successfully identify even minor and low mass mergers even at these redshifts. Our model achieves overall accuracy, purity, and completeness of over 73%. We show that some galaxy mergers can only be identified from certain observation angles, leading to a potential upper limit in overall accuracy. Using Grad-CAMs and UMAPs, we more deeply examine the performance and observe a visible gradient in the latent space with stellar mass and specific star formation rate, but no visible gradient with merger mass ratio or merger stage.

Schechter, Aimee L. [U. Colorado, Boulder]↗

Toward Complete Merger Identification at Cosmic Noon with Deep Learning

As we enter the era of large imaging surveys such as $\textit{Roman}$, Rubin, and $\textit{Euclid}$, a deeper understanding of potential biases and selection effects in optical astronomical catalogs created with the use of ML-based methods is paramount. This work focuses on a deeper understanding of the performance and limitations of deep learning-based classifiers as tools for galaxy merger identification. We train a ResNet18 model on mock Hubble Space Telescope CANDELS images from the IllustrisTNG50 simulation. Our focus is on a more challenging classification of galaxy mergers and nonmergers at higher redshifts $1

Schechter, Aimee [Colorado U.] (ORCID:000000017120↗

Online and Offline Identification of False Data Injection Attacks in Battery Sensors Using a Single Particle Model

The cells in battery energy storage systems are monitored, protected, and controlled by battery management systems whose sensors are susceptible to cyberattacks. False data injection attacks (FDIAs) targeting batteries’ voltage sensors affect cell protection functions and the estimation of critical battery states like the state of charge (SoC). Inaccurate SoC estimation could result in battery overcharging and over discharging, which can have disastrous consequences on grid operations. This paper proposes a three-pronged online and offline method to detect, identify, and classify FDIAs corrupting the voltage sensors of a battery stack. To accurately model the dynamics of the series-connected cells a single particle model is used and to estimate the SoC, the unscented Kalman filter is employed. FDIA detection, identification, and classification was accomplished using a tuned cumulative sum (CUSUM) algorithm, which was compared with a baseline method, the chi-squared error detector. Online simulations and offline batch simulations were performed to determine the effectiveness of the proposed approach. Throughout the batch simulations, the CUSUM algorithm detected attacks, with no false positives, in 99.83% of cases, identified the corrupted sensor in 97% of cases, and determined if the attack was positively or negatively biased in 97% of cases.

25 ENERGY STORAGE↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗

The DESI Early Data Release white dwarf catalogue

The Early Data Release (EDR) of the Dark Energy Spectroscopic Instrument (DESI) comprises spectroscopy obtained from 2020 December 14 to 2021 June 10. White dwarfs were targeted by DESI both as calibration sources and as science targets and were selected based on Gaia photometry and astrometry. Here, we present the DESI EDR white dwarf catalogue, which includes 2706 spectroscopically confirmed white dwarfs of which approximately 60 per cent have been spectroscopically observed for the first time, as well as 66 white dwarf binary systems. We provide spectral classifications for all white dwarfs, and discuss their distribution within the Gaia Hertzsprung–Russell diagram. We provide atmospheric parameters derived from spectroscopic and photometric fits for white dwarfs with pure hydrogen or helium photospheres, a mixture of those two, and white dwarfs displaying carbon features in their spectra. We also discuss the less abundant systems in the sample, such as those with magnetic fields, and cataclysmic variables. The DESI EDR white dwarf sample is significantly less biased than the sample observed by the Sloan Digital Sky Survey, which is skewed to bluer and therefore hotter white dwarfs, making DESI more complete and suitable for performing statistical studies of white dwarfs.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Artificial Intelligence with Liquid Argon Neutrino Experiments (Technical Report)

The grant allowed two main contributions: 1) The development of a first successful demonstration of the employment of Optimal Transport in liquid argon time projection chamber neutrino detectors. Optimal Transport, used in other contexts and specifically with LHC calorimetric data, was adapted to address a key particle identification challenge in LArTPCs: the separation of pi0 backgrounds from single-electrons produced in charged-current electron neutrino interactions. The work, leveraging ML methods such as k-nearest-neighbor (kNN) and support-vector-machine (SVM), showed an increase in background rejection of a factor of two or more. Work is now ongoing to incorporate this development in physics analyses for LArTPC experiments and more broadly expand the use of OT in LArTPC detectors including DUNE. This work was done in collaboration with the phenomenology group led by Nathaniel Craig at UCSB. 2) The deployment of NuGraph2, a graph neural network developed for LArTPC reconstruction, in the MicroBooNE experiment. NuGraph2 uses novel graph-neural-network methods on the rather simple LArTPC inputs of reconstructed hits, greatly simplifying the workflow compared to the use of waveform or signal-deconvolved wire ROIs. The network performed particle classification and was shown to address many challenging problems in LArTPC imaging including track-shower separation and the identification of protons and charged pions from primary muons. Our group collaborated with Giuseppe Cerati (FNAL scientist) who is one of the core developers of NuGraph2 to integrate this tool in MicroBooNE’s analysis framework. This consisted in tow key contributions: a) Studying performance on real data, which came with several months of iterations because the MC-trained version of the network was found to show significant bias that our group investigated and addressed. b) Integrating the output hit labeling of NuGraph2 into the existing particle tracking and shower reconstruction code. As a result of this work led by our team NuGraph2 is now enabling a suite of new analyses which benefit from enhanced capabilities and thus broader physics reach. The grant supported primarily the salary of UCSB graduate student Chuyue “Michaelia” Fang as well as partial summer salary support for PI Caratelli. Some funds were used for travel by Michaelia to ML related schools and conferences.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A typological framework of non-floodplain wetlands for global collaborative research and sustainable use

Non-floodplain wetlands (NFWs) are important but vulnerable inland freshwater systems that are receiving increased attention and protection worldwide. However, a lack of consistent terminology, incohesive research objectives, and inherent heterogeneity in existing knowledge hinder cross-regional information sharing and global collaboration. To address this challenge and facilitate future management decisions, we synthesized recent work to understand the state of NFW science and explore new opportunities for research and sustainable NFW use globally. Results from our synthesis show that although NFWs have been widely studied across all continents, regional biases exist in the literature. We hypothesize these biases in the literature stem from terminology rather than real geographical bias around existence and functionality. To confirm this observation, we explored a set of geographically representative NFW regions around the world and characteristics of research focal areas. We conclude that there is more that unites NFW research and management efforts than we might otherwise appreciate. Furthermore, opportunities for cross-regional information sharing and global collaboration exist, but a unified terminology will be needed, as will a focus on wetland functionality. Based on these findings, we discuss four pathways that aid in better collaboration, including improved cohesion in classification and terminology, and unified approaches to modeling and simulation. In turn, legislative objectives must be informed by science to drive conservation and management priorities. Finally, an educational pathway serves to integrate the measures and to promote new technologies that aid in our collective understanding of NFWs. Our resulting framework from NFW synthesis serves to encourage interdisciplinary collaboration and sustainable use and conservation of wetland systems globally.

54 ENVIRONMENTAL SCIENCES↗

Robust Measurement of Stellar Streams around the Milky Way: Correcting Spatially Variable Observational Selection Effects in Optical Imaging Surveys

Observations of density variations in stellar streams are a promising probe of low-mass dark matter substructure in the Milky Way. However, survey systematics such as variations in seeing and sky brightness can also induce artificial fluctuations in the observed densities of known stellar streams. These variations arise because survey conditions affect both object detection and star–galaxy misclassification rates. To mitigate these effects, we use Balrog synthetic source injections in the Dark Energy Survey (DES) Y3 data to calculate detection rate variations and classification rates as functions of survey properties. We show that these rates are nearly separable with respect to survey properties and can be estimated with sufficient statistics from the synthetic catalogs. Applying these corrections reduces the standard deviation of relative detection rates across the DES footprint by a factor of 5, and our corrections significantly change the inferred linear density of the Phoenix stream when including faint objects. Additionally, for artificial streams with DES-like survey properties we are able to recover density power spectra with reduced bias. We also find that uncorrected power-spectrum results for Legacy Survey of Space and Time (LSST)-like data can be around 5 times more biased, highlighting the need for such corrections in future ground-based surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Uranium Oxide Synthetic Pathway Discernment through Unsupervised Morphological Analysis

We present a novel unsupervised machine learning method for quantitative representation of scanning electron micrographs and its applications and performance for nuclear forensic analysis of uranium ore concentrates. The method uses a vector quantizing variational autoencoder followed by a histogram operation to encode a micrograph into a single dimensional representation, called the latent vector. The method requires no extant labeling of the data and can be applied over large datasets of micrographs with minimal human interaction. The representations generated are broadly descriptive of each micrograph and the microstructure of the material imaged. In the case of uranium ore concentrate analysis, the representations were amenable to processing reagent and ore concentrate species classification with accuracy of 81:8%, which is competitive with state-of-the-art supervised networks. The representations were also used to classify previously unseen processing routes, were able to classify imaging parameters such as magnification (to 76:0% accuracy), were able to classify fine grained process parameters such as calcining temperature (to 74:4% accuracy), and their informatic properties indicate that they are generally descriptive of the image represented. This method can be applied across microstructure analysis fields to perform quantitative analysis without the need for labor intensive and possibly biased human analysis.

Scanning Electron Microscopy, Vector Quantizing Va↗

A Template-based Approach to the Photometric Classification of SN 1991bg-like Supernovae in the SDSS-II Supernova Survey

The use of SNe Ia to measure cosmological parameters has grown significantly over the past two decades. However, there exists a significant diversity in the SN Ia population that is not well understood. Overluminous SN 1991T-like and subluminous SN 1991bg-like objects are two characteristic examples of peculiar SNe. The identification and classification of such objects is an important step in studying what makes them unique from the remaining SN population. With the upcoming Vera C. Rubin Observatory promising on the order of a million new SNe over a 10 year survey, spectroscopic classifications will be possible for only a small subset of observed targets. As such, photometric classification has become an increasingly important concern in preparing for the next generation of astronomical surveys. Using observations from the Sloan Digital Sky Survey II (SDSS-II) SN Survey, we apply here an empirically based classification technique targeted at the identification of SN 1991bg-like SNe in photometric data sets. By performing dedicated fits to photometric data in the rest-frame redder and bluer bandpasses, we classify 16 previously unidentified 91bg-like SNe. Using SDSS-II host galaxy measurements, we find that these SNe are preferentially found in host galaxies with an older average stellar age than the hosts of normal SNe Ia. We also find that these SNe are found at a further physical distance from the center of their host galaxies. We find no statistically significant bias in host galaxy mass or specific star formation rate for these targets.

Astronomy & Astrophysics↗

Event Classifications on DNE2 Main Experiment Data using a Convolutional Neural Network Ensemble

The Dynamic Networks (DN) Experiment for FY24 (DNE2) is an experiment within DN with the goal of quantitatively evaluating the effectiveness of solutions developed so far by various researchers under the Low Yield Nuclear Monitoring (LYNM) program using a shared set of metrics and datasets. A key component of this experiment is the mimicking of a signature processing pipeline, and comparing currently accepted and standard-use processing methods to more state-of-the-art processes developed under DN. In this work, we focus specifically on the Event Characterization (EC) Focus Area (FA) of the pipeline, where a seismic event’s magnitude, yield and class are identified. We use Deep Learning (DL) to classify the type of events being processed as either earthquakes (EQs) or explosions (EXs) for three iterations of experiment datasets. The model is noticeably more confident and accurate in classifying explosions than earthquakes, reflecting a known shortcoming of the model, that being of a bias towards predicting explosions over earthquakes in the west coast due to training data biases.

97 MATHEMATICS AND COMPUTING↗

Group-Invariant Quantum Machine Learning

Quantum machine learning (QML) models are aimed at learning from data encoded in quantum states. Recently, it has been shown that models with little to no inductive biases (i.e., with no assumptions about the problem embedded in the model) are likely to have trainability and generalization issues, especially for large problem sizes. As such, it is fundamental to develop schemes that encode as much information as available about the problem at hand. In this work we present a simple, yet powerful, framework where the underlying invariances in the data are used to build QML models that, by construction, respect those symmetries. These so-called group-invariant models produce outputs that remain invariant under the action of any element of the symmetry group $\mathfrak{G}$ associated with the dataset. We present theoretical results underpinning the design of $\mathfrak{G}$-invariant models, and exemplify their application through several paradigmatic QML classification tasks, including cases when $\mathfrak{G}$ is a continuous Lie group and also when it is a discrete symmetry group. Notably, our framework allows us to recover, in an elegant way, several well-known algorithms for the literature, as well as to discover new ones. Taken together, we expect that our results will help pave the way towards a more geometric and group-theoretic approach to QML model design.

97 MATHEMATICS AND COMPUTING↗