Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Galaxy morphological classification catalogue of the Dark Energy Survey Year 3 data with convolutional neural networks

ABSTRACT We present in this paper one of the largest galaxy morphological classification catalogues to date, including over 20 million galaxies, using the Dark Energy Survey (DES) Year 3 data based on convolutional neural networks (CNNs). Monochromatic i-band DES images with linear, logarithmic, and gradient scales, matched with debiased visual classifications from the Galaxy Zoo 1 (GZ1) catalogue, are used to train our CNN models. With a training set including bright galaxies (16 ≤ i < 18) at low redshift (z < 0.25), we furthermore investigate the limit of the accuracy of our predictions applied to galaxies at fainter magnitude and at higher redshifts. Our final catalogue covers magnitudes 16 ≤ i < 21, and redshifts z < 1.0, and provides predicted probabilities to two galaxy types – ellipticals and spirals (disc galaxies). Our CNN classifications reveal an accuracy of over 99 per cent for bright galaxies when comparing with the GZ1 classifications (i < 18). For fainter galaxies, the visual classification carried out by three of the co-authors shows that the CNN classifier correctly categorizes discy galaxies with rounder and blurred features, which humans often incorrectly visually classify as ellipticals. As a part of the validation, we carry out one of the largest examinations of non-parametric methods, including ∼100 ,000 galaxies with the same coverage of magnitude and redshift as the training set from our catalogue. We find that the Gini coefficient is the best single parameter discriminator between ellipticals and spirals for this data set.

79 ASTRONOMY AND ASTROPHYSICS↗

Use of Convolutional Neural Network Image Classification and High-Speed Ion Probe Data Toward Real-Time Detonation Characterization in a Water-Cooled Rotating Detonation Engine

As rotating detonation engines (RDEs) progress in maturity, the importance of monitoring advancements toward development of active control becomes more critical. Experimental RDE data processing at time scales which satisfy real-time diagnostics will likely require the use of machine learning. This study aims to develop and deploy a novel real-time monitoring technique capable of determining detonation wave number, direction, frequency, and individual wave speeds throughout experimental RDE operational windows. To do so, the diagnostic integrates image classification by a convolutional neural network (CNN) and ionization current signal analysis. Wave mode identification through single-image CNN classification bypasses the need to evaluate sequential images and offers instantaneous identification of the wave mode present in the RDE annulus. Here, real-time processing speeds are achieved due to low data volumes required by the methodology, namely one short-exposure image and a short window of sensor data to generate each diagnostic output. The diagnostic acquires live data using a modified experimental setup alongside Pylon and PyDAQmx libraries within a python data acquisition environment. Lab-deployed diagnostic results are presented across varying wave modes, operating conditions, and data quality, currently executed at 3–4 Hz with a variety of iteration speed optimization options to be considered as future work. These speeds exceed that of conventional techniques and offer a proven structure for real-time RDE monitoring. The demonstrated ability to analyze detonation wave presence and behavior during RDE operation will certainly play a vital role in the development of RDE active control, necessary for RDE technology maturation toward industrial integration.

42 ENGINEERING↗

Repository scale classification and decomposition of tandem mass spectral data

Various studies have shown associations between molecular features and phenotypes of biological samples. These studies, however, focus on a single phenotype per study and are not applicable to repository scale metabolomics data. Here we report MetSummarizer, a method for predicting (i) the biological phenotypes of environmental and host-oriented samples, and (ii) the raw ingredient composition of complex mixtures. We show that the aggregation of various metabolomic datasets can improve the accuracy of predictions. Since these datasets have been collected using different standards at various laboratories, in order to get unbiased results it is crucial to detect and discard standard-specific features during the classification step. We further report high accuracy in prediction of the raw ingredient composition of complex foods from the Global Foodomics Project.

59 BASIC BIOLOGICAL SCIENCES↗

EE-SMOTE: An oversampling method in conjunction with information entropy for imbalanced learning

Imbalanced learning attracts great attention in various research fields. Existing literature-reported methodologies in imbalanced learning have shown drawbacks including over-generation or noisy/wrong samples generations. This paper presents EE-SMOTE, an oversampling technique based on information entropy, to support the imbalance classifications. Specifically, we propose a metric, Eigen-Entropy (EE), to identify homogenous samples from minority classes for oversampling technique, specifically, SMOTE to reach data balances for classification. Experiments on public dataset and real-world datasets demonstrate the efficacy and effectiveness of the proposed EE-SMOTE in imbalanced learning.

Huang, Jiajing↗

Grid Edge Waveform Analytics Framework for Event Detection and Classification

This paper provides a grid edge waveform analytics framework for power system event detection and classification in the local as well as in the wide area. This framework overviews data excellence for event detection and classification. The data excellence describes the data acquisition process and requirements, data processing, data quality, and data integrity. Power system event detection in the local area based on different features such as energy-based, cyclostationary approach, template matching, and wavelet transform are also discussed. Furthermore, local area event detection and classification using approaches such as statistical, signal processing, artificial intelligence, and hybrid are also discussed. Moreover, an overview of wide-area event detection and classification along with several other aspects such as wide-area events, wide-area event detection approaches, event location and system performance, event pattern recognition, inter-area oscillation, and wide-area frequency response under variable deployment of inverter-based resources are also provided. The proposed framework is the first step toward the goal of developing appropriate tools and methodologies to detect and classify local as well as wide-area events using waveform analytics. The appropriate event detection and classification framework development is especially important now as more and more grid edge devices with communication capabilities are being deployed in the modern power grid than ever before.

Bhusal, Narayan↗

Data driven discovery and quantification of hyperspectral leaf reflectance phenotypes across a maize diversity panel

Abstract Estimates of plant traits derived from hyperspectral reflectance data have the potential to efficiently substitute for traits, which are time or labor intensive to manually score. Typical workflows for estimating plant traits from hyperspectral reflectance data employ supervised classification models that can require substantial ground truth datasets for training. We explore the potential of an unsupervised approach, autoencoders, to extract meaningful traits from plant hyperspectral reflectance data using measurements of the reflectance of 2151 individual wavelengths of light from the leaves of maize ( Zea mays ) plants harvested from 1658 field plots in a replicated field trial. A subset of autoencoder‐derived variables exhibited significant repeatability, indicating that a substantial proportion of the total variance in these variables was explained by difference between maize genotypes, while other autoencoder variables appear to capture variation resulting from changes in leaf reflectance between different batches of data collection. Several of the repeatable latent variables were significantly correlated with other traits scored from the same maize field experiment, including one autoencoder‐derived latent variable (LV8) that predicted plant chlorophyll content modestly better than a supervised model trained on the same data. In at least one case, genome‐wide association study hits for variation in autoencoder‐derived variables were proximal to genes with known or plausible links to leaf phenotypes expected to alter hyperspectral reflectance. In aggregate, these results suggest that an unsupervised, autoencoder‐based approach can identify meaningful and genetically controlled variation in high‐dimensional, high‐throughput phenotyping data and link identified variables back to known plant traits of interest.

Tross, Michael C.↗

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United ↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

A Taxonomic Classification Approach for Global Spatio-temporal Data

The World Bank, World Health Organization, and other major vendors collectively provide thousands of global time series datasets that focus on issues of the environment, public health, economics, violence, education, and national security. Sorting these data into meaningful information requires the use of data mining techniques to cluster trends into an orderly and manageable number of cases. The World SpatioTemporal Analytics and Mapping (WSTAMP) project database (wstamp.ornl.gov) was developed to spatiotemporally harmonize global vendor data (23,300+ attributes, 200+ locations, 50+ years). Within the WSTAMP analytical environment, Dynamic Time Warping (DTW) has been a highly effective data-driven approach for clustering and mapping these time series into national spatiotemporal behavior maps. Two significant properties have surfaced from this work. First, several recognizable cluster patterns have emerged and persist across a range of locations, attributes, and time frames (e.g., increasing, decreasing, rebounding, peak, oscillating). Secondly, practitioners engaging WSTAMP have noted the explanatory and anticipatory value of these patterns and articulated particular interest in detecting them within the spatiotemporal cube. This need was addressed by shifting DTW-based clustering from an open ended, data-driven implementation to a taxonomic pattern matching approach. This paper presents the method including implementation strategies for visualization and human computer interaction and applies the approach to a sample data set and concludes with next steps.

Stewart, Robert↗

Feature Extraction: Improving Remote Sensor Classification of Non-Proliferation

This research focuses on developing algorithms for nuclear non-proliferation detection using remote sensor modeling. To improve the performance of classification models, we implemented a data pipeline with feature extraction. This pipeline takes raw data and transforms it into smaller data points called features that still describe the model. Improving this classification works towards the departments of energy’s missions of ensuring American’s security and prosperity by creating technology that addresses nuclear challenges. To conduct this analysis, we used the Python programming language and some key packages, including tsfresh and TSFEL. Originally tsfresh was selected because it has the most statistical features out of all the packages. Later TSFEL was incorporated due to the additional features it can extract from data, such as temporal and spectral. However, feature extraction becomes challenging in the presence of missing values. In this case, two additional Python packages were added to our workflow, NumPy and pandas, allowing for the feature extraction process to handle unknown values. Our data pipeline was tested on data collected from a simulation that describes the process state of a physical example. The results show the pipeline’s capability to consume and extract a total 17 features from tabular data. Future work includes producing classifications using decision tree-based models such as XGBoost and improving data collection by analyzing feature importance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deep Multimodal Networks for M-type Star Classification with Paired Spectrum and Photometric Image

Abstract Traditional stellar classification methods include spectral and photometric classification separately. Although satisfactory results can be achieved, the accuracy could be improved. In this paper, we pioneer a novel approach to deeply fuse the spectra and photometric images of the sources in an advanced multimodal network to enhance the model’s discriminatory ability. We use Transformer as the fusion module and apply a spectrum–image contrastive loss function to enhance the consistency of the spectrum and photometric image of the same source in two different feature spaces. We perform M-type stellar subtype classification on two data sets with high and low signal-to-noise ratio (S/N) spectra and corresponding photometric images, and the F1-score achieves 95.65% and 90.84%, respectively. In our experiments, we prove that our model effectively utilizes the information from photometric images and is more accurate than advanced spectrum and photometric image classifiers. Our contributions can be summarized as follows: (1) We propose an innovative idea for stellar classification that allows the model to simultaneously consider information from spectra and photometric images. (2) We discover the challenge of fusing low-S/N spectra and photometric images in the Transformer and provide a solution. (3) The effectiveness of Transformer for spectral classification is discussed for the first time and will inspire more Transformer-based spectral classification models.

Astronomy & Astrophysics↗

When Spectral Modeling Meets Convolutional Networks: A Method for Discovering Reionization-era Lensed Quasars in Multiband Imaging Data

Over the last two decades, around 300 quasars have been discovered at z ≳ 6, yet only one has been identified as being strongly gravitationally lensed. We explore a new approach—enlarging the permitted spectral parameter space, while introducing a new spatial geometry veto criterion—which is implemented via image-based deep learning. We first apply this approach to a systematic search for reionization-era lensed quasars, using data from the Dark Energy Survey, the Visible and Infrared Survey Telescope for Astronomy Hemisphere Survey, and the Wide-field Infrared Survey Explorer. Our search method consists of two main parts: (i) the preselection of the candidates, based on their spectral energy distributions (SEDs), using catalog-level photometry; and (ii) relative probability calculations of the candidates being a lens or some contaminant, utilizing a convolutional neural network (CNN) classification. The training data sets are constructed by painting deflected point-source lights over actual galaxy images, to generate realistic galaxy–quasar lens models, optimized to find systems with small image separations, i.e., Einstein radii of θ E ≤ 1''. Visual inspection is then performed for sources with CNN scores of P lens > 0.1, which leads us to obtain 36 newly selected lens candidates, which are awaiting spectroscopic confirmation. These findings show that automated SED modeling and deep learning pipelines, supported by modest human input, are a promising route for detecting strong lenses from large catalogs, which can overcome the veto limitations of primarily dropout-based SED selection approaches.

High-redshift galaxies↗

Self-Supervised Cloud Classification

Abstract Low-level marine clouds play a pivotal role in Earth’s weather and climate through their interactions with radiation, heat and moisture transport, and the hydrological cycle. These interactions depend on a range of dynamical and microphysical processes that result in a broad diversity of cloud types and spatial structures, and a comprehensive understanding of cloud morphology is critical for continued improvement of our atmospheric modeling and prediction capabilities moving forward. Deep learning has recently accelerated our ability to study clouds using satellite remote sensing, and machine learning classifiers have enabled detailed studies of cloud morphology. A major limitation of deep learning approaches to this problem, however, is the large number of hand-labeled samples that are required for training. This work applies a recently developed self-supervised learning scheme to train a deep convolutional neural network (CNN) to map marine cloud imagery to vector embeddings that capture information about mesoscale cloud morphology and can be used for satellite image classification. The model is evaluated against existing cloud classification datasets and several use cases are demonstrated, including training cloud classifiers with very few labeled samples, interrogation of the CNN’s learned internal feature representations, cross-instrument application, and resilience against sensor calibration drift and changing scene brightness. The self-supervised approach learns meaningful internal representations of cloud structures and achieves comparable classification accuracy to supervised deep learning methods without the expense of creating large hand-annotated training datasets. Significance Statement Marine clouds heavily influence Earth’s weather and climate, and improved understanding of marine clouds is required to improve our atmospheric modeling capabilities and physical understanding of the atmosphere. Recently, deep learning has emerged as a powerful research tool that can be used to identify and study specific marine cloud types in the vast number of images collected by Earth-observing satellites. While powerful, these approaches require hand-labeling of training data, which is prohibitively time intensive. This study evaluates a recently developed self-supervised deep learning method that does not require human-labeled training data for processing images of clouds. We show that the trained algorithm performs competitively with algorithms trained on hand-labeled data for image classification tasks. We also discuss potential downstream uses and demonstrate some exciting features of the approach including application to multiple satellite instruments, resilience against changing image brightness, and its learned internal representations of cloud types. The self-supervised technique removes one of the major hurdles for applying deep learning to very large atmospheric datasets.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript on residence time distribution simulation in two 10-kilometer long river sections

This data package is associated with the publication “On the Transferability of Residence Time Distributions in Two 10-km Long River Sections with Similar Hydromorphic Units” submitted to the Journal of Hydrology (Bao et al. 2024).Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface, along with their residence time distributions (RTDs) in the subsurface, is crucial for managing water quality and ecosystem health in dynamic river corridors. However, directly simulating high-spatial resolution HEFs and RTDs can be a time-consuming process, particularly for watershed-scale modeling. Efficient surrogate models that link RTDs to hydromorphic units (HUs) may serve as alternatives for simulating RTDs in large-scale models. One common concern with these surrogate models, however, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this, we evaluated the HEFs and the resulting RTD-HU relationships for two 10-kilometer-long river corridors along the Columbia River, using a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework that we previously developed. Applying this framework to the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. This data package includes the model inputs files and the simulation results data. This data package contains 10 folders. The modeling simulation results data are in the folders 100H_pt_data and 300area_pt_data, for the study domain Hanford 100H and 300 area respectively. The remaining eight folders contain the scripts and data to generate the manuscript figures. The file-level metadata file (Bao_2024_Residence_Time_Distribution _flmd.csv) includes a list of all files contained in this data package and descriptions for each. The data dictionary file (Bao_2024_Residence_Time_Distribution _dd.csv) includes column header definitions and units of all tabular files.

54 ENVIRONMENTAL SCIENCES↗

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models↗

Characterizing Long COVID: Deep Phenotype of a Complex Condition

Background: Numerous publications describe the clinical manifestations of post-acute sequelae of SARS-CoV-2 (PASC or "long COVID"), but they are difficult to integrate because of heterogeneous methods and the lack of a standard for denoting the many phenotypic manifestations. Patient-led studies are of particular importance for understanding the natural history of COVID-19, but integration is hampered because they often use different terms to describe the same symptom or condition. This significant disparity in patient versus clinical characterization motivated the proposed ontological approach to specifying manifestations, which will improve capture and integration of future long COVID studies. Methods: The Human Phenotype Ontology (HPO) is a widely used standard for exchange and analysis of phenotypic abnormalities in human disease but has not yet been applied to the analysis of COVID-19. Funding: We identified 303 articles published before April 29, 2021, curated 59 relevant manuscripts that described clinical manifestations in 81 cohorts three weeks or more following acute COVID-19, and mapped 287 unique clinical findings to HPO terms. We present layperson synonyms and definitions that can be used to link patient self-report questionnaires to standard medical terminology. Long COVID clinical manifestations are not assessed consistently across studies, and most manifestations have been reported with a wide range of synonyms by different authors. Across at least 10 cohorts, authors reported 31 unique clinical features corresponding to HPO terms; the most commonly reported feature was Fatigue (median 45.1%) and the least commonly reported was Nausea (median 3.9%), but the reported percentages varied widely between studies. Interpretation: Translating long COVID manifestations into computable HPO terms will improve analysis, data capture, and classification of long COVID patients. If researchers, clinicians, and patients share a common language, then studies can be compared/pooled more effectively. Furthermore, mapping lay terminology to HPO will help patients assist clinicians and researchers in creating phenotypic characterizations that are computationally accessible, thereby improving the stratification, diagnosis, and treatment of long COVID.

60 APPLIED LIFE SCIENCES↗