Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Genetic Algorithm for Hyperparameter Optimization in Gaussian Process Modeling

A genetic algorithm is developed and applied to optimize hyperparameters of convolutional recursively determined dual neural network-Gaussian process (NNGP) kernels. As a specific application of the combined GPNN-GA algorithm, it is applied to image classification in publicly available data of Hyper Suprime-Cam Subaru Strategic Program. Matthews correlation coefficient is calculated based on results of binary star-galaxy classification and used as a fitting function of the GA module of the algorithm. The simulation results confirm significant improvement of the classification accuracy with optimized hyperparameters.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric redshift-aided classification using ensemble learning

We present SHEEP, a new machine learning approach to the classic problem of astronomical source classification, which combines the outputs from the XGBoost, LightGBM, and CatBoost learning algorithms to create stronger classifiers. A novel step in our pipeline is that prior to performing the classification, SHEEP first estimates photometric redshifts, which are then placed into the data set as an additional feature for classification model training; this results in significant improvements in the subsequent classification performance. SHEEP contains two distinct classification methodologies: (i) Multi-class and (ii) one versus all with correction by a meta-learner. We demonstrate the performance of SHEEP for the classification of stars, galaxies, and quasars using a data set composed of SDSS and WISE photometry of 3.5 million astronomical sources. The resulting F1 -scores are as follows: 0.992 for galaxies; 0.967 for quasars; and 0.985 for stars. In terms of the F1-scores for the three classes, SHEEP is found to outperform a recent RandomForest-based classification approach using an essentially identical data set. Our methodology also facilitates model and data set explainability via feature importances; it also allows the selection of sources whose uncertain classifications may make them interesting sources for follow-up observations.

79 ASTRONOMY AND ASTROPHYSICS↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

ANN↗

Automated Stellar Spectra Classification with Ensemble Convolutional Neural Network

Large sky survey telescopes have produced a tremendous amount of astronomical data, including spectra. Machine learning methods must be employed to automatically process the spectral data obtained by these telescopes. Classification of stellar spectra by applying deep learning is an important research direction for the automatic classification of high-dimensional celestial spectra. In this paper, a robust ensemble convolutional neural network (ECNN) was designed and applied to improve the classification accuracy of massive stellar spectra from the Sloan digital sky survey. We designed six classifiers which consist six different convolutional neural networks (CNN), respectively, to recognize the spectra in DR16. Then, according the cross-entropy testing error of the spectra at different signal-to-noise ratios, we integrate the results of different classifiers in an ensemble learning way to improve the effect of classification. The experimental result proved that our one-dimensional ECNN strategy could achieve 95.0% accuracy in the classification task of the stellar spectra, a level of accuracy that exceeds that of the classical principal component analysis and support vector machine model.

79 ASTRONOMY AND ASTROPHYSICS↗

Improving ProtoDUNE pion cross-section measurements with NuGraph Michel-electron tagging

Understanding hadron-argon interactions is essential for precise neutrino energy reconstruction and final-state interaction modeling in liquid-argon time projection chamber (LArTPC) experiments such as DUNE. In particular, pion absorption and charge-exchange processes constitute significant sources of systematic uncertainty in neutrino oscillation measurements. ProtoDUNE-SP, a large-scale LArTPC prototype operated at the CERN Neutrino Platform and exposed to charged-particle test beams in the few-GeV range, enables direct measurements of these processes. This work focuses on the measurement of differential cross sections for pion absorption and charge exchange using the 2 GeV/c pion beam data from the ProtoDUNE-SP run. A key component of this analysis is the identification of Michel electrons from $\pi \rightarrow \mu \rightarrow e$ decay chains, which helps separate different interaction topologies and improves background rejection. Michel electron identification will also assist in reliably calibrating the electromagnetic response in ProtoDUNE-SP data and for the future DUNE detectors. In this analysis, we apply NuGraph to identify Michel electrons. NuGraph is a graph neural network that models detector hits as nodes connected by spatial and temporal edges for particle and topology classification in LArTPC detectors. We first benchmark NuGraph’s Michel electron classification performance using ICEBERG data, a small-scale LArTPC prototype used for DUNE electronics and reconstruction development, and then transfer the approach to ProtoDUNE-SP. This poster presents the analysis strategy, NuGraph-based classification studies, and discusses how these developments are expected to improve the pion cross-section measurement.

Razafinime, Soamasina Herilala [Cincinnati U.] (OR↗

Outage Cause Classification of Power Distribution Systems with Machine Learning and Real-World Data

Power distribution systems are geographically dispersed by nature. It may be affected by various factors, such as vegetation, weather, animal and human behaviors. Present response procedures to an outage event massively rely on expert experience and thus tend to be time-consuming. Automatic outage event detection and classification will help to reduce the responding and restoration time. However, this issue is less addressed with existing research done in this area. In this applied research, a set of waveform pre-processing techniques are first proposed to prepare the waveform data for being used as inputs to the classification algorithm. Further, a machine learning-based algorithm is proposed to classify the outage events according to their root causes, e.g. tree contact, animal contact, lightning, etc. Available data include three phase current & voltage waveforms and contextual information during the distribution system outages. The proposed machine learning algorithm takes the current and voltage waveforms as direct inputs in search of features that humans are unable to capture. Real data provided by a distribution company in the East Tennessee region is used to test the proposed pre-processing techniques and the classification algorithm.

Sun, Haoyuan↗

Review of multi-faceted morphologic signatures of actinide process materials for nuclear forensic science

Particle morphology is an emerging signature that has the potential to identify the processing history of unknown nuclear materials. Using readily available scanning electron microscopes (SEM), the morphology of nearly any solid material can be measured within hours. Coupled with robust image analysis and classification methods, the morphological features can be quantified and support identification of the processing history of unknown nuclear materials. The viability of this signature depends on developing databases of morphological features, coupled with a rapid data analysis and accurate classification process. With developed reference methods, datasets, and throughputs, morphological analysis can be applied within days to (i) interdicted bulk nuclear materials (gram to kilogram quantities), and (ii) trace amounts of nuclear materials detected on swipes or environmental samples. In conclusion, this review aims to develop validated and verified analytical strategies for morphological analysis relevant to nuclear forensics.

36 MATERIALS SCIENCE↗

End-to-end jet classification of boosted top quarks with the CMS open data

Here we describe a novel application of the end-to-end deep learning technique to the task of discriminating top quark-initiated jets from those originating from the hadronization of a light quark or a gluon. The end-to-end deep learning technique uses low-level detector representation of high-energy collision event as inputs to deep learning algorithms. In this study, we use low-level detector information from the simulated Compact Muon Solenoid (CMS) open data samples to construct the top jet classifiers. To optimize classifier performance we progressively add low-level information from the CMS tracking detector, including pixel detector reconstructed hits and impact parameters, and demonstrate the value of additional tracking information even when no new spatial structures are added. Relying only on calorimeter energy deposits and reconstructed pixel detector hits, the end-to-end classifier achieves an area under the receiver operator curve (AUC) score of 0.975 ± 0.002 for the task of classifying boosted top quark jets. After adding derived track quantities, the classifier AUC score increases to 0.9824 ± 0.0013, serving as the first performance benchmark for these CMS open data samples.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Square Kilometre Array Science Data Challenge 1: analysis and results

ABSTRACT As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques developed to exploit the rich scientific value of the data. To drive forward the development of efficient and accurate analysis methods, we are designing a series of data challenges that will provide the scientific community with high-quality data sets for testing and evaluating new techniques. In this paper, we present a description and results from the first such Science Data Challenge 1 (SDC1). Based on SKA MID continuum simulated observations and covering three frequencies (560, 1400, and 9200 MHz) at three depths (8, 100, and 1000 h), SDC1 asked participants to apply source detection, characterization, and classification methods to simulated data. The challenge opened in 2018 November, with nine teams submitting results by the deadline of 2019 April. In this work, we analyse the results for eight of those teams, showcasing the variety of approaches that can be successfully used to find, characterize, and classify sources in a deep, crowded field. The results also demonstrate the importance of building domain knowledge and expertise on this kind of analysis to obtain the best performance. As high-resolution observations begin revealing the true complexity of the sky, one of the outstanding challenges emerging from this analysis is the ability to deal with highly resolved and complex sources as effectively as the unresolved source population.

Bonaldi, A.↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Learning to identify electrons

In this report we investigate whether state-of-the-art classification features commonly used to distinguish electrons from jet backgrounds in collider experiments are overlooking valuable information. A deep convolutional neural network analysis of electromagnetic and hadronic calorimeter deposits is compared to the performance of typical features, revealing a ≈ 5% gap which indicates that these lower-level data do contain untapped classification power. To reveal the nature of this unused information, we use a recently developed technique to map the deep network into a space of physically interpretable observables. We identify two simple calorimeter observables which are not typically used for electron identification, but which mimic the decisions of the convolutional network and nearly close the performance gap.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data-Driven Whole-Genome Clustering to Detect Geospatial, Temporal, and Functional Trends in SARS-CoV-2 Evolution

Current methods for defining SARS-CoV-2 lineages ignore the vast majority of the SARS-CoV-2 genome. We develop and apply an exhaustive vector comparison method that directly compares all known SARS-CoV-2 genome sequences to produce novel lineage classifications. We utilize data-driven models that (i) accurately capture the complex interactions across the set of all known SARS-CoV-2 genomes, (ii) scale to leadership-class computing systems, and (iii) enable tracking how such strains evolve geospatially over time. We show that during the height of the original Omicron surge, countries across Europe, Asia, and the Americas had a spatially asynchronous distribution of Omicron sub-strains. Moreover, neighboring countries were often dominated by either different clusters of the same variant or different variants altogether throughout the pandemic. Analyses of this kind may suggest a different pattern of epidemiological risk than was understood from conventional data, as well as produce actionable insights and transform our ability to prepare for and respond to current and future biological threats.

Jacobson, Daniel↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗