Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Square Kilometre Array Science Data Challenge 1: analysis and results

ABSTRACT As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques developed to exploit the rich scientific value of the data. To drive forward the development of efficient and accurate analysis methods, we are designing a series of data challenges that will provide the scientific community with high-quality data sets for testing and evaluating new techniques. In this paper, we present a description and results from the first such Science Data Challenge 1 (SDC1). Based on SKA MID continuum simulated observations and covering three frequencies (560, 1400, and 9200 MHz) at three depths (8, 100, and 1000 h), SDC1 asked participants to apply source detection, characterization, and classification methods to simulated data. The challenge opened in 2018 November, with nine teams submitting results by the deadline of 2019 April. In this work, we analyse the results for eight of those teams, showcasing the variety of approaches that can be successfully used to find, characterize, and classify sources in a deep, crowded field. The results also demonstrate the importance of building domain knowledge and expertise on this kind of analysis to obtain the best performance. As high-resolution observations begin revealing the true complexity of the sky, one of the outstanding challenges emerging from this analysis is the ability to deal with highly resolved and complex sources as effectively as the unresolved source population.

Bonaldi, A.↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Learning to identify electrons

In this report we investigate whether state-of-the-art classification features commonly used to distinguish electrons from jet backgrounds in collider experiments are overlooking valuable information. A deep convolutional neural network analysis of electromagnetic and hadronic calorimeter deposits is compared to the performance of typical features, revealing a ≈ 5% gap which indicates that these lower-level data do contain untapped classification power. To reveal the nature of this unused information, we use a recently developed technique to map the deep network into a space of physically interpretable observables. We identify two simple calorimeter observables which are not typically used for electron identification, but which mimic the decisions of the convolutional network and nearly close the performance gap.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data-Driven Whole-Genome Clustering to Detect Geospatial, Temporal, and Functional Trends in SARS-CoV-2 Evolution

Current methods for defining SARS-CoV-2 lineages ignore the vast majority of the SARS-CoV-2 genome. We develop and apply an exhaustive vector comparison method that directly compares all known SARS-CoV-2 genome sequences to produce novel lineage classifications. We utilize data-driven models that (i) accurately capture the complex interactions across the set of all known SARS-CoV-2 genomes, (ii) scale to leadership-class computing systems, and (iii) enable tracking how such strains evolve geospatially over time. We show that during the height of the original Omicron surge, countries across Europe, Asia, and the Americas had a spatially asynchronous distribution of Omicron sub-strains. Moreover, neighboring countries were often dominated by either different clusters of the same variant or different variants altogether throughout the pandemic. Analyses of this kind may suggest a different pattern of epidemiological risk than was understood from conventional data, as well as produce actionable insights and transform our ability to prepare for and respond to current and future biological threats.

Jacobson, Daniel↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Automatic Detection and Classification of Radio Galaxy Images by Deep Learning

Abstract Surveys conducted by radio astronomy observatories, such as SKA, MeerKAT, Very Large Array, and ASKAP, have generated massive astronomical images containing radio galaxies (RGs). This generation of massive RG images has imposed strict requirements on the detection and classification of RGs and makes manual classification and detection increasingly difficult, even impossible. Rapid classification and detection of images of different types of RGs help astronomers make full use of the observed astronomical image data for further processing and analysis. The classification of FRI and FRII is relatively easy, and there are more studies and literature on them at present, but FR0 and FRI are similar, so it is difficult to distinguish them. It poses a greater challenge to image processing. At present, deep learning has made breakthrough progress in the field of image analysis and processing and has preliminary applications in astronomical data processing. Compared with classification algorithms that can only classify galaxies, object detection algorithms that can locate and classify RGs simultaneously are preferred. In target detection algorithms, YOLOv5 has outstanding advantages in the classification and positioning of small targets. Therefore, we propose a deep-learning method based on an improved YOLOv5 object detection model that makes full use of multisource data, combining FIRST radio with SDSS optical image data, and realizes the automatic detection of FR0, FRI, and FRII RGs. The innovation of our work is that on the basis of the original YOLOv5 object detection model, we introduce the SE Net attention mechanism, increase the number of preset anchors, adjust the network structure of the feature pyramid, and modify the network structure, thereby allowing our model to demonstrate galaxy classification and position detection effects. Our improved model produces satisfactory results, as evidenced by experiments. Overall, the mean average precision (mAP@0.5) of our improved model on the test set reaches 89.4%, which can determine the position (R.A. and decl.) and automatically detect and classify FR0s, FRIs, and FRIIs. Our work contributes to astronomy because it allows astronomers to locate FR0, FRI, and FRII galaxies in a relatively short time and can be further combined with other astronomically generated data to study the properties of these galaxies. The target detection model can also help astronomers find FR0s, FRIs, and FRIIs in future surveys and build a large-scale star RG catalog. Moreover, our work is also useful for the detection of other types of galaxies.

Astronomy & Astrophysics↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Star–Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

Abstract We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam Subaru Strategic Program data. In our experiments, we evaluate the performance of a principal component analysis embedded GP predictive model against other machine-learning algorithms, including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

cTULIP: application of a human-based RNA-seq primary tumor classification tool for cross-species primary tumor classification in canine

The domestic dog, Canis familiaris, is quickly gaining traction as an advantageous model for use in the study of cancer, one of the leading causes of death worldwide. Naturally occurring canine cancers share clinical, histological, and molecular characteristics with the corresponding human diseases. In this study, we take a deep-learning approach to test how similar the gene expression profile of canine glioma and bladder cancer (BLCA) tumors are to the corresponding human tumors. We likewise develop a tool for identifying misclassified or outlier samples in large canine oncological datasets, analogous to that which was developed for human datasets. We test a number of machine learning algorithms and found that a convolutional neural network outperformed logistic regression and random forest approaches. We use a recently developed RNA-seq-based convolutional neural network, TULIP, to test the robustness of a human-data-trained primary tumor classification tool on cross-species primary tumor prediction. Our study ultimately highlights the molecular similarities between canine and human BLCA and glioma tumors, showing that protein-coding one-to-one homologs shared between humans and canines, are sufficient to distinguish between BLCA and gliomas. The results of this study indicate that using protein-coding one-to-one homologs as the features in the input layer of TULIP performs good primary tumor prediction in both humans and canines. Furthermore, our analysis shows that our selected features also contain the majority of features with known clinical relevance in BLCA and gliomas. Our success in using a human-data-trained model for cross-species primary tumor prediction also sheds light on the conservation of oncological pathways in humans and canines, further underscoring the importance of the canine model system in the study of human disease.

60 APPLIED LIFE SCIENCES↗

Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation

To protect tropical forest biodiversity, we need to be able to detect it reliably, cheaply, and at scale. Automated detection of sound producing animals from passively recorded soundscapes via machine-learning approaches is a promising technique towards this goal, but it is constrained by the necessity of large training data sets. Using soundscapes from a tropical forest in Borneo and a Convolutional Neural Network model (CNN), we investigate i) the minimum viable training data set size for accurate prediction of call types (‘sonotypes’), and ii) the extent to which data augmentation and transfer learning can overcome the issue of small and imbalanced training data sets. We found that even relatively high sample sizes (>80 per sonotype) lead to mediocre accuracy, which however improved significantly with data augmentation and transfer learning, including at extremely small sample sizes (3 per sonotype), regardless of taxonomic group or call characteristics. Neither transfer learning nor data augmentation alone achieved high accuracy. Our results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape-based projects with many rare species. Retraining our open-source model requires only basic programming skills which makes it possible for individual conservation initiatives to match their local context, in order to enable more evidence-informed management of biodiversity.

54 ENVIRONMENTAL SCIENCES↗

A Compound Data Poisoning Technique with Significant Adversarial Effects on Transformer-based Sentiment Classification Tasks

Transformer-based models have demonstrated much success in various natural language processing tasks. However, they are often vulnerable to adversarial attacks, such as data poisoning, which can intentionally fool the model into generating incorrect results. In this article, we present a novel, compound variant of a data poisoning attack on a transformer-based model that maximizes the poisoning effect while minimizing the scope of poisoning. Here we do so by combining the established data poisoning technique (label flipping) with a novel adversarial artifact selection and insertion technique aimed at minimizing detectability and the scope of the poisoning footprint. We find that by using a combination of these two techniques, we achieve a state-of-the-art attack success rate of approximately 90% while poisoning only 0.5% of the original training set, thus minimizing the scope and detectability of the poisoning action. These findings have the potential to advance the development of better data poisoning detection methods.

97 MATHEMATICS AND COMPUTING↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

PERSIANN-CCS-CDR, a 3-hourly 0.04° global precipitation climate data record for heavy precipitation studies

Accurate long-term global precipitation estimates, especially for heavy precipitation rates, at fine spatial and temporal resolutions is vital for a wide variety of climatological studies. Most of the available operational precipitation estimation datasets provide either high spatial resolution with short-term duration estimates or lower spatial resolution with long-term duration estimates. Furthermore, previous research has stressed that most of the available satellite-based precipitation products show poor performance for capturing extreme events at high temporal resolution. Therefore, there is a need for a precipitation product that reliably detects heavy precipitation rates with fine spatiotemporal resolution and a longer period of record. Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks-Cloud Classification System-Climate Data Record (PERSIANN-CCS-CDR) is designed to address these limitations. This dataset provides precipitation estimates at 0.04° spatial and 3-hourly temporal resolutions from 1983 to present over the global domain of 60°S to 60°N. Evaluations of PERSIANN-CCS-CDR and PERSIANN-CDR against gauge and radar observations show the better performance of PERSIANN-CCS-CDR in representing the spatiotemporal resolution, magnitude, and spatial distribution patterns of precipitation, especially for extreme events.

54 ENVIRONMENTAL SCIENCES↗

Pose Classification Using Three-Dimensional Atomic Structure-Based Neural Networks Applied to Ion Channel–Ligand Docking

The identification of promising lead compounds showing pharmacological activities toward a biological target is essential in early stage drug discovery. With the recent increase in available small-molecule databases, virtual high-throughput screening using physics-based molecular docking has emerged as an essential tool in assisting fast and cost-efficient lead discovery and optimization. However, the best scored docking poses are often suboptimal, resulting in incorrect screening and chemical property calculation. We address the pose classification problem by leveraging data-driven machine learning approaches to identify correct docking poses from AutoDock Vina and Glide screens. To enable effective classification of docking poses, we present two convolutional neural network approaches: a three-dimensional convolutional neural network (3D-CNN) and an attention-based point cloud network (PCN) trained on the PDBbind refined set. We demonstrate the effectiveness of our proposed classifiers on multiple evaluation data sets including the standard PDBbind CASF-2016 benchmark data set and various compound libraries with structurally different protein targets including an ion channel data set extracted from Protein Data Bank (PDB) and an in-house KCa3.1 inhibitor data set. Our experiments show that excluding false positive docking poses using the proposed classifiers improves virtual high-throughput screening to identify novel molecules against each target protein compared to the initial screen based on the docking scores.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification and Photometric Classification of Extragalactic Transients in the Vera C. Rubin Observatory’s Data Preview 1

The Vera C. Rubin Observatory will soon survey the southern sky, delivering a depth and sky coverage that is unprecedented in time-domain astronomy. As part of commissioning, Data Preview 1 (DP1) has been released. It comprises a Legacy Survey of Space and Time (LSST) Commissioning Camera observing campaign between 2024 November and December with multiband imaging of seven fields, covering roughly 0.4 deg 2 each, providing a first glimpse into the data products that will become available once the LSST begins. In this work, we search three fields for extragalactic transients. We identify eight new likely supernovae (SNe), and three known ones from a sample of 369,644 difference image analysis objects. Photometric classification using Superphot+ assigns subclasses with >95% confidence to only one SN Ia and one SN II in this sample. Our findings are in agreement with SN detection rate predictions of 15 ± 4 SNe from simulations using simsurvey. The SN detection rate in the data is possibly affected by the lack of suitable templates. Nevertheless, this work demonstrates the quality of the data products delivered in DP1 and indicates that the Rubin Observatory’s LSST is well placed to fulfill its discovery potential in time-domain astronomy.

Freeburn, James [University of North Carolina, Cha↗