Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

DeepAdversaries: examining the robustness of deep learning models for galaxy morphology classification

With increased adoption of supervised deep learning methods for work with cosmological survey data, the assessment of data perturbation effects (that can naturally occur in the data processing and analysis pipelines) and the development of methods that increase model robustness are increasingly important. In the context of morphological classification of galaxies, we study the effects of perturbations in imaging data. In particular, we examine the consequences of using neural networks when training on baseline data and testing on perturbed data. We consider perturbations associated with two primary sources: (a) increased observational noise as represented by higher levels of Poisson noise and (b) data processing noise incurred by steps such as image compression or telescope errors as represented by one-pixel adversarial attacks. We also test the efficacy of domain adaptation techniques in mitigating the perturbation-driven errors. We use classification accuracy, latent space visualizations, and latent space distance to assess model robustness in the face of these perturbations. For deep learning models without domain adaptation, we find that processing pixel-level errors easily flip the classification into an incorrect class and that higher observational noise makes the model trained on low-noise data unable to classify galaxy morphologies. On the other hand, we show that training with domain adaptation improves model robustness and mitigates the effects of these perturbations, improving the classification accuracy up to 23% on data with higher observational noise. Domain adaptation also increases up to a factor of ${\approx}2.3$ the latent space distance between the baseline and the incorrectly classified one-pixel perturbed image, making the model more robust to inadvertent perturbations. Successful development and implementation of methods that increase model robustness in astronomical survey pipelines will help pave the way for many more uses of deep learning for astronomy.

79 ASTRONOMY AND ASTROPHYSICS↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting Search Task Difficulty through a Discrete‐Time Action Log Representation on Spectrum Kernel

ABSTRACT Predicting perceived difficulty on a web search task is an open problem in the interactive information retrieval field. A common approach to tackle it, is through features obtained from full search sessions, which are then used to train classification models. In this poster we attempt to predict perceived task difficulty at different stages of the search process. To do so, we use the spectrum kernel for support vector machine (SVM) classification. Our preliminary results suggest that by using behavioral data from the first query segment, it is possible to provide timely classifications of whether a search task is perceived as hard or easy.

Gacitúa, Daniel↗

Deep Learning Classification of Cheatgrass Invasion in the Western United States Using Biophysical and Remote Sensing Data

Cheatgrass (Bromus tectorum) invasion is driving an emerging cycle of increased fire frequency and irreversible loss of wildlife habitat in the western US. Yet, detailed spatial information about its occurrence is still lacking for much of its presumably invaded range. Deep learning (DL) has demonstrated success for remote sensing applications but is less tested on more challenging tasks like identifying biological invasions using sub-pixel phenomena. We compare two DL architectures and the more conventional Random Forest and Logistic Regression methods to improve upon a previous effort to map cheatgrass occurrence at >2% canopy cover. High-dimensional sets of biophysical, MODIS, and Landsat-7 ETM+ predictor variables are also compared to evaluate different multi-modal data strategies. All model configurations improved results relative to the case study and accuracy generally improved by combining data from both sensors with biophysical data. Cheatgrass occurrence is mapped at 30 m ground sample distance (GSD) with an estimated 78.1% accuracy, compared to 250-m GSD and 71% map accuracy in the case study. Furthermore, DL is shown to be competitive with well-established machine learning methods in a limited data regime, suggesting it can be an effective tool for mapping biological invasions and more broadly for multi-modal remote sensing applications.

54 ENVIRONMENTAL SCIENCES↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PECAN2

PECAN2, Pose Classification with 3D Atomic Network 2, presents an innovative approach to pose classification in computational modeling, particularly in the context of molecular docking. Traditional methods like docking, which rely on physics-based calculations, are prone to inaccuracies in predicting binding poses during experimental testing. While machine learning (ML) approaches, such as ML-driven pose classification, have been introduced to address these issues, they often depend on the availability of crystal structure data. In contrast, PECAN2 introduces a novel pose classification method that operates independently of crystal structures. Instead, it utilizes a 3D atomic neural network with Point Cloud Network (PCN) to establish correlations between docking scores and experimental data. This approach marks a departure from previous studies that heavily relied on crystal structures for labeling. The key innovation lies in the ability of PECAN2 to enhance the performance of molecular docking by filtering out false positives and false negatives through the correlation between docking scores and experimental data.

Shim, Heesung↗

Deep convolutional neural networks for multi-scale time-series classification and application to tokamak disruption prediction using raw, high temporal resolution diagnostic data

In this paper we discuss recent advances in deep convolutional neural networks (CNN) for sequence learning, which allow identifying long-range, multi-scale phenomena in long sequences, such as those found in fusion plasmas. We point out several benefits of these deep CNN architectures, such as not requiring experts such as physicists to hand-craft input data features, the ability to capture longer range dependencies compared to the more common sequence neural networks (recurrent neural networks like long short-term memory (LSTM) networks), and the comparative computational efficiency. We apply this neural network architecture to the popular problem of disruption prediction in fusion energy tokamaks, utilizing raw data from a single diagnostic, the Electron Cyclotron Emission imaging (ECEi) diagnostic from the DIII-D tokamak. Initial results trained on a large ECEi dataset show promise, achieving an F 1 -score of ~91% on individual time-slices using only the ECEi data. This indicates the ECEi diagnostic by itself can be sensitive to a number of pre-disruption markers useful for predicting disruptions on timescales not only for mitigation but also avoidance. Future opportunities for utilizing these deep CNN architectures with fusion data are outlined, including impact of recent upgrades to the ECEi diagnostic.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deep learning model to detect various synchrophasor data anomalies

High-density synchrophasors provide valuable information for power grid situational awareness, operation and control. Unfortunately, due to factors including communication instability and hardware failure, their data quality can be greatly deteriorated by anomalies. Since the anomalies can impact the performance of the synchrophasor applications, it is of paramount significance to propose a model to detect anomalies in synchrophasor. In this study, a convolutional neural network model is established to detect and classify the anomalies in the synchrophasor measurements. Additionally, four types of anomalies observed in actual synchrophasors including erroneous patterns, random spikes, missing points and high-frequency interferences are considered in this study. The proposed model is extensively evaluated via field-collected measurements from the synchrophasor network in Jiangsu grid, China. The superior performance of the proposed model indicates the great potential of using deep learning for the detection of abnormal synchrophasor measurements.

42 ENGINEERING↗

Energy Efficient Streaming Time Series Classification with Attentive Power Iteration

Efficiently processing time series data streams in real-time on resource-constrained devices offers significant advantages in terms of enhanced computational energy efficiency and reduced time-related risks. We introduce an innovative streaming time series classification network that utilizes attentive power iteration, enabling real-time processing on resource-constrained devices. Our model continuously updates a compact representation of the entire time series, enhancing classification accuracy while conserving energy and processing time. Notably, it excels in streaming scenarios without requiring complete time series access, enabling swift decisions. Experimental results show that our approach excels in classification accuracy and energy efficiency, with over 70% less consumption and threefold faster task completion than benchmarks. This work advances real-time responsiveness, energy conservation, and operational effectiveness for constrained devices, contributing to optimizing various applications.

97 MATHEMATICS AND COMPUTING↗

Mapping Vegetation at Species Level with High-Resolution Multispectral and Lidar Data Over a Large Spatial Area: A Case Study with Kudzu

Mapping vegetation species is critical to facilitate related quantitative assessment, and mapping invasive plants is important to enhance monitoring and management activities. Integrating high-resolution multispectral remote-sensing (RS) images and lidar (light detection and ranging) point clouds can provide robust features for vegetation mapping. However, using multiple sources of high-resolution RS data for vegetation mapping on a large spatial scale can be both computationally and sampling intensive. Here, we designed a two-step classification workflow to potentially decrease computational cost and sampling effort and to increase classification accuracy by integrating multispectral and lidar data in order to derive spectral, textural, and structural features for mapping target vegetation species. We used this workflow to classify kudzu, an aggressive invasive vine, in the entire Knox County (1362 km2) of Tennessee (U.S.). Object-based image analysis was conducted in the workflow. The first-step classification used 320 kudzu samples and extensive, coarsely labeled samples (based on national land cover) to generate an overprediction map of kudzu using random forest (RF). For the second step, 350 samples were randomly extracted from the overpredicted kudzu and labeled manually for the final prediction using RF and support vector machine (SVM). Computationally intensive features were only used for the second-step classification. SVM had constantly better accuracy than RF, and the producer’s accuracy, user’s accuracy, and Kappa for the SVM model on kudzu were 0.94, 0.96, and 0.90, respectively. SVM predicted 1010 kudzu patches covering 1.29 km2 in Knox County. We found the sample size of kudzu used for algorithm training impacted the accuracy and number of kudzu predicted. The proposed workflow could also improve sampling efficiency and specificity. Our workflow had much higher accuracy than the traditional method conducted in this research, and could be easily implemented to map kudzu in other regions as well as map other vegetation species.

59 BASIC BIOLOGICAL SCIENCES↗

Applying novel analytical tools for analyzing multidimensional secondary organic aerosol measurements

In the atmosphere, secondary organic aerosols (SOA) are often the major components of fine particulate matter and interact with clouds and radiation. SOA comprises a mixture of thousands of organic compounds. There is tremendous complexity and uncertainty in understanding SOA formation, since it is formed by oxidation and gas to particle conversion of a variety of sources: natural biogenic, anthropogenic (vehicles, cooking coal combustion) and biomass burning. The Aerosol Mass Spectrometer (AMS) produces multidimensional chemical information about SOA but analyzing this data to understand SOA sources relies on time consuming analyses (~months to years) such as the positive matrix factorization (PMF). PMF also becomes difficult for aircraft data where signal to noise ratio is weaker. There is a critical need to develop fast machine learning techniques that can analytically provide information about SOA sources using AMS data on the same timescales as the data is being collected (~minutes). We apply a machine learning supervised classification approach: the multinomial logistic regression to rapidly classify AMS data obtained from aircraft measurements.

47 OTHER INSTRUMENTATION↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

Evaluation of WSR-88D Level III and MRMS Rainfall Estimates Against Rain Gauge Observations at the Savannah River Site

Rainfall data for the Savannah River Site (SRS) has been historically measured by rain gauges. These instruments serve as ground truth for most climatological and weather applications; however, gauge measurements are prone to errors or biases under certain weather conditions. Rainfall estimates from radar reflectivity values have been developed and improved over the years and serve as an alternative or supplement for gauge measurements. This study compares measurements from tipping bucket rain gauges with rainfall estimates from the NOAA NWS WSR-88D Level III hourly rainfall product and the Multi-Radar Multi-Sensor (MRMS) gauge-corrected precipitation estimate. Results show good agreement between radar derived amounts and ground measurements, with MRMS values showing correlation coefficients of 0.60-0.95 and RMSE values of less than 1.0 cm (0.4 in). The WSR-88D Level III estimates result in correlation coefficients of 0.46-0.86 and RMSE values of less than 1.3 cm (0.5 in). Few outliers are observed for each data pair and are evaluated against precipitation classification products (WSR-88D Hybrid Hydrometeor Classification product and MRMS Precipitation Flag product). The rain gauges used in this study are not part of the Hydrometeorological Automated Data System (HADS) network used to correct the MRMS estimates and therefore, results of this work provide an independent validation of the MRMS gauge correction scheme.

54 ENVIRONMENTAL SCIENCES↗

An Overview of the Usefulness of Machine Learning Techniques on Network Packet Data

Understanding the health and behavior of a computer network allows for better network efficiency and security. We present an overview of various machine learning techniques for classifying network packet data via packet metadata. While some classical machine learning approaches achieve reasonable results, the most accurate classification can be achieved with deep learning. On the four data sets studied herein, a basic deep learning model achieved at or near 100\% classification accuracy. We also propose a method for determining variable importance as a means for potential transfer learning applications to classifying yet unseen network packet data.

97 MATHEMATICS AND COMPUTING↗

Persistent Classification: Understanding Adversarial Attacks by Studying Decision Boundary Dynamics

ABSTRACT There are a number of hypotheses underlying the existence of adversarial examples for classification problems. These include the high‐dimensionality of the data, the high codimension in the ambient space of the data manifolds of interest, and that the structure of machine learning models may encourage classifiers to develop decision boundaries close to data points. This article proposes a new framework for studying adversarial examples that does not depend directly on the distance to the decision boundary. Similarly to the smoothed classifier literature, we define a (natural or adversarial) data point to be ( γ , σ)‐stable if the probability of the same classification is at least for points sampled in a Gaussian neighborhood of the point with a given standard deviation . We focus on studying the differences between persistence metrics along interpolants of natural and adversarial points. We show that adversarial examples have significantly lower persistence than natural examples for large neural networks in the context of the MNIST and ImageNet datasets. We connect this lack of persistence with decision boundary geometry by measuring angles of interpolants with respect to decision boundaries. Finally, we connect this approach with robustness by developing a manifold alignment gradient metric and demonstrating the increase in robustness that can be achieved when training with the addition of this metric.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES↗

Optical Photometric Indicators of Galaxy Cluster Relaxation

Abstract The most dynamically relaxed clusters of galaxies play a special role in cosmological studies as well as astrophysical studies of the intracluster medium (ICM) and active galactic nucleus feedback. While high-spatial-resolution imaging of the morphology of the ICM has long been the gold standard for establishing a cluster’s dynamical state, such data are not available from current or planned surveys, and thus require separate, pointed follow-up observations. With optical and/or near-IR photometric imaging, and red-sequence cluster finding results from those data, expected to be ubiquitously available for clusters discovered in upcoming optical and millimeter-wavelength surveys, it is worth asking how effectively photometric data alone can identify relaxed cluster candidates, before investing in, e.g., high-resolution X-ray observations. Here we assess the ability of several simple photometric measurements, based on the redMaPPer cluster finder run on Sloan Digital Sky Survey data, to reproduce X-ray classifications of dynamical state for an X-ray selected sample of massive clusters. We find that two simple metrics contrasting the bright central galaxy (BCG) to other cluster members can identify a complete sample of relaxed clusters with a purity of ∼40% in our data set. Including minimal ICM information in the form of a center position increases the purity to ∼60%. However, all three metrics depend critically on correctly identifying the BCG, which is presently a challenge for optical red-sequence cluster finders.

79 ASTRONOMY AND ASTROPHYSICS↗

‘Flux+Mutability’: a conditional generative approach to one-class classification and anomaly detection

Abstract Anomaly Detection is becoming increasingly popular within the experimental physics community. At experiments such as the Large Hadron Collider, anomaly detection is growing in interest for finding new physics beyond the Standard Model. This paper details the implementation of a novel Machine Learning architecture, called Flux+Mutability, which combines cutting-edge conditional generative models with clustering algorithms. In the ‘flux’ stage we learn the distribution of a reference class. The ‘mutability’ stage at inference addresses if data significantly deviates from the reference class. We demonstrate the validity of our approach and its connection to multiple problems spanning from one-class classification to anomaly detection. In particular, we apply our method to the isolation of neutral showers in an electromagnetic calorimeter and show its performance in detecting anomalous dijets events from standard QCD background. This approach limits assumptions on the reference sample and remains agnostic to the complementary class of objects of a given problem. We describe the possibility of dynamically generating a reference population and defining selection criteria via quantile cuts. Remarkably this flexible architecture can be deployed for a wide range of problems, and applications like multi-class classification or data quality control are left for further exploration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗