Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quantum cluster algorithm for data classification

Abstract We present a quantum algorithm for data classification based on the nearest-neighbor learning algorithm. The classification algorithm is divided into two steps: Firstly, data in the same class is divided into smaller groups with sublabels assisting building boundaries between data with different labels. Secondly we construct a quantum circuit for classification that contains multi control gates. The algorithm is easy to implement and efficient in predicting the labels of test data. To illustrate the power and efficiency of this approach, we construct the phase transition diagram for the metal-insulator transition of VO 2 , using limited trained experimental data, where VO 2 is a typical strongly correlated electron materials, and the metallic-insulating phase transition has drawn much attention in condensed matter physics. Moreover, we demonstrate our algorithm on the classification of randomly generated data and the classification of entanglement for various Werner states, where the training sets can not be divided by a single curve, instead, more than one curves are required to separate them apart perfectly. Our preliminary result shows considerable potential for various classification problems, particularly for constructing different phases in materials.

97 MATHEMATICS AND COMPUTING↗

Towards AI Based Data Classification for Decision Making During Testing

During the development of high-consequence items, test systems should be capable of differentiating between test failures resulting from narrowly missing requirements versus those indicating potentially catastrophic faults. In many instances, classifying the data corresponds to simply identifying whether measured waveforms have approximately the anticipated shape. Cast in this light, the problem reduces to converting raw data into a form optimal for use with neural network classifiers. This manuscript investigates different means of representing raw data for image classification. Raw data plots and Short Time Fourier Transform (STFT) spectrograms are classified by both custom built, small-scale, Convolution Neural Networks (CNN) and open-source, multi-million parameter, pre-trained deep CNNs. In the case of time varying frequency content, the STFTs provide images with greater detail and can be accurately classified with simpler networks. This requires less memory and runs faster than classifying the raw data using the more sophisticated options—making STFTs optimal for applications with memory constraints. STFTs are not a panacea. In some cases the time-domain signal contains useful information that should not be discarded. Rather than using raw data or STFTs, the images can be constructed from both by using red and green channels of an RGB image to visualize the real and imaginary components of the transform, with the raw data occupying the blue channel.

97 MATHEMATICS AND COMPUTING↗

Countering Weapons of Mass Destruction Office (CWMD) Data Categorization Study: Chemical, Biological, Radiological, and Nuclear (CBRN) Detection Device Data

Pacific Northwest National Laboratory (PNNL) seeks to address critical questions related to chemical, biological, radiological, and nuclear (CBRN) detection devices. This research aims to enhance the security and understanding of these devices by investigating various aspects of their identification, communication, and functionality. The primary focus is on network security, malware detection, device identification, and intelligence gathering. CBRN data can be categorized in various ways depending on the purpose of CBRN detection devices and the specific context of the applications for analysis. Criteria that can be used to assist in this effort include but are not limited to data type, data protocol, source/destination, application, time, security, and content. This study will inform additional paths for data classification, data profiling, data mapping, and data modeling. This will help the Countering Weapons of Mass Destruction Office (CWMD) better understand their data and make informed decisions based on the insights gained from this study and their application. The CBRN Data Categorization study will include the identification of 5–10 different CBRN detection devices with unique characteristics for assessing and analyzing the data that is being produced by and transmitted from these devices.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Enhancing Electron Microscopy Image Classification Using Data Augmentation

Manual labeling for machine learning tasks such as image classification is tedious and labor-intensive; as a result, scientific datasets suitable for deep learning applications are scarce and limited. While data augmentation techniques have shown promise for extending image datasets, very little work has been done to understand the impact of combining multiple augmentation methods sequentially or the limits of their effectiveness when combined. Our work addresses this gap by examining how standard and combinatorial data augmentation affects the performance of machine learning models when trained on small datasets for label classification tasks. For our analysis, we generate single, double and quadruple-augmented datasets for a microscopy image classification task using six standard augmentation methods, and compare the resultant improvements observed in binary classification accuracy with three standard image classification models (DenseNet169, MobileNetV2, ResNet101V2). Our experiments show a non-monotonic relationship between the number of simultaneous augmentation methods and classification accuracy, indicating that there is a trade-off between the degree of augmentation and the model performance. These findings suggest that the optimal number of augmentation methods will vary by domain and use case. We also find that the order in which augmentation methods are applied to a limited dataset matters when combining augmentation schemes, with our use case showing performance differences up to 2.6% when the augmentation order is reversed for double-augmented datasets. Our work offers insights to the limits of data augmentation when working on image classification tasks with limited datasets.

Welsman, Jordan A↗

Transparent application-layer/os deeper packet inspector

A computer-implemented method of deep packet inspection (DPI) in a network is provided. The method comprises collecting data packets comprising a number of traffic flows from a number of devices via a number of traffic taps and classifying each traffic flow according to data about network protocol layers of the packets comprising the traffic flow. Application layer metadata is extracted from the packets. Traffic flow classification data and the extracted metadata are ingested into a data cluster and normalized. The normalized classification data and extracted metadata is then correlated to other data sets.

Urias, Vincent↗

Grassmannian Diffusion Maps--Based Dimension Reduction and Classification for High-Dimensional Data

This work introduces the Grassmannian diffusion maps (GDMaps), a novel nonlinear dimensionality reduction technique that defines the affinity between points through their representation as low-dimensional subspaces corresponding to points on the Grassmann manifold. Here, the method is designed for applications, such as image recognition and data-based classification of constrained high-dimensional data where each data point itself is a high-dimensional object (i.e., a large matrix) that can be compactly represented in a lower-dimensional subspace. The GDMaps is composed of two stages. The first is a pointwise linear dimensionality reduction wherein each high-dimensional object is mapped onto the Grassmann manifold representing the low-dimensional subspace on which it resides. The second stage is a multipoint nonlinear kernel-based dimension reduction using diffusion maps to identify the subspace structure of the points on the Grassmann manifold. To this end, an appropriate Grassmannian kernel is used to construct the transition matrix of a random walk on a graph connecting points on the Grassmann manifold. Spectral analysis of the transition matrix yields low-dimensional Grassmannian diffusion coordinates embedding the data into a low-dimensional reproducing kernel Hilbert space. Further, a novel data classification/recognition technique is developed based on the construction of an overcomplete dictionary of reduced dimension whose atoms are given by the Grassmannian diffusion coordinates. Three examples are considered. First, a "toy" example shows that the GDMaps can identify an appropriate parametrization of structured points on the unit sphere. The second example demonstrates the ability of the GDMaps to revealing the intrinsic subspace structure of high-dimensional random field data. In the last ex- ample, a face recognition problem is solved considering face images subject to varying illumination conditions, changes in face expressions, and occurrence of occlusions. The technique presented high recognition rates (i.e., 95% in the best case) using a fraction of the data required by conventional methods.

42 ENGINEERING↗

Assurance by Design for Cyber Physical Data-Driven Systems

Currently, Cyber Physical Data-Driven Systems (CPDDS) employ machine learning for the classification, data fusion, and control of our nation’s infrastructure, such as the power grid, transportation networks (e.g., fuel distribution, air traffic control), and DoD long-duration collaborative autonomous platforms including unmanned underwater, ground, surface, space, and aerial systems. Many CPDDSs are system-of-systems that should be designed to communicate over disadvantaged networks. It is important to assure that the CPDDSs are resilient against physical and cyber threats by design. Additionally, their design should tolerate misclassification errors resulting from natural and/or adversarial distribution shifts within their data driven components. The all-domain nature of the problem of assuring the design of CPDDSs requires a multi-disciplinary perspective as outlined in this chapter.

Chikkagoudar, Satish↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Pixel Anomaly Detection Tool : a user-friendly GUI for classifying detector frames using machine-learning approaches

Data collection at X-ray free electron lasers has particular experimental challenges, such as continuous sample delivery or the use of novel ultrafast high-dynamic-range gain-switching X-ray detectors. This can result in a multitude of data artefacts, which can be detrimental to accurately determining structure-factor amplitudes for serial crystallography or single-particle imaging experiments. Here, a new data-classification tool is reported that offers a variety of machine-learning algorithms to sort data trained either on manual data sorting by the user or by profile fitting the intensity distribution on the detector based on the experiment. This is integrated into an easy-to-use graphical user interface, specifically designed to support the detectors, file formats and software available at most X-ray free electron laser facilities. The highly modular design makes the tool easily expandable to comply with other X-ray sources and detectors, and the supervised learning approach enables even the novice user to sort data containing unwanted artefacts or perform routine data-analysis tasks such as hit finding during an experiment, without needing to write code.

47 OTHER INSTRUMENTATION↗

Exploring Legacy Document Management Systems to Increase Access to Project Data

The current Electronic Data Management System (EDMS) makes it difficult to efficiently access information in critical time frames. This project explored ways to allow Environmental Support and Services staff the ability to efficiently access documents with necessary historic information while leveraging existing data classification methods and the capabilities of EDMS.

99 GENERAL AND MISCELLANEOUS↗

SAIL Field Campaign X-Band Precipitation Radar Seasonal Surface Hydrometeor Phase Classification Value-Added Product Report

The Surface Atmosphere Integrated Field Laboratory (SAIL) campaign generated an unprecedented, multisensory record of cold-season precipitation in the Upper Colorado River Basin (UCRB). To complement the Surface Quantitative Precipitation Estimation (SQUIRE) snow water equivalent (SWE) product derived from the Colorado State University (CSU) X-band radar, a hydrometeor phase classification data set was developed to systematically identify phase of the near-surface precipitation. The algorithm uses Corrected Moments in Antenna Coordinates (CMAC) fields with semi-supervised classification (SSC) and fuzzy-logic hydrometeor classification (FHC) methods and maps hydrometeor classes to a compact set of phase classes suitable for Quantitative Precipitation Estimation (QPE) and hydrologic applications. We used SQUIRE grid projections to mitigate beam-blockage issues by selecting the near-surface gates for each beam as well as to make it directly usable for the QPE products.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Quantum-Classical Neural Networks

Deep learning is one of the most successful and far-reaching strategies used in machine learning today. However, the scale and utility of neural networks is still greatly limited by the current hardware used to train them. These concerns have become increasingly pressing as conventional computers are soon expected to approach the physical limitations that will slow their performance improvements in the near future. For these reasons, scientists have begun to explore alternative computing platforms, like quantum computers, for training neural networks. In recent years, variational quantum circuits have emerged as one of the most successful approaches to quantum deep learning on noisy intermediate scale quantum devices. We propose a hybrid quantum-classical neural network architecture where each neuron is a variational quantum circuit. We empirically analyze the performance of this hybrid neural network on a series of binary classification data sets using a simulated IBM universal quantum computer and a state-of-the-art IBM universal quantum computer. On the simulated hardware, we observe that the hybrid neural network achieves around 10% higher classification accuracy and 20% better minimization of the cost function than an individual variational quantum circuit. On the quantum hardware, we observe that each model only performs well when the qubit and gate count is sufficiently small.

Arthur, Davis↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Fast decay of classification error in variational quantum circuits

Variational quantum circuits (VQCs) have shown great potential in near-term applications. However, the discriminative power of a VQC, in connection to its circuit architecture and depth, is not understood. To unleash the genuine discriminative power of a VQC, we propose a VQC system with the optimal classical post-processing—maximum-likelihood estimation on measuring all VQC output qubits. Via extensive numerical simulations, we find that the error of VQC quantum data classification typically decays exponentially with the circuit depth, when the VQC architecture is extensive—the number of gates does not shrink with the circuit depth. This fast error suppression ends at the saturation towards the ultimate Helstrom limit of quantum state discrimination. On the other hand, non-extensive VQCs such as quantum convolutional neural networks are sub-optimal and fail to achieve the Helstrom limit, demonstrating a trade-off between ansatz complexity and classification performance in general. To achieve the best performance for a given VQC, the optimal classical post-processing is crucial even for a binary classification problem. To simplify VQCs for near-term implementations, we find that utilizing the symmetry of the input properly can improve the performance, while oversimplification can lead to degradation.

Zhang, Bingzhi↗