Making the Most of Data: Feature Engineering for Applied Supervised Machine Learning
DOE Data Days, Livermore, CA, June 1-3, 2022
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
DOE Data Days, Livermore, CA, June 1-3, 2022
Explore the source record for details and available documents.
The modern industrial environment is equipping myriads of smart manufacturing machines where the state of each device can be monitored continuously. Such monitoring can help identify possible future failures and develop a cost-effective maintenance plan. However, it is a daunting task to perform early detection with low false positives and negatives from the huge volume of collected data. This requires developing a holistic machine learning framework to address the issues in condition monitoring of high priority components and develop efficient techniques to detect anomalies that can detect and possibly localize the faulty components. This paper presents a comparative analysis of recent machine learning approaches for robust, cost-effective anomaly detection in cyber-physical systems. While detection has been extensively studied, very few researchers have analyzed the localization of the anomalies. We show that supervised learning outperforms unsupervised algorithms. For supervised cases, we achieve near-perfect accuracy of 98% (specifically for tree-based algorithms). In contrast, the best-case accuracy in the unsupervised cases was 63%—the area under the receiver operating characteristic curve (AUC) exhibits similar outcomes as an additional metric.
Abstract As a critical component of coherent X-ray diffraction imaging (CDI), phase retrieval has been extensively applied in X-ray structural science to recover the 3D morphological information inside measured particles. Despite meeting all the oversampling requirements of Sayre and Shannon, current phase retrieval approaches still have trouble achieving a unique inversion of experimental data in the presence of noise. Here, we propose to overcome this limitation by incorporating a 3D Machine Learning (ML) model combining (optional) supervised learning with transfer learning. The trained ML model can rapidly provide an immediate result with high accuracy which could benefit real-time experiments, and the predicted result can be further refined with transfer learning. More significantly, the proposed ML model can be used without any prior training to learn the missing phases of an image based on minimization of an appropriate ‘loss function’ alone. We demonstrate significantly improved performance with experimental Bragg CDI data over traditional iterative phase retrieval algorithms.
Self-contrastive learning (SCL), a self-supervised learning method, has been shown to improve image and signal classifier accuracies and reduce the training time for neural communications receivers. In particular, prior work has shown that SCL applied as a pre-training step can improve simulated performance of OFDM in 3GPP TDL channel models by reducing the training time of the downstream classification task (demodulation and demapping). In this work a practical implementation demonstrating SCL pre-training using software defined radios (SDRs) is proposed.
Light-matter interaction optimization in complex nanophotonic structures is a critical step towards the tailored performance of photonic devices. The increasing complexity of such systems requires new optimization strategies beyond intuitive methods. For example, in disordered photonic structures, the spatial distribution of energy densities has large random fluctuations due to the interference of multiply scattered electromagnetic waves, even though the statistically averaged spatial profiles of the transmission eigenchannels are universal. Classification of these eigenchannels for a single configuration based on visualization of intensity distributions is difficult. However, successful classification could provide vital information about disordered nanophotonic structures. Emerging methods in machine learning have enabled new investigations into optimized photonic structures. In this work, we combine intensity distributions of the transmission eigenchannels and the transmitted speckle-like intensity patterns to classify the eigenchannels of a single configuration of disordered photonic structures using machine learning techniques. Specifically, we leverage supervised learning methods, such as decision trees and fully connected neural networks, to achieve classification of these transmission eigenchannels based on their intensity distributions with an accuracy greater than 99%, even with a dataset including photonic devices of various disorder strengths. Simultaneous classification of the transmission eigenchannels and the relative disorder strength of the nanophotonic structure is also possible. Our results open new directions for machine learning assisted speckle-based metrology and demonstrate a novel approach to classifying nanophotonic structures based on their electromagnetic field distributions. These insights can be of paramount importance for optimizing light-matter interactions at the nanoscale.
Solving nonlinear optimal power flow (OPF) problem is computationally expensive, and poses scalability challenges for power distribution networks. An alternative to solving the original nonlinear OPF is the linear approximated OPF models. Although, these linear approximated OPF models are fast, the resulting solutions may result in significant optimality gap. Lately, the application of machine learning (ML) methods in successfully solving the nonlinear OPF has been reported. These methods learn and estimate the nonlinear control policies using a purely data-driven approach. In this paper, we propose an approach to complements the ML based approach to solving OPF using solutions from known linearized OPF model. Specifically, we use supervised learning to map the solutions of linear OPF to nonlinear control variables. Unlike, the traditional ML based methods for OPF that approximate the full distribution feeder model using function approximation, our approach uses a two-node approximation of radial networks. The proposed approach is validated using IEEE 123 bus test system for OPF solutions obtained using the nonlinear OPF models.
Background: Currently, the identification of infectious disease re-emergence is performed without describing specific quantitative criteria that can be used to identify re-emergence events consistently. This practice may lead to ineffective mitigation. In addition, identification of factors contributing to local disease re-emergence and assessment of global disease re-emergence require access to data about disease incidence and a large number of factors at the local level for the entire world. This paper presents Re-emerging Disease Alert (RED Alert), a web-based tool designed to help public health officials detect and understand infectious disease re-emergence. Objective: Our objective is to bring together a variety of disease-related data and analytics needed to help public health analysts answer the following 3 primary questions for detecting and understanding disease re-emergence: Is there a potential disease re-emergence at the local (country) level? What are the potential contributing factors for this re-emergence? Is there a potential for global re-emergence? Methods: We collected and cleaned disease-related data (eg, case counts, vaccination rates, and indicators related to disease transmission) from several data sources including the World Health Organization (WHO), Pan American Health Organization (PAHO), World Bank, and Gideon. We combined these data with machine learning and visual analytics into a tool called RED Alert to detect re-emergence for the following 4 diseases: measles, cholera, dengue, and yellow fever. We evaluated the performance of the machine learning models for re-emergence detection and reviewed the output of the tool through a number of case studies. Results: Our supervised learning models were able to identify 82%-90% of the local re-emergence events, although with 18%-31% (except 46% for dengue) false positives. This is consistent with our goal of identifying all possible re-emergences while allowing some false positives. The review of the web-based tool through case studies showed that local re-emergence detection was possible and that the tool provided actionable information about potential factors contributing to the local disease re-emergence and trends in global disease re-emergence. Conclusions: To the best of our knowledge, this is the first tool that focuses specifically on disease re-emergence and addresses the important challenges mentioned above.
It is of imperative interests for regional transmission organizations (RTOs) to effectively extract daily load profiles at transmission buses, which remains a gap in existing technology paradigm. This digest proposes an explicit yet efficient linear estimator, to disaggregate metered load profiles at buses with significant behind-the-meter (BTM) solar generations in a data driven manner. The proposed estimator is based on utility zonal load profiles and proxy solar irradiance profiles, which in reality is the aggregated waveform at each transmission bus and equivalent to the mix of summed load profiles minus actual BTM solar generation. To overcome technical challenges in the lack of “ground truth” and validate the performance of supervised learning algorithms, we propose semi-supervised mechanisms with parameter tuning, and leverage the unique characteristics of zero-crossing points in BTM solar peaking behaviors.
The ROI-finder software is being developed for use by several Microscopy Group beamlines at Argonne National Laboratory, including 2-ID microprobes and 9-ID-B Bionanoprobe which use multi-scale scanning fluorescence microscopy to acquire elemental maps (multi-modal image data). Microscopy experiments require scan of samples at a coarse resolution followed by ROI identification using feature detection based on domain expertise. Finer resolution scans are then conducted based on identified ROI. The decision-making process based on domain expertise will be difficult to perform for faster data rates and much larger sampling volumes anticipated after APS-U necessitating the need for the ROI-finder software. The ROI- finder detects regions of interest through a continuous learning process, starting with a unsupervised representation learning and improving its recommendations through supervised learning and an interactive tool for user annotation. The scope of ongoing development efforts includes the integration of image registration module to correlate optical and X-ray images, extraction of feature morphology as well as elemental signatures in the image space and incorporation of beamtime streaming data by the scanning probe via EPICS.
In many non-canonical data science scenarios, obtaining, detecting, attributing, and annotating enough high-quality training data is the primary barrier to developing highly effective models. Moreover, in many problems that are not sufficiently defined or constrained, manually developing a training dataset can often overlook interesting phenomena that should be included. To this end, we have developed and demonstrated an iterative self-supervised learning procedure, whereby models are successfully trained and applied to new data to extract new training examples that are added to the corpus of training data. Successive generations of classifiers are then trained on this augmented corpus. Using low-frequency acoustic data collected by a network of infrasound sensors deployed around the High Flux Isotope Reactor and Radiochemical Engineering Development Center at Oak Ridge National Laboratory, we test the viability of our proposed approach to develop a powerful classifier with the goal of identifying vehicles from continuously streamed data and differentiating these from other sources of noise such as tools, people, airplanes, and wind. Using a small collection of exhaustively manually labeled data, we test several implementation details of the procedure and demonstrate its success regardless of the fidelity of the initial model used to seed the iterative procedure. Finally, we demonstrate the method’s ability to update a model to accommodate changes in the data-generating distribution encountered during long-term persistent data collection.
Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.
We introduce a physics guided data-driven method for image-based multi-material decomposition for dual-energy computed tomography (CT) scans. The method is demonstrated for CT scans of virtual human phantoms containing more than two types of tissues. The method is a physics-driven supervised learning technique. We take advantage of the mass attenuation coefficient of dense materials compared to that of muscle tissues to perform a preliminary extraction of the dense material from the images using unsupervised methods. We then perform supervised deep learning on the images processed by the extracted dense material to obtain the final multi-material tissue map. The method is demonstrated on simulated breast models with calcifications as the dense material placed amongst the muscle tissues. The physics-guided machine learning method accurately decomposes the various tissues from input images, achieving a normalized root-mean-squared error of 2.75%.
A method of supervised machine learning-based spectrum analysis information, using a neural network trained with spectrum information, to identify a specified feature of a given material, a system for supervised machine learning-based spectrum analysis, and a method of training a neural network to analyze spectrum data. The method of supervised machine learning-base spectrum analysis comprises inputting into the neural network spectrum data obtained from a sample of the given material; and the neural network processing the spectrum data, in accordance with the training of the neural network, and outputting one or more values for the specified feature of the sample of the material. In an embodiment, the training set of data includes x-ray absorption spectroscopy data for the given material. In an embodiment, the training set of data includes electron energy loss spectra (EELS) data.
A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.
Abstract Graph learning, when used as a semi-supervised learning (SSL) method, performs well for classification tasks with a low label rate. We provide a graph-based batch active learning pipeline for pixel/patch neighborhood multi- or hyperspectral image segmentation. Our batch active learning approach selects a collection of unlabeled pixels that satisfy a graph local maximum constraint for the active learning acquisition function that determines the relative importance of each pixel to the classification. This work builds on recent advances in the design of novel active learning acquisition functions (e.g., the Model Change approach in arXiv:2110.07739) while adding important further developments including patch-neighborhood image analysis and batch active learning methods to further increase the accuracy and greatly increase the computational efficiency of these methods. In addition to improvements in the accuracy, our approach can greatly reduce the number of labeled pixels needed to achieve the same level of the accuracy based on randomly selected labeled pixels.
Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.
In physical networks trained using supervised learning, physical parameters are adjusted to produce desired responses to inputs. An example is an electrical contrastive local learning network of nodes connected by edges that adjust their conductances during training. When an edge conductance changes, it upsets the current balance of every node. In response, physics adjusts the node voltages to minimize the dissipated power. Learning in these systems is therefore a coupled double-optimization process, in which the network descends both a cost landscape in the high-dimensional space of edge conductances and a physical landscape—the power dissipation—in the high-dimensional space of node voltages. Because of this coupling, the physical landscape of a trained network contains information about the learned task. Here, we derive a structure-function relation for trained tunable networks and demonstrate that all the physical information relevant to the trained input-output relation can be captured by a tuning susceptibility, an experimentally measurable quantity. We supplement our theoretical results with simulations to show that the tuning susceptibility is correlated with functional importance and that we can extract physical insight into how the system performs the task from the conductances of highly susceptible edges. Our analysis is general and can be applied directly to mechanical networks, such as networks trained for protein-inspired function such as allostery.