Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Cy-Phy ADS: Cyber Physical Anomaly Detection Framework for EV Charging Systems

Today’s large-scale Electric Vehicle (EV) infrastructures are heavily dependent on information communication technologies to maintain their operation and to support communication within sub-system components as well as the outside world. These technologies are vulnerable to various cyber and physical threats. Timely identification and mitigation of these threats are critical for improving human safety, avoiding economic losses, and preventing catastrophic system failures. By addressing this, our work presents a ResNet Autoencoder (AE) based Cyber-Physical Anomaly Detection System (Cy-Phy ADS) for detecting anomalies in EV Controller Area Network (CAN) protocol communication. It consists of four main components: Cyber-Physical Feature Extractor, ResNet AE-based Anomaly Detection Framework, Cyber-Physical Health Metric (CPHM), and Visualization Dashboard. The presented framework was trained and tested using CAN data collected from the EV charging system testbed at the Idaho National Laboratory. The presented Cy-Phy ADS compared against six widely used unsupervised anomaly detection algorithms: One Class Support Vector Machine (OCSVM), Variational Autoencoder (VAE), LSTM Autoencoder (LSTM AE), Isolation Forest (IForest), Principle Component Analysis (PCA) and Local Outlier Factor (LOF). Here the presented approach showed the highest accuracy among the compared methods. Further, the proposed approach showed comparable performance in terms of precision, F1, and False positive rate. It also showed the lowest training and inference time compared to the neural network-based baseline algorithms compared against with. Additionally, the Cy-Phy ADS has advantages such as unsupervised training, the ability to provide a holistic metric for system health characterization, and non-linear feature extraction.

99 GENERAL AND MISCELLANEOUS↗

Lowering of Tc in Van Der Waals Layered Materials Under In-Plane Strain

The dependence of electromechanical behavior on strain in ferroelectric materials can be leveraged as parameter to tune ferroelectric properties such as the Curie temperature. For van der Waals materials, a unique opportunity arises because of wrinkling, bubbling, and Moiré phenomena accessible due to structural properties inherent to the van der Waals gap. Here, we use piezoresponse force microscopy and unsupervised machine learning methods to gain insight into the ferroelectric properties of layered CuInP2S6 where local areas are strained in-plane due to a partial delamination, resulting in a topographic bubble feature. We observe significant differences between strained and unstrained areas in piezoresponse images as well as voltage spectroscopy, during which strained areas show a sigmoid-shaped response usually associated with the response measured around the Curie temperature, indicating a lowering of the Curie temperature under tensile strain. These results suggest that strain engineering might be used to further increase the functionality of CuInP2S6 through locally modifying ferroelectric properties on the micro- and nanoscale.

36 MATERIALS SCIENCE↗

REC protein family expansion by the emergence of a new signaling pathway

This report presents multi-genome evidence that REC protein family expansion occurs when the emergence of new pathways gives rise to functional discordance. Specificity between residues in REC domain containing response regulators with paired histidine kinases is under negative purifying selection, constrained by the presence of other bacterial two-component systems signaling cascades that share sequence and structural identity. Presuming that the two-component systems can evolve by neutral amino acid changes (neutral drift) when purifying evolutionary constraints are relaxed, how might the REC protein family expand by amino acid changes when these constraints remain intact? Using an unsupervised machine learning approach to observe the sequence landscape of REC domains across long phylogenetic distances, we find that within-gene recombination, a subcategory of gene conversion, switched the effector domain and, consequently, the regulatory context of a duplicated response regulator from transcriptional regulation by σ54 to that by σ70. We determined that the recombined response regulator diverged from its parent by episodic diversifying selection and neutral drift. Functional experiments of the parent of recombined response regulators in a model Pseudomonas putida KT2440 model system revealed that the parent and recombined response regulators sense and respond to different carboxylic acids. Finally, a residue-switching experiment using structural predictions and functional characterization suggests that the new residues in the recombined regulator could form a new interaction interface and mediate condition-specific phosphotransfer. Overall, our study finds that genetic perturbations can create conditions of functional discordance, whereby the REC protein family can evolve by episodic diversifying selection.

59 BASIC BIOLOGICAL SCIENCES↗

Generalized Canonical Polyadic Tensor Decomposition

Tensor decomposition is a fundamental unsupervised machine learning method in data science, with applications including network analysis and sensor data processing. This work develops a generalized canonical polyadic (GCP) low-rank tensor decomposition that allows other loss functions besides squared error. For instance, we can use logistic loss or Kullback--Leibler divergence, enabling tensor decomposition for binary or count data. We present a variety of statistically motivated loss functions for various scenarios. We provide a generalized framework for computing gradients and handling missing data that enables the use of standard optimization methods for fitting the model. Furthermore, we demonstrate the flexibility of the GCP decomposition on several real-world examples including interactions in a social network, neural activity in a mouse, and monthly rainfall measurements in India.

97 MATHEMATICS AND COMPUTING↗

Search for Beyond the Standard Model physics with anomaly detection in multilepton final states in pp collisions at s=13TeV with the ATLAS detector

A model-agnostic search for Beyond the Standard Model physics is presented, targeting final states with at least four light leptons (electrons or muons). The search regions are separated by event topology and unsupervised machine learning is used to identify anomalous events in the full 140 fb-1$$^{-1}$$ of proton–proton collision data collected with the ATLAS detector during Run 2. No significant excess above the Standard Model background expectation is observed. Model-agnostic limits are presented in each topology, along with limits on several benchmark models including vector-like leptons, wino-like charginos and neutralinos, or smuons. Limits are set on the flavourful vector-like lepton model for the first time.

Aad, G↗

Enhancing the hunt for new phenomena in dijet final states using anomaly detection filters at the high-luminosity large Hadron Collider

In the realm of dijet searches in high-energy physics, a significant challenge has emerged: with experiments producing more and more data, the traditional methods of using analytic functions to describe dijet mass spectra start to fail. Here, to address this, we suggest the application of an anomaly detection approach to eliminate less interesting background events based on event final states. This method not only bypasses the limitations of conventional background models but also significantly enhances our ability to detect potential signals of new physics. Through simulations that mimic the conditions of the upcoming high-luminosity large Hadron collider, we demonstrate the strength and efficiency of this approach in dealing with large data volumes. The integration of unsupervised machine learning into our experimental framework paves the way for a promising avenue to unveil hidden physics discoveries within the overwhelming influx of data.

47 OTHER INSTRUMENTATION↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

LCA

The Locally Competitive Algorithm (LCA) is a dynamical sparse solver that uses only local computations, allowing for massively parallel implementations on compatible neuromorphic architectures such as Intel's Loihi research chip. In this invention, we show how unsupervised dictionary learning with spiking LCA can be implemented on GPUs and Intel's Loihi research chip.

Parpart, Gavin↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Anomaly Detection and Identification Using a Leave-One-Variable-Out Method

At nuclear power plants (NPPs), anomaly detection and identification (i.e., determining the causes of anomalies) are important tasks for ensuring the safe and efficient operation of NPPs. These tasks are currently labor-intensive and costly, and are made more difficult by the size and complexity of NPP systems. An alternative approach to conducting these tasks is to automate them, such as via the reconstruction-based contribution method, which is a well-researched unsupervised machine learning method that uses a data-driven model of anomaly-free behavior to detect events and then identify each variable’s contributions to those events. The present effort developed a novel contribution approach that utilized a leave-one-variable-out (LOVO) model, with which each variable is predicted using all the other variables. The novelty lay in transforming this model into a reconstruction model and modifying the identification algorithm to work with the new reconstruction model. To evaluate this method in a controlled environment, a synthetic dataset based on spring-mass-damper (SMD) systems (commonly found in mechanical engineering references) was used, with known anomalies introduced into the system. The proposed method successfully detected the anomalies and afforded insights into their causes, thus enabling the appropriate identifications to be made.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Performance of Compact Pulsed Thermal Imaging System for In-Service Applications. Pulsed thermal tomography nondestructive examination of additively manufactured reactor materials and components

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of complex topology nuclear reactor parts from high-strength corrosion resistance alloys, such as stainless steel and Inconel. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which has the capability of melting metallic powder and net shaping the structures with relatively high precision. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures due to creep in high temperature nuclear reactor environment. Currently, there exist limited capabilities to evaluate actual AM structures nondestructively. Pulsed Thermography (PT) imaging provides a capability for non-destructive evaluation (NDE) of sub-surface defects in arbitrary size structures. The PT method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PT method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. The data cube of PT measurements consists of surface temperature taken at sequential time intervals T(x,y,t). Material defects can be detected either by analyzing the thermograms T(x,y,t) data cube, or by using thermal tomography (TT) algorithm to obtain 3D spatial reconstruction of thermal effusivity e(x,y,z). To reduce the cost and enable in-service NDE in spatially constrained environment, it is highly desirable to develop PT with compact and inexpensive IR camera. Following initial qualification of an AM component for deployment in a nuclear reactor, a compact PT system can also be used for in-service nondestructive evaluation (NDE) applications. However, data cube obtained with PT based on compact IR camera suffers from strong thermal noises and loss of features due to relatively low sampling rate. In this report we describe two unsupervised machine learning (ML) algorithms for enhancement of PT images obtained with compact IR camera. In one approach, we introduce Sparse Coding Discrete Cosine Transform (SC/DCT) algorithm to remove additive white Gaussian noise (AWGN) from spatial thermal effusivity reconstructions. In another approach we introduce a Spatial Temporal Denoised Thermal Source Separation (STDTSS) ML algorithm to process thermograms. The STDTSS algorithm consists of spatial and temporal denoising using Gaussian and Savitzky–Golay filtering, followed by the matrix decomposition using Principal Component Analysis (PCA), and Independent Component Analysis (ICA) to automatically detect flaws. In the work described in this report, we constructed a compact PT system using a relatively small and low-cost FLIR A65 camera, consisting on uncooled microbolometer detector. Performance of SC/DCT algorithm was demonstrated on enhancing TT images of Inconel 718 AM plate. Performance of the STDTSS methods was investigated using thermography data obtained from imaging stainless steel 316L specimens produced with LPBF method with imprinted calibrated porosity defects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Performance Validation of Pulsed Thermal Imaging System for In-Service Applications

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of complex topology nuclear reactor parts from high-strength corrosion resistance alloys, such as stainless steel and Inconel. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which has the capability of melting metallic powder and net shaping the structures with relatively high precision. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures in high temperature nuclear reactor environment. Currently, there exist limited capabilities to evaluate actual AM structures non-destructively. Pulsed Thermography Imaging (PTI) provides a capability for non-destructive evaluation (NDE) of subsurface defects in arbitrary size structures. The PTI method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PTI method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. Following initial qualification of an AM component for deployment in a nuclear reactor, a PTI system can also be used for in-service nondestructive evaluation (NDE) applications. In this report, we describe recent progress in enhancing PTI capabilities in detecting microscopic defects in metallic specimens. SS316 and IN718 specimens were developed with a pattern of subsurface calibrated flat bottom hole (FBH) defects with diameters from 500µm to 200µm. FBH’s were created with EDM (electron discharge machining) drill. PTI imaging data was processed Spatial Temporal Denoised Thermal Source Separation (STDTSS) unsupervised machine learning (ML) algorithm. We show that defects as small as 200µm in SS316 and IN718 can be detected with STDTSS algorithm. To the best of our knowledge, these are the smallest detected defects which are reported in literature.

42 ENGINEERING↗

Pulsed Thermal Tomography Nondestructive Examination of Additively Manufactured Reactor Materials and Components. Third Annual Progress Report

Additive manufacturing (AM) of high-strength corrosion resistance alloys for nuclear energy applications, such as stainless steel and Inconel, is currently based on laser powder bed fusion (LPBF) process. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Probability of crack initiation at the pore depends on size, shape, and orientation of the defect. Pulsed Infrared Thermography Imaging (PIT) provides a capability for non-destructive evaluation (NDE) of sub-surface defects in arbitrary size structures. The PIT method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PIT method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. Following initial qualification of an AM component for deployment in a nuclear reactor, a PIT system can also be used for in-service nondestructive evaluation (NDE) applications. In this report, we describe recent progress in enhancing PIT capabilities in detecting microscopic subsurface defects in metals, and classifying shapes and orientation of pores in thermal images. For detection of microscopic defects in PIT imaging data, we have developed Spatial Temporal Denoised Thermal Source Separation (STDTSS) unsupervised machine learning (ML) image processing algorithm. We show that flat bottom hole (FBH) defects as small as 200µm in SS316 and IN718 specimens, can be detected with STDTSS algorithm. To the best of our knowledge, these are the smallest detected defects which are reported in literature. For classification of defects shapes, we have previously developed thermal tomography (TT) algorithm to obtain depth reconstructions of material defects from data cube of sequentially recorded surface temperatures. However, interpretation of TT images is non-trivial because of blurring with increasing depth. To address this challenge, we have developed a deep learning convolutional neural network (CNN) to classify size and orientation subsurface defects in simulated TT images.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Detection of Anomalies in Gamma Background Radiation Data with K-Means and Self-Organizing Map Clustering Algorithms (Consortium on Nuclear Security Technologies (CONNECT) Q1 Report)

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. The challenge is that spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

61 RADIATION PROTECTION AND DOSIMETRY↗