Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Design of a Robust Memristive Spiking Neuromorphic System with Unsupervised Learning in Hardware

Spiking neural networks (SNN) offer a power efficient, biologically plausible learning paradigm by encoding information into spikes. The discovery of the memristor has accelerated the progress of spiking neuromorphic systems, as the intrinsic plasticity of the device makes it an ideal candidate to mimic a biological synapse. Despite providing a nanoscale form factor, non-volatility, and low-power operation, memristors suffer from device-level non-idealities, which impact system-level performance. To address these issues, this article presents a memristive crossbar-based neuromorphic system using unsupervised learning with twin-memristor synapses, fully digital pulse width modulated spike-timing-dependent plasticity, and homeostasis neurons. Additionally, the implemented single-layer SNN was applied to a pattern-recognition task of classifying handwritten-digits. The performance of the system was analyzed by varying design parameters such as number of training epochs, neurons, and capacitors. Furthermore, the impact of memristor device non-idealities, such as device-switching mismatch, aging, failure, and process variations, were investigated and the resilience of the proposed system was demonstrated.

97 MATHEMATICS AND COMPUTING↗

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗

Classification of Photovoltaic Failures with Hidden Markov Modeling, an Unsupervised Statistical Approach

Failure detection methods are of significant interest for photovoltaic (PV) site operators to help reduce gaps between expected and observed energy generation. Current approaches for field-based fault detection, however, rely on multiple data inputs and can suffer from interpretability issues. In contrast, this work offers an unsupervised statistical approach that leverages hidden Markov models (HMM) to identify failures occurring at PV sites. Using performance index data from 104 sites across the United States, individual PV-HMM models are trained and evaluated for failure detection and transition probabilities. This analysis indicates that the trained PV-HMM models have the highest probability of remaining in their current state (87.1% to 93.5%), whereas the transition probability from normal to failure (6.5%) is lower than the transition from failure to normal (12.9%) states. A comparison of these patterns using both threshold levels and operations and maintenance (O&M) tickets indicate high precision rates of PV-HMMs (median = 82.4%) across all of the sites. Although additional work is needed to assess sensitivities, the PV-HMM methodology demonstrates significant potential for real-time failure detection as well as extensions into predictive maintenance capabilities for PV.

classification↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Virtual Sensing with Unsupervised Image-to-Image Translation

Earth observing satellites carrying multi-spectral sensors are widely used to monitor the physical and biological states of the atmosphere, land, and oceans. These satellites have different vantage points above the earth and different spectral imaging bands resulting in inconsistent imagery from one to another. This presents challenges in building downstream applications. What if we could generate synthetic bands for existing satellites from the union of all domains? We tackle the problem of generating synthetic spectral imagery for multispectral sensors as an unsupervised image-to-image translation problem with partial labels and introduce a novel shared spectral reconstruction loss. Simulated experiments performed by dropping one or more spectral bands show that cross-domain reconstruction outperforms measurements obtained from a second vantage point. On a downstream cloud detection task, we show that generating synthetic bands with our model improves segmentation performance beyond our baseline. Our proposed approach enables synchronization of multispectral data and provides a basis for more homogeneous remote sensing datasets.

Geostationary satellites↗

Unsupervised Learning for Equitable DER Control: Preprint

In the context of managing distributed energy resources (DERs) within distribution networks (DNs), this work focuses on the task of developing local controllers. We propose an unsupervised learning framework to train functions that can closely approximate optimal power flow (OPF) solutions. The primary aim is to establish specific conditions under which these learned functions can collectively guide the network towards desired configurations asymptotically, leveraging an incremental control approach. The flexibility of the proposed methodology allows to integrate fairness-driven components into the cost function associated with the OPF problem. This addition seeks to mitigate power curtailment disparities among DERs, thereby promoting equitable power injections across the network. To demonstrate the effectiveness of the proposed approach, power flow simulations are conducted using the IEEE 37-bus feeder. The findings not only showcase the guaranteed system stability but also underscore its improved overall performance.

asymptotic stability↗

Challenges for unsupervised anomaly detection in particle physics

Anomaly detection relies on designing a score to determine whether a particular event is uncharacteristic of a given background distribution. One way to define a score is to use autoencoders, which rely on the ability to reconstruct certain types of data (background) but not others (signals). In this paper, we study some challenges associated with variational autoencoders, such as the dependence on hyperparameters and the metric used, in the context of anomalous signal (top and W) jets in a QCD background. We find that the hyperparameter choices strongly affect the network performance and that the optimal parameters for one signal are non-optimal for another. In exploring the networks, we uncover a connection between the latent space of a variational autoencoder trained using mean-squared-error and the optimal transport distances within the dataset. We then show that optimal transport distances to representative events in the background dataset can be used directly for anomaly detection, with performance comparable to the autoencoders. Whether using autoencoders or optimal transport distances for anomaly detection, we find that the choices that best represent the background are not necessarily best for signal identification. These challenges with unsupervised anomaly detection bolster the case for additional exploration of semi-supervised or alternative approaches.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Informed unsupervised machine learning analysis of dislocation microstructure from high-resolution differential aperture X-ray structural microscopy data

This study leverages high-resolution differential-aperture X-ray structural microscopy (DAXM) to probe the local dislocation structure in deformed 304L-stainless steel at small strain, by measuring the lattice rotation and deviatoric elastic strain with a sub-micron resolution. For a single grain in a polycrystalline specimen, the measured lattice rotation field over the measured volume exhibited a multimodal distribution while the deviatoric elastic strain showed a single-mode distribution. An unsupervised Cauchy mixture machine learning model was developed to resolve the multimodal distribution of the lattice rotation. By mapping the lattice rotation data associated with each Cauchy peak in the model back onto the measured volume, we identify contiguous regions of the crystal rotated near the average values corresponding to the peaks of the overall rotation distribution. These regions represent the grain subdivision in the microstructure. Finally, the dislocation density tensor was also computed and its norm was laid over the rotation field to detect the subgrain boundaries. This step provided a validation of the Cauchy mixture model for the analysis of the lattice rotation distribution. The current study highlights the integration of advanced X-ray microscopy techniques with data-driven analysis methods to uncover detailed microstructure scales in deformed crystals.

Machine learning; Lattice rotation; High-energy X-↗

Unsupervised Azimuth Estimation of Solar Arrays in Low-Resolution Satellite Imagery through Semantic Segmentation and Hough Transform

This paper explains the use of a convolutional neural network (CNN) to segment solar panels in a satellite image containing solar arrays, and extract associated metadata from the arrays. A novel unsupervised technique is introduced to estimate the azimuth of each individual solar panel from the predicted mask of the convolutional neural network. This pipeline was developed with the aim of extracting necessary metadata for a solar installation, using only a set of latitude–longitude coordinates. Azimuth prediction results for 669 individual solar installations associated with 387 sites located across the United States are provided. A mean average error and median average error of 21.65 degrees and 1.0 degrees were obtained, respectively, when predicting the azimuth of the solar fleet data set, with about 80% of the results within an error of zero degrees of the ground truth azimuth value and about 85% within an error of 25 degrees. The predicted azimuth was then used to estimate the energy conversion of the solar arrays. Results show a 90.9 and 90.6 R-squared value for estimating alternating current (AC) and direct current (DC) energy, respectively, and a mean absolute percentage error (MAPE) of 1.70% in estimating the alternating current (AC) energy using the fully automated algorithm.

14 SOLAR ENERGY↗

Unsupervised anomaly detection in MeV ultrafast electron diffraction

MeV ultrafast electron diffraction (MUED) is a pump-probe technique used to study the dynamic structural evolution of materials. An ultrashort laser pulse triggers structural changes, which are then probed by an ultrashort relativistic electron beam. To overcome low signal-to-noise ratios, diffraction patterns are averaged over thousands of shots. However, shot-to-shot instabilities in the electron beam can distort individual patterns, introducing uncertainty. Improving MUED accuracy requires detecting and removing these anomalous patterns from large datasets. In this work, we developed a fully unsupervised methodology for the detection of anomalous diffraction patterns. Using a convolutional autoencoder, we calculate the reconstruction mean squared error of the diffraction patterns. Based on the statistical analysis of this error, we provide the user an estimation of the probability that the pattern is normal, which also allows a posterior visual inspection of the images that are difficult to classify. This method has been trained with only 100 diffraction patterns and tested on 1521 patterns, resulting in a false positive rate between 0.2% and 0.4%, with a training time of 10 s per image and a test time of about 1 s per image. Here, the proposed methodology can also be applied to other diffraction techniques in which large datasets are collected that include faulty images due to instrumental instabilities.

43 PARTICLE ACCELERATORS↗

Discovery of peculiar radio morphologies with ASKAP using unsupervised machine learning

Abstract We present a set of peculiar radio sources detected using an unsupervised machine learning method. We use data from the Australian Square Kilometre Array Pathfinder (ASKAP) telescope to train a self-organizing map (SOM). The radio maps from three ASKAP surveys, Evolutionary Map of Universe pilot survey (EMU-PS), Deep Investigation of Neutral Gas Origins pilot survey (DINGO), and Survey With ASKAP of GAMA-09 + X-ray (SWAG-X), are used to search for the rarest or unknown radio morphologies. We use an extension of the SOM algorithm that implements rotation and flipping invariance on astronomical sources. The SOM is trained using the images of all ‘complex’ radio sources in the EMU-PS which we define as all sources catalogued as ‘multi-component’. The trained SOM is then used to estimate a similarity score for complex sources in all surveys. We select 0.5% of the sources that are most complex according to the similarity metric and visually examine them to find the rarest radio morphologies. Among these, we find two new odd radio circle (ORC) candidates and five other peculiar morphologies. We discuss multiwavelength properties and the optical/infrared counterparts of selected peculiar sources. In addition, we present examples of conventional radio morphologies including: diffuse emission from galaxy clusters, and resolved, bent-tailed, and FR-I and FR-II type radio galaxies. We discuss the overdense environment that may be the reason behind the circular shape of ORC candidates.

Astronomy & Astrophysics↗

Analysis of Interpretable Data Representations for 4D-STEM Using Unsupervised Learning

Abstract Understanding the structure of materials is crucial for engineering devices and materials with enhanced performance. Four-dimensional scanning transmission electron microscopy (4D-STEM) is capable of mapping nanometer-scale local crystallographic structure over micron-scale field of views. However, 4D-STEM datasets can contain tens of thousands of images from a wide variety of material structures, making it difficult to automate detection and classification of structures. Traditional automated analysis pipelines for 4D-STEM focus on supervised approaches, which require prior knowledge of the material structure and cannot describe anomalous or deviant structures. In this article, a pipeline for engineering 4D-STEM feature representations for unsupervised clustering using non-negative matrix factorization (NMF) is introduced. Each feature is evaluated using NMF and results are presented for both simulated and experimental data. It is shown that some data representations more reliably identify overlapping grains. Additionally, real space refinement is applied to identify spatially distinct sample regions, allowing for size and shape analysis to be performed. This work lays the foundation for improved analysis of nanoscale structural features in materials that deviate from expected crystallographic arrangement using 4D-STEM.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Multiscale Reactive Model for 1,3,5-Triamino-2,4,6-trinitrobenzene Inferred by Reactive MD Simulations and Unsupervised Learning

When high-energy-density materials are subjected to thermal or mechanical insults at extreme conditions (shock loading), a coupled response between the thermo-mechanical and chemical behaviors is systematically induced. Herein we develop a reaction model for the fast chemistry of 1,3,5-triamino-2,4,6-trinitrobenzene (TATB) at the mesoscopic scale, where the chemical behavior is determined by underlying microscopic reactive simulations. The slow carbon cluster formation is not discussed in the present work. All-atom reactive molecular dynamics (MD) simulations are performed with the ReaxFF potential, and a reduced-order chemical kinetics model for TATB is fitted to isothermal and adiabatic simulations of single crystal chemical decomposition. Unsupervised machine learning techniques based on non-negative matrix factorization are applied to MD trajectories to model the decomposition kinetics of TATB in terms of a four-component model. The associated heats of reaction are fit to the temperature evolution from adiabatic decomposition trajectories. Using a chemical species analysis, we show that non-negative matrix factorization captures the main chemical decomposition steps of TATB and provides an accurate estimation of their evolution with temperature. The final analytical formulation, coupled to a diffusion term, is incorporated into a continuum formalism, and simulation results are compared one-to-one against MD simulations of 1D reaction propagation along different crystallographic directions and with different initial temperatures. A good agreement is found for both the temporal and spatial evolution of the temperature field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

4D-STEM Coupled with Unsupervised Machine Learning to Reveal at Large-Scale the Microstructural Evolution in Li- and Mn-Rich Cathodes

Li- and Mn-rich (LMR) layered oxides are known to exhibit a thin surface reconstruction layer, which grows during electrochemical cycling in a manner that depends on exposed crystallographic facets, cycling conditions, and electrolyte chemistry. Direct characterization of this layer has traditionally relied on high-resolution electron microscopy, which is inherently limited to small fields of view. Here, we employ four-dimensional scanning transmission electron microscopy (4D-STEM) combined with unsupervised machine-learning clustering to quantitatively map phase distributions over large areas and track their evolution in LMR cathodes during electrochemical aging. Our results show that the surface reconstruction layer consists predominantly of a rocksalt phase, whose thickness varies across different facets following activation cycling and becomes substantially thicker and more uniform during calendar aging. In contrast, a spinel-like phase is observed within the particle bulk. Large-area phase mapping and correlative high-resolution imaging reveal that this spinel-like phase preferentially nucleates at bulk crystallographic defects, including boundaries between 60°-rotated layered domains and associated mixed-phase regions, rather than exclusively at the particle surface. Our findings establish a mechanistic distinction between surface-driven rocksalt formation and bulk-defect-mediated spinel nucleation while demonstrating the unique capability of 4D-STEM to provide statistically robust, mesoscale insight into complex phase-evolution processes in LMR cathodes.

4D-STEM↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

LHC physics dataset for unsupervised New Physics detection at 40 MHz

In the particle detectors at the Large Hadron Collider, hundreds of millions of proton-proton collisions are produced every second. If one could store the whole data stream produced in these collisions, tens of terabytes of data would be written to disk every second. The general-purpose experiments ATLAS and CMS reduce this overwhelming data volume to a sustainable level, by deciding in real-time whether each collision event should be kept for further analysis or be discarded. We introduce a dataset of proton collision events that emulates a typical data stream collected by such a real-time processing system, pre-filtered by requiring the presence of at least one electron or muon. This dataset could be used to develop novel event selection strategies and assess their sensitivity to new phenomena. In particular, we intend to stimulate a community-based effort towards the design of novel algorithms for performing unsupervised new physics detection, customized to fit the bandwidth, latency and computational resource constraints of the real-time event selection system of a typical particle detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗