Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

GraphEM

An unsupervised segmentation method for structural analysis of Scanning Transmission Electron Microscopy (STEM) images

Pope, Jenna↗

Subseasonal Representation and Predictability of North American Weather Regimes Using Cluster Analysis

Abstract This study focuses on assessing the representation and predictability of North American weather regimes, which are persistent large-scale atmospheric patterns, in a set of initialized subseasonal reforecasts created using the Community Earth System Model, version 2 (CESM2). The k -means clustering was used to extract four key North American (10°–70°N, 150°–40°W) weather regimes within ERA5 reanalysis, which were used to interpret CESM2 subseasonal forecast performance. Results show that CESM2 can recreate the climatology of the four main North American weather regimes with skill but exhibits biases during later lead times with overoccurrence of the West Coast high regime and underoccurrence of the Greenland high and Alaskan ridge regimes. Overall, the West Coast high and Pacific trough regimes exhibited higher predictability within CESM2, partly related to El Niño. Despite biases, several reforecasts were skillful and exhibited high predictability during later lead times, which could be partly attributed to skillful representation of the atmosphere from the tropics to extratropics upstream of North America. The high predictability at the subseasonal time scale of these case-study examples was manifested as an “ensemble realignment,” in which most ensemble members agreed on a prediction despite ensemble trajectory dispersion during earlier lead times. Weather regimes were also shown to project distinct temperature and precipitation anomalies across North America that largely agree with observational products. This study further demonstrates that unsupervised learning methods can be used to uncover sources and limits of subseasonal predictability, along with systematic biases present in numerical prediction systems. Significance Statement North American weather regimes are large-scale atmospheric patterns that can persist for several days. Their skillful subseasonal (2 weeks or greater) prediction can provide valuable lead time to prepare for temperature and precipitation anomalies that can stress energy and water resources. The purpose of this study was to assess the climatological representation and subseasonal predictability of four key North American weather regimes using a research subseasonal prediction system and clustering analysis. We found that the Pacific trough and West Coast high regimes exhibited higher predictability than other regimes and that skillful representation of conditions across the tropics and extratropics can increase predictability during later lead times. Future work will quantify causal pathways associated with high predictability.

58 GEOSCIENCES↗

Machine Learning for Geothermal Resource Exploration in the Tularosa Basin, New Mexico

Geothermal energy is considered an essential renewable resource to generate flexible electricity. Geothermal resource assessments conducted by the U.S. Geological Survey showed that the southwestern basins in the U.S. have a significant geothermal potential for meeting domestic electricity demand. Within these southwestern basins, play fairway analysis (PFA), funded by the U.S. Department of Energy’s (DOE) Geothermal Technologies Office, identified that the Tularosa Basin in New Mexico has significant geothermal potential. This short communication paper presents a machine learning (ML) methodology for curating and analyzing the PFA data from the DOE’s geothermal data repository. The proposed approach to identify potential geothermal sites in the Tularosa Basin is based on an unsupervised ML method called non-negative matrix factorization with custom k-means clustering. This methodology is available in our open-source ML framework, GeoThermalCloud (GTC). Using this GTC framework, we discover prospective geothermal locations and find key parameters defining these prospects. Our ML analysis found that these prospects are consistent with the existing Tularosa Basin’s PFA studies. This instills confidence in our GTC framework to accelerate geothermal exploration and resource development, which is generally time-consuming.

15 GEOTHERMAL ENERGY↗

Anomaly Detection and Identification Using a Leave-One-Variable-Out Method

At nuclear power plants (NPPs), anomaly detection and identification (i.e., determining the causes of anomalies) are important tasks for ensuring the safe and efficient operation of NPPs. These tasks are currently labor-intensive and costly, and are made more difficult by the size and complexity of NPP systems. An alternative approach to conducting these tasks is to automate them, such as via the reconstruction-based contribution method, which is a well-researched unsupervised machine learning method that uses a data-driven model of anomaly-free behavior to detect events and then identify each variable’s contributions to those events. The present effort developed a novel contribution approach that utilized a leave-one-variable-out (LOVO) model, with which each variable is predicted using all the other variables. The novelty lay in transforming this model into a reconstruction model and modifying the identification algorithm to work with the new reconstruction model. To evaluate this method in a controlled environment, a synthetic dataset based on spring-mass-damper (SMD) systems (commonly found in mechanical engineering references) was used, with known anomalies introduced into the system. The proposed method successfully detected the anomalies and afforded insights into their causes, thus enabling the appropriate identifications to be made.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Methods for rapid identification of anomalous layers in laser powder bed fusion

In situ process monitoring is a key requirement for increased industry acceptance of powder bed Additive Manufacturing. As sensing technologies increase in maturity, attention must also be given to effective data exploration techniques. These data are often high-resolution and multi-modal, with each build consisting of thousands of layers. Here, in this paper, the authors propose two methods enabling users to rapidly identify layers of interest within a build. Both methods leverage results from deep learning based segmentations of in situ powder bed images. The first method is an unsupervised “reverse layer search” algorithm while the second method uses supervised machine learning.

36 MATERIALS SCIENCE↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

Model-agnostic search for dijet resonances with anomalous jet substructure in proton–proton collisions at $\sqrt{s}$ = 13 TeV

This paper presents a model-agnostic search for narrow resonances in the dijet final state in the mass range 1.8-6 TeV. The signal is assumed to produce jets with substructure atypical of jets initiated by light quarks or gluons, with minimal additional assumptions. Search regions are obtained by utilizing multivariate machine-learning methods to select jets with anomalous substructure. A collection of complementary anomaly detection methods - based on unsupervised, weakly supervised, and semisupervised algorithms - are used in order to maximize the sensitivity to unknown new physics signatures. These algorithms are applied to data corresponding to an integrated luminosity of 138 fb -1 , recorded by the CMS experiment at the LHC, at a center-of-mass energy of 13 TeV. No significant excesses above background expectations are seen. Exclusion limits are derived on the production cross section of benchmark signal models varying in resonance mass, jet mass, and jet substructure. Many of these signatures have not been previously sought, making several of the limits reported on the corresponding benchmark models the first ever. When compared to benchmark inclusive and substructure-based search strategies, the anomaly detection methods are found to significantly enhance the sensitivity to a variety of models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Scoring of Morphological Changes in Images of Pentaerythritol Tetranitrate

Recent advances in characterization techniques that generate large datasets of material microstructure images require robust, automated image-processing. We applied an unsupervised anomaly detection method called feature anomaly detection system (FADS) to automatically detect and quantify microstructure changes in images of the explosive pentaerythritol tetranitrate (PETN) aged at various temperatures. We demonstrated the FADS approach on two-dimensional images extracted from computed tomography scans, but the same technique can be readily applied to other imaging modalities. FADS calculates anomaly scores on the basis of differences in filter activations of nominal and test data in pretrained convolutional neural networks. The FADS scores successfully differentiated between pristine PETN and PETN aged at a temperature where material coarsening occurred. Morphological metric analysis of segmented images verified observed trends in FADS scores as a function of aging temperature and aging time, specifically by calculating volume fractions, specific boundary lengths, two-point correlation functions, and local thicknesses. Here, the FADS technique has two important advantages compared to traditional morphological analysis: First, it uses grayscale images as input, rather than images that are segmented to separate the appropriate phases; and second, FADS scores capture any type of changes among image sets, rather than requiring prior knowledge or selection of a relevant set of metrics.

Accelerated aging↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

The mass profiles of dwarf galaxies from Dark Energy Survey lensing

We present a novel approach to extracting dwarf galaxies from photometric data to measure their average halo mass profile with weak lensing. We characterize their stellar mass and redshift distributions with a spectroscopic calibration sample. By combining the ${\sim} 5000\,\mathrm{deg}^2$ multiband photometry from the Dark Energy Survey and redshifts from the Satellites Around Galactic Analogs Survey with an unsupervised machine learning method, we select a low-mass galaxy sample spanning redshifts $z\lt 0.3$ and divide it into three mass bins. From low to high median mass, the bins contain [146 420, 330 146, 275 028] galaxies and have median stellar masses of $\log _{10}(M_*/\text{M}_\odot)=\left[8.52\substack{+0.57 -0.76},\, 9.02\substack{+0.50 -0.64},\, 9.49\substack{+0.50 -0.58}\right]$ . We measure the stacked excess surface mass density profiles, $\Delta \Sigma (R)$, of these galaxies using galaxy–galaxy lensing with a signal-to-noise ratio of [14, 23, 28]. Through a simulation-based forward-modelling approach, we fit the measurements to constrain the stellar-to-halo mass relation and find the median halo mass of these samples to be $\log _{10}(M_{\rm halo}/\text{M}_\odot)$ = [$10.67\substack{+0.2 -0.4}$, $11.01\substack{+0.14 -0.27}$, $11.40\substack{+0.08 -0.15}$]. The cold dark matter profiles are consistent with NFW (Navarro, Frenk, and White) profiles over scales ${\lesssim} 0.15 \, {h}^{-1}$ Mpc. We find that ${\sim} 20$ per cent of the dwarf galaxy sample are satellites. This is the first measurement of the halo profiles and masses of such a comprehensive, low-mass galaxy sample. The techniques presented here pave the way for extracting and analysing even lower mass dwarf galaxies and for more finely splitting galaxies by their properties with future photometric and spectroscopic survey data.

dark matter↗

Automatic Point Cloud Building Envelope Segmentation (AutoCuBES)

The Auto-CuBES algorithm is based on unsupervised machine learning that automatically labels 3D point cloud data and reduces the time spent in manual segmentation. The algorithm can process high-resolution point clouds and generate a wire-frame building envelope model with a small set of calibration parameters. The algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters.

Puente, BryanMaldonado↗

Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) using Machine Learning

Modern retrofit construction practices use 3D point cloud data of the building envelope to obtain the as-built dimensions. However, manual segmentation by a trained professional is required to identify and measure window openings, door openings, and other architectural features, making the use of 3D point clouds labor-intensive. In this study, the Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) algorithm is described, which can significantly reduce the time spent during point cloud segmentation. The Auto-CuBES algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters. Additionally, Auto-CuBES generates a heat map of each facade indicating non-planar characteristics that are crucial for the optimization of connections used in overclad envelope retrofits. With a scan resolution of 3 mm, the resulting window dimensions showed a mean absolute error of 4.2 mm compared to manual laser measurements.

Maldonado Puente, Bryan↗

Seismic Characterization of the Blue Mountain Geothermal Field

Subsurface characterization is crucial for geothermal energy exploration and production. Yet hydrothermal reservoirs usually reside in highly fractured and faulted zones where accurate characterization is very challenging because of low signal-to-noise ratios of land seismic data and lack of coherent reflection signals. We perform an active-source seismic characterization for the Blue Mountain geothermal field in Nevada using active seismic data to reveal the elastic medium property complexity and fault distribution at this field. We first employ an unsupervised machine learning method to attenuate groundroll and near-surface guided-wave noise and enhance coherent reflection and scattering signals from noisy seismic data. We then build a smooth initial P-wave velocity model based on an existing magnetotellurics survey result, and use 3D first-arrival traveltime tomography to refine the initial velocity model. We then derive a set of elastic wave velocities and anisotropic parameters using elastic full-waveform inversion, and obtain PP and PS images using elastic reverse-time migration. We identify major faults by analyzing the variations of seismic velocities and anisotropy parameters, and reveal mid- to small-scale faults by applying a supervised machine learning method to the seismic migration images. Our characterization reveals complex velocity heterogeneities and anisotropies, as well as faults, with a high spatial resolution. These results can provide valuable information for optimal placement of future injection and production wells to increase geothermal energy production at the Blue Mountain geothermal power plant.

58 GEOSCIENCES↗