Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Comparing weak- and unsupervised methods for resonant anomaly detection

Abstract Anomaly detection techniques are growing in importance at the Large Hadron Collider (LHC), motivated by the increasing need to search for new physics in a model-agnostic way. In this work, we provide a detailed comparative study between a well-studied unsupervised method called the autoencoder (AE) and a weakly-supervised approach based on the Classification Without Labels (CWoLa) technique. We examine the ability of the two methods to identify a new physics signal at different cross sections in a fully hadronic resonance search. By construction, the AE classification performance is independent of the amount of injected signal. In contrast, the CWoLa performance improves with increasing signal abundance. When integrating these approaches with a complete background estimate, we find that the two methods have complementary sensitivity. In particular, CWoLa is effective at finding diverse and moderately rare signals while the AE can provide sensitivity to very rare signals, but only with certain topologies. We therefore demonstrate that both techniques are complementary and can be used together for anomaly detection at the LHC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Merope: Unsupervised Learning for Network Operations (Merope) v1

The ability to analyze network flow characteristics is critical to understand network patterns that can help us understand how our networks are behaving. Merope is a collection of many unsupervised learning approaches and visualizers to allow networks to be more visual and analytical.

Kiran, Mariam↗

Unsupervised acoustic detection of fatigue-induced damage modes from wind turbine blades

This paper proposes a new in-situ damage detection approach for wind turbine blades, which leverages blade-internal non-stationary acoustic pressure fluctuations caused by the mechanical loading as the main source of excitation. This acoustic excitation was leveraged for the detection of fatigue-related damage modes on a full-scale wind turbine blade undergoing edgewise fatigue testing. An unsupervised, data-driven structural health monitoring strategy was developed to learn the normal cavity-internal acoustic sequences generated by the blade’s load cycles and to detect damage-related anomalies in the context of those sequences. A linear cepstral-coefficient based feature set was used to characterize the cavity-internal acoustics and LSTM-autoencoders were trained to accurately reconstruct healthy-case sequences. The reconstruction error was then used to characterize anomalous acoustic patterns within the blade cavity. The technique was able to detect a damage event earlier than a strain-based system by 120,000 load cycles.

17 WIND ENERGY↗

Identifying Vehicle Signals in Continuous Seismic Data Using Unsupervised Machine-Learning Techniques

Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Trajectory Optimization via Unsupervised Probabilistic Learning On Manifolds

This report investigates the use of unsupervised probabilistic learning techniques for the analysis of hypersonic trajectories. The algorithm first extracts the intrinsic structure in the data via a diffusion map approach. Using the diffusion coordinates on the graph of training samples, the probabilistic framework augments the original data with samples that are statistically consistent with the original set. The augmented samples are then used to construct conditional statistics that are ultimately assembled in a path-planing algorithm. In this framework the controls are determined stage by stage during the flight to adapt to changing mission objectives in real-time. A 3DOF model was employed to generate optimal hypersonic trajectories that comprise the training datasets. The diffusion map algorithm identfied that data resides on manifolds of much lower dimensionality compared to the high-dimensional state space that describes each trajectory. In addition to the path-planing worflow we also propose an algorithm that utilizes the diffusion map coordinates along the manifold to label and possibly remove outlier samples from the training data. This algorithm can be used to both identify edge cases for further analysis as well as to remove them from the training set to create a more robust set of samples to be used for the path-planing process.

42 ENGINEERING↗

Unsupervised Learning Based Interaction Force Model for Nonspherical Particles in Incompressible Flows

This project provides a neural network-based interaction force model for gas-solid flows from low to intermediate Reynolds numbers and concentration, which can be linked to MFiX-DEM. We have constructed a database of the interaction force between the irregular-shaped particles using a spherical harmonic method and the fluid phase based on the particle-resolved direct numerical simulation (PR-DNS) with immersed boundary-based gas kinetic scheme. Unsupervised learning method, i.e., variational auto-encoder (VAE) has been applied to extract the primitive shape factors determining the drag force, lifting forces, and torque. The interaction force model has been trained and validated with a simple but effective multi-layer feed-forward neural network: multi-layer perceptron (MLP), which will be concatenated after the encoder of the previously trained VAE for geometry feature extraction for single, irregular particles. We have trained transpose convolutional neural networks with the PR-DNS data to predict the velocity and pressure gradient of the single particle systems and utilized them to calculate drag force of multi-particle systems. This model can provide high computational efficiency because it does not require collecting multiparticle system data from PR-DNS.

99 GENERAL AND MISCELLANEOUS↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

Poster presented at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. This poster highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Uncertainty Quantification in CO2 Trapping Mechanisms: A Case Study of PUNQ-S3 Reservoir Model Using Representative Geological Realizations and Unsupervised Machine Learning

Evaluating uncertainty in CO2 injection projections often requires numerous high-resolution geological realizations (GRs) which, although effective, are computationally demanding. This study proposes the use of representative geological realizations (RGRs) as an efficient approach to capture the uncertainty range of the full set while reducing computational costs. A predetermined number of RGRs is selected using an integrated unsupervised machine learning (UML) framework, which includes Euclidean distance measurement, multidimensional scaling (MDS), and a deterministic K-means (DK-means) clustering algorithm. In the context of the intricate 3D aquifer CO2 storage model, PUNQ-S3, these algorithms are utilized. The UML methodology selects five RGRs from a pool of 25 possibilities (20% of the total), taking into account the reservoir quality index (RQI) as a static parameter of the reservoir. To determine the credibility of these RGRs, their simulation results are scrutinized through the application of the Kolmogorov–Smirnov (KS) test, which analyzes the distribution of the output. In this assessment, 40 CO2 injection wells cover the entire reservoir alongside the full set. The end-point simulation results indicate that the CO2 structural, residual, and solubility trapping within the RGRs and full set follow the same distribution. Simulating five RGRs alongside the full set of 25 GRs over 200 years, involving 10 years of CO2 injection, reveals consistently similar trapping distribution patterns, with an average value of Dmax of 0.21 remaining lower than Dcritical (0.66). Using this methodology, computational expenses related to scenario testing and development planning for CO2 storage reservoirs in the presence of geological uncertainties can be substantially reduced.

Mahjour, Seyed Kourosh↗

Reclassification of ASFV into 7 Biotypes Using Unsupervised Machine Learning

In 2007, an outbreak of African swine fever (ASF), a deadly disease of domestic swine and wild boar caused by the African swine fever virus (ASFV), occurred in Georgia and has since spread globally. Historically, ASFV was classified into 25 different genotypes. However, a newly proposed system recategorized all ASFV isolates into 6 genotypes exclusively using the predicted protein sequences of p72. However, ASFV has a large genome that encodes between 150–200 genes, and classifications using a single gene are insufficient and misleading, as strains encoding an identical p72 often have significant mutations in other areas of the genome. We present here a new classification of ASFV based on comparisons performed considering the entire encoded proteome. A curated database consisting of the protein sequences predicted to be encoded by 220 reannotated ASFV genomes was analyzed for similarity between homologous protein sequences. Weights were applied to the protein identity matrices and averaged to generate a genome-genome identity matrix that was then analyzed by an unsupervised machine learning algorithm, DBSCAN, to separate the genomes into distinct clusters. We conclude that all available ASFV genomes can be classified into 7 distinct biotypes.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Analysis of TRGBs (CATs) from Unsupervised, Multi-halo-field Measurements: Contrast is Key

The tip of the red giant branch (TRGB) is an apparent discontinuity of the luminosity function (LF) due to the end of the red giant evolutionary phase and is used to measure distances in the local universe. In practice, tip localization via edge detection response (EDR) relies on several methods applied on a case-by-case basis. It is hard to evaluate how individual choices affect a distance estimation using only a single host field while also avoiding confirmation bias. To devise a standardized approach, we compare unsupervised, algorithmic analyses of the TRGB in multiple halo fields per galaxy. We first optimize methods for the lowest field-to-field dispersion, including spatial filtering, smoothing, and weighting of LF, color band selection, and tip selection based on the number of likely RGB stars and the ratio of stars below versus above the tip (R). We find R, which we call the tip contrast, to be the most important indicator of the quality of EDR measurements; higher R selection can decrease field-to-field dispersion. Further, since R is found to correlate with the age or metallicity of the stellar population based on theoretical modeling, it might result in a displacement of the detected tip magnitude. We find a tip-contrast relation with a slope of -0.023 ± 0.0046 mag/ratio, an ~5σ result that can be used to correct these variations in the detections. When using TRGB to establish a distance ladder, consistent TRGB standardization using tip-contrast relation across rungs is vital to make robust cosmological measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Hunting for Polluted White Dwarfs and Other Treasures with Gaia XP Spectra and Unsupervised Machine Learning

White dwarfs (WDs) polluted by exoplanetary material provide the unprecedented opportunity to directly observe the interiors of exoplanets. However, spectroscopic surveys are often limited by brightness constraints, and WDs tend to be very faint, making detections of large populations of polluted WDs difficult. In this paper, we aim to increase considerably the number of WDs with multiple metals in their atmospheres. Using 96,134 WDs with Gaia DR3 BP/RP (XP) spectra, we constructed a 2D map using an unsupervised machine-learning technique called Uniform Manifold Approximation and Projection (UMAP) to organize the WDs into identifiable spectral regions. The polluted WDs are among the distinct spectral groups identified in our map. We have shown that this selection method could potentially increase the number of known WDs with five or more metal species in their atmospheres by an order of magnitude. Such systems are essential for characterizing exoplanet diversity and geology.

79 ASTRONOMY AND ASTROPHYSICS↗

Autoencoders on FPGAs for real-time, unsupervised new physics detection at 40 MHz at the Large Hadron Collider

In this paper, we show how to adapt and deploy anomaly detection algorithms based on deep autoencoders, for the unsupervised detection of new physics signatures in the extremely challenging environment of a real-time event selection system at the Large Hadron Collider (LHC). We demonstrate that new physics signatures can be enhanced by three orders of magnitude, while staying within the strict latency and resource constraints of a typical LHC event filtering system. This would allow for collecting datasets potentially enriched with high-purity contributions from new physics processes. Through per-layer, highly parallel implementations of network layers, support for autoencoder-specific losses on FPGAs and latent space based inference, we demonstrate that anomaly detection can be performed in as little as $80\,$ns using less than 3% of the logic resources in the Xilinx Virtex VU9P FPGA. Opening the way to real-life applications of this idea during the next data-taking campaign of the LHC.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Unsupervised Resource Allocation with Graph Neural Networks

We present an approach for maximizing a global utility function by learning how to allocate resources in an unsupervised way. We expect interactions between allocation targets to be important and therefore propose to learn the reward structure for near-optimal allocation policies with a GNN. By relaxing the resource constraint, we can employ gradient-based optimization in contrast to more standard evolutionary algorithms. Our algorithm is motivated by a problem in modern astronomy, where one needs to select-based on limited initial information-among $10^9$ galaxies those whose detailed measurement will lead to optimal inference of the composition of the universe. Our technique presents a way of flexibly learning an allocation strategy by only requiring forward simulators for the physics of interest and the measurement process. We anticipate that our technique will also find applications in a range of resource allocation problems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Unsupervised, supervised and reinforced learning via spiking computation

The present invention relates to unsupervised, supervised and reinforced learning via spiking computation. The neural network comprises a plurality of neural modules. Each neural module comprises multiple digital neurons such that each neuron in a neural module has a corresponding neuron in another neural module. An interconnection network comprising a plurality of edges interconnects the plurality of neural modules. Each edge interconnects a first neural module to a second neural module, and each edge comprises a weighted synaptic connection between every neuron in the first neural module and a corresponding neuron in the second neural module.

Modha, Dharmendra S.↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

This is the conference paper accompanying a poster presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. This work highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Unsupervised Clustering and Supervised Regression Learning to Select High Temperature Oxidation-Resistant Materials

High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.

Romedenne, Marie [ORNL] (ORCID:0000000317936561)↗