Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Rapid subsurface analysis of frequency-domain thermoreflectance images with K-means clustering

K-means clustering analysis is applied to frequency-domain thermoreflectance (FDTR) hyperspectral image data to rapidly screen the spatial distribution of thermophysical properties at material interfaces. Performing FDTR while raster scanning a sample consisting of 8.6 μm of doped-silicon (Si) bonded to a doped-Si substrate identifies spatial variation in the subsurface bond quality. Routine thermal analysis at select pixels quantifies this variation in bond quality and allows assignment of bonded, partially bonded, and unbonded regions. Performing this same routine thermal analysis across the entire map, however, becomes too computationally demanding for rapid screening of bond quality. To address this, K-means clustering was used to reduce the dimensionality of the dataset from more than 20 000 pixel spectra to just K = 3 component spectra. The three component spectra were then used to express every pixel in the image through a least-squares minimized linear combination providing continuous interpolation between the components across spatially varying features, e.g., bonded to unbonded transition regions. Fitting the component spectra to the thermal model, thermal properties for each K cluster are extracted and then distributed according to the weighting established by the regressed linear combination. Thermophysical property maps are then constructed and capture significant variation in bond quality over 25 μm length scales. The use of K-means clustering to achieve these thermal property maps results in a 74-fold speed improvement over explicit fitting of every pixel.

36 MATERIALS SCIENCE↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

Balanced k -means clustering on an adiabatic quantum computer

Adiabatic quantum computers are a promising platform for efficiently solving challenging optimization problems. Therefore, many are interested in using these computers to train computationally expensive machine learning models. We present a quantum approach to solving the balanced k-means clustering training problem on the D-Wave 2000Q adiabatic quantum computer. In order to do this, we formulate the training problem as a quadratic unconstrained binary optimization (QUBO) problem. Unlike existing classical algorithms, our QUBO formulation targets the global solution to the balanced k-means model. We test our approach on a number of small problems and observe that despite the theoretical benefits of the QUBO formulation, the clustering solution obtained by a modern quantum computer is usually inferior to the solution obtained by the best classical clustering algorithms. Nevertheless, the solutions provided by the quantum computer do exhibit some promising characteristics. We also perform a scalability study to estimate the run time of our approach on large problems using future quantum hardware. Finally, as a final proof of concept, we used the quantum approach to cluster random subsets of the Iris benchmark data set.

97 MATHEMATICS AND COMPUTING↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

Optimal Electrification Using Renewable Energies: Microgrid Installation Model with Combined Mixture k-Means Clustering Algorithm, Mixed Integer Linear Programming, and Onsset Method

Optimal planning and design of microgrids are priorities in the electrification of off-grid areas. Indeed, in one of the Sustainable Development Goals (SDG 7), the UN recommends universal access to electricity for all at the lowest cost. Several optimization methods with different strategies have been proposed in the literature as ways to achieve this goal. This paper proposes a microgrid installation and planning model based on a combination of several techniques. The programming language Python 3.10 was used in conjunction with machine learning techniques such as unsupervised learning based on K-means clustering and deterministic optimization methods based on mixed linear programming. These methods were complemented by the open-source spatial method for optimal electrification planning: onsset. Four levels of study were carried out. The first level consisted of simulating the model obtained with a cluster, which is considered based on the elbow and k-means clustering method as a case study. The second level involved sizing the microgrid with a capacity of 40 kW and optimizing all the resources available on site. The example of the different resources in the Togo case was considered. At the third level, the work consisted of proposing an optimal connection model for the microgrid based on voltage stability constraints and considering, above all, the capacity limit of the source substation. Finally, the fourth level involved a planning study of electrification strategies based mainly on microgrids according to the study scenario. The results of the first level of study enabled us to obtain an optimal location for the centroid of the cluster under consideration, according to the different load positions of this cluster. Then, the results of the second level of study were used to highlight the optimal resources obtained and proposed by the optimization model formulated based on the various technology costs, such as investment, maintenance, and operating costs, which were based on the technical limits of the various technologies. In these results, solar systems account for 80% of the maximum load considered, compared to 7.5% for wind systems and 12.5% for battery systems. Next, an optimal microgrid connection model was proposed based on the constraints of a voltage stability limit estimated to be 10% of the maximum voltage drop. The results obtained for the third level of study enabled us to present selective results for load nodes in relation to the source station node. Finally, the last results made it possible to plan electrification using different network technologies and systems in the short and long term. The case study of Togo was taken into account. The various results obtained from the different techniques provide the necessary leads for a feasibility study for optimal electrification of off-grid areas using microgrid systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data–Driven Velocity Model Evaluation Using K–Means Clustering

In this work, we develop a data-driven clustering method to evaluate a velocity model using surface wave velocity dispersion. This is done by first computing theoretical dispersion curves for 1-D velocity profiles of all the grid locations and then splitting the resulting dispersion curves into a certain number of groups via the K-means clustering. The observed dispersion curves are also clustered following the same procedure and the velocity model is assessed by comparing the spatial patterns obtained for the observed and synthetic data sets. The method is applied to evaluate two community velocity models in southern California, CVM-S4.26 and CVM-H15.1, using phase velocity maps derived for 3–16 s Rayleigh waves. We found a good correlation in the spatial distribution of clusters between the result of CVM-S4.26 and that of the observed data, suggesting that the CVM-S4.26 fits the observed dispersion maps better than the CVM-H15.1 in terms of features extracted from the clustering analysis.

58 GEOSCIENCES↗

Discovering hidden geothermal signatures using non-negative matrix factorization with customized k-means clustering

Discovery of hidden geothermal resources is challenging. It requires the mining of large datasets with diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used play fairway analysis approach typically incorporates subject-matter expertise to analyze regional data to estimate geothermal characteristics and favorability. We demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset from southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. Several of these systems are not well characterized because of insufficient existing data and limited past explorative work. This study discovers hidden patterns and relations in the SWNM geothermal dataset to improve our understanding of the regional hydrothermal conditions and energy-production favorability. This understanding is obtained by applying an unsupervised ML algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden signatures characterizing analyzed datasets, (2) the optimal number of these signatures, (3) the dominant data attributes associated with each signature, and (4) the spatial distribution of the extracted signatures. Here, in this study, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, and geothermal attributes at 44 locations in SWNM. Using NMFk, we find data patterns and identify the spatial associations of hydrothermal signatures within two physiographic provinces (Colorado Plateau and Basin and Range) and two sub-regions of these provinces (the Mogollon-Datil volcanic field and the Rio Grande rift) in SWNM. The ML algorithm extracted five hydrothermal signatures in the SWNM datasets that differentiate between low (<90°C) and medium (90-150°C)-temperature hydrothermal systems. The algorithm also suggests that the Rio Grande rift and northern Mogollon-Datil volcanic field are the most favorable regions for future geothermal resource discovery. NMFk also identified critical attributes to identify medium-temperature hydrothermal systems in the study area. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored regions. The code to execute the performed analyses as well as the corresponding data can be found at https://github.com/SmartTensors/GeoThermalCloud.jl.

15 GEOTHERMAL ENERGY↗

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Differentially Private K -Means Clustering Applied to Meter Data Analysis and Synthesis

The proliferation of smart meters has resulted in a large amount of data being generated. It is increasingly apparent that methods are required for allowing a variety of stakeholders to leverage the data in a manner that preserves the privacy of the consumers. The sector is scrambling to define policies, such as the so called ‘15/15 rule’, to respond to the need. However, the current policies fail to adequately guarantee privacy. Here, in this paper, we address the problem of allowing third parties to apply K-means clustering, obtaining customer labels and centroids for a set of load time series by applying the framework of differential privacy. We leverage the method to design an algorithm that generates differentially private synthetic load data consistent with the labeled data. We test our algorithm’s utility by answering summary statistics such as average daily load profiles for a 2-dimensional synthetic dataset and a real-world power load dataset.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Detection of Anomalies in Gamma Background Radiation Data with K-Means and Self-Organizing Map Clustering Algorithms (Consortium on Nuclear Security Technologies (CONNECT) Q1 Report)

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. The challenge is that spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

61 RADIATION PROTECTION AND DOSIMETRY↗

Presentation: Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning: An overview of methods and results

This presentation is a culmination of work which has occurred over the course of a 10-week internship. Anomaly detection methods must be both robust enough to detect subtle anomalies yet not so sensitive to report false positives, which would result significant loss of revenue. Methods currently being developed for autonomous systems are often pursuing a Digital Twin method, which will look at the entire system and model it as a whole. This presentation, however, focuses less on direct application to an NPP, rather acting as a proof of concept for the methods developed. For the project, we look to develop methods to analyze steady-state data and report anomalies.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Coreset Clustering on Small Quantum Computers

Many quantum algorithms for machine learning require access to classical data in superposition. However, for many natural data sets and algorithms, the overhead required to load the data set in superposition can erase any potential quantum speedup over classical algorithms. Recent work by Harrow introduces a new paradigm in hybrid quantum-classical computing to address this issue, relying on coresets to minimize the data loading overhead of quantum algorithms. We investigated using this paradigm to perform k-means clustering on near-term quantum computers, by casting it as a QAOA optimization instance over a small coreset. We used numerical simulations to compare the performance of this approach to classical k-means clustering. We were able to find data sets with which coresets work well relative to random sampling and where QAOA could potentially outperform standard k-means on a coreset. However, finding data sets where both coresets and QAOA work well—which is necessary for a quantum advantage over k-means on the entire data set—appears to be challenging.

42 ENGINEERING↗