Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

Coreset Clustering on Small Quantum Computers

Many quantum algorithms for machine learning require access to classical data in superposition. However, for many natural data sets and algorithms, the overhead required to load the data set in superposition can erase any potential quantum speedup over classical algorithms. Recent work by Harrow introduces a new paradigm in hybrid quantum-classical computing to address this issue, relying on coresets to minimize the data loading overhead of quantum algorithms. We investigated using this paradigm to perform k-means clustering on near-term quantum computers, by casting it as a QAOA optimization instance over a small coreset. We used numerical simulations to compare the performance of this approach to classical k-means clustering. We were able to find data sets with which coresets work well relative to random sampling and where QAOA could potentially outperform standard k-means on a coreset. However, finding data sets where both coresets and QAOA work well—which is necessary for a quantum advantage over k-means on the entire data set—appears to be challenging.

42 ENGINEERING↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Characterization of Fuel Cladding Chemical Interaction on a High Burnup U-10Zr Metallic Fuel via Electron Energy Loss Spectroscopy Enhanced by Machine Learning

Fuel cladding chemical interaction (FCCI) is one of the main performance limiting factors for metallic nuclear fuels. The interaction destabilizes the martensitic microstructure and deteriorates mechanical properties of HT-9 cladding. The detection of low atomic number elements (Z<10) and overlapping of elemental peaks can be problematic in interpreting energy dispersive X-ray spectroscopy (EDS) data. Electron energy loss spectroscopy (EELS) provides precise elemental edge energy values and can detect elements with a low atomic number. This work utilizes EELS to study the distribution of lanthanides and light elements at the interaction region. The sample was prepared from the FCCI region of a U-10Zr (wt.%) solid fuel with HT-9 cladding, irradiated to a burnup of 13.2 at.%. Processing the EELS data included three major steps: 1) enhance the signal to noise ratio by denoising the spectrum with principal component analysis (PCA) method, removing background and performing deconvolution; 2) identify chemical elements with core energy loss edges; 3) confirm different phases using a popular machine learning method, K-means. This work presents qualitative assessment of lanthanides and light elements like carbon (C) and oxygen (O) enhanced by the application of machine learning algorithms. By comparing with EDS elemental maps, EELS provides higher resolution chemical maps, reveals the distribution of carbon at the interaction region supporting the formation of zirconium carbide, a rind-like microstructure feature that was proposed to mitigate the chemical interaction. Furthermore, the plasmon peak map was also found to indicate an energy shift associated with the formation of phases/compounds. K-means clustering method was used on the processed electron energy loss (EEL) spectrum to automatically reveal different phases. The resulting clustered maps from K-means clustering align well with elemental maps confirming certain phases, especially Fe-Ce and Zr-C, in the FCCI region.

EELS↗

Decay of enveloped SARS-CoV-2 and non-enveloped PMMoV RNA in raw sewage from university dormitories

Although severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) RNA has been frequently detected in sewage from many university dormitories to inform public health decisions during the COVID-19 pandemic, a clear understanding of SARS-CoV-2 RNA persistence in site-specific raw sewage is still lacking. To investigate the SARS-CoV-2 RNA persistence, a field trial was conducted in the University of Tennessee dormitories raw sewage, similar to municipal wastewater. The decay of enveloped SARS-CoV-2 RNA and non-enveloped Pepper mild mottle virus (PMMoV) RNA was investigated by reverse transcription-quantitative polymerase chain reaction (RT-qPCR) in raw sewage at 4°C and 20°C. Temperature, followed by the concentration level of SARS-CoV-2 RNA, was the most significant factors that influenced the first-order decay rate constants (k) of SARS-CoV-2 RNA. The mean k values of SARS-CoV-2 RNA were 0.094 day –1 at 4°C and 0.261 day –1 at 20°C. At high-, medium-, and low-concentration levels of SARS-CoV-2 RNA, the mean k values were 0.367, 0.169, and 0.091 day –1 , respectively. Furthermore, there was a statistical difference between the decay of enveloped SARS-CoV-2 and non-enveloped PMMoV RNA at different temperature conditions. The first decay rates for both temperatures were statistically comparable for SARS-CoV-2 RNA, which showed sensitivity to elevated temperatures but not for PMMoV RNA. This study provides evidence for the persistence of viral RNA in site-specific raw sewage at different temperature conditions and concentration levels.

59 BASIC BIOLOGICAL SCIENCES↗

High-Throughput Nanoindentation Mapping of Additively Manufactured T91 Steel

Here, this work aims to adapt nanoindentation mapping combined with a k-means algorithm as a high-throughput technique to study the nano-scale spatial changes in mechanical properties for a heterogeneous material. This technique can also classify the individual data points based on their properties. Hundreds to thousands of indents were performed on additively manufactured T91 at room temperature, 300°C, 400°C, and 500°C across a square area with a side length of 120 μm to 400 μm. From this data, the hardness and reduced modulus at each point could be calculated and mapped. Using k-means clustering, we were able to arrange the data into three or four clusters corresponding roughly to the ferritic and martensitic phases as well as one or two intermediate clusters sampling both the phases. The hardness of these two phases appears to be quite stable as a function of temperature. Nanoindentation mapping and the k-means algorithm can therefore be used to rapidly assess the feasibility of heterogeneous materials under extreme conditions, such as nuclear reactor steels.

36 MATERIALS SCIENCE↗

Machine Learning-Driven Quantification of CO2 Plume Dynamics at Illinois Basin Decatur Project Sites Using Microseismic Data

This study utilizes machine learning to quantify CO2 plume extents by analyzing microseismic data from the Illinois Basin Decatur Project (IBDP). Leveraging a unique dataset of well logs, microseismic records, and CO2 injection metrics, this work aims to predict the temporal evolution of subsurface CO2 saturation plumes. The findings illustrate that machine learning can predict plume dynamics, revealing vertical clustering of microseismic events over distinct time periods within certain proximities to the injection well, consistent with an invasion percolation model. The buoyant CO2 plume partially trapped within sandstone intervals periodically breaches localized barriers or baffles, which act as leaky seals and impede vertical migration until buoyancy overcomes gravity and capillary forces, leading to breakthroughs along vertical zones of weakness. Between different unsupervised clustering techniques, K-Means and DBSCAN were applied and analyzed in detail, where K-means outperformed DBSCAN in this specific study by indicating the combination of the highest Silhouette Score and the lowest Davies–Bouldin Index. The predictive capability of machine learning models in quantifying CO2 saturation plume extension is significant for real-time monitoring and management of CO2 sequestration sites. The models exhibit high accuracy, validated against physical models and injection data from the IBDP, reinforcing the viability of CO2 geological sequestration as a climate change mitigation strategy and enhancing advanced tools for safe management of these operations.

Iyegbekedo, Ikponmwosa↗

Machine learning to identify geologic factors associated with production in geothermal fields: a case-study using 3D geologic data, Brady geothermal field, Nevada

Abstract In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid-flow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fluid-flow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k -means clustering (NMF k ), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.

58 GEOSCIENCES↗

Multi-channel Imager Algorithm (MIA): A novel cloud-top phase classification algorithm

The current Geostationary Operational Environmental Satellites (GOES-16 and 17) cloud-top phase classification algorithm is based primarily on empirical thresholds at multiple wavelengths that have varying absorption capabilities for water and ice. The performance of current GOES-16 cloud-top phase product largely depends on the accuracy of the selection of reflectance ratios. Here this study aims at presenting a novel cloud-top phase classification algorithm (the Multi-channel Imager Algorithm, MIA) that provides a more judicious selection of relationships between channels using a supervised K-mean clustering method on multi-channel Red-Green-Blue images. The K-mean clustering method works analogously to how human eyes separate different colors in a microphysical color rendering set of satellite images, which differentiates water, ice and unclassified thin clouds. For water phase, cloud-top temperature information is used to further distinguish supercooled water. To evaluate the performance of the MIA, an extensive comparison with Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP), Moderate Resolution Imaging Spectroradiometer, and current GOES-16 cloud-top phase products is conducted, using CALIOP as the benchmark. Compared to the current GOES-16 cloud-top phase product, MIA demonstrates a substantial improvement in phase classification, where hit rate increases from 69% to 76% over the Continental United States and 58% to 66% over the full disk domain.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Enabled Quantitative Analysis of Optically Obscure Scratches on Nickel-Plated Additively Manufactured (AM) Samples

Additively manufactured metal components often have rough and uneven surfaces, necessitating post-processing and surface polishing. Hardness is a critical characteristic that affects overall component properties, including wear. This study employed K-means unsupervised machine learning to explore the relationship between the relative surface hardness and scratch width of electroless nickel plating on additively manufactured composite components. The Taguchi design of experiment (TDOE) L9 orthogonal array facilitated experimentation with various factors and levels. Initially, a digital light microscope was used for 3D surface mapping and scratch width quantification. However, the microscope struggled with the reflections from the shiny Ni-plating and scatter from small scratches. To overcome this, a scanning electron microscope (SEM) generated grayscale images and 3D height maps of the scratched Ni-plating, thus enabling the precise characterization of scratch widths. Optical identification of the scratch regions and quantification were accomplished using Python code with a K-means machine-learning clustering algorithm. The TDOE yielded distinct Ni-plating hardness levels for the nine samples, while an increased scratch force showed a non-linear impact on scratch widths. The enhanced surface quality resulting from Ni coatings will have significant implications in various industrial applications, and it will play a pivotal role in future metal and alloy surface engineering.

36 MATERIALS SCIENCE↗

Intelligent Energy Optimizer for Residential Buildings

Demand-side management in the buildings is essential for meeting grid flexibility needs in a highly renewable energy scenario. Appliance load monitoring helps decision making for demand-side management by providing the information on operation status/power consumption from different appliances in the buildings. Nonintrusive load monitoring (NILM) is an attractive option for appliance load monitoring using because it has lower cost for sensors and helps mitigate privacy concerns. In this study, the team used an event detection technique followed by two different methods for event classification. The results from k-means clustering showed that the events from a single appliance are often distributed in multiple clusters. Thus, the unsupervised method of NILM using k-means clustering used in this study was not very suitable for load disaggregation. The results from NILM showed that the F1 score for event classification was 0.77 for a heat pump water heater and very low for other appliances using the rule-based classification.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Clustering of Software Vulnerabilities using Self Organizing Map (SOM)

The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.

Panchal, Khyati↗

Spatio-Temporal Surrogates for Interaction of a Jet with High Explosives: Part II - Clustering Extremely High-Dimensional Grid-Based Data

Building an accurate surrogate model for the spatio-temporal outputs of a computer simulation is a challenging task. A simple approach to improve the accuracy of the surrogate is to cluster the outputs based on similarity and build a separate surrogate model for each cluster. This clustering is relatively straightforward when the output at each time step is of moderate size. However, when the spatial domain is represented by a large number of grid points, numbering in the millions, the clustering of the data becomes more challenging. In this report, we consider output data from simulations of a jet interacting with high explosives. These data are available on spatial domains of different sizes, at grid points that vary in their spatial coordinates, and in a format that distributes the output across multiple files at each time step of the simulation. We first describe how we bring these data into a consistent format prior to clustering. Borrowing the idea of random projections from data mining, we reduce the dimension of our data by a factor of thousand, making it possible to use the iterative k-means method for clustering. We show how we can use the randomness of both the random projections, and the choice of initial centroids in k-means clustering, to determine the number of clusters in our data set. Our approach makes clustering of extremely high dimensional data tractable, generating meaningful cluster assignments for our problem, despite the approximation introduced in the random projections.

97 MATHEMATICS AND COMPUTING↗