Toolpath Planning for Multiple Build Points using K-Means Clustering
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.
Many quantum algorithms for machine learning require access to classical data in superposition. However, for many natural data sets and algorithms, the overhead required to load the data set in superposition can erase any potential quantum speedup over classical algorithms. Recent work by Harrow introduces a new paradigm in hybrid quantum-classical computing to address this issue, relying on coresets to minimize the data loading overhead of quantum algorithms. We investigated using this paradigm to perform k-means clustering on near-term quantum computers, by casting it as a QAOA optimization instance over a small coreset. We used numerical simulations to compare the performance of this approach to classical k-means clustering. We were able to find data sets with which coresets work well relative to random sampling and where QAOA could potentially outperform standard k-means on a coreset. However, finding data sets where both coresets and QAOA work well—which is necessary for a quantum advantage over k-means on the entire data set—appears to be challenging.
In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.
Fuel cladding chemical interaction (FCCI) is one of the main performance limiting factors for metallic nuclear fuels. The interaction destabilizes the martensitic microstructure and deteriorates mechanical properties of HT-9 cladding. The detection of low atomic number elements (Z<10) and overlapping of elemental peaks can be problematic in interpreting energy dispersive X-ray spectroscopy (EDS) data. Electron energy loss spectroscopy (EELS) provides precise elemental edge energy values and can detect elements with a low atomic number. This work utilizes EELS to study the distribution of lanthanides and light elements at the interaction region. The sample was prepared from the FCCI region of a U-10Zr (wt.%) solid fuel with HT-9 cladding, irradiated to a burnup of 13.2 at.%. Processing the EELS data included three major steps: 1) enhance the signal to noise ratio by denoising the spectrum with principal component analysis (PCA) method, removing background and performing deconvolution; 2) identify chemical elements with core energy loss edges; 3) confirm different phases using a popular machine learning method, K-means. This work presents qualitative assessment of lanthanides and light elements like carbon (C) and oxygen (O) enhanced by the application of machine learning algorithms. By comparing with EDS elemental maps, EELS provides higher resolution chemical maps, reveals the distribution of carbon at the interaction region supporting the formation of zirconium carbide, a rind-like microstructure feature that was proposed to mitigate the chemical interaction. Furthermore, the plasmon peak map was also found to indicate an energy shift associated with the formation of phases/compounds. K-means clustering method was used on the processed electron energy loss (EEL) spectrum to automatically reveal different phases. The resulting clustered maps from K-means clustering align well with elemental maps confirming certain phases, especially Fe-Ce and Zr-C, in the FCCI region.
Although severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) RNA has been frequently detected in sewage from many university dormitories to inform public health decisions during the COVID-19 pandemic, a clear understanding of SARS-CoV-2 RNA persistence in site-specific raw sewage is still lacking. To investigate the SARS-CoV-2 RNA persistence, a field trial was conducted in the University of Tennessee dormitories raw sewage, similar to municipal wastewater. The decay of enveloped SARS-CoV-2 RNA and non-enveloped Pepper mild mottle virus (PMMoV) RNA was investigated by reverse transcription-quantitative polymerase chain reaction (RT-qPCR) in raw sewage at 4°C and 20°C. Temperature, followed by the concentration level of SARS-CoV-2 RNA, was the most significant factors that influenced the first-order decay rate constants (k) of SARS-CoV-2 RNA. The mean k values of SARS-CoV-2 RNA were 0.094 day –1 at 4°C and 0.261 day –1 at 20°C. At high-, medium-, and low-concentration levels of SARS-CoV-2 RNA, the mean k values were 0.367, 0.169, and 0.091 day –1 , respectively. Furthermore, there was a statistical difference between the decay of enveloped SARS-CoV-2 and non-enveloped PMMoV RNA at different temperature conditions. The first decay rates for both temperatures were statistically comparable for SARS-CoV-2 RNA, which showed sensitivity to elevated temperatures but not for PMMoV RNA. This study provides evidence for the persistence of viral RNA in site-specific raw sewage at different temperature conditions and concentration levels.
Here, this work aims to adapt nanoindentation mapping combined with a k-means algorithm as a high-throughput technique to study the nano-scale spatial changes in mechanical properties for a heterogeneous material. This technique can also classify the individual data points based on their properties. Hundreds to thousands of indents were performed on additively manufactured T91 at room temperature, 300°C, 400°C, and 500°C across a square area with a side length of 120 μm to 400 μm. From this data, the hardness and reduced modulus at each point could be calculated and mapped. Using k-means clustering, we were able to arrange the data into three or four clusters corresponding roughly to the ferritic and martensitic phases as well as one or two intermediate clusters sampling both the phases. The hardness of these two phases appears to be quite stable as a function of temperature. Nanoindentation mapping and the k-means algorithm can therefore be used to rapidly assess the feasibility of heterogeneous materials under extreme conditions, such as nuclear reactor steels.
This study utilizes machine learning to quantify CO2 plume extents by analyzing microseismic data from the Illinois Basin Decatur Project (IBDP). Leveraging a unique dataset of well logs, microseismic records, and CO2 injection metrics, this work aims to predict the temporal evolution of subsurface CO2 saturation plumes. The findings illustrate that machine learning can predict plume dynamics, revealing vertical clustering of microseismic events over distinct time periods within certain proximities to the injection well, consistent with an invasion percolation model. The buoyant CO2 plume partially trapped within sandstone intervals periodically breaches localized barriers or baffles, which act as leaky seals and impede vertical migration until buoyancy overcomes gravity and capillary forces, leading to breakthroughs along vertical zones of weakness. Between different unsupervised clustering techniques, K-Means and DBSCAN were applied and analyzed in detail, where K-means outperformed DBSCAN in this specific study by indicating the combination of the highest Silhouette Score and the lowest Davies–Bouldin Index. The predictive capability of machine learning models in quantifying CO2 saturation plume extension is significant for real-time monitoring and management of CO2 sequestration sites. The models exhibit high accuracy, validated against physical models and injection data from the IBDP, reinforcing the viability of CO2 geological sequestration as a climate change mitigation strategy and enhancing advanced tools for safe management of these operations.
Abstract In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid-flow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fluid-flow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k -means clustering (NMF k ), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.
The current Geostationary Operational Environmental Satellites (GOES-16 and 17) cloud-top phase classification algorithm is based primarily on empirical thresholds at multiple wavelengths that have varying absorption capabilities for water and ice. The performance of current GOES-16 cloud-top phase product largely depends on the accuracy of the selection of reflectance ratios. Here this study aims at presenting a novel cloud-top phase classification algorithm (the Multi-channel Imager Algorithm, MIA) that provides a more judicious selection of relationships between channels using a supervised K-mean clustering method on multi-channel Red-Green-Blue images. The K-mean clustering method works analogously to how human eyes separate different colors in a microphysical color rendering set of satellite images, which differentiates water, ice and unclassified thin clouds. For water phase, cloud-top temperature information is used to further distinguish supercooled water. To evaluate the performance of the MIA, an extensive comparison with Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP), Moderate Resolution Imaging Spectroradiometer, and current GOES-16 cloud-top phase products is conducted, using CALIOP as the benchmark. Compared to the current GOES-16 cloud-top phase product, MIA demonstrates a substantial improvement in phase classification, where hit rate increases from 69% to 76% over the Continental United States and 58% to 66% over the full disk domain.
Additively manufactured metal components often have rough and uneven surfaces, necessitating post-processing and surface polishing. Hardness is a critical characteristic that affects overall component properties, including wear. This study employed K-means unsupervised machine learning to explore the relationship between the relative surface hardness and scratch width of electroless nickel plating on additively manufactured composite components. The Taguchi design of experiment (TDOE) L9 orthogonal array facilitated experimentation with various factors and levels. Initially, a digital light microscope was used for 3D surface mapping and scratch width quantification. However, the microscope struggled with the reflections from the shiny Ni-plating and scatter from small scratches. To overcome this, a scanning electron microscope (SEM) generated grayscale images and 3D height maps of the scratched Ni-plating, thus enabling the precise characterization of scratch widths. Optical identification of the scratch regions and quantification were accomplished using Python code with a K-means machine-learning clustering algorithm. The TDOE yielded distinct Ni-plating hardness levels for the nine samples, while an increased scratch force showed a non-linear impact on scratch widths. The enhanced surface quality resulting from Ni coatings will have significant implications in various industrial applications, and it will play a pivotal role in future metal and alloy surface engineering.
Demand-side management in the buildings is essential for meeting grid flexibility needs in a highly renewable energy scenario. Appliance load monitoring helps decision making for demand-side management by providing the information on operation status/power consumption from different appliances in the buildings. Nonintrusive load monitoring (NILM) is an attractive option for appliance load monitoring using because it has lower cost for sensors and helps mitigate privacy concerns. In this study, the team used an event detection technique followed by two different methods for event classification. The results from k-means clustering showed that the events from a single appliance are often distributed in multiple clusters. Thus, the unsupervised method of NILM using k-means clustering used in this study was not very suitable for load disaggregation. The results from NILM showed that the F1 score for event classification was 0.77 for a heat pump water heater and very low for other appliances using the rule-based classification.
The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.
Building an accurate surrogate model for the spatio-temporal outputs of a computer simulation is a challenging task. A simple approach to improve the accuracy of the surrogate is to cluster the outputs based on similarity and build a separate surrogate model for each cluster. This clustering is relatively straightforward when the output at each time step is of moderate size. However, when the spatial domain is represented by a large number of grid points, numbering in the millions, the clustering of the data becomes more challenging. In this report, we consider output data from simulations of a jet interacting with high explosives. These data are available on spatial domains of different sizes, at grid points that vary in their spatial coordinates, and in a format that distributes the output across multiple files at each time step of the simulation. We first describe how we bring these data into a consistent format prior to clustering. Borrowing the idea of random projections from data mining, we reduce the dimension of our data by a factor of thousand, making it possible to use the iterative k-means method for clustering. We show how we can use the randomness of both the random projections, and the choice of initial centroids in k-means clustering, to determine the number of clusters in our data set. Our approach makes clustering of extremely high dimensional data tractable, generating meaningful cluster assignments for our problem, despite the approximation introduced in the random projections.
Light detection and ranging (LiDAR) measurements of isolated wakes generated by wind turbines installed at an onshore wind farm are leveraged to characterize the variability of the wake mean velocity and turbulence intensity during typical operations, which encompass a breadth of atmospheric stability regimes and rotor thrust coefficients. The LiDAR measurements are clustered through the k-means algorithm, which enables identifying the most representative realizations of wind turbine wakes while avoiding the imposition of thresholds for the various wind and turbine parameters. Considering the large number of LiDAR samples collected to probe the wake velocity field, the dimensionality of the experimental dataset is reduced by projecting the LiDAR data on an intelligently truncated basis obtained with the proper orthogonal decomposition (POD). The coefficients of only five physics-informed POD modes are then injected in the k-means algorithm for clustering the LiDAR dataset. The analysis of the clustered LiDAR data and the associated supervisory control and data acquisition and meteorological data enables the study of the variability of the wake velocity deficit, wake extent, and wake-added turbulence intensity for different thrust coefficients of the turbine rotor and regimes of atmospheric stability. Furthermore, the cluster analysis of the LiDAR data allows for the identification of systematic off-design operations with a certain yaw misalignment of the turbine rotor with the mean wind direction.
Quantum machine learning (QML) algorithms have obtained great relevance in the machine learning (ML) field due to the promise of quantum speedups when performing basic linear algebra subroutines (BLAS), a fundamental element in most ML algorithms. By making use of BLAS operations, we propose, implement and analyze a quantum k-means (qk-means) algorithm with a low time complexity of O(NKlog(D)I/C) to apply it to the fundamental problem of discriminating quantum states at readout. Discriminating quantum states allows the identification of quantum states |0⟩ and |1⟩ from low-level in-phase and quadrature signal (IQ) data, and can be done using custom ML models. In order to reduce dependency on a classical computer, we use the qk-means to perform state discrimination on the IBMQ Bogota device and managed to find assignment fidelities of up to 98.7% that were only marginally lower than that of the k-means algorithm. We also performed a cross-talk benchmark on the quantum device by applying both algorithms to perform state discrimination on a combination of quantum states and using Pearson Correlation coefficients and assignment fidelities of discrimination results to conclude on the presence of cross-talk on qubits. Evidence shows cross-talk in the (1, 2) and (2, 3) neighboring qubit couples for the analyzed device.
Abstract Training machine learning models on classical computers is usually a time and compute intensive process. With Moore’s law nearing its inevitable end and an ever-increasing demand for large-scale data analysis using machine learning, we must leverage non-conventional computing paradigms like quantum computing to train machine learning models efficiently. Adiabatic quantum computers can approximately solve NP-hard problems, such as the quadratic unconstrained binary optimization (QUBO), faster than classical computers. Since many machine learning problems are also NP-hard, we believe adiabatic quantum computers might be instrumental in training machine learning models efficiently in the post Moore’s law era. In order to solve problems on adiabatic quantum computers, they must be formulated as QUBO problems, which is very challenging. In this paper, we formulate the training problems of three machine learning models—linear regression, support vector machine (SVM) and balanced k-means clustering—as QUBO problems, making them conducive to be trained on adiabatic quantum computers. We also analyze the computational complexities of our formulations and compare them to corresponding state-of-the-art classical approaches. We show that the time and space complexities of our formulations are better (in case of SVM and balanced k-means clustering) or equivalent (in case of linear regression) to their classical counterparts.
Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.