Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Using Machine Learning to Infer Material Properties of Debris Fragments from X-ray Images in the DebriSat Project

The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.

Saik Anam Siam↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Integrated Land Suitability Assessment for Depots Siting in a Sustainable Biomass Supply Chain

A sustainable biomass supply chain would require not only an effective and fluid transportation system with a reduced carbon footprint and costs, but also good soil characteristics ensuring durable biomass feedstock presence. Unlike existing approaches that fail to account for ecological factors, this work integrates ecological as well as economic factors for developing sustainable supply chain development. For feedstock to be sustainably supplied, it necessitates adequate environmental conditions, which need to be captured in supply chain analysis. Using geospatial data and heuristics, we present an integrated framework that models biomass production suitability, capturing the economic aspect via transportation network analysis and the environmental aspect via ecological indicators. Production suitability is estimated using scores, considering both ecological factors and road transportation networks. These factors include land cover/crop rotation, slope, soil properties (productivity, soil texture, and erodibility factor) and water availability. This scoring determines the spatial distribution of depots with priority to fields scoring the highest. Two methods for depot selection are presented using graph theory and a clustering algorithm to benefit from contextualized insights from both and potentially gain a more comprehensive understanding of biomass supply chain designs. Graph theory, via the clustering coefficient, helps determine dense areas in the network and indicate the most appropriate location for a depot. Clustering algorithm, via K-means, helps form clusters and determine the depot location at the center of these clusters. An application of this innovative concept is performed on a case study in the US South Atlantic, in the Piedmont region, determining distance traveled and depot locations, with implications on supply chain design. The findings from this study show that a more decentralized depot-based supply chain design with 3depots, obtained using the graph theory method, can be more economical and environmentally friendly compared to a design obtained from the clustering algorithm method with 2 depots. In the former, the distance from fields to depots totals 801,031,476 miles, while in the latter, it adds up to 1,037,606,072 miles, which represents about 30% more distance covered for feedstock transportation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Signature extension through the application of cluster matching algorithms to determine appropriate signature transformations

Signature extension is intended to increase the space-time range over which a set of training statistics can be used to classify data without significant loss of recognition accuracy. A first cluster matching algorithm MASC (Multiplicative and Additive Signature Correction) was developed at the Environmental Research Institute of Michigan to test the concept of using associations between training and recognition area cluster statistics to define an average signature transformation. A more recent signature extension module CROP-A (Cluster Regression Ordered on Principal Axis) has shown evidence of making significant associations between training and recognition area cluster statistics, with the clusters to be matched being selected automatically by the algorithm.

Lambeck, P. F.↗

Automated thematic mapping and change detection of ERTS-A images

The author has identified the following significant results. In the first part of the investigation, spatial and spectral features were developed which were employed to automatically recognize terrain features through a clustering algorithm. In this part of the investigation, the size of the cell which is the number of digital picture elements used for computing the spatial and spectral features was varied. It was determined that the accuracy of terrain recognition decreases slowly as the cell size is reduced and coincides with increased cluster diffuseness. It was also proven that a cell size of 17 x 17 pixels when used with the clustering algorithm results in high recognition rates for major terrain classes. ERTS-1 data from five diverse geographic regions of the United States were processed through the clustering algorithm with 17 x 17 pixel cells. Simple land use maps were produced and the average terrain recognition accuracy was 82 percent.

Gramenopoulos, N.↗

Alpha Shape Analysis (ASA) Framework for Post- Clustering Property Determination in Atom Probe Tomographic Data

While application of clustering algorithms to atom probe tomography data have enabled quantification of solute clusters in terms of number density, size, and subcomposition there exist other properties (e.g., volume, surface area, and composition) that are better determined by defining an interface between the cluster and the surrounding matrix. The limitation in composition results from an ion selection step where the expected matrix ion types are omitted from the cluster search algorithm to enhance the contrast between the matrix and cluster and to reduce the complexity of the search. Previously, composition determination within solute clusters has utilized a secondary envelopment and erosion step on top of conventional methods such as maximum separation. In this work, we present a novel stochastic method that combines the particle identification fidelity of a conventional clustering algorithm with the analytical flexibility of mesh-based approaches through the generation of alpha shapes for each identified cluster. Furthermore, the corresponding mesh accounts for concave components of the clusters and determines the volume and surface area of the clusters; additionally, the mesh boundary is utilized to update the total composition according to the internal ions.

42 ENGINEERING↗

Metric DBSCAN

SAND2025-11725O Metric DBSCAN is an implementation of the popular DBSCAN clustering algorithm that works in general metric spaces. DBSCAN is a clustering algorithm, a fundamental building block in machine learning. It takes a set of objects and, given some notion of distance, identifies coherent groups of objects. With Metric DBSCAN, users can provide an arbitrary function to compute distance. Nearly all existing implementations of DBSCAN restrict distance to one of a few formulations. Metric DBScan accomplishes this cleanly and efficiently. The Python source code is on Github. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Automated clustering-based workload characterization

The demands placed on the mass storage systems at various federal agencies and national laboratories are continuously increasing in intensity. This forces system managers to constantly monitor the system, evaluate the demand placed on it, and tune it appropriately using either heuristics based on experience or analytic models. Performance models require an accurate workload characterization. This can be a laborious and time consuming process. It became evident from our experience that a tool is necessary to automate the workload characterization process. This paper presents the design and discusses the implementation of a tool for workload characterization of mass storage systems. The main features of the tool discussed here are: (1)Automatic support for peak-period determination. Histograms of system activity are generated and presented to the user for peak-period determination; (2) Automatic clustering analysis. The data collected from the mass storage system logs is clustered using clustering algorithms and tightness measures to limit the number of generated clusters; (3) Reporting of varied file statistics. The tool computes several statistics on file sizes such as average, standard deviation, minimum, maximum, frequency, as well as average transfer time. These statistics are given on a per cluster basis; (4) Portability. The tool can easily be used to characterize the workload in mass storage systems of different vendors. The user needs to specify through a simple log description language how the a specific log should be interpreted. The rest of this paper is organized as follows. Section two presents basic concepts in workload characterization as they apply to mass storage systems. Section three describes clustering algorithms and tightness measures. The following section presents the architecture of the tool. Section five presents some results of workload characterization using the tool.Finally, section six presents some concluding remarks.

Pentakalos, Odysseas I.↗

When is Constrained Clustering Beneficial, and Why?

Several researchers have shown that constraints can improve the results of a variety of clustering algorithms. However, there can be a large variation in this improvement, even for a fixed number of constraints for a given data set. We present the first attempt to provide insight into this phenomenon by characterizing two constraint set properties: informativeness and coherence. We show that these measures can help explain why some constraint sets are more beneficial to clustering algorithms than others. Since they can be computed prior to clustering, these measures can aid in deciding which constraints to use in practice.

clustering algorithms↗

A Dissimilarity Measure for Clustering High- and Infinite Dimensional Data that Satisfies the Triangle Inequality

The cosine or correlation measures of similarity used to cluster high dimensional data are interpreted as projections, and the orthogonal components are used to define a complementary dissimilarity measure to form a similarity-dissimilarity measure pair. Using a geometrical approach, a number of properties of this pair is established. This approach is also extended to general inner-product spaces of any dimension. These properties include the triangle inequality for the defined dissimilarity measure, error estimates for the triangle inequality and bounds on both measures that can be obtained with a few floating-point operations from previously computed values of the measures. The bounds and error estimates for the similarity and dissimilarity measures can be used to reduce the computational complexity of clustering algorithms and enhance their scalability, and the triangle inequality allows the design of clustering algorithms for high dimensional distributed data.

Socolovsky, Eduardo A.↗

Reducing Earth Topography Resolution for SMAP Mission Ground Tracks Using K-Means Clustering

The K-means clustering algorithm is used to reduce Earth topography resolution for the SMAP mission ground tracks. As SMAP propagates in orbit, knowledge of the radar antenna footprints on Earth is required for the antenna misalignment calibration. Each antenna footprint contains a latitude and longitude location pair on the Earth surface. There are 400 pairs in one data set for the calibration model. It is computationally expensive to calculate corresponding Earth elevation for these data pairs. Thus, the antenna footprint resolution is reduced. Similar topographical data pairs are grouped together with the K-means clustering algorithm. The resolution is reduced to the mean of each topographical cluster called the cluster centroid. The corresponding Earth elevation for each cluster centroid is assigned to the entire group. Results show that 400 data points are reduced to 60 while still maintaining algorithm performance and computational efficiency. In this work, sensitivity analysis is also performed to show a trade-off between algorithm performance versus computational efficiency as the number of cluster centroids and algorithm iterations are increased.

ground tracks↗

Robust clustering of the local Milky Way stellar kinematic substructures with Gaia eDR3

Understanding local stellar kinematic substructures in the solar neighbourhood helps build a complete picture of the formation of the Milky Way, as well as an empirical phase space distribution of dark matter that would inform detection experiments. We apply the clustering algorithm HDBSCAN on the Gaia early third data release to identify a list of stable clusters in velocity space and action-angle space by taking into account the measurement uncertainties and studying the stability of the clustering results. We find 1405 (497) stars in 23 (6) robust clusters in velocity space (action-angle space) that are consistently not associated with noise. We discuss the kinematic properties of these structures and study whether many of the small clusters belong to a similar larger cluster based on their chemical abundances. They are attributed to the known structures: the Gaia Sausage-Enceladus, the Helmi Stream, and globular cluster NGC 3201 are found in both spaces, while NGC 104 and the thick disc (Sequoia) are identified in velocity space (action-angle space). Although we do not identify any new structures, we find that the HDBSCAN member selection of already known structures is unstable to input kinematics of the stars when resampled within their uncertainties. We therefore present the stable subset of local kinematic structures, which are consistently identified by the clustering algorithm, and emphasize the need to take into account error propagation during both the manual and automated identification of stellar structures, both for existing ones as well as future discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

Automated thematic mapping and change detection of ERTS-A images

The author has identified the following significant results. A diffraction pattern analysis of MSS images led to the development of spatial signatures for farm land, urban areas and mountains. Four spatial features are employed to describe the spatial characteristics of image cells in the digital data. Three spectral features are combined with the spatial features to form a seven dimensional vector describing each cell. Then, the classification of the feature vectors is accomplished by using the maximum likelihood criterion. It was determined that the recognition accuracy with the maximum likelihood criterion depends on the statistics of the feature vectors. It was also determined that for a given geographic area the statistics of the classes remain invariable for a period of a month, but vary substantially between seasons. Three ERTS-1 images from the Phoenix, Arizona area were processed, and recognition rates between 85% and 100% were obtained for the terrain classes of desert, farms, mountains, and urban areas. To eliminate the need for training data, a new clustering algorithm has been developed. Seven ERTS-1 images from four test sites have been processed through the clustering algorithm, and high recognition rates have been achieved for all terrain classes.

Gramenopoulos, N.↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

Operational Dynamic Configuration Analysis

Sectors may combine or split within areas of specialization in response to changing traffic patterns. This method of managing capacity and controller workload could be made more flexible by dynamically modifying sector boundaries. Much work has been done on methods for dynamically creating new sector boundaries [1-5]. Many assessments of dynamic configuration methods assume the current day baseline configuration remains fixed [6-7]. A challenging question is how to select a dynamic configuration baseline to assess potential benefits of proposed dynamic configuration concepts. Bloem used operational sector reconfigurations as a baseline [8]. The main difficulty is that operational reconfiguration data is noisy. Reconfigurations often occur frequently to accommodate staff training or breaks, or to complete a more complicated reconfiguration through a rapid sequence of simpler reconfigurations. Gupta quantified a few aspects of airspace boundary changes from this data [9]. Most of these metrics are unique to sector combining operations and not applicable to more flexible dynamic configuration concepts. To better understand what sort of reconfigurations are acceptable or beneficial, more configuration change metrics should be developed and their distribution in current practice should be computed. This paper proposes a method to select a simple sequence of configurations among operational configurations to serve as a dynamic configuration baseline for future dynamic configuration concept assessments. New configuration change metrics are applied to the operational data to establish current day thresholds for these metrics. These thresholds are then corroborated, refined, or dismissed based on airspace practitioner feedback. The dynamic configuration baseline selection method uses a k-means clustering algorithm to select the sequence of configurations and trigger times from a given day of operational sector combination data. The clustering algorithm selects a simplified schedule containing k configurations based on stability score of the sector combinations among the raw operational configurations. In addition, the number of the selected configurations is determined based on balance between accuracy and assessment complexity.

Lai, Chok Fung↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Evaluation of large area crop estimation techniques

The performance of the USDAs EDITOR system is evaluated. The system processes Landsat imagery and estimates crop hectarage for large areas based on a regression estimator developed on a sample with known ground truth. It is found that use of multitemporal data over unitemporal significantly improves the hectarage estimates, and a 15% reduction in the r-squared of the regression occurs when independent and jackknifed test sets are used to evaluate the performance of the estimator. When an alternative clustering algorithm, CLASSY, is substituted for the current EDITOR method, estimator performance is improved with reduced need for analyst decisions. It is recommended that the CLASSY clustering algorithm and some form of jackknifing be implemented on EDITOR.

Amis, M. L.↗

Cluster Analysis of IRIS Spectroscopic Line Profiles and SDO/AIA EUV Emission in Observations and RMHD Simulations of the Solar Atmosphere

Spatially-resolved observations from the IRIS and SDO/AIA satellites, especially when coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar chromosphere, transition region, and corona. However, the complexity of the data makes understanding the observations and modeling results difficult. In this work, we apply unsupervised clustering algorithms for analysis of observational and synthetic chromospheric Mg II h&k 2796Å&2803Å and transition region C II 1334Å&1335Å line profiles observed by IRIS, and extreme ultraviolet (EUV) emission observed by SDO/AIA, for various types of problems. The synthetic line profiles are computed for simulations of the quiescent solar atmosphere (using the StellarBox and RH1.5 codes). The K-Means clustering algorithm is applied, and the selection of an optimal number of clusters is supported by the average silhouette width technique. We discuss applications of the line profile clustering method to 1) visualization of computational and observational spectroscopic imaging data; 2) understanding of evolutionary trends and behavior patterns of quiet Sun emission and during solar flares; and 3) recognition of heating events and shock waves.

Sadykov, Viacheslav↗