Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

LANDSAT-4 image data quality analysis

Seven heterogeneous areas within the Des Moines, Iowa area test site were selected to define candidate spectral training classes using a clustering algorithm. In addition to the 91 cluster (nonsupervised) classes, three supervised training classes were defined and subsequently included in the training statistics file. The identity of all 94 candidate classes were determined using available reference data. Through analysis of the interclass separabilities, the original 94 candidate training classes were reduced to 42 spectrally separable final classes. The minimum average transformed divergence values for the 42 spectral classes and for the best subsets of TM spectral bands are shown in a table.

Anuta, P. E.↗

LANDSAT-4 MSS and TM Spectral Class Comparison and Coherent Noise Analysis

A detailed spectral analysis is conducted of thematic mapper and MSS data for an area near Des Moines, Iowa. Data are utilized from 7 blocks distributed throughout the area which included agricultural, forest, suburan, urban, and water scene types. The blocks are processed using a clustering algorithm to produce up to 18 cluster groupings for each block. Each cluster class is then identified with a ground-cover class using aerial photography and maps of the area. The clusters from each of the 7 blocks are inspected with regard to separability, mean, and variances. The separability measure used in the transformed divergence function or processor measures the statistical distance between classes based on class means and covariance matrices. The measure has a maximum value of 2,000 and the minimum of 0. Spectrally, very close classes will typically have values as low as 50 to 500.

Anuta, P. E.↗

Cooperative Clustering Techniques Applied to Contact Graph Routing

Routing in the space internet has to face many unique challenges - from unplanned disconnections and interruptions to predictable intermittent connectivity due to high network mobility and long propagation delays. NASA’s current approach to such routing is Contact Graph Routing (CGR), using a graph formed of prescheduled communication contacts to compute routes through the network. While this approach manages to tackle issues of connectivity and propagation delays, it is a global approach that requires continuous knowledge of the entire network. In a potential future Solar Space Internet (SSI) such an approach on its own cannot scale to large networks with thousands of members. In this presentation we propose clustering as a solution to CGR scalability. Clustering has been used in many networking problems as a way to subdivide the network and allow for localized routing and better scalability. Using techniques from graph theory and game theory, we explore various existing clustering algorithms and adapt them to the Contact Graph Routing setting. Finally, we propose a way to combine multiple algorithms to create a Delay Tolerant Clustering Protocol.

Yael Kirkpatrick↗

Cooperative Clustering Techniques For Space Network Scalability

Routing in the space internet must face many unique challenges - from unplanned disconnections and interruptions to predictable intermittent connectivity due to high network mobility and long propagation delays. NASA’s current approach to such routing is Contact Graph Routing (CGR), using a graph formed of prescheduled communication contacts to compute routes through the network. While this approach manages to tackle issues of connectivity and propagation delays, it is a global approach that requires continuous knowledge of the entire network. In a potential future Solar Space Internet (SSI) such an approach on its own cannot scale to large networks with thousands of members. In this paper we propose clustering as a solution to CGR scalability. Clustering has been used in many networking problems as a way to subdivide the network and allow for localized routing and better scalability. Using techniques from graph theory and game theory, we explore various existing clustering algorithms and adapt them to the Contact Graph Routing setting. We propose a way to combine multiple algorithms to create a Delay Tolerant Clustering Protocol (DTCP). In addition, we explore the underlying networking mechanisms such as multicast, neighbor discovery, and software defined networking that may be used to enable DTCP.

Delay Tolerant Networking↗

Real Time Intelligent Target Detection and Analysis with Machine Vision

We present an algorithm for detecting a specified set of targets for an Automatic Target Recognition (ATR) application. ATR involves processing images for detecting, classifying, and tracking targets embedded in a background scene. We address the problem of discriminating between targets and nontarget objects in a scene by evaluating 40x40 image blocks belonging to an image. Each image block is first projected onto a set of templates specifically designed to separate images of targets embedded in a typical background scene from those background images without targets. These filters are found using directed principal component analysis which maximally separates the two groups. The projected images are then clustered into one of n classes based on a minimum distance to a set of n cluster prototypes. These cluster prototypes have previously been identified using a modified clustering algorithm based on prior sensed data. Each projected image pattern is then fed into the associated cluster's trained neural network for classification. A detailed description of our algorithm will be given in this paper. We outline our methodology for designing the templates, describe our modified clustering algorithm, and provide details on the neural network classifiers. Evaluation of the overall algorithm demonstrates that our detection rates approach 96% with a false positive rate of less than 0.03%.

Howard, Ayanna↗

Determining the Number of Clusters in a Data Set Without Graphical Interpretation

Cluster analysis is a data mining technique that is meant ot simplify the process of classifying data points. The basic clustering process requires an input of data points and the number of clusters wanted. The clustering algorithm will then pick starting C points for the clusters, which can be either random spatial points or random data points. It then assigns each data point to the nearest C point where "nearest usually means Euclidean distance, but some algorithms use another criterion. The next step is determining whether the clustering arrangement this found is within a certain tolerance. If it falls within this tolerance, the process ends. Otherwise the C points are adjusted based on how many data points are in each cluster, and the steps repeat until the algorithm converges,

Aguirre, Nathan S.↗

Lightning Strike Distance Distribution Beyond a Preexisting Lightning Area

The 45th Weather Squadron (45 WS) asked the Applied Meteorology Unit (AMU) to review the 30-year-old, lightning stand-off distances of 5 nautical miles (nmi) for applicability to today's operations. This was based on the realization that previous lightning strike distance studies did not match how 45 WS issues lightning warnings (Roeder, 2008). The previous lightning distance studies were from the point of origin of the lightning or the average starting location that would tend to be in the core of the thunderstorm. However, the 45 WS issues lightning warnings based on the edge of a preexisting lightning area. Before beginning the AMU project, it took several years to develop a method to calculate a distance distribution beyond a preexisting area (Roeder, 2015). The AMU pulled Lightning Detection and Ranging (LDAR) sensor data from 1/1/2013 to 12/31/2013. This dataset consisted of 37 million individual source data points from the LDAR sensors. Only sources within 50 km north, south, east or west of the LDAR grid center were included in the dataset. This limited the use of LDAR data to that with the greatest accuracy of source detection and increased data processing. Points were grouped into flashes based on spatial and temporal criteria. Based on the sensitivity analysis the AMU performed on the flash clustering algorithm, a time value of 0.3 seconds was found to model flashes adequately. Distance parameters were tested from 1,500 to 7,500 meters (m) in 500 m increments. Distance parameters of both 3,000 m and 4,000 m produced results in the plotting tool that were most representative of the physical behavior of lightning. Thus statistics were gathered for the most representative of these spatial and temporal criteria on the flash size and the polygon expansion distance in order to find the correct data distributions. The best fit curves for the LDAR polygon expansion frequency vs. distances for both the 3 kilometer (km) and 4 km distance threshold values were exponential decay functions and had R2 values of > 0.998, indicating good model fits. The equations of the best fit curves were then used to calculate a desired safety radius of 4 nmi for either 3 km or 4 km distance threshold criteria. The AMU analysis concludes the safe reduction of the 5 nmi lightning warning circles to 4 nmi should improve the operational impact by 36% if based on distance from the center of the property area being protected. If based on the edge of the property being protected, then the reduction is 4.5 nmi to 4 nmi and the operational impact is decreased by 21%. For the 6 nmi lightning warning circles, the recommended 4 nmi stand-off distance will result in a safe reduction of operational impact of 31% if based on the center of the area being protected, or 16% if based on the edge of the property.

Flash clustering algorithm↗

Cluster analysis based on dimensional information with applications to feature selection and classification

A new clustering algorithm is presented that is based on dimensional information. The algorithm includes an inherent feature selection criterion, which is discussed. Further, a heuristic method for choosing the proper number of intervals for a frequency distribution histogram, a feature necessary for the algorithm, is presented. The algorithm, although usable as a stand-alone clustering technique, is then utilized as a global approximator. Local clustering techniques and configuration of a global-local scheme are discussed, and finally the complete global-local and feature selector configuration is shown in application to a real-time adaptive classification scheme for the analysis of remote sensed multispectral scanner data.

Eigen, D. J.↗

A comparison of unsupervised classification procedures on LANDSAT MSS data for an area of complex surface conditions in Basilicata, Southern Italy

Two unsupervised classification procedures were applied to ratioed and unratioed LANDSAT multispectral scanner data of an area of spatially complex vegetation and terrain. An objective accuracy assessment was undertaken on each classification and comparison was made of the classification accuracies. The two unsupervised procedures use the same clustering algorithm. By on procedure the entire area is clustered and by the other a representative sample of the area is clustered and the resulting statistics are extrapolated to the remaining area using a maximum likelihood classifier. Explanation is given of the major steps in the classification procedures including image preprocessing; classification; interpretation of cluster classes; and accuracy assessment. Of the four classifications undertaken, the monocluster block approach on the unratioed data gave the highest accuracy of 80% for five coarse cover classes. This accuracy was increased to 84% by applying a 3 x 3 contextual filter to the classified image. A detailed description and partial explanation is provided for the major misclassification. The classification of the unratioed data produced higher percentage accuracies than for the ratioed data and the monocluster block approach gave higher accuracies than clustering the entire area. The moncluster block approach was additionally the most economical in terms of computing time.

Justice, C.↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

Clustering with Missing Values: No Imputation Required

Clustering algorithms can identify groups in large data sets, such as star catalogs and hyperspectral images. In general, clustering methods cannot analyze items that have missing data values. Common solutions either fill in the missing values (imputation) or ignore the missing data (marginalization). Imputed values are treated as just as reliable as the truly observed data, but they are only as good as the assumptions used to create them. In contrast, we present a method for encoding partially observed features as a set of supplemental soft constraints and introduce the KSC algorithm, which incorporates constraints into the clustering process. In experiments on artificial data and data from the Sloan Digital Sky Survey, we show that soft constraints are an effective way to enable clustering with missing values.

constraints↗

Semi-Supervised Data Summarization: Using Spectral Libraries to Improve Hyperspectral Clustering

Hyperspectral imagers produce very large images, with each pixel recorded at hundreds or thousands of different wavelengths. The ability to automatically generate summaries of these data sets enables several important applications, such as quickly browsing through a large image repository or determining the best use of a limited bandwidth link (e.g., determining which images are most critical for full transmission). Clustering algorithms can be used to generate these summaries, but traditional clustering methods make decisions based only on the information contained in the data set. In contrast, we present a new method that additionally leverages existing spectral libraries to identify materials that are likely to be present in the image target area. We find that this approach simultaneously reduces runtime and produces summaries that are more relevant to science goals.

Wagstaff, K. L.↗

Optimization of a Non-traditional Unsupervised Classification Approach for Land Cover Analysis

The conditions under which a hybrid of clustering and canonical analysis for image classification produce optimum results were analyzed. The approach involves generation of classes by clustering for input to canonical analysis. The importance of the number of clusters input and the effect of other parameters of the clustering algorithm (ISOCLS) were examined. The approach derives its final result by clustering the canonically transformed data. Therefore the importance of number of clusters requested in this final stage was also examined. The effect of these variables were studied in terms of the average separability (as measured by transformed divergence) of the final clusters, the transformation matrices resulting from different numbers of input classes, and the accuracy of the final classifications. The research was performed with LANDSAT MSS data over the Hazleton/Berwick Pennsylvania area. Final classifications were compared pixel by pixel with an existing geographic information system to provide an indication of their accuracy.

Boyd, R. K.↗

Cloud-Precipitation Hybrid Regimes and their Projection onto IMERG Precipitation Data

We extend and enhance the concept of the Cloud Regimes (CRs) developed from two-dimensional joint histograms of cloud optical thickness and cloud top pressure from the Moderate Resolution Imaging Spectroradiometer (MODIS), by adding precipitation information in order to better understand cloud-precipitation relationships. Taking advantage of the high-resolution Integrated Multi-satellitE Retrievals for GPM (IMERG) precipitation dataset, cloud-precipitation “hybrid” regimes are derived by implementing the k-means clustering algorithm with advanced initialization and objective measures to determine the most optimal clusters. By expressing precipitation rates within 1-degree grid cell as histograms and making choices on the relative weight of cloud and precipitation, we could obtain several editions of hybrid cloud-precipitation regimes (CPRs), and examine their characteristics. In the deep tropics, when precipitation is weighted weakly, the cloud part of the hybrid entroids resembles the centroid of cloud-only regimes, but still tightens the cloud-precipitation relationship by decreasing the precipitation variability of each regime. As precipitation weight progressively increases, the shape of the cloudy part of the hybrid centroids becomes blunter, while the precipitation part of the centroids sharpens. In the case where cloud and precipitation are weighted equally, the CPRs representing high clouds with intermediate to heavy precipitation exhibit distinct features in the precipitation parts of the centroids, which allows us to project them onto the 30-minly IMERG domain. Such a projection can be used to overcome the temporal sparseness of MODIS cloud observations, which leads to great application potential for various convection-focused studies, including diurnal cycle analysis.

cloud-precipitation↗

The JSC clustering program ISOCLS and its applications

The clustering program ISOCLS developed at the Johnson Space Center, Houston, Texas, has been extensively used in the pattern analysis and classification of remote sensor data collected by aircraft and by the Earth Resources Technology Satellite ERTS-1. This paper discusses the theory behind this clustering algorithm. Several new ideas that have been incorporated in ISOCLS are discussed. Among these are the novel philosophy of operation behind the procedure, which assumes that a population (i.e., a class or a cluster) can be treated as the union of an appropriate number of subpopulations, and the termination of the clustering program by a 'chaining algorithm.' Finally, this paper reports the results of the application of ISOCLS to an investigation on rangeland vegetation mapping using ERTS-1 data.

Kan, E. P.↗

Investigation of correlation classification techniques

A two-step classification algorithm for processing multispectral scanner data was developed and tested. The first step is a single pass clustering algorithm that assigns each pixel, based on its spectral signature, to a particular cluster. The output of that step is a cluster tape in which a single integer is associated with each pixel. The cluster tape is used as the input to the second step, where ground truth information is used to classify each cluster using an iterative method of potentials. Once the clusters have been assigned to classes the cluster tape is read pixel-by-pixel and an output tape is produced in which each pixel is assigned to its proper class. In addition to the digital classification programs, a method of using correlation clustering to process multispectral scanner data in real time by means of an interactive color video display is also described.

Haskell, R. E.↗

A Census of Young Stellar Objects in Two Line-of-Sight Star-Forming Regions Toward IRAS 22147+5948 in the Outer Galaxy

Context. Star formation in the outer Galaxy, namely, outside of the Solar circle, has not been extensively studied in part due to the low CO brightness of the molecular clouds linked with the negative metallicity gradient. Recent infrared surveys provide an overview of dust emission in large sections of the Galaxy, but they suffer from cloud confusion and poor spatial resolution at far-infrared wavelengths. Aims. We aim to develop a methodology to identify and classify young stellar objects (YSOs) in star-forming regions in the outer Galaxy and use it to resolve a long-standing disparity in terms of the distance and evolutionary status of IRAS 22147+5948. Methods. We used a support vector machine learning algorithm to complement standard color–color and color–magnitude diagrams in our search for YSOs in the IRAS 22147 region, based on publicly available data from the Spitzer Mapping of the Outer Galaxy survey. The agglomerative hierarchical clustering algorithm was used to identify clusters. Then the physical properties of individual YSOs were calculated. The distances were determined using CO 1–0 from the Five College Radio Astronomy Observatory survey. Results. We identified 13 Class I and 13 Class II YSO candidates using the color–color diagrams, along with an additional 2 and 21 sources, respectively, using the applied machine learning techniques. The spectral energy distributions of 23 sources were modeled with a star and a passive disk, corresponding to Class II objects. The models of three sources include envelopes that are typical for Class I objects. The objects were grouped into two clusters located at a distance of 2:2 kpc and 5 clusters at 5:6 kpc. The spatial extent of CO, radio continuum, and dust emission confirms the origin of YSOs in two distinct star-forming regions along a similar line of sight. Conclusions. The outer Galaxy may serve as a unique laboratory for exploring star formation across environments, on the condition that complementary methods and ancillary data are used to properly account for cloud confusion and distance uncertainties.

Agata Karska↗

Investigation of the application of remote sensing technology to environmental monitoring

Activities and results are reported of a project to investigate the application of remote sensing technology developed for the LACIE, AgRISTARS, Forestry and other NASA remote sensing projects for the environmental monitoring of strip mining, industrial pollution, and acid rain. Following a remote sensing workshop for EPA personnel, the EOD clustering algorithm CLASSY was selected for evaluation by EPA as a possible candidate technology. LANDSAT data acquired for a North Dakota test sight was clustered in order to compare CLASSY with other algorithms.

Rader, M. L.↗