Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fatigue Crack and Porosity Measurement in Composite Materials by Thermographic and Ultrasonic Methods

Many nondestructive methods exist for the detection of localized material anomalies in an otherwise good composite structure. The problem arises when the material system as a whole has degraded during service or was improperly manufactured. Porosity and intra-ply microcracking are two such conditions that in unlined composite pressure vessels can be very troublesome to detect and when linked through the thickness can be critical to mission success. These leak paths may lead to loss of pressure/propellant, increased risk of explosion and possible cryo-pumping. Research sought nondestructive methods for quantifying porosity and microcracking in composite tankage. Both thermographic and resonance ultrasound methods have been utilized with artificial neural network and statistical approaches to analyze the data. Resonant ultrasound spectroscopy provides measurements, which are sensitive to fine details in the materials character, such as micro-cracking and porosity. Here, the higher frequency (shorter wavelength) components of the signal train provide more significant interaction with the defects causing the spectral characteristics to shift toward lower amplitudes at the higher frequencies. As the density of the defects increases more interactions occur and more drastic amplitude changes are observed. From a thermal perspective, the higher the defect density the lower the through thickness thermal diffusivity will be. Utilizing a point heat source, and thermographically recording the heat profile with time, diffusivity calculations can be made which in turn can be related to the relative quality of the material. Preliminary experiments to verify the measurable effect on the resonance spectrum of the ultrasonic data to detect microcracking and for porosity detection thermographically are presented. Methods involving supervised and unsupervised artificial neural networks as well as other clustering algorithms are developed for signal identification.

Walker, James L.↗

Machine Learning for Biological Trajectory Classification Applications

Machine-learning techniques, including clustering algorithms, support vector machines and hidden Markov models, are applied to the task of classifying trajectories of moving keratocyte cells. The different algorithms axe compared to each other as well as to expert and non-expert test persons, using concepts from signal-detection theory. The algorithms performed very well as compared to humans, suggesting a robust tool for trajectory classification in biological applications.

Sbalzarini, Ivo F.↗

INDUCTIVE SYSTEM HEALTH MONITORING WITH STATISTICAL METRICS

Model-based reasoning is a powerful method for performing system monitoring and diagnosis. Building models for model-based reasoning is often a difficult and time consuming process. The Inductive Monitoring System (IMS) software was developed to provide a technique to automatically produce health monitoring knowledge bases for systems that are either difficult to model (simulate) with a computer or which require computer models that are too complex to use for real time monitoring. IMS processes nominal data sets collected either directly from the system or from simulations to build a knowledge base that can be used to detect anomalous behavior in the system. Machine learning and data mining techniques are used to characterize typical system behavior by extracting general classes of nominal data from archived data sets. In particular, a clustering algorithm forms groups of nominal values for sets of related parameters. This establishes constraints on those parameter values that should hold during nominal operation. During monitoring, IMS provides a statistically weighted measure of the deviation of current system behavior from the established normal baseline. If the deviation increases beyond the expected level, an anomaly is suspected, prompting further investigation by an operator or automated system. IMS has shown potential to be an effective, low cost technique to produce system monitoring capability for a variety of applications. We describe the training and system health monitoring techniques of IMS. We also present the application of IMS to a data set from the Space Shuttle Columbia STS-107 flight. IMS was able to detect an anomaly in the launch telemetry shortly after a foam impact damaged Columbia's thermal protection system.

Iverson, David L.↗

How Much Global Burned Area Can Be Forecast on Seasonal Time Scales Using Sea Surface Temperatures?

Large-scale sea surface temperature (SST) patterns influence the interannual variability of burned area in many regions by means of climate controls on fuel continuity, amount, and moisture content. Some of the variability in burned area is predictable on seasonal timescales because fuel characteristics respond to the cumulative effects of climate prior to the onset of the fire season. Here we systematically evaluated the degree to which annual burned area from the Global Fire Emissions Database version 4 with small fires (GFED4s) can be predicted using SSTs from 14 different ocean regions. We found that about 48 of global burned area can be forecast with a correlation coefficient that is significant at a p < 0.01 level using a single ocean climate index (OCI) 3 or more months prior to the month of peak burning. Continental regions where burned area had a higher degree of predictability included equatorial Asia, where 92% of the burned area exceeded the correlation threshold, and Central America, where 86% of the burned area exceeded this threshold. Pacific Ocean indices describing the El Nino-Southern Oscillation were more important than indices from other ocean basins, accounting for about 1/3 of the total predictable global burned area. A model that combined two indices from different oceans considerably improved model performance, suggesting that fires in many regions respond to forcing from more than one ocean basin. Using OCI-burned area relationships and a clustering algorithm, we identified 12 hotspot regions in which fires had a consistent response to SST patterns. Annual burned area in these regions can be predicted with moderate confidence levels, suggesting operational forecasts may be possible with the aim of improving ecosystem management.

Chen, Yang↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Supporting Responsible Machine Learning in Heliophysics

Over the last decade, Heliophysics researchers have increasingly adopted a variety of machine learning methods such as artificial neural networks, decision trees, and clustering algorithms into their workflow. Adoption of these advanced data science methods had quickly outpaced institutional response, but many professional organizations such as the European Commission, the National Aeronautics and Space Administration (NASA), and the American Geophysical Union have now issued (or will soon issue) standards for artificial intelligence and machine learning that will impact scientific research. These standards add further (necessary) burdens on the individual researcher who must now prepare the public release of data and code in addition to traditional paper writing. Support for these is not reflected in the current state of institutional support, community practices, or governance systems. We examine here some of these principles and how our institutions and community can promote their successful adoption within the Heliophysics discipline.

Machine learning↗

How Have Hydrological Extremes Changed Over the Past 20 Years?

Severe floods and droughts, including their back-to-back occurrences (weather whiplash), have been increasing in frequency and severity around the world. Improved understanding of systematic changes in hydrological extremes is essential for preparation and adaptation. In this study, we identified and quantified extreme wet and dry events globally by applying a clustering algorithm to terrestrial water storage (TWS) data from the Gravity Recovery and Climate Experiment (GRACE) and GRACE Follow-On (FO). The most intense events, ranked using an intensity metric, often reflect impacts of large-scale oceanic oscillations such as El Niño–Southern Oscillation and consequences of climate change. The severity of both wet and dry events, represented by standardized TWS anomalies, increased significantly in most cases, likely associated with intensification of wet and dry weather regimes in a warmer world, and consequently, exhibited strongest correlation with global temperature. In the Dry climate, the number of wet events decreased while the number of dry events increased significantly, suggesting a drying trend that may be attributed to climate variability and possible increases in irrigation and reliance on groundwater. In the Continental climate where temperature has risen faster than global average, dry events increased significantly. Characteristics of extreme events often showed strong correlations with global temperature, especially when averaged over all climates. These results suggest changes in hydrological extremes and underscore the importance of quantifying total water storage changes when studying hydrological extremes. Extending the GRACE/FO record, which spans 2002 to the present, is essential to continuously tracking changes in TWS and hydrological extremes.

Bailing Li↗

Multiyear Dry Periods in Southern Africa

Characteristics and physical features related to low precipitation across many years in Southern Africa that lead to societal disruptions are diagnosed using observed analyses and an ensemble of historical coupled climate model simulations during 1921 to 2014. Four regions are evaluated, as identified through a hierarchical clustering algorithm applied to the Standardized Precipitation Index (SPI) during the October–April precipitation season. Although dryness spanning many October–April occurs periodically in each region, they seldom occur simultaneously, consistent with largely insignificant SPI cross-correlations between them. However, characteristics relevant to low precipitation across many years are generalizable between the four regions, including the serial persistence of October–April precipitation, the likelihood of consecutive dry October–April, and the likelihood of dry October–April in temporal extents of up to 10 consecutive such 7-month seasons. Systematic precipitation persistence is not a feature in any of the four Southern Africa regions, as serial correlations of October–April SPI are not statistically significant at any time lags. It follows that there is an exponential-folding decay in the likelihood of consecutive October–April for various SPI thresholds and that there is a large spread in the likelihood of low October–April SPI across many years. In terms of physical features, low October–April SPI in each Southern Africa region is closely related to local atmospheric circulations; however, they are not as closely related to sea surface temperatures (SSTs). These results suggest that dryness spanning many years is determined primarily by persistent local circulations related to atmospheric variability and to a lesser extent variability related to SST anomalies, including the El Niño–Southern Oscillation.

subtropical Indian Ocean dipole↗

Characterizing the California Current System through Sea Surface Temperature and Salinity

Characterizing temperature and salinity (T-S) conditions is a standard framework in oceanography to identify and describe deep water masses and their dynamics. At the surface, this practice is hindered by multiple air–sea–land processes impacting T-S properties at shorter time scales than can easily be monitored. Now, however, the unsurpassed spatial and temporal coverage and resolution achieved with satellite sea surface temperature (SST) and salinity (SSS) allow us to use these variables to investigate the variability of surface processes at climate-relevant scales. In this work, we use SSS and SST data, aggregated into domains using a cluster algorithm over a T-S diagram, to describe the surface characteristics of the California Current System (CCS), validating them with in situ data from uncrewed Saildrone vessels. Despite biases and uncertainties in SSS and SST values in highly dynamic coastal areas, this T-S framework has proven useful in describing CCS regional surface properties and their variability in the past and in real time, at novel scales. This analysis also shows the capacity of remote sensing data for investigating variability in land–air–sea interactions not previously possible due to limited in situ data.

Marisol García-Reyes↗

Development of a Genetic Algorithm to Automate Clustering of a Dependency Structure Matrix

Much technology assessment and organization design data exists in Microsoft Excel spreadsheets. Tools are needed to put this data into a form that can be used by design managers to make design decisions. One need is to cluster data that is highly coupled. Tools such as the Dependency Structure Matrix (DSM) and a Genetic Algorithm (GA) can be of great benefit. However, no tool currently combines the DSM and a GA to solve the clustering problem. This paper describes a new software tool that interfaces a GA written as an Excel macro with a DSM in spreadsheet format. The results of several test cases are included to demonstrate how well this new tool works.

Rogers, James L.↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

DHARMA - Discriminant hyperplane abstracting residuals minimization algorithm for separating clusters with fuzzy boundaries

Learning of discriminant hyperplanes in imperfectly supervised or unsupervised training sample sets with unreliably labeled samples along the fuzzy joint boundaries between sample clusters is discussed, with the discriminant hyperplane designed to be a least-squares fit to the unreliably labeled data points. (Samples along the fuzzy boundary jump back and forth from one cluster to the other in recursive cluster stabilization and are considered unreliably labeled.) Minimization of the distances of these unreliably labeled samples from the hyperplanes does not sacrifice the ability to discriminate between classes represented by reliably labeled subsets of samples. An equivalent unconstrained linear inequality problem is formulated and algorithms for its solution are indicated. Landsat earth sensing data were used in confirming the validity and computational feasibility of the approach, which should be useful in deriving discriminant hyperplanes separating clusters with fuzzy boundaries, given supervised training sample sets with unreliably labeled boundary samples.

Dasarathy, B. V.↗

Filamentary galaxy clustering - A mapping algorithm

A simple and objective algorithm is presented which not only accurately identifies the filamentary structures in the Shane-Wirtanen galaxy count catalog, but also finds a set of visually less impressive filaments in a static hierarchical model of the clustering conducted by Soneira and Peebles (1978). The statistical properties of the elements in the model, while very similar to those in the data, show a significant excess of long and bright filaments in the data relative to the model. Two possible interpretations of these results are presented and discussed.

Gott, J. R., III↗

The void spectrum in two-dimensional numerical simulations of gravitational clustering

An algorithm for deriving a spectrum of void sizes from two-dimensional high-resolution numerical simulations of gravitational clustering is tested, and it is verified that it produces the correct results where those results can be anticipated. The method is used to study the growth of voids as clustering proceeds. It is found that the most stable indicator of the characteristic void 'size' in the simulations is the mean fractional area covered by voids of diameter d, in a density field smoothed at its correlation length. Very accurate scaling behavior is found in power-law numerical models as they evolve. Eventually, this scaling breaks down as the nonlinearity reaches larger scales. It is shown that this breakdown is a manifestation of the undesirable effect of boundary conditions on simulations, even with the very large dynamic range possible here. A simple criterion is suggested for deciding when simulations with modest large-scale power may systematically underestimate the frequency of larger voids.

Kauffmann, Guinevere↗

AHIMSA - Ad hoc histogram information measure sensing algorithm for feature selection in the context of histogram inspired clustering techniques

An algorithm is proposed for dimensionality reduction in the context of clustering techniques based on histogram analysis. The approach is based on an evaluation of the hills and valleys in the unidimensional histograms along the different features and provides an economical means of assessing the significance of the features in a nonparametric unsupervised data environment. The method has relevance to remote sensing applications.

Dasarathy, B. V.↗

An algorithm for spatial heirarchy clustering

A method for utilizing both spectral and spatial redundancy in compacting and preclassifying images is presented. In multispectral satellite images, a high correlation exists between neighboring image points which tend to occupy dense and restricted regions of the feature space. The image is divided into windows of the same size where the clustering is made. The classes obtained in several neighboring windows are clustered, and then again successively clustered until only one region corresponding to the whole image is obtained. By employing this algorithm only a few points are considered in each clustering, thus reducing computational effort. The method is illustrated as applied to LANDSAT images.

Dejesusparada, N.↗

An unsupervised classification approach for analysis of Landsat data to monitor land reclamation in Belmont county, Ohio

Two unsupervised classification procedures for analyzing Landsat data used to monitor land reclamation in a surface mining area in east central Ohio are compared for agreement with data collected from the corresponding locations on the ground. One procedure is based on a traditional unsupervised-clustering/maximum-likelihood algorithm sequence that assumes spectral groupings in the Landsat data in n-dimensional space; the other is based on a nontraditional unsupervised-clustering/canonical-transformation/clustering algorithm sequence that not only assumes spectral groupings in n-dimensional space but also includes an additional feature-extraction technique. It is found that the nontraditional procedure provides an appreciable improvement in spectral groupings and apparently increases the level of accuracy in the classification of land cover categories.

Brumfield, J. O.↗

Performance tests of signature extension algorithms

Comparative tests were performed on seven signature extension algorithms to evaluate their effectiveness in correcting for changes in atmospheric haze and sun angle in a LANDSAT scene. Four of the algorithms were cluster matching, and two were maximum likelihood algorithms. The seventh algorithm determined the haze level in both training and recognition segments and used a set of tables calculated from an atmospheric model to determine the affine transformation that corrects the training signatures for changes in sun angle and haze level. Three of the algorithms were tested on a simulated data set, and all of the algorithms were tested on consecutive-day data.

Abotteen, R. A.↗