Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Failure Mode Identification Through Clustering Analysis

Research has shown that nearly 80% of the costs and problems are created in product development and that cost and quality are essentially designed into products in the conceptual stage. Currently, failure identification procedures (such as FMEA (Failure Modes and Effects Analysis), FMECA (Failure Modes, Effects and Criticality Analysis) and FTA (Fault Tree Analysis)) and design of experiments are being used for quality control and for the detection of potential failure modes during the detail design stage or post-product launch. Though all of these methods have their own advantages, they do not give information as to what are the predominant failures that a designer should focus on while designing a product. This work uses a functional approach to identify failure modes, which hypothesizes that similarities exist between different failure modes based on the functionality of the product/component. In this paper, a statistical clustering procedure is proposed to retrieve information on the set of predominant failures that a function experiences. The various stages of the methodology are illustrated using a hypothetical design example.

Arunajadai, Srikesh G.↗

Unsupervised classification of remote multispectral sensing data

The new unsupervised classification technique for classifying multispectral remote sensing data which can be either from the multispectral scanner or digitized color-separation aerial photographs consists of two parts: (a) a sequential statistical clustering which is a one-pass sequential variance analysis and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. Applications of the technique using an IBM-7094 computer on multispectral data sets over Purdue's Flight Line C-1 and the Yellowstone National Park test site have been accomplished. Comparisons between the classification maps by the unsupervised technique and the supervised maximum liklihood technique indicate that the classification accuracies are in agreement.

Su, M. Y.↗

Nineteen hundred seventy three significant accomplishments

Data collected by the Skylab remote sensing satellites was used to develop applications techniques and to combine automatic data classification with statistical clustering methods. Continuing research was concentrated in the correlation and registration of data products and in the definition of the atmospheric effects on remote sensing. The causes of errors encountered in the automated classification of agricultural data are identified. Other applications in forestry, geography, environmental geology, and land use are discussed.

Source record↗

Analysis of the Tanana River Basin using LANDSAT data

Digital image classification techniques were used to classify land cover/resource information in the Tanana River Basin of Alaska. Portions of four scenes of LANDSAT digital data were analyzed using computer systems at Ames Research Center in an unsupervised approach to derive cluster statistics. The spectral classes were identified using the IDIMS display and color infrared photography. Classification errors were corrected using stratification procedures. The classification scheme resulted in the following eleven categories; sedimented/shallow water, clear/deep water, coniferous forest, mixed forest, deciduous forest, shrub and grass, bog, alpine tundra, barrens, snow and ice, and cultural features. Color coded maps and acreage summaries of the major land cover categories were generated for selected USGS quadrangles (1:250,000) which lie within the drainage basin. The project was completed within six months.

Morrissey, L. A.↗

MEASURE: An integrated data-analysis and model identification facility

The first phase of the development of MEASURE, an integrated data analysis and model identification facility is described. The facility takes system activity data as input and produces as output representative behavioral models of the system in near real time. In addition a wide range of statistical characteristics of the measured system are also available. The usage of the system is illustrated on data collected via software instrumentation of a network of SUN workstations at the University of Illinois. Initially, statistical clustering is used to identify high density regions of resource-usage in a given environment. The identified regions form the states for building a state-transition model to evaluate system and program performance in real time. The model is then solved to obtain useful parameters such as the response-time distribution and the mean waiting time in each state. A graphical interface which displays the identified models and their characteristics (with real time updates) was also developed. The results provide an understanding of the resource-usage in the system under various workload conditions. This work is targeted for a testbed of UNIX workstations with the initial phase ported to SUN workstations on the NASA, Ames Research Center Advanced Automation Testbed.

Singh, Jaidip↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

An Artificial Intelligence Classification Tool and Its Application to Gamma-Ray Bursts

Despite being the most energetic phenomenon in the known universe, the astrophysics of gamma-ray bursts (GRBs) has still proven difficult to understand. It has only been within the past five years that the GRB distance scale has been firmly established, on the basis of a few dozen bursts with x-ray, optical, and radio afterglows. The afterglows indicate source redshifts of z=1 to z=5, total energy outputs of roughly 10(exp 52) ergs, and energy confined to the far x-ray to near gamma-ray regime of the electromagnetic spectrum. The multi-wavelength afterglow observations have thus far provided more insight on the nature of the GRB mechanism than the GRB observations; far more papers have been written about the few observed gamma-ray burst afterglows in the past few years than about the thousands of detected gamma-ray bursts. One reason the GRB central engine is still so poorly understood is that GRBs have complex, overlapping characteristics that do not appear to be produced by one homogeneous process. At least two subclasses have been found on the basis of duration, spectral hardness, and fluence (time integrated flux); Class 1 bursts are softer, longer, and brighter than Class 2 bursts (with two second durations indicating a rough division). A third GRB subclass, overlapping the other two, has been identified using statistical clustering techniques; Class 3 bursts are intermediate between Class 1 and Class 2 bursts in brightness and duration, but are softer than Class 1 bursts. We are developing a tool to aid scientists in the study of GRB properties. In the process of developing this tool, we are building a large gamma-ray burst classification database. We are also scientifically analyzing some GRB data as we develop the tool. Tool development thus proceeds in tandem with the dataset for which it is being designed. The tool invokes a modified KDD (Knowledge Discovery in Databases) process, which is described as follows.

Hakkila, Jon↗

Radiation-induced gene expression in the nematode Caenorhabditis elegans

We used the nematode C. elegans to characterize the genotoxic and cytotoxic effects of ionizing radiation in a simple animal model emphasizing the unique effects of charged particle radiation. Here we demonstrate by RT-PCR differential display and whole genome microarray hybridization experiments that gamma rays, accelerated protons and iron ions at the same physical dose lead to unique transcription profiles. 599 of 17871 genes analyzed (3.4%) showed differential expression 3 hrs after exposure to 3 Gy of radiation. 193 were up-regulated, 406 were down-regulated and 90% were affected only by a single species of radiation. A novel statistical clustering technique identified the regulatory relationships between the radiation-modulated genes and showed that genes affected by each radiation species were associated with unique regulatory clusters. This suggests that independent homeostatic mechanisms are activated in response to radiation exposure as a function of track structure or ionization density.

Non-NASA Center↗

Dumb-bell galaxies in southern clusters: Catalog and preliminary statistical results

The dominant galaxy of a rich cluster is often an object whose formation and evolution is closely connected to the dynamics of the cluster itself. Hoessel (1980) and Schneider et al. (1983) estimate that 50 percent of the dominant galaxies are either of the dumb-bell type or have companions at projected distances less than 20 kpc, which is far in excess of the number expected from chance projection (see also Rood and Leir 1979). Presently there is no complete sample of these objects, with the exception of the listing of dumb-bell galaxies in BM type I and I-II clusters in the Abell statistical sample of Rood and Leir (1979). Recent dynamical studies of dumb-bell galaxies in clusters (Valentijn and Casertano, 1988) still suffer from inhomogeneity of the sample. The fact that it is a mixture of optically and radio selected objects may have introduced an unknown biases, for instance if the probability of radio emission is enhanced by the presence of close companions (Stocke, 1978, Heckman et al. 1985, Vettolani and Gregorini 1988) a bias could be present in their velocity distribution. However, this situation is bound to improve: a new sample of Abell clusters in the Southern Hemisphere has been constructed (Abell et al., 1988 hereafter ACO), which has several advantages over the original northern catalog. The plate material (IIIaJ plates) is of better quality and reaches fainter magnitudes. This makes it possible to classify the cluster types with a higher degree of accuracy, as well as to fainter magnitudes. The authors therefore decided to reconsider the whole problem constructing a new sample of dumb-bell galaxies homogeneously selected from the ACO survey. Details of the classification criteria are given.

Vettolani, G.↗

Statistical simulations of clusters of galaxies

By comparing observed and simulated rich clusters of galaxies, it is shown that the observed clusters actually possess physical cores. The accuracy with which core radii can be determined is found. It is also shown that the observations of density profiles of galaxies in the clusters give no significant evidence for a dynamical reason as the cause of the anomalously close resemblance found previously between such density profiles and the isothermal gas distribution.

Avni, Y.↗

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗

Statistical association of QSO's with foreground galaxy clusters

We report a statistically significant overdensity of high redshift quasi-stellar objects (QSO's) in the directions of foreground galaxy clusters. QSO's are taken from the Large Bright QSO Survey (LBQS) between 1.4 less than or equal z less than or equal 2.2 with a limiting magnitude of m(sub B) = 18.5. Foreground clusters are regions within 6 Zwicky radii of small Zwicky clusters at a characteristic redshift of about z approximately = 0.2, covering about 40% of the total area surveyed (304 sq. deg). The overdensity, defined as the ratio of the number density of QSO's in the directions of clusters ('association QSO's) to that in the remainder of the fields ('background QSO's), is equal to 1.7, and formally differs from unity at 4.7 sigma significance. The observed overdensity probably is not due to statistical variation in QSO density, intrinsic QSO-QSO and/or cluster-cluster autocorrelations, or patchy Galactic obscuration. We thus interpret this observation as being due to statistical gravitational lensing of background QSO's by galaxy clusters. However, this amplitude of overdensity behind clusters cannot be accounted for in any cluster lensing model if the background QSO number-magnitude counts are similar to the intrinsic (unlensed) counts, and is implausible in any conventional model of cosmic mass distribution.

Rodrigues-Williams, Liliya L.↗

Statistical Issues in Galaxy Cluster Cosmology

The number and growth of massive galaxy clusters are sensitive probes of cosmological structure formation. Surveys at various wavelengths can detect clusters to high redshift, but the fact that cluster mass is not directly observable complicates matters, requiring us to simultaneously constrain scaling relations of observable signals with mass. The problem can be cast as one of regression, in which the data set is truncated, the (cosmology-dependent) underlying population must be modeled, and strong, complex correlations between measurements often exist. Simulations of cosmological structure formation provide a robust prediction for the number of clusters in the Universe as a function of mass and redshift (the mass function), but they cannot reliably predict the observables used to detect clusters in sky surveys (e.g. X-ray luminosity). Consequently, observers must constrain observable-mass scaling relations using additional data, and use the scaling relation model in conjunction with the mass function to predict the number of clusters as a function of redshift and luminosity.

Galaxy↗

Evolution of massive stars in very young clusters and associations

Statistics concerning the stellar content of young galactic clusters and associations which show well defined main sequence turnups have been analyzed in order to derive information about stellar evolution in high-mass galaxies. The analytical approach is semiempirical and uses natural spectroscopic groups of stars on the H-R diagram together with the stars' apparent magnitudes. The new approach does not depend on absolute luminosities and requires only the most basic elements of stellar evolution theory. The following conclusions are offered on the basis of the statistical analysis: (1) O-tupe main-sequence stars evolve to a spectral type of B1 during core hydrogen burning; (2) most O-type blue stragglers are newly formed massive stars burning core hydrogen; (3) supergiants lying redward of the main-sequence turnup are burning core helium; and most Wolf-Rayet stars are burning core helium and originally had masses greater than 30-40 solar mass. The statistics of the natural spectroscopic stars in young galactic clusters and associations are given in a table.

Stothers, R. B.↗

Clusters and cycles in the cosmic ray age distributions of meteorites

Statistically significant clusters in the cosmic ray exposure age distributions of some groups of iron and stone meteorites were observed, suggesting epochs of enhanced collision and breakups. Fourier analyses of the age distributions of chondrites reveal no significant periods, nor does the same analysis when applied to iron meteorite clusters.

Woodard, M. F.↗

Statistics of arcs in clusters of galaxies

Samples of gravitational lens events in clusters show many large arcs compared to arclets, relative to what can be obtained by idealized singular lens models. We describe the probability of image magnification for point sources and for simple but more realistic gravitational lensing models that include a finite core size and an ellipticity. In addition, we explore the changes in the probability distribution of image magnifications, distortions, and angular extents for sources of different sizes as the parameters of the lenses are varied. A finite core in spherically symmetric lens models introduces a discontinuity in the probability distribution at which the relative number of highly magnified images is increased. In elliptical lenses, this discontinuity and its effect are replaced by a continuous increase in the probability of obtaining high-magnification images relative to singular spherically symmetric models. We also find that the finite size of the source causes a further increase in the expected number of images just below the maximum possible magnification.

Bergmann, Anton G.↗

Statistics of Experiments on Cluster Formation and Transport in a Gravitational Field

Metastable state relaxation in a gravitational field is investigated in the case of non-critical binary solutions. A relaxation description is presented in terms of the time-dependent Ginzburg-Landau formalism for a non-conserved order parameter. A new ansatz for solution of the corresponding partial nonlinear stochastic differential equation is discussed. It is proved that, for the supersaturated solution under consideration, the metastable state relaxation in a gravitational field leads to formation of solute concentration gradients due to the sedimentation of subcritical solute clusters. The pure discussion of the possible methods to compare theoretical results and experimental data related to solute sedimentation in a gravitational field is presented. It is shown that in order to describe these experiments it is necessary to deal both with the value of the solute concentration gradient and with its formation rate. The stochastic nature of the sedimentation process is shown.

Izmailov, Alexander F.↗