Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

MEASURE: An integrated data-analysis and model identification facility

The first phase of the development of MEASURE, an integrated data analysis and model identification facility is described. The facility takes system activity data as input and produces as output representative behavioral models of the system in near real time. In addition a wide range of statistical characteristics of the measured system are also available. The usage of the system is illustrated on data collected via software instrumentation of a network of SUN workstations at the University of Illinois. Initially, statistical clustering is used to identify high density regions of resource-usage in a given environment. The identified regions form the states for building a state-transition model to evaluate system and program performance in real time. The model is then solved to obtain useful parameters such as the response-time distribution and the mean waiting time in each state. A graphical interface which displays the identified models and their characteristics (with real time updates) was also developed. The results provide an understanding of the resource-usage in the system under various workload conditions. This work is targeted for a testbed of UNIX workstations with the initial phase ported to SUN workstations on the NASA, Ames Research Center Advanced Automation Testbed.

Singh, Jaidip↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

An Artificial Intelligence Classification Tool and Its Application to Gamma-Ray Bursts

Despite being the most energetic phenomenon in the known universe, the astrophysics of gamma-ray bursts (GRBs) has still proven difficult to understand. It has only been within the past five years that the GRB distance scale has been firmly established, on the basis of a few dozen bursts with x-ray, optical, and radio afterglows. The afterglows indicate source redshifts of z=1 to z=5, total energy outputs of roughly 10(exp 52) ergs, and energy confined to the far x-ray to near gamma-ray regime of the electromagnetic spectrum. The multi-wavelength afterglow observations have thus far provided more insight on the nature of the GRB mechanism than the GRB observations; far more papers have been written about the few observed gamma-ray burst afterglows in the past few years than about the thousands of detected gamma-ray bursts. One reason the GRB central engine is still so poorly understood is that GRBs have complex, overlapping characteristics that do not appear to be produced by one homogeneous process. At least two subclasses have been found on the basis of duration, spectral hardness, and fluence (time integrated flux); Class 1 bursts are softer, longer, and brighter than Class 2 bursts (with two second durations indicating a rough division). A third GRB subclass, overlapping the other two, has been identified using statistical clustering techniques; Class 3 bursts are intermediate between Class 1 and Class 2 bursts in brightness and duration, but are softer than Class 1 bursts. We are developing a tool to aid scientists in the study of GRB properties. In the process of developing this tool, we are building a large gamma-ray burst classification database. We are also scientifically analyzing some GRB data as we develop the tool. Tool development thus proceeds in tandem with the dataset for which it is being designed. The tool invokes a modified KDD (Knowledge Discovery in Databases) process, which is described as follows.

Hakkila, Jon↗

Radiation-induced gene expression in the nematode Caenorhabditis elegans

We used the nematode C. elegans to characterize the genotoxic and cytotoxic effects of ionizing radiation in a simple animal model emphasizing the unique effects of charged particle radiation. Here we demonstrate by RT-PCR differential display and whole genome microarray hybridization experiments that gamma rays, accelerated protons and iron ions at the same physical dose lead to unique transcription profiles. 599 of 17871 genes analyzed (3.4%) showed differential expression 3 hrs after exposure to 3 Gy of radiation. 193 were up-regulated, 406 were down-regulated and 90% were affected only by a single species of radiation. A novel statistical clustering technique identified the regulatory relationships between the radiation-modulated genes and showed that genes affected by each radiation species were associated with unique regulatory clusters. This suggests that independent homeostatic mechanisms are activated in response to radiation exposure as a function of track structure or ionization density.

Non-NASA Center↗

Nearest neighbour distributions: New statistical measures for cosmological clustering

ABSTRACT The use of summary statistics beyond the two-point correlation function to analyse the non-Gaussian clustering on small scales, and thereby, increasing the sensitivity to the underlying cosmological parameters, is an active field of research in cosmology. In this paper, we explore a set of new summary statistics – the k-Nearest Neighbour Cumulative Distribution Functions (kNN-CDF). This is the empirical cumulative distribution function of distances from a set of volume-filling, Poisson distributed random points to the k-nearest data points, and is sensitive to all connected N-point correlations in the data. The kNN-CDF can be used to measure counts in cell, void probability distributions, and higher N-point correlation functions, all using the same formalism exploiting fast searches with spatial tree data structures. We demonstrate how it can be computed efficiently from various data sets – both discrete points, and the generalization for continuous fields. We use data from a large suite of N-body simulations to explore the sensitivity of this new statistic to various cosmological parameters, compared to the two-point correlation function, while using the same range of scales. We demonstrate that the use of kNN-CDF improves the constraints on the cosmological parameters by more than a factor of 2 when applied to the clustering of dark matter in the range of scales between 10 and $40\, h^{-1}\, {\rm Mpc}$. We also show that relative improvement is even greater when applied on the same scales to the clustering of haloes in the simulations at a fixed number density, both in real space, as well as in redshift space. Since the kNN-CDF are sensitive to all higher order connected correlation functions in the data, the gains over traditional two-point analyses are expected to grow as progressively smaller scales are included in the analysis of cosmological data, provided the higher order correlation functions are sensitive to cosmology on the scales of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Dumb-bell galaxies in southern clusters: Catalog and preliminary statistical results

The dominant galaxy of a rich cluster is often an object whose formation and evolution is closely connected to the dynamics of the cluster itself. Hoessel (1980) and Schneider et al. (1983) estimate that 50 percent of the dominant galaxies are either of the dumb-bell type or have companions at projected distances less than 20 kpc, which is far in excess of the number expected from chance projection (see also Rood and Leir 1979). Presently there is no complete sample of these objects, with the exception of the listing of dumb-bell galaxies in BM type I and I-II clusters in the Abell statistical sample of Rood and Leir (1979). Recent dynamical studies of dumb-bell galaxies in clusters (Valentijn and Casertano, 1988) still suffer from inhomogeneity of the sample. The fact that it is a mixture of optically and radio selected objects may have introduced an unknown biases, for instance if the probability of radio emission is enhanced by the presence of close companions (Stocke, 1978, Heckman et al. 1985, Vettolani and Gregorini 1988) a bias could be present in their velocity distribution. However, this situation is bound to improve: a new sample of Abell clusters in the Southern Hemisphere has been constructed (Abell et al., 1988 hereafter ACO), which has several advantages over the original northern catalog. The plate material (IIIaJ plates) is of better quality and reaches fainter magnitudes. This makes it possible to classify the cluster types with a higher degree of accuracy, as well as to fainter magnitudes. The authors therefore decided to reconsider the whole problem constructing a new sample of dumb-bell galaxies homogeneously selected from the ACO survey. Details of the classification criteria are given.

Vettolani, G.↗

Statistical simulations of clusters of galaxies

By comparing observed and simulated rich clusters of galaxies, it is shown that the observed clusters actually possess physical cores. The accuracy with which core radii can be determined is found. It is also shown that the observations of density profiles of galaxies in the clusters give no significant evidence for a dynamical reason as the cause of the anomalously close resemblance found previously between such density profiles and the isothermal gas distribution.

Avni, Y.↗

Detection of the significant impact of source clustering on higher order statistics with DES Year 3 weak gravitational lensing data

We measure the impact of source galaxy clustering on higher order summary statistics of weak gravitational lensing data. By comparing simulated data with galaxies that either trace or do not trace the underlying density field, we show that this effect can exceed measurement uncertainties for common higher order statistics for certain analysis choices. We evaluate the impact on different weak lensing observables, finding that third moments and wavelet phase harmonics are more affected than peak count statistics. Using Dark Energy Survey (DES) Year 3 (Y3) data, we construct null tests for the source-clustering-free case, finding a p-value of p = 4 × 10 −3 (2.6σ) using third-order map moments and p = 3 × 10 −11 (6.5σ) using wavelet phase harmonics. The impact of source clustering on cosmological inference can be either included in the model or minimized through ad hoc procedures (e.g. scale cuts). We verify that the procedures adopted in existing DES Y3 cosmological analyses were sufficient to render this effect negligible. Failing to account for source clustering can significantly impact cosmological inference from higher order gravitational lensing statistics, e.g. higher order N-point functions, wavelet-moment observables, and deep learning or field-level summary statistics of weak lensing maps.

79 ASTRONOMY AND ASTROPHYSICS↗

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗

Statistical association of QSO's with foreground galaxy clusters

We report a statistically significant overdensity of high redshift quasi-stellar objects (QSO's) in the directions of foreground galaxy clusters. QSO's are taken from the Large Bright QSO Survey (LBQS) between 1.4 less than or equal z less than or equal 2.2 with a limiting magnitude of m(sub B) = 18.5. Foreground clusters are regions within 6 Zwicky radii of small Zwicky clusters at a characteristic redshift of about z approximately = 0.2, covering about 40% of the total area surveyed (304 sq. deg). The overdensity, defined as the ratio of the number density of QSO's in the directions of clusters ('association QSO's) to that in the remainder of the fields ('background QSO's), is equal to 1.7, and formally differs from unity at 4.7 sigma significance. The observed overdensity probably is not due to statistical variation in QSO density, intrinsic QSO-QSO and/or cluster-cluster autocorrelations, or patchy Galactic obscuration. We thus interpret this observation as being due to statistical gravitational lensing of background QSO's by galaxy clusters. However, this amplitude of overdensity behind clusters cannot be accounted for in any cluster lensing model if the background QSO number-magnitude counts are similar to the intrinsic (unlensed) counts, and is implausible in any conventional model of cosmic mass distribution.

Rodrigues-Williams, Liliya L.↗

Statistical Issues in Galaxy Cluster Cosmology

The number and growth of massive galaxy clusters are sensitive probes of cosmological structure formation. Surveys at various wavelengths can detect clusters to high redshift, but the fact that cluster mass is not directly observable complicates matters, requiring us to simultaneously constrain scaling relations of observable signals with mass. The problem can be cast as one of regression, in which the data set is truncated, the (cosmology-dependent) underlying population must be modeled, and strong, complex correlations between measurements often exist. Simulations of cosmological structure formation provide a robust prediction for the number of clusters in the Universe as a function of mass and redshift (the mass function), but they cannot reliably predict the observables used to detect clusters in sky surveys (e.g. X-ray luminosity). Consequently, observers must constrain observable-mass scaling relations using additional data, and use the scaling relation model in conjunction with the mass function to predict the number of clusters as a function of redshift and luminosity.

Galaxy↗

Toward Accurate Modeling of Galaxy Clustering on Small Scales: Constraining the Galaxy-halo Connection with Optimal Statistics

Applying halo models to analyze the small-scale clustering of galaxies is a proven method for characterizing the connection between galaxies and their host halos. Such works are often plagued by systematic errors or limited to clustering statistics that can be predicted analytically. In this work, we employ a numerical mock-based modeling procedure to examine the clustering of Sloan Digital Sky Survey DR7 galaxies. We apply a standard halo occupation distribution (HOD) model to dark matter only simulations with a ΛCDM cosmology. To constrain the theoreStical models, we utilize a combination of galaxy number density and selected scales of the projected correlation function, redshift-space correlation function, group multiplicity function, average group velocity dispersion, mark correlation function, and counts-in-cells statistics. We design an algorithm to choose an optimal combination of measurements that yields tight and accurate constraints on our model parameters. Compared to previous work using fewer clustering statistics, we find a significant improvement in the constraints on all parameters of our halo model for two different luminosity-threshold galaxy samples. Most interestingly, we obtain unprecedented high-precision constraints on the scatter in the relationship between galaxy luminosity and halo mass. However, our best-fit model results in significant tension (>4σ) for both samples, indicating the need to add second-order features to the standard HOD model. To guarantee the robustness of these results, we perform an extensive analysis of the systematic and statistical errors in our modeling procedure, including a first of its kind study of the sensitivity of our constraints to changes in the halo mass function due to baryonic physics.

79 ASTRONOMY AND ASTROPHYSICS↗

2D k -th nearest neighbour statistics: a highly informative probe of galaxy clustering

ABSTRACT Beyond standard summary statistics are necessary to summarize the rich information on non-linear scales in the era of precision galaxy clustering measurements. For the first time, we introduce the 2D k-th nearest neighbour (kNN) statistics as a summary statistic for discrete galaxy fields. This is a direct generalization of the standard 1D kNN by disentangling the projected galaxy distribution from the redshift-space distortion signature along the line-of-sight. We further introduce two different flavours of 2D kNNs that trace different aspects of the galaxy field: the standard flavour which tabulates the distances between galaxies and random query points, and a ‘DD’ flavour that tabulates the distances between galaxies and galaxies. We showcase the 2D kNNs’ strong constraining power both through theoretical arguments and by testing on realistic galaxy mocks. Theoretically, we show that 2D kNNs are computationally efficient and directly generate other statistics such as the popular two-point correlation function (2PCF), voids probability function, and counts-in-cell statistics. In a more practical test, we apply the 2D kNN statistics to simulated galaxy mocks that fold in a large range of observational realism and recover parameters of the underlying extended halo occupation distribution (HOD) model that includes velocity bias and galaxy assembly bias. We find unbiased and significantly tighter constraints on all aspects of the HOD model with the 2D kNNs, both compared to the standard 1D kNN, and the classical redshift-space 2PCF.

79 ASTRONOMY AND ASTROPHYSICS↗

Hydrated Anions: From Clusters to Bulk Solution with Quasi-Chemical Theory

The interactions of hydrated ions with molecular and macromolecular solution and interface partners are strong on a chemical energy scale. Here we recount the foremost ab initio theory for the evaluation of the hydration free energies of ions, namely, quasi-chemical theory (QCT). We focus on anions, particularly halides but also the hydroxide anion, because they have been outstanding challenges for all theories. For example, this work supports understanding the high selectivity for F – over Cl – in fluoride-selective ion channels despite the identical charge and the size similarity of these ions. QCT is built by the identification of inner-shell clusters, separate treatment of those clusters, and then the integration of those results into the broader-scale solution environment. Recent work has focused on a close comparison with mass-spectrometric measurements of ion-hydration equilibria. We delineate how ab initio molecular dynamics (AIMD) calculations on ion-hydration clusters, elementary statistical thermodynamics, and electronic structure calculations on cluster structures sampled from the AIMD calculations obtain just the free energies extracted from the cluster experiments. That theory–experiment comparison has not been attempted before the work discussed here, but the agreement is excellent with moderate computational effort. This agreement reinforces both theory and experiment and provides a numerically accurate inner-shell contribution to QCT. The inner-shell complexes involving heavier halides display strikingly asymmetric hydration clusters. Asymmetric hydration structures can be problematic for the evaluation of the QCT outer-shell contribution with the polarizable continuum model (PCM). Nevertheless, QCT provides a favorable setting for the exploitation of PCM when the inner-shell material shields the ion from the outer solution environment. For the more asymmetrically hydrated, and thus less effectively shielded, heavier halide ions clustered with waters, the PCM is less satisfactory. We therefore investigate an inverse procedure in which the inner-shell structures are sampled from readily available AIMD calculations on the bulk solutions. This inverse procedure is a remarkable improvement; our final results are in close agreement with a standard tabulation of hydration free energies, and the final composite results are independent of the coordination number on the chemical energy scale of relevance, as they should be. Finally, a comparison of anion hydration structure in clusters and bulk solutions from AIMD simulations emphasize some differences: the asymmetries of bulk solution inner-shell structures are moderated compared with clusters but are still present, and inner hydration shells fill to slightly higher average coordination numbers in bulk solution than in clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evolution of massive stars in very young clusters and associations

Statistics concerning the stellar content of young galactic clusters and associations which show well defined main sequence turnups have been analyzed in order to derive information about stellar evolution in high-mass galaxies. The analytical approach is semiempirical and uses natural spectroscopic groups of stars on the H-R diagram together with the stars' apparent magnitudes. The new approach does not depend on absolute luminosities and requires only the most basic elements of stellar evolution theory. The following conclusions are offered on the basis of the statistical analysis: (1) O-tupe main-sequence stars evolve to a spectral type of B1 during core hydrogen burning; (2) most O-type blue stragglers are newly formed massive stars burning core hydrogen; (3) supergiants lying redward of the main-sequence turnup are burning core helium; and most Wolf-Rayet stars are burning core helium and originally had masses greater than 30-40 solar mass. The statistics of the natural spectroscopic stars in young galactic clusters and associations are given in a table.

Stothers, R. B.↗

Clusters and cycles in the cosmic ray age distributions of meteorites

Statistically significant clusters in the cosmic ray exposure age distributions of some groups of iron and stone meteorites were observed, suggesting epochs of enhanced collision and breakups. Fourier analyses of the age distributions of chondrites reveal no significant periods, nor does the same analysis when applied to iron meteorite clusters.

Woodard, M. F.↗

Statistics of arcs in clusters of galaxies

Samples of gravitational lens events in clusters show many large arcs compared to arclets, relative to what can be obtained by idealized singular lens models. We describe the probability of image magnification for point sources and for simple but more realistic gravitational lensing models that include a finite core size and an ellipticity. In addition, we explore the changes in the probability distribution of image magnifications, distortions, and angular extents for sources of different sizes as the parameters of the lenses are varied. A finite core in spherically symmetric lens models introduces a discontinuity in the probability distribution at which the relative number of highly magnified images is increased. In elliptical lenses, this discontinuity and its effect are replaced by a continuous increase in the probability of obtaining high-magnification images relative to singular spherically symmetric models. We also find that the finite size of the source causes a further increase in the expected number of images just below the maximum possible magnification.

Bergmann, Anton G.↗