Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Transport in the Subtropical Lowermost Stratosphere during CRYSTAL-FACE

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NO(y)) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations, and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and lastly by the extent of convective influence, potentially related to the latitude of convective injection [Dessler and Sherwuud, 2004]. We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and non-local events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Transport in the Subtropical Lowermost Stratosphere during the Cirrus Regional Study of Tropical Anvils and Cirrus Layers-Florida Area Cirrus Experiment

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NOy) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and last by the extent of convective influence, potentially related to the latitude of convective injection (Dessler and Sherwood, 2004). We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and nonlocal events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Suppressing the sample variance of DESI-like galaxy clustering with fast simulations

Ongoing and upcoming galaxy redshift surveys, such as the Dark Energy Spectroscopic Instrument (DESI) survey, will observe vast regions of sky and a wide range of redshifts. In order to model the observations and address various systematic uncertainties, N-body simulations are routinely adopted, however, the number of large simulations with sufficiently high mass resolution is usually limited by available computing time. Therefore, achieving a simulation volume with the effective statistical errors significantly smaller than those of the observations becomes prohibitively expensive. In this study, we apply the Convergence Acceleration by Regression and Pooling (CARPool) method to mitigate the sample variance of the DESI-like galaxy clustering in the AbacusSummit simulations, with the assistance of the quasi-N-body simulations FastPM. Based on the halo occupation distribution (HOD) models, we construct different FastPM galaxy catalogs, including the luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars, with their number densities and two-point clustering statistics well matched to those of AbacusSummit. We also employ the same initial conditions between AbacusSummit and FastPM to achieve high cross-correlation, as it is useful in effectively suppressing the variance. Our method of reducing noise in clustering is equivalent to performing a simulation with volume larger by a factor of 5 and 4 for LRGs and ELGs, respectively. We also mitigate the standard deviation of the LRG bispectrum with the triangular configurations k 2 = 2k 1 = 0.2 h Mpc -1 by a factor of 1.6. With smaller sample variance on galaxy clustering, we are able to constrain the baryon acoustic oscillations (BAO) scale parameters to higher precision. The CARPool method will be beneficial to better constrain the theoretical systematics of BAO, redshift space distortions (RSD) and primordial non-Gaussianity (NG).

79 ASTRONOMY AND ASTROPHYSICS↗

Non-gaussian statistics of pencil beam surveys

We study the effect of the non-Gaussian clustering of galaxies on the statistics of pencil beam surveys. We derive the probability from the power spectrum peaks by means of Edgeworth expansion and find that the higher order moments of the galaxy distribution play a dominant role. The probability of obtaining the 128 Mpc/h periodicity found in pencil beam surveys is raised by more than one order of magnitude, up to 1%. Further data are needed to decide if non-Gaussian distribution alone is sufficient to explain the 128 Mpc/h periodicity, or if extra large-scale power is necessary.

Amendola, Luca↗

Cluster cosmology without cluster finding

ABSTRACT We propose that observations of supermassive galaxies contain cosmological statistical constraining power similar to conventional cluster cosmology, and we provide promising indications that the associated systematic errors are comparably easier to control. We consider a fiducial spectroscopic and stellar mass complete sample of galaxies drawn from the Dark Energy Spectroscopic Instrument (DESI) and forecast how constraints on Ωm–σ8 from this sample will compare with those from number counts of clusters based on richness λ. At fixed number density, we find that massive galaxies offer similar constraints to galaxy clusters. However, a mass-complete galaxy sample from DESI has the potential to probe lower halo masses than standard optical cluster samples (which are typically limited to λ ≳ 20 and Mhalo ≳ 1013.5 M⊙ h−1); additionally, it is straightforward to cleanly measure projected galaxy clustering wp for such a DESI sample, which we show can substantially improve the constraining power on Ωm. We also compare the constraining power of M*-limited samples to those from larger but mass-incomplete samples [e.g. the DESI Bright Galaxy Survey (BGS) sample]; relative to a lower number density M*-limited samples, we find that a BGS-like sample improves statistical constraints by 60 per cent for Ωm and 40 per cent for σ8, but this uses small-scale information that will be harder to model for BGS. Our initial assessment of the systematics associated with supermassive galaxy cosmology yields promising results. The proposed samples have a ∼10 per cent satellite fraction, but we show that cosmological constraints may be robust to the impact of satellites. These findings motivate future work to realize the potential of supermassive galaxies to probe lower halo masses than richness-based clusters and to potentially avoid persistent systematics associated with optical cluster finding.

(cosmology): large-scale structure of Universe↗

Do the major axes of rich clusters of galaxies point toward their neighbors?

The major axis orientation of rich clusters of galaxies, determined from an analysis of X-ray images, is used to investigate whether these clusters point toward their nearest neighbors. No statistical significance is found for a pointing effect between clusters and their nearest neighbors in either X-ray, optical, or combined X-ray and optical samples. Using updated redshifts and permitting nonstatistical sample Abell clusters as nearest neighbors does not affect this conclusion. The lack of statistical significance for a pointing effect favors hierarchical models in which galaxies form first, followed by clusters and then superclusters. For clusters with well-defined X-ray orientations, it is found that cluster position angles determined from X-ray and optical data are in general agreement.

Ulmer, M. P.↗

Spectroscopic quantification of projection effects in the SDSS redMaPPer galaxy cluster catalogue

ABSTRACT Projection effects, whereby galaxies along the line of sight to a galaxy cluster are mistakenly associated with the cluster halo, present a significant challenge for optical cluster cosmology. We use statistically representative spectral coverage of luminous galaxies to investigate how projection effects impact the low-redshift limit of the Sloan Digital Sky Survey (SDSS) redMaPPer galaxy cluster catalogue. Spectroscopic redshifts enable us to differentiate true cluster members from false positives and determine the fraction of candidate cluster members viewed in projection. Our main results can be summarized as follows: first, we show that a simple double-Gaussian model can be used to describe the distribution of line-of-sight velocities in the redMaPPer sample; secondly, the incidence of projection effects is substantial, accounting for ∼16 per cent of the weighted richness for the lowest richness objects; thirdly, projection effects are a strong function of richness, with the contribution in the highest richness bin being several times smaller than for low-richness objects; fourthly, our measurement has a similar amplitude to state-of-the-art models, but finds a steeper dependence of projection effects on richness than these models; and fifthly, the slope of the observed velocity dispersion–richness relation, corrected for projection effects, implies an approximately linear relationship between the true, three-dimensional halo mass and three-dimensional richness. Our results provide a robust, empirical description of the impact of projection effects on the SDSS redMaPPer cluster sample and exemplify the synergies between optical imaging and spectroscopic data for studies of galaxy cluster astrophysics and cosmology.

79 ASTRONOMY AND ASTROPHYSICS↗

A comparison of unsupervised classification procedures on LANDSAT MSS data for an area of complex surface conditions in Basilicata, Southern Italy

Two unsupervised classification procedures were applied to ratioed and unratioed LANDSAT multispectral scanner data of an area of spatially complex vegetation and terrain. An objective accuracy assessment was undertaken on each classification and comparison was made of the classification accuracies. The two unsupervised procedures use the same clustering algorithm. By on procedure the entire area is clustered and by the other a representative sample of the area is clustered and the resulting statistics are extrapolated to the remaining area using a maximum likelihood classifier. Explanation is given of the major steps in the classification procedures including image preprocessing; classification; interpretation of cluster classes; and accuracy assessment. Of the four classifications undertaken, the monocluster block approach on the unratioed data gave the highest accuracy of 80% for five coarse cover classes. This accuracy was increased to 84% by applying a 3 x 3 contextual filter to the classified image. A detailed description and partial explanation is provided for the major misclassification. The classification of the unratioed data produced higher percentage accuracies than for the ratioed data and the monocluster block approach gave higher accuracies than clustering the entire area. The moncluster block approach was additionally the most economical in terms of computing time.

Justice, C.↗

Superclustering and the large-scale structure of the Universe

A complete redshift survey of nearby rich galactic clusters is analyzed in terms of a spatial correlation function, the growth of supercluster, the presence of superclusters around the Bootes void and the discovery of a 300 Mpc void. Abell's (1958) statistical sampling of rich clusters is used for the study. Attention is given to the dependence of the correlation function on the richness of the system, with finding that chances are greater of finding rich neighbors closer together than poor neighbors. Samples of redshifts of 4 or less correlate well with clusters of redshift 5 or more. The size of intergalactic voids is proportional to the surrounding galactic excesses. Finally, the superclusters exist on scales of about 100 Mpc/h, which is larger than expected.

Bahcall, N. A.↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Deconvolute individual genomes from metagenome sequences through short read clustering

Metagenome assembly from short next-generation sequencing data is a challenging process due to its large scale and computational complexity. Clustering short reads by species before assembly offers a unique opportunity for parallel downstream assembly of genomes with individualized optimization. However, current read clustering methods suffer either false negative (under-clustering) or false positive (over-clustering) problems. Here we extended our previous read clustering software, SpaRC, by exploiting statistics derived from multiple samples in a dataset to reduce the under-clustering problem. Using synthetic and real-world datasets we demonstrated that this method has the potential to cluster almost all of the short reads from genomes with sufficient sequencing coverage. The improved read clustering in turn leads to improved downstream genome assembly quality.

59 BASIC BIOLOGICAL SCIENCES↗

The spatial correlation function of rich clusters of galaxies

A series of objective statistical estimators is applied to directly study the three-dimensional distribution of rich clusters, using a recently completed redshift sample of 104 Abell clusters out to distance class four and galactic latitude of 30 deg or more. The sample and the relevant positional and redshift selection effects are discussed, and the two-point angular correlation function and nearest-neighbor distribution of clusters in the sample are determined, as well as in the larger D = 5+6 sample and in several subsamples of different richness classes and latitude zones. The redshift correlation function of clusters for various angular separations on the sky is examined as a test for the reality of the observed angular correlations. The two-point spatial correlation function of clusters in the sample is determined, and the consistency among the three observed correlation functions is tested. A model that describes all the observed correlations is presented.

Bahcall, N. A.↗

On possible associations of quasi-stellar objects and radio galaxies with rich clusters of galaxies

A comparison of the cataloged coordinates of QSOs and of 3CR galaxies with those of Abell's rich clusters of galaxies yields the following results: (1) There is no statistically significant evidence that high-redshift QSOs lie preferentially close to Abell clusters. (2) There is no statistically significant evidence that low-redshift QSOs lie preferentially close to Abell clusters. (3) There is (as is well known) highly significant evidence that 3CR galaxies lie preferentially close to Abell clusters. (4) The distributions of angular separations between QSOs and clusters and between 3CR galaxies and clusters differ at statistically significant (but not highly significant) levels. In view of these results, a generic relationship between low-redshift QSOs and radio galaxies seems questionable. This result for the QSOs is also entirely consistent with the idea that they are 'local' objects with redshifts of noncosmological origin.

Roberts, D. H.↗

The locations of X-ray sources in globular clusters

The SAS 3 X-ray observatory has been used to measure the positions of the X-ray sources in the globular clusters NGC 1851, 6441, 6624, 6712, and 7078. The derived positions have error circles with 90% confidence radii of 10 or 30 arcsec, which, in each case, include the optical center of the cluster. The results of a statistical analysis demonstrate that the X-ray sources are more concentrated toward the cluster centers than the visible stars and are therefore more massive than those stars.

Jernigan, J. G.↗

Galaxy Cluster Contribution to the Diffuse Extragalactic Ultraviolet Background

The diffuse ultraviolet background radiation has been mapped over most of the sky with 2′ resolution using data from the Galaxy Evolution Explorer survey. We utilize this map to study the correlation between the UV background and clusters of galaxies discovered via the Sunyaev–Zeldovich effect in the Planck survey. We use only high Galactic latitude (|b|>60{sup ∘}) galaxy clusters to avoid contamination by Galactic foregrounds, and we only analyze clusters with a measured redshift. This leaves us with a sample of 142 clusters over the redshift range of 0.02 ≤ z ≤ 0.72, which we further subdivide into four redshift bins. In analyzing our stacked samples binned by redshift, we find evidence for a central excess of UV background light compared to local backgrounds for clusters with z < 0.3. We then stacked these z < 0.3 clusters to find a statistically significant excess of 12 ± 2.3 photon cm{sup −2} s{sup −1} sr{sup −1} Å{sup −1} over the median of ∼380 photon cm{sup −2} s{sup −1} sr{sup −1} Å{sup −1} measured around random blank fields. We measure the stacked radial profile of these clusters, and find that the excess UV radiation decays to the level of the background at a radius of ∼1 Mpc, roughly consistent with the maximum radial extent of the clusters. Analysis of possible physical processes contributing to the excess UV brightness indicates that non-thermal emission from relativistic electrons in the intracluster medium and faint, unresolved UV emission from cluster member galaxies and intracluster light are likely the dominant contributors.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey Year 3 results: Cosmology from combined galaxy clustering and lensing validation on cosmological simulations

In this work, we present a validation of the Dark Energy Survey Year 3 (DES Y3) 3×2-point analysis choices by testing them on Buzzard2.0, a new suite of cosmological simulations that is tailored for the testing and validation of combined galaxy clustering and weak-lensing analyses. We show that the buzzard2.0 simulations accurately reproduce many important aspects of the DES Y3 data, including photometric redshift and magnitude distributions, and the relevant set of two-point clustering and weak-lensing statistics. We then show that our model for the 3×2-point data vector is accurate enough to recover the true cosmology in simulated surveys assuming the true redshift distributions for our source and lens samples, demonstrating robustness to uncertainties in the modeling of the nonlinear matter power spectrum, nonlinear galaxy bias, and higher-order lensing corrections. Additionally, we demonstrate for the first time that our photometric redshift calibration methodology, including information from photometry, spectroscopy, clustering cross-correlations, and galaxy–galaxy lensing ratios, is accurate enough to recover the true cosmology in simulated surveys in the presence of realistic photometric redshift uncertainties.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-Parent Clustering Algorithms from Stochastic Grammar Data Models

We introduce a statistical data model and an associated optimization-based clustering algorithm which allows data vectors to belong to zero, one or several "parent" clusters. For each data vector the algorithm makes a discrete decision among these alternatives. Thus, a recursive version of this algorithm would place data clusters in a Directed Acyclic Graph rather than a tree. We test the algorithm with synthetic data generated according to the statistical data model. We also illustrate the algorithm using real data from large-scale gene expression assays.

Mjoisness, Eric↗

A framework to measure the properties of intergalactic metal systems with two-point flux statistics

ABSTRACT The abundance, temperature, and clustering of metals in the intergalactic medium are important parameters for understanding their cosmic evolution and quantifying their impact on cosmological analysis with the Ly α forest. The properties of these systems are typically measured from individual quasar spectra redward of the quasar’s Ly α emission line, yet that approach may provide biased results due to selection effects. We present an alternative approach to measure these properties in an unbiased manner with the two-point statistics commonly employed to quantify large-scale structure. Our model treats the observed flux of a large sample of quasar spectra as a continuous field and describes the one-dimensional, two-point statistics of this field with three parameters per ion: the abundance (column density distribution), temperature (Doppler parameter), and clustering (cloud–cloud correlation function). We demonstrate this approach on multiple ions (e.g. ${\rm C\, \small {\rm IV}}$ , ${\rm Si\, \small {\rm IV}}$ , and ${\rm Mg\, \small {\rm II}}$ ) with early data from the Dark Energy Spectroscopic Instrument (DESI) and high-resolution spectra from the literature. Our initial results show some evidence that the ${\rm C\, \small {\rm IV}}$ abundance is higher than previous measurements and evidence for abundance evolution over time. The first full year of DESI observations will have over an order of magnitude more quasar spectra than this study. In a future paper, we will use those data to measure the growth of clustering and its impact on the Ly α forest, as well as test other DESI analysis infrastructure such as the pipeline noise estimates and the resolution matrix.

79 ASTRONOMY AND ASTROPHYSICS↗