SEARCH · Engineering Papers
Results for “Bayesian clustering”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
SPT clusters with DES and HST weak lensing. I. Cluster lensing and Bayesian population modeling of multiwavelength cluster datasets
We present a Bayesian population modeling method to analyze the abundance of galaxy clusters identified by the South Pole Telescope (SPT) with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). We discuss and validate the modeling choices with a particular focus on a robust, weak-lensing-based mass calibration using DES data. For the DES Year 3 data, we report a systematic uncertainty in weak-lensing mass calibration that increases from 1% at z = 0.25 to 10% at z = 0.95 , to which we add 2% in quadrature to account for uncertainties in the impact of baryonic effects. We implement an analysis pipeline that joins the cluster abundance likelihood with a multiobservable likelihood for the Sunyaev-Zel’dovich effect, optical richness, and weak-lensing measurements for each individual cluster. We validate that our analysis pipeline can recover unbiased cosmological constraints by analyzing mocks that closely resemble the cluster sample extracted from the SPT-SZ, SPTpol ECS, and SPTpol 500d surveys and the DES Year 3 and HST-39 weak-lensing datasets. This work represents a crucial prerequisite for the subsequent cosmological analysis of the real dataset.
Evaluation of the procedure 1A component of the 1980 US/Canada wheat and barley exploratory experiment
Several techniques which use clusters generated by a new clustering algorithm, CLASSY, are proposed as alternatives to random sampling to obtain greater precision in crop proportion estimation: (1) Proportional Allocation/relative count estimator (PA/RCE) uses proportional allocation of dots to clusters on the basis of cluster size and a relative count cluster level estimate; (2) Proportional Allocation/Bayes Estimator (PA/BE) uses proportional allocation of dots to clusters and a Bayesian cluster-level estimate; and (3) Bayes Sequential Allocation/Bayesian Estimator (BSA/BE) uses sequential allocation of dots to clusters and a Bayesian cluster level estimate. Clustering in an effective method in making proportion estimates. It is estimated that, to obtain the same precision with random sampling as obtained by the proportional sampling of 50 dots with an unbiased estimator, samples of 85 or 166 would need to be taken if dot sets with AI labels (integrated procedure) or ground truth labels, respectively were input. Dot reallocation provides dot sets that are unbiased. It is recommended that these proportion estimation techniques are maintained, particularly the PA/BE because it provides the greatest precision.
Exploring the contamination of the DES-Y1 cluster sample with SPT-SZ selected clusters
ABSTRACT We perform a cross validation of the cluster catalogue selected by the red-sequence Matched-filter Probabilistic Percolation algorithm (redMaPPer) in Dark Energy Survey year 1 (DES-Y1) data by matching it with the Sunyaev–Zel’dovich effect (SZE) selected cluster catalogue from the South Pole Telescope SPT-SZ survey. Of the 1005 redMaPPer selected clusters with measured richness $\hat{\lambda }\gt 40$ in the joint footprint, 207 are confirmed by SPT-SZ. Using the mass information from the SZE signal, we calibrate the richness–mass relation using a Bayesian cluster population model. We find a mass trend λ ∝ MB consistent with a linear relation (B ∼ 1), no significant redshift evolution and an intrinsic scatter in richness of σλ = 0.22 ± 0.06. By considering two error models, we explore the impact of projection effects on the richness–mass modelling, confirming that such effects are not detectable at the current level of systematic uncertainties. At low richness SPT-SZ confirms fewer redMaPPer clusters than expected. We interpret this richness dependent deficit in confirmed systems as due to the increased presence at low richness of low-mass objects not correctly accounted for by our richness-mass scatter model, which we call contaminants. At a richness $\hat{\lambda }=40$, this population makes up ${\gt}12{{\ \rm per\ cent}}$ (97.5 percentile) of the total population. Extrapolating this to a measured richness $\hat{\lambda }=20$ yields ${\gt}22{{\ \rm per\ cent}}$ (97.5 percentile). With these contamination fractions, the predicted redMaPPer number counts in different plausible cosmologies are compatible with the measured abundance. The presence of such a population is also a plausible explanation for the different mass trends (B ∼ 0.75) obtained from mass calibration using purely optically selected clusters. The mean mass from stacked weak lensing (WL) measurements suggests that these low-mass contaminants are galaxy groups with masses ∼3–5 × 1013 M⊙ which are beyond the sensitivity of current SZE and X-ray surveys but a natural target for SPT-3G and eROSITA.
Tool Support for Parametric Analysis of Large Software Simulation Systems
The analysis of large and complex parameterized software systems, e.g., systems simulation in aerospace, is very complicated and time-consuming due to the large parameter space, and the complex, highly coupled nonlinear nature of the different system components. Thus, such systems are generally validated only in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. We have addressed the factors deterring such an analysis with a tool to support envelope assessment: we utilize a combination of advanced Monte Carlo generation with n-factor combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. Additional test-cases, automatically generated from models (e.g., UML, Simulink, Stateflow) improve the coverage. The distributed test runs of the software system produce vast amounts of data, making manual analysis impossible. Our tool automatically analyzes the generated data through a combination of unsupervised Bayesian clustering techniques (AutoBayes) and supervised learning of critical parameter ranges using the treatment learner TAR3. The tool has been developed around the Trick simulation environment, which is widely used within NASA. We will present this tool with a GN&C (Guidance, Navigation and Control) simulation of a small satellite system.
Redshift inference from the combination of galaxy colours and clustering in a hierarchical Bayesian model – Application to realistic N -body simulations
ABSTRACT Photometric galaxy surveys constitute a powerful cosmological probe but rely on the accurate characterization of their redshift distributions using only broad-band imaging, and can be very sensitive to incomplete or biased priors used for redshift calibration. A hierarchical Bayesian model has recently been developed to estimate those from the robust combination of prior information, photometry of single galaxies, and the information contained in the galaxy clustering against a well-characterized tracer population. In this work, we extend the method so that it can be applied to real data, developing some necessary new extensions to it, especially in the treatment of galaxy clustering information, and we test it on realistic simulations. After marginalizing over the mapping between the clustering estimator and the actual density distribution of the sample galaxies, and using prior information from a small patch of the survey, we find the incorporation of clustering information with photo-z’s tightens the redshift posteriors and overcomes biases in the prior that mimic those happening in spectroscopic samples. The method presented here uses all the information at hand to reduce prior biases and incompleteness. Even in cases where we artificially bias the spectroscopic sample to induce a shift in mean redshift of $\Delta \bar{z} \approx 0.05,$ the final biases in the posterior are $\Delta \bar{z} \lesssim 0.003.$ This robustness to flaws in the redshift prior or training samples would constitute a milestone for the control of redshift systematic uncertainties in future weak lensing analyses.
Proactive Frequency Stability Scheme Based on Bayesian Filters and Spectral Clustering
Not Available
Copacabana: a probabilistic membership assignment method for galaxy clusters
Cosmological analyses using galaxy clusters in optical/near-infrared photometric surveys require robust characterization of their galaxy content. Precisely determining which galaxies belong to a cluster is crucial. In this paper, we present the COlor Probabilistic Assignment of Clusters And BAyesiaN Analysis (Copacabana) algorithm. Copacabana computes membership probabilities for all galaxies within an aperture centred on the cluster using photometric redshifts, colours, and projected radial probability density functions. We use simulations to validate Copacabana and we show that it achieves up to 89 per cent membership accuracy with a mild dependence on photometric redshift uncertainties and choice of aperture size. We find that the precision of the photometric redshifts has the largest impact on the determination of the membership probabilities followed by the choice of the cluster aperture size. We also quantify how much these uncertainties in the membership probabilities affect the stellar mass–cluster mass scaling relation, a relation that directly impacts cosmology. Using the sum of the stellar masses weighted by membership probabilities (μ * ) as the observable, we find that Copacabana can reach an accuracy of 0.06 dex in the measurement of the scaling relation at low redshift for a Legacy Survey of Space and Time type survey. These results indicate the potential of Copacabana and μ * to be used in cosmological analyses of optically selected clusters in the future.
A Multi-Armed Bayesian Ordinal Outcome Utility-Based Sequential Trial with a Pairwise Null Clustering Prior
A multi-armed trial based on ordinal outcomes is proposed that leverages a flexible non-proportional odds cumulative logit model and numerical utility scores for each outcome to determine treatment optimality. This trial design uses a Bayesian clustering prior on the treatment effects that encourages the pairwise null hypothesis of no differences between treatments. A group sequential design is proposed to determine which treatments are clinically different with an adaptive decision boundary that becomes more aggressive as the sample size or clinical significance grows, or the number of active treatments decreases. A simulation study is conducted for 3 and 5 treatment arms, which shows that the design has superior operating characteristics (family wise error rate, generalized power, average sample size) compared to utility designs that do not allow clustering, a frequentist proportional odds model, or a permutation test based on empirical mean utilities.
Precision calibration of calorimeter signals in the ATLAS experiment using an uncertainty-aware neural network
The ATLAS experiment at the Large Hadron Collider explores the use of modern neural networks for a multi-dimensional calibration of its calorimeter signal defined by clusters of topologically connected cells (topo-clusters). The Bayesian neural network (BNN) approach not only yields a continuous and smooth calibration function that improves performance relative to the standard calibration but also provides uncertainties on the calibrated energies for each topo-cluster. The results obtained by using a trained BNN are compared to the standard local hadronic calibration and to a calibration provided by training a deep neural network. The uncertainties predicted by the BNN are interpreted in the context of a fractional contribution to the systematic uncertainties of the trained calibration. They are also compared to uncertainty predictions obtained from an alternative estimator employing repulsive ensembles.
Protein Conformational States—A First Principles Bayesian Method
Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naïve Bayes classifier from the machine learning community for use on atom-to-atom pairwise contacts. The result is an unsupervised learning algorithm that samples a ‘distribution’ over potential classification schemes. We apply the classifier to a series of test structures and one real protein, showing that it identifies the conformational transition with >95% accuracy in most cases. A nontrivial feature of our adaptation is a new connection to information entropy that allows us to vary the level of structural detail without spoiling the categorization. This is confirmed by comparing results as the number of atoms and time-samples are varied over 1.5 orders of magnitude. Further, the method’s derivation from Bayesian analysis on the set of inter-atomic contacts makes it easy to understand and extend to more complex cases.
Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis
Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.
Alleviating prior dependencies for DESI DR1 clustering fits through reparameterization
Bayesian analyses of the full-shape clustering of Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) exhibit prior-volume projection effects, whereby weakly constrained nuisance parameters of the Effective Field Theory of Large Scale Structure (EFTofLSS) shift marginalized cosmological posteriors away from the posterior maximum. We reanalyze DESI DR1 power spectrum multipoles using two complementary mitigation strategies: (i) nonlinear orthogonalization to decorrelate nuisance and cosmological parameter priors, and (ii) a fully reparameterization-invariant Jeffreys prior over all EFTofLSS coefficients, evaluated on-the-fly via closed-form Jacobians. Including data from DESI, Big-Bang Nuclesynthesis and a constraint on $n_{\mathrm{s}}$, baseline priors lead to multi-$σ$ projection in the Hubble parameter $H_{0}$ and dark energy equation of state parameters $w_{0}$ and $w_{a}$; the Jeffreys prior successfully recenters these posteriors to enclose the maximum a posteriori estimate within the 68% credible regions, demonstrating clear mitigation of projection effects for these late-time expansion parameters. A hybrid Jeffreys+baseline-Gaussian configuration controls residual over-broad tails in the physical cold dark matter density $ω_{\mathrm{c}}$ while preserving the volume correction, and is our favoured approach. We compare the credible intervals derived using our methodology to those obtained using Halo Occupation Distribution (HOD)-informed priors and to confidence intervals derived using frequentist profile likelihood analyses, finding agreement in both central values and degeneracy directions in the $w_{0}$--$w_{a}$ plane. This demonstrates that, once projection effects are properly controlled, we can make robust inferences about the late-time cosmological expansion independent of the statistical framework adopted.
Bayesian prior construction for uncertainty quantification in first-principles statistical mechanics
First-principles statistical mechanics enables the prediction of thermodynamic and kinetic properties of materials, but is computationally expensive. Many approaches require surrogate models to calculate energies within Monte Carlo or molecular dynamics simulations. Inexpensive surrogates such as cluster expansions enable otherwise intractable calculations by interpolating data from higher accuracy methods, such as Density Functional Theory (DFT). Surrogate models introduce uncertainty into downstream calculations, in addition to any uncertainty inherent to DFT calculations. Bayesian frameworks address this by quantifying uncertainty and incorporating expert knowledge through priors. However, constructing effective priors remains challenging. This work introduces and describes practical strategies for building Bayesian cluster expansions, focusing on basis truncation, hyperparameter selection, and ground state replication. We analyze multiple basis truncation schemes, compare cross-validation to the evidence-approximation for hyperparameter optimization, and provide methods to find and enforce ground-state-preserving models through priors. Additionally, we compare the uncertainties between different approximations to DFT (LDA, PBE, SCAN) against the uncertainty introduced with the use of cluster expansion surrogate models. These approaches are demonstrated on the BCC Li x Mg 1-x and Li x Al 1-x alloys, which are both of interest for solid-state Li batteries. Our results provide guidelines for constructing and utilizing Bayesian cluster expansions, thereby improving the transparency of materials modeling. Furthermore, the approaches and insights developed in this work can be transferred to a wide range of cluster expansion surrogate models, including the atomic cluster expansion and related machine-learned interatomic potential architectures.
Population Subdivision in the Gopher Frog (Rana capito) across the Fragmented Longleaf Pine-Wiregrass Savanna of the Southeastern USA
Delineating genetically distinct population segments of threatened species and quantifying population connectivity are important steps in developing effective conservation and management strategies aimed at preventing extinction. The gopher frog (Rana capito) is a xeric-adapted, pond-breeding species endemic to the Gulf and Atlantic coastal plains of the southeastern United States. This species has experienced extensive habitat loss and fragmentation in the formerly widespread longleaf pine-wiregrass savanna where it lives, resulting in individual abundance declines and population extinctions throughout its range. We used individual-based clustering methods along with Bayesian inference of historical migration based on almost 1500 multilocus microsatellite genotypes to examine genetic structure in this taxon. Clustering analyses identified panhandle and peninsular populations in Florida as distinct genetic clusters separated by the Aucilla River, consistent with the division between the Coastal Plain and peninsular mitochondrial lineages, respectively. Analysis of historical migration indicated an east–west population divergence event followed by immigration to the east. Together, our results indicate that the genetically distinct Coastal Plain and peninsular Florida lineages should be considered separately for conservation and management purposes.
Atacama Cosmology Telescope measurements of a large sample of candidates from the Massive and Distant Clusters of WISE Survey: Sunyaev-Zeldovich effect confirmation of MaDCoWS candidates using ACT
Context. Galaxy clusters are an important tool for cosmology, and their detection and characterization are key goals for current and future surveys. Using data from the Wide-field Infrared Survey Explorer (WISE), the Massive and Distant Clusters of WISE Survey (MaDCoWS) located 2839 significant galaxy overdensities at redshifts 0.7 . z . 1.5, which included extensive follow-up imaging from the Spitzer Space Telescope to determine cluster richnesses. Concurrently, the Atacama Cosmology Telescope (ACT) has produced large area millimeter-wave maps in three frequency bands along with a large catalog of Sunyaev-Zeldovich (SZ)-selected clusters as part of its Data Release 5 (DR5). Aims. We aim to verify and characterize MaDCoWS clusters using measurements of, or limits on, their thermal SZ effect signatures. We also use these detections to establish the scaling relation between SZ mass and the MaDCoWS-defined richness. Methods. Using the maps and cluster catalog from DR5, we explore the scaling between SZ mass and cluster richness. We do this by comparing cataloged detections and extracting individual and stacked SZ signals from the MaDCoWS cluster locations. We use complementary radio survey data from the Very Large Array, submillimeter data from Herschel, and ACT 224 GHz data to assess the impact of contaminating sources on the SZ signals from both ACT and MaDCoWS clusters. We use a hierarchical Bayesian model to fit the mass-richness scaling relation, allowing for clusters to be drawn from two populations: one, a Gaussian centered on the mass-richness relation, and the other, a Gaussian centered on zero SZ signal. Results. We find that MaDCoWS clusters have submillimeter contamination that is consistent with a gray-body spectrum, while the ACT clusters are consistent with no submillimeter emission on average. Additionally, the intrinsic radio intensities of ACT clusters are lower than those of MaDCoWS clusters, even when the ACT clusters are restricted to the same redshift range as the MaDCoWS clusters. We find the best-fit ACT SZ mass versus MaDCoWS richness scaling relation has a slope of p1 = 1.84+0.15 −0.14, where the slope is defined as M ∝ λ p1 15 and λ15 is the richness. We also find that the ACT SZ signals for a significant fraction (∼57%) of the MaDCoWS sample can statistically be described as being drawn from a noise-like distribution, indicating that the candidates are possibly dominated by low-mass and unvirialized systems that are below the mass limit of the ACT sample. Further, we note that a large portion of the optically confirmed ACT clusters located in the same volume of the sky as MaDCoWS are not selected by MaDCoWS, indicating that the MaDCoWS sample is not complete with respect to SZ selection. Finally, we find that the radio loud fraction of MaDCoWS clusters increases with richness, while we find no evidence that the submillimeter emission of the MaDCoWS clusters evolves with richness. Conclusions. We conclude that the original MaDCoWS selection function is not well defined and, as such, reiterate the MaDCoWS collaboration’s recommendation that the sample is suited for probing cluster and galaxy evolution, but not cosmological analyses. We find a best-fit mass-richness relation slope that agrees with the published MaDCoWS preliminary results. Additionally, we find that while the approximate level of infill of the ACT and MaDCoWS cluster SZ signals (1–2%) is subdominant to other sources of uncertainty for current generation experiments, characterizing and removing this bias will be critical for next-generation experiments hoping to constrain cluster masses at the sub-percent level.