Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Improving Dark Energy Constraints Using Low-Redshift Large-Scale Structures

The primary goal of this project was to improve constraints on dark energy measurements by improving our ability to extract cosmological information from low redshift large-scale structures. PI Clowe's project's primary aim was to reduce the bias in measurements of the masses of clusters of galaxies to a level where the evolution of the cluster mass function can be used in the Vera Rubin Observatory's Legacy Survey of Space and Time Dark Energy Science Collaboration survey to improve the accuracy of the measurement of dark energy and other cosmological parameters. Co-PI Seo's project studied observational systematics affecting large-scale clustering of galaxies, which will be used to improve dark energy constraints from the Dark Energy Spectroscopic Instrument (DESI). The cluster lensing project employed a series of simulations and observations of clusters of galaxies to test numerous potential systematic errors in cluster mass measurements using weak gravitational lensing as the accuracy of current weak lensing measurements are more than order of magnitude worse than what is required to use clusters of galaxies for accurate determination of dark energy parameters. PI Clowe and group developed and analyzed simulations to test for, and correct biases introduced in, the weak lensing measurement process. Finally, PI Clowe and group developed a method of detecting clusters using galaxy overdensities and applied the method to the BLISS and DES surveys. The success of spectroscopic dark energy mission such as the extended Baryon Oscillation Spectroscopic Survey (eBOSS) and the Dark Energy Spectroscopic Instrument (DESI) will depend on a thorough understanding of various observational systematics in the target density fluctuations that would give rise to spurious, non-cosmological signals. PI Seo and group developed a deep learning, artificial neural network (ANN) technique that modeled and mitigated such effects, aimed at deriving more robust galaxy clustering signals not only for the baryon acoustic oscillation feature and redshift-space distortions but also for primordial non-Gaussianity constraint.

79 ASTRONOMY AND ASTROPHYSICS↗

CCCP and MENeaCS: (updated) weak-lensing masses for 100 galaxy clusters

ABSTRACT Large area surveys continue to increase the samples of galaxy clusters that can be used to constrain cosmological parameters, provided that the masses of the clusters are measured robustly. To improve the calibration of cluster masses using weak gravitational lensing we present new results for 48 clusters at 0.05 < z < 0.15, observed as part of the Multi Epoch Nearby Cluster Survey, and re-evaluate the mass estimates for 52 clusters from the Canadian Cluster Comparison Project. Updated high-fidelity photometric redshift catalogues of reference deep fields are used in combination with advances in shape measurements and state-of-the-art cluster simulations, yielding an average systematic uncertainty in the lensing signal below 5 per cent, similar to the statistical uncertainty for our cluster sample. We derive a scaling relation with Planck measurements for the full sample and find a bias in the Planck masses of 1 − b = 0.84 ± 0.04 (stat) ±0.05 (syst). We find no statistically significant trend of the mass bias with redshift or cluster mass, but find that different selections could change the bias by up to 0.07. We find a gas fraction of 0.139 ± 0.014 (stat) for eight relaxed clusters in our sample, which can also be used to infer cosmological parameters.

79 ASTRONOMY AND ASTROPHYSICS↗

Methods to Calculate Electronic Excited-State Dynamics for Molecules on Large Metal Clusters with Many States: Ensuring Fast Overlap Calculations and a Robust Choice of Phase

Here, we present an efficient set of methods for propagating excited-state dynamics involving a large number of configuration interaction singles (CIS) or Tamm-Dancoff approximation (TDA) single-reference excited states. Specifically, (i) following Head-Gordon et al., we implement an exact evaluation of the overlap of singly-excited CIS/TDA electronic states at different nuclear geometries using a biorthogonal basis and (ii) we employ a unified protocol for choosing the correct phase for each adiabat at each geometry. For many-electron systems, the combination of these techniques significantly reduces the computational cost of integrating the electronic Schrodinger equation and imposes minimal overhead on top of the underlying electronic structure calculation. As a demonstration, we calculate the electronic excited-state dynamics for a hydrogen molecule scattering off a silver metal cluster, focusing on high-lying excited states, where many electrons can be excited collectively and crossings are plentiful. Interestingly, we find that the high-lying, plasmon-like collective excitation spectrum changes with nuclear dynamics, highlighting the need to simulate non-adiabatic nuclear dynamics and plasmonic excitations simultaneously. In the future, the combination of methods presented here should help theorists build a mechanistic understanding of plasmon-assisted charge transfer and excitation energy relaxation processes near a nanoparticle or metal surface.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust sampling for weak lensing and clustering analyses with the Dark Energy Survey

Recent cosmological analyses rely on the ability to accurately sample from high-dimensional posterior distributions. A variety of algorithms have been applied in the field, but justification of the particular sampler choice and settings is often lacking. Here, we investigate three such samplers to motivate and validate the algorithm and settings used for the Dark Energy Survey (DES) analyses of the first 3 yr (Y3) of data from combined measurements of weak lensing and galaxy clustering. We employ the full DES Year 1 likelihood alongside a much faster approximate likelihood, which enables us to assess the outcomes from each sampler choice and demonstrate the robustness of our full results. We find that the ellipsoidal nested sampling algorithm multinest reports inconsistent estimates of the Bayesian evidence and somewhat narrower parameter credible intervals than the sliced nested sampling implemented in polychord. We compare the findings from multinest and polychord with parameter inference from the Metropolis–Hastings algorithm, finding good agreement. We determine that polychord provides a good balance of speed and robustness for posterior and evidence estimation, and recommend different settings for testing purposes and final chains for analyses with DES Y3 data. Our methodology can readily be reproduced to obtain suitable sampler settings for future surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

Scaling Subspace-Driven Approaches Using Information Fusion

In this work, we seek to exploit the deep structure of multi-modal data to robustly exploit the group subspace distribution of the information using the Convolutional Neural Networks (CNNs) formalism. Upon unfolding the set of subspaces constituting each data modality, and learning their corresponding encoders, an optimized integration of the generated inherent information is carried out to yield a characterization of various classes. Referred to as deep Multimodal Robust Group Subspace Clustering (DRoGSuRe), this approach is compared against the independently developed state-of-the-art approach named Deep Multimodal Subspace Clustering (DMSC). Experiments on different multimodal datasets show that our approach is competitive and more robust in the presence of noise.

Ghanem, Sally↗

Numerical experiments on the clustering of galaxies

Consistent and robust growth rates for disturbances which lead to galaxy clustering are obtainable with a precision of 1-2 percent, in numerical experiments that encompass such conditions as expansion, nonexpansion, and parameter variations. The experiments have given attention to the dominant physical processes of gravitational clustering in an expanding universe of conventional matter, and are based on n-body integrations for 100,000 particles responding self-consistently to forces of self-gravitation with periodic boundary conditions. Observed structures of the scale of galaxy clusters and superclusters are most easily described in terms of matter swept away from growing empty regions. The result of this process has a cellular appearance which resembles clustering of the scale of large voids and superclusters.

Miller, R. H.↗

Unveiling Synergetic Photocatalytic Activity from Heterometallic Ti/Ce Clusters

Titanium-oxo clusters, with their robust structure and suitable optical and electronic properties, have been widely investigated as photocatalysts. Heterometallic Ti/M-oxo clusters provide additional tunability and functionality, which enable systematic structure-activity investigations to elucidate reaction mechanisms and improve catalyst de-sign. Incorporating cerium into Ti-oxo clusters can provide additional redox (Ce IV /Ce III ) and oxygen harvesting ability, but to date only limited number of structurally defined titanium-cerium (Ti/Ce) clusters have been reported due to their synthetic challenges. Herein, we report the synthesis and photocatalytic properties of two structurally defined Ti/Ce-oxo clusters, Ti 8 Ce 2 (BA) 16 and Ti 9 Ce 4 (BA) 20 , as well as a TiCe-BA cluster with a calculated formula of Ti 20 Ce 9 O 36 (BA) 42 . Photocatalytic study of these clusters demonstrates that the amount of Ce 3+ species greatly impacts its photocatalytic oxidation performance, and their superior photocatalytic reactivity towards aerobic alcohol oxidation can be contributed to the synergistic effects of the multiple radical species generated upon light absorption. Furthermore, this work represents a significant milestone in the construction of stable Ti/Ce-oxo clusters, enriching the current library of known heterometallic Ti/M-oxo clusters, and providing a series of crystalline materials with great promise of photo-luminescence and photovoltaic chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancements in cerium/titanium metal-organic frameworks: Unparalleled stability in CO oxidation

Due to the excellent catalytic properties of Ce-based materials, the development of thermally stable metal-organic frameworks (MOFs) based on Ce-oxo clusters has attracted significant attention but remains challenging. In this work, we report the synthesis of an unreported Ce 4 Ti 2 -TMA (Ce IV 4 Ti IV 2 O 4 (OH) 4 (C(CH 3 ) 3 COO) 12 ·3H 2 O·3MeCN) cluster, which serves as an ideal source for the assembly of robust Ce/Ti-MOFs. Using this cluster, we constructed two isostructural MOFs, denoted as NU-3000 and NU-3001. Single-crystal X-ray diffraction analysis confirms these MOFs as mesoporous structures with 12-coordinated Ce 3 Ti 3 nodes. Furthermore, structural analysis reveals a plane triangular node structure that likely contributes to the excellent thermal stability of these MOFs. Finally, both MOFs show catalytic activity toward high-temperature (250°C) CO oxidation and maintain significant porosity, emphasizing the thermal stability of these materials under practical catalytic conditions. Furthermore, the straightforward synthesis of thermally robust Ce/Ti-MOFs from the Ce 4 Ti 2 -TMA cluster will pave the way for future Ce/Ti-MOF-based catalyst development.

36 MATERIALS SCIENCE↗

Brightest cluster galaxies are statistically special from z = 0.3 to z = 1

ABSTRACT We study brightest cluster galaxies (BCGs) in ∼5000 galaxy clusters from the Hyper Suprime-Cam (HSC) Subaru Strategic Program. The sample is selected over an area of 830 deg2 and is uniformly distributed in redshift over the range of z = 0.3−1.0. The clusters have stellar masses in the range of 1011.8−1012.9M⊙. We compare the stellar mass of the BCGs in each cluster to what we would expect if their masses were drawn from the mass distribution of the other member galaxies of the clusters. The BCGs are found to be ‘special’, in the sense that they are not consistent with being a statistical extreme of the mass distribution of other cluster galaxies. This result is robust over the full range of cluster stellar masses and redshifts in the sample, indicating that BCGs are special up to a redshift of z = 1.0. However, BCGs with a large separation from the centre of the cluster are found to be consistent with being statistical extremes of the cluster member mass distribution. We discuss the implications of these findings for BCG formation scenarios.

Dalal, Roohi (ORCID:0000000279989899)↗

Detection of spatial clustering in the 1000 richest SDSS DR8 redMaPPer clusters with nearest neighbor distributions

ABSTRACT Distances to the k-nearest-neighbor (kNN) data points from volume-filling query points are a sensitive probe of spatial clustering. Here, we present the first application of kNN summary statistics to observational clustering measurement, using the 1000 richest redMaPPer clusters (0.1 ≤ z ≤ 0.3) from the SDSS DR8 catalog. A clustering signal is defined as a difference in the cumulative distribution functions (CDFs) of kNN distances from fixed query points to the observed clusters versus a set of unclustered random points. We find that the k = 1, 2-NN CDFs of redMaPPer deviate significantly from the randoms’ across scales of 35 to 155 Mpc, which is a robust signature of clustering. In addition to kNN, we also measure the two-point correlation function for the same set of redMaPPer clusters versus random points, which shows a noisier and less significant clustering signal within the same radial scales. Quantitatively, the χ2 distribution for both the kNN-CDFs and the two-point correlation function measured on the randoms peak at χ2 ∼ 50 (null hypothesis), whereas the kNN-CDFs (χ2 ∼ 300, p = 1.54 × 10−36) pick up a much more significant clustering signal than the two-point function (χ2 ∼ 100, p = 1.16 × 10−6) when measured on redMaPPer. Finally, the measured 3NN and 4NN CDFs deviate from the predicted k = 3, 4-NN CDFs assuming an ideal Gaussian field, indicating a non-Gaussian clustering signal for redMaPPer clusters, although its origin might not be cosmological due to observational systematics. Therefore, kNN serves as a more sensitive probe of clustering complementary to the two point correlation function, providing a novel approach for constraining cosmology and galaxy–halo connection.

79 ASTRONOMY AND ASTROPHYSICS↗

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

SPT-3G D1: Compton-$y$ maps using data from the SPT-3G and Planck surveys

We present thermal Sunyaev-Zel'dovich (tSZ) Compton-$y$ parameter maps constructed from two years (2019-2020) of observations with the South Pole Telescope (SPT) third-generation camera, SPT-3G, combined with data from the Planck satellite. Using a linear combination (LC) pipeline, we obtain a suite of reconstructions that explore different trade-offs between statistical sensitivity and suppression of astrophysical contaminants, including minimum-variance, CMB-deprojected, and CIB-deprojected $y$-maps. We validate these maps through different statistical techniques such as auto- and cross-power spectra with large-scale structure tracers as well as stacking on cluster locations. These tests are used to understand the balance between noise and astrophysical foreground residuals (such as the CIB) in combination with the recovery of the tSZ signal for different maps. For example, results from stacking at the location of clusters confirm the robustness of the recovered tSZ signal over the $\sim 1500\: {\rm deg}^2$ SPT-3G survey field used in this analysis. The high-resolution and low-noise maps produced here provide an important cosmological tool for future studies, including measurements of the Compton-$y$ map power spectrum, cross-correlations with other tracers of the large-scale structure, detailed modeling of cluster pressure profiles, and study of the thermodynamic state of the baryons in the Universe.

Maniyar, A. S. [Harvard-Smithsonian Ctr. Astrophys↗

Constraining Cosmology with Simulation-based inference and Optical Galaxy Cluster Abundance

We test the robustness of simulation-based inference (SBI) in the context of cosmological parameter estimation from galaxy cluster counts and masses in simulated optical datasets. We construct ``simulations'' using analytical models for the galaxy cluster halo mass function (HMF) and for the observed richness (number of observed member galaxies) to train and test the SBI method. We compare the SBI parameter posterior samples to those from an MCMC analysis that uses the same analytical models to construct predictions of the observed data vector. The two methods exhibit comparable performance, with reliable constraints derived for the primary cosmological parameters, ($\Omega_m$ and $\sigma_8$), and richness-mass relation parameters. We also perform out-of-domain tests with observables constructed from galaxy cluster-sized halos in the Quijote simulations. Again, the SBI and MCMC results have comparable posteriors, with similar uncertainties and biases. Unsurprisingly, upon evaluating the SBI method on thousands of simulated data vectors that span the parameter space, SBI exhibits worsened posterior calibration metrics in the out-of-domain application. We note that such calibration tests with MCMC is less computationally feasible and highlight the potential use of SBI to stress-test limitations of analytical models, such as in the use for constructing models for inference with MCMC.

79 ASTRONOMY AND ASTROPHYSICS↗

Source identification by non-negative matrix factorization combined with semi-supervised clustering

Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.

97 MATHEMATICS AND COMPUTING↗

Source identification by non-negative matrix factorization combined with semi-supervised clustering

Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.

Alexandrov, Boian S.↗

HLA-Clus: HLA class I clustering based on 3D structure

In a previous paper, we classified populated HLA class I alleles into supertypes and subtypes based on the similarity of 3D landscape of peptide binding grooves, using newly defined structure distance metric and hierarchical clustering approach. Compared to other approaches, our method achieves higher correlation with peptide binding specificity, intra-cluster similarity (cohesion), and robustness. Here we introduce HLA-Clus, a Python package for clustering HLA Class I alleles using the method we developed recently and describe additional features including a new nearest neighbor clustering method that facilitates clustering based on user-defined criteria. The HLA-Clus pipeline includes three stages: First, HLA Class I structural models are coarse grained and transformed into clouds of labeled points. Second, similarities between alleles are determined using a newly defined structure distance metric that accounts for spatial and physicochemical similarities. Finally, alleles are clustered via hierarchical or nearest-neighbor approaches. We also interfaced HLA-Clus with the peptide:HLA affinity predictor MHCnuggets. By using the nearest neighbor clustering method to select optimal allele-specific deep learning models in MHCnuggets, the average accuracy of peptide binding prediction of rare alleles was improved. The HLA-Clus package offers a solution for characterizing the peptide binding specificities of a large number of HLA alleles. This method can be applied in HLA functional studies, such as the development of peptide affinity predictors, disease association studies, and HLA matching for grafting. HLA-Clus is freely available at our GitHub repository (https://github.com/yshen25/HLA-Clus).

59 BASIC BIOLOGICAL SCIENCES↗