Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Robust sampling for weak lensing and clustering analyses with the Dark Energy Survey

Recent cosmological analyses rely on the ability to accurately sample from high-dimensional posterior distributions. A variety of algorithms have been applied in the field, but justification of the particular sampler choice and settings is often lacking. Here, we investigate three such samplers to motivate and validate the algorithm and settings used for the Dark Energy Survey (DES) analyses of the first 3 yr (Y3) of data from combined measurements of weak lensing and galaxy clustering. We employ the full DES Year 1 likelihood alongside a much faster approximate likelihood, which enables us to assess the outcomes from each sampler choice and demonstrate the robustness of our full results. We find that the ellipsoidal nested sampling algorithm multinest reports inconsistent estimates of the Bayesian evidence and somewhat narrower parameter credible intervals than the sliced nested sampling implemented in polychord. We compare the findings from multinest and polychord with parameter inference from the Metropolis–Hastings algorithm, finding good agreement. We determine that polychord provides a good balance of speed and robustness for posterior and evidence estimation, and recommend different settings for testing purposes and final chains for analyses with DES Y3 data. Our methodology can readily be reproduced to obtain suitable sampler settings for future surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Robustness of topological insulating phase against vacancy, vacancy cluster, and grain boundary bulk defects

One distinguished property of the topological insulator (TI) is its robust quantized edge conductance against edge defect. However, this robustness, underlined by the topological principle of bulk-boundary correspondence, is conditioned by assuming a perfect bulk. Here, we investigate the robustness of the TI phase against bulk defects, including vacancy (VA), vacancy cluster (VC), and grain boundary (GB), instead of edge defect. In this work, based on a tight-binding model analysis, we show that a two-dimensional (2D) TI phase, as characterized by a nonzero spin Bott index, will vanish beyond a critical VA concentration ($n^\text{c}_\text{v}$). Generally, $n^\text{c}_\text{v}$ decreases monotonically with the decreasing topological gap induced by spin-orbit coupling. Interestingly, the $n^\text{c}_\text{v}$ to destroy the topological order, namely, the robustness of the TI phase, is shown to be increased by the presence of VCs but decreased by GBs. As a specific example of a large-gap 2D TI, we further show that the surface-supported monolayer Bi can sustain a nontrivial topology up to $n^\text{c}_\text{v}$ ~ 17%, based on a density-functional theory–Wannier-function calculation. Our findings should provide useful guidance for future experimental studies of effects of defects on TIs.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Scaling Subspace-Driven Approaches Using Information Fusion

In this work, we seek to exploit the deep structure of multi-modal data to robustly exploit the group subspace distribution of the information using the Convolutional Neural Networks (CNNs) formalism. Upon unfolding the set of subspaces constituting each data modality, and learning their corresponding encoders, an optimized integration of the generated inherent information is carried out to yield a characterization of various classes. Referred to as deep Multimodal Robust Group Subspace Clustering (DRoGSuRe), this approach is compared against the independently developed state-of-the-art approach named Deep Multimodal Subspace Clustering (DMSC). Experiments on different multimodal datasets show that our approach is competitive and more robust in the presence of noise.

Ghanem, Sally↗

Unveiling Synergetic Photocatalytic Activity from Heterometallic Ti/Ce Clusters

Titanium-oxo clusters, with their robust structure and suitable optical and electronic properties, have been widely investigated as photocatalysts. Heterometallic Ti/M-oxo clusters provide additional tunability and functionality, which enable systematic structure-activity investigations to elucidate reaction mechanisms and improve catalyst de-sign. Incorporating cerium into Ti-oxo clusters can provide additional redox (Ce IV /Ce III ) and oxygen harvesting ability, but to date only limited number of structurally defined titanium-cerium (Ti/Ce) clusters have been reported due to their synthetic challenges. Herein, we report the synthesis and photocatalytic properties of two structurally defined Ti/Ce-oxo clusters, Ti 8 Ce 2 (BA) 16 and Ti 9 Ce 4 (BA) 20 , as well as a TiCe-BA cluster with a calculated formula of Ti 20 Ce 9 O 36 (BA) 42 . Photocatalytic study of these clusters demonstrates that the amount of Ce 3+ species greatly impacts its photocatalytic oxidation performance, and their superior photocatalytic reactivity towards aerobic alcohol oxidation can be contributed to the synergistic effects of the multiple radical species generated upon light absorption. Furthermore, this work represents a significant milestone in the construction of stable Ti/Ce-oxo clusters, enriching the current library of known heterometallic Ti/M-oxo clusters, and providing a series of crystalline materials with great promise of photo-luminescence and photovoltaic chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancements in cerium/titanium metal-organic frameworks: Unparalleled stability in CO oxidation

Due to the excellent catalytic properties of Ce-based materials, the development of thermally stable metal-organic frameworks (MOFs) based on Ce-oxo clusters has attracted significant attention but remains challenging. In this work, we report the synthesis of an unreported Ce 4 Ti 2 -TMA (Ce IV 4 Ti IV 2 O 4 (OH) 4 (C(CH 3 ) 3 COO) 12 ·3H 2 O·3MeCN) cluster, which serves as an ideal source for the assembly of robust Ce/Ti-MOFs. Using this cluster, we constructed two isostructural MOFs, denoted as NU-3000 and NU-3001. Single-crystal X-ray diffraction analysis confirms these MOFs as mesoporous structures with 12-coordinated Ce 3 Ti 3 nodes. Furthermore, structural analysis reveals a plane triangular node structure that likely contributes to the excellent thermal stability of these MOFs. Finally, both MOFs show catalytic activity toward high-temperature (250°C) CO oxidation and maintain significant porosity, emphasizing the thermal stability of these materials under practical catalytic conditions. Furthermore, the straightforward synthesis of thermally robust Ce/Ti-MOFs from the Ce 4 Ti 2 -TMA cluster will pave the way for future Ce/Ti-MOF-based catalyst development.

36 MATERIALS SCIENCE↗

Brightest cluster galaxies are statistically special from z = 0.3 to z = 1

ABSTRACT We study brightest cluster galaxies (BCGs) in ∼5000 galaxy clusters from the Hyper Suprime-Cam (HSC) Subaru Strategic Program. The sample is selected over an area of 830 deg2 and is uniformly distributed in redshift over the range of z = 0.3−1.0. The clusters have stellar masses in the range of 1011.8−1012.9M⊙. We compare the stellar mass of the BCGs in each cluster to what we would expect if their masses were drawn from the mass distribution of the other member galaxies of the clusters. The BCGs are found to be ‘special’, in the sense that they are not consistent with being a statistical extreme of the mass distribution of other cluster galaxies. This result is robust over the full range of cluster stellar masses and redshifts in the sample, indicating that BCGs are special up to a redshift of z = 1.0. However, BCGs with a large separation from the centre of the cluster are found to be consistent with being statistical extremes of the cluster member mass distribution. We discuss the implications of these findings for BCG formation scenarios.

Dalal, Roohi (ORCID:0000000279989899)↗

Detection of spatial clustering in the 1000 richest SDSS DR8 redMaPPer clusters with nearest neighbor distributions

ABSTRACT Distances to the k-nearest-neighbor (kNN) data points from volume-filling query points are a sensitive probe of spatial clustering. Here, we present the first application of kNN summary statistics to observational clustering measurement, using the 1000 richest redMaPPer clusters (0.1 ≤ z ≤ 0.3) from the SDSS DR8 catalog. A clustering signal is defined as a difference in the cumulative distribution functions (CDFs) of kNN distances from fixed query points to the observed clusters versus a set of unclustered random points. We find that the k = 1, 2-NN CDFs of redMaPPer deviate significantly from the randoms’ across scales of 35 to 155 Mpc, which is a robust signature of clustering. In addition to kNN, we also measure the two-point correlation function for the same set of redMaPPer clusters versus random points, which shows a noisier and less significant clustering signal within the same radial scales. Quantitatively, the χ2 distribution for both the kNN-CDFs and the two-point correlation function measured on the randoms peak at χ2 ∼ 50 (null hypothesis), whereas the kNN-CDFs (χ2 ∼ 300, p = 1.54 × 10−36) pick up a much more significant clustering signal than the two-point function (χ2 ∼ 100, p = 1.16 × 10−6) when measured on redMaPPer. Finally, the measured 3NN and 4NN CDFs deviate from the predicted k = 3, 4-NN CDFs assuming an ideal Gaussian field, indicating a non-Gaussian clustering signal for redMaPPer clusters, although its origin might not be cosmological due to observational systematics. Therefore, kNN serves as a more sensitive probe of clustering complementary to the two point correlation function, providing a novel approach for constraining cosmology and galaxy–halo connection.

79 ASTRONOMY AND ASTROPHYSICS↗

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

SPT-3G D1: Compton-$y$ maps using data from the SPT-3G and Planck surveys

We present thermal Sunyaev-Zel'dovich (tSZ) Compton-$y$ parameter maps constructed from two years (2019-2020) of observations with the South Pole Telescope (SPT) third-generation camera, SPT-3G, combined with data from the Planck satellite. Using a linear combination (LC) pipeline, we obtain a suite of reconstructions that explore different trade-offs between statistical sensitivity and suppression of astrophysical contaminants, including minimum-variance, CMB-deprojected, and CIB-deprojected $y$-maps. We validate these maps through different statistical techniques such as auto- and cross-power spectra with large-scale structure tracers as well as stacking on cluster locations. These tests are used to understand the balance between noise and astrophysical foreground residuals (such as the CIB) in combination with the recovery of the tSZ signal for different maps. For example, results from stacking at the location of clusters confirm the robustness of the recovered tSZ signal over the $\sim 1500\: {\rm deg}^2$ SPT-3G survey field used in this analysis. The high-resolution and low-noise maps produced here provide an important cosmological tool for future studies, including measurements of the Compton-$y$ map power spectrum, cross-correlations with other tracers of the large-scale structure, detailed modeling of cluster pressure profiles, and study of the thermodynamic state of the baryons in the Universe.

Maniyar, A. S. [Harvard-Smithsonian Ctr. Astrophys↗

Constraining Cosmology with Simulation-based inference and Optical Galaxy Cluster Abundance

We test the robustness of simulation-based inference (SBI) in the context of cosmological parameter estimation from galaxy cluster counts and masses in simulated optical datasets. We construct ``simulations'' using analytical models for the galaxy cluster halo mass function (HMF) and for the observed richness (number of observed member galaxies) to train and test the SBI method. We compare the SBI parameter posterior samples to those from an MCMC analysis that uses the same analytical models to construct predictions of the observed data vector. The two methods exhibit comparable performance, with reliable constraints derived for the primary cosmological parameters, ($\Omega_m$ and $\sigma_8$), and richness-mass relation parameters. We also perform out-of-domain tests with observables constructed from galaxy cluster-sized halos in the Quijote simulations. Again, the SBI and MCMC results have comparable posteriors, with similar uncertainties and biases. Unsurprisingly, upon evaluating the SBI method on thousands of simulated data vectors that span the parameter space, SBI exhibits worsened posterior calibration metrics in the out-of-domain application. We note that such calibration tests with MCMC is less computationally feasible and highlight the potential use of SBI to stress-test limitations of analytical models, such as in the use for constructing models for inference with MCMC.

79 ASTRONOMY AND ASTROPHYSICS↗

Source identification by non-negative matrix factorization combined with semi-supervised clustering

Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.

97 MATHEMATICS AND COMPUTING↗

Source identification by non-negative matrix factorization combined with semi-supervised clustering

Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.

Alexandrov, Boian S.↗

HLA-Clus: HLA class I clustering based on 3D structure

In a previous paper, we classified populated HLA class I alleles into supertypes and subtypes based on the similarity of 3D landscape of peptide binding grooves, using newly defined structure distance metric and hierarchical clustering approach. Compared to other approaches, our method achieves higher correlation with peptide binding specificity, intra-cluster similarity (cohesion), and robustness. Here we introduce HLA-Clus, a Python package for clustering HLA Class I alleles using the method we developed recently and describe additional features including a new nearest neighbor clustering method that facilitates clustering based on user-defined criteria. The HLA-Clus pipeline includes three stages: First, HLA Class I structural models are coarse grained and transformed into clouds of labeled points. Second, similarities between alleles are determined using a newly defined structure distance metric that accounts for spatial and physicochemical similarities. Finally, alleles are clustered via hierarchical or nearest-neighbor approaches. We also interfaced HLA-Clus with the peptide:HLA affinity predictor MHCnuggets. By using the nearest neighbor clustering method to select optimal allele-specific deep learning models in MHCnuggets, the average accuracy of peptide binding prediction of rare alleles was improved. The HLA-Clus package offers a solution for characterizing the peptide binding specificities of a large number of HLA alleles. This method can be applied in HLA functional studies, such as the development of peptide affinity predictors, disease association studies, and HLA matching for grafting. HLA-Clus is freely available at our GitHub repository (https://github.com/yshen25/HLA-Clus).

59 BASIC BIOLOGICAL SCIENCES↗

Multiresolution Quantum Chemistry: Nonlinear Response Properties at the Basis Set Limit

We benchmark the accuracy of Dunning correlation-consistent Gaussian basis sets for computing frequencydependent second-order hyperpolarizabilities relevant to second-harmonic generation (SHG), using multiresolution analysis (MRA) as a reference. Basis set errors are analyzed using a unit-sphere representation of the effective hyperpolarizability vector, enabling direct assessment of directional error structure. We introduce a relative RMS total error metric that integrates directional deviations over the unit sphere and complement it with signed projection errors that distinguish over- and underestimation. Unsupervised clustering based on these signed directional metrics reveals four distinct convergence behaviors across a set of 68 molecules. Unitsphere visualizations of representative systems show that basis set errors are often highly anisotropic and localized along specific bond directions, even when global error measures appear small. Doubly augmented basis sets consistently outperform singly augmented ones, and core-polarization functions are required for uniform convergence in second-row systems. Overall, this work demonstrates that directional analysis combined with clustering provides a robust framework for understanding basis set convergence in nonlinear optical response properties.

Basis sets↗

Isolated and H 2 -reduced Anderson clusters catalyse low-temperature hydrogenation of CO 2 to methanol

CO 2 hydrogenation, especially to methanol, is crucial to establishing sustainable closed-loop systems for carbon utilization. However, the difficulties of CO 2 activation at low temperatures and the ambiguity of structure–activity correlations are obstacles to reducing the energy consumption of the hydrogenation process. Here we report that molecularly defined Anderson PtMo 6 O 24 clusters, sited within a robust metal–organic framework, are catalytic for low-temperature CO 2 hydrogenation. The performance of the cluster showed no signs of decay in either its activity or methanol selectivity over 3,600 h at 180 °C. It also achieves a per-pass yield exceeding that of state-of-the-art heterogeneous catalysts under similar conditions. Combined in situ spectroscopy and density functional theory calculations demonstrated that CH 3 OH formation is dominated by the reverse water–gas shift and subsequent CO* hydrogenation pathway, while the HCOO* pathway may serve as a supplementary route. The well-defined cluster structure offers an ideal model for elucidating structure–activity correlations and opens exciting avenues for the rational design of high-activity, low-temperature catalysts for CO 2 hydrogenation.

heterogeneous catalysis↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

SPT clusters with DES and HST weak lensing. II. Cosmological constraints from the abundance of massive halos

We present cosmological constraints from the abundance of galaxy clusters selected via the thermal Sunyaev-Zel’dovich (SZ) effect in South Pole Telescope (SPT) data with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). The cluster sample is constructed from the combined SPT-SZ, SPTpol ECS, and SPTpol 500d surveys, and comprises 1,005 confirmed clusters in the redshift range 0.25–1.78 over a total sky area of 5200 deg 2 . We use DES Year 3 weak-lensing data for 688 clusters with redshifts 𝑧 < 0.95 and HST weak-lensing data for 39 clusters with 0.6 < 𝑧 < 1.7. The weak-lensing measurements enable robust mass measurements of sample clusters and allow us to empirically constrain the SZ observable-mass relation without having to make strong assumptions about, e.g., the hydrodynamical state of the clusters. For a flat Λ⁢ CDM cosmology, and marginalizing over the sum of massive neutrinos, we measure Ω m = 0.286 ± 0.032, 𝜎 8 = 0.817 ± 0.026, and the parameter combination 𝜎 8 ⁢(Ω m /0.3) 0.25 = 0.805 ± 0.016. Our measurement of 𝑆 8 ≡ 𝜎 8 ⁢$\sqrt{Ω_{m}/0.3}$ = 0.795 ± 0.029 and the constraint from Planck CMB anisotropies (2018 TT, TE, EE+lowE) differ by 1.1⁢𝜎. In combination with that Planck dataset, we place a 95% upper limit on the sum of neutrino masses ∑𝑚 𝜈 < 0.18 eV. When additionally allowing the dark energy equation of state parameter 𝑤 to vary, we obtain 𝑤 = −1.45 ± 0.31 from our cluster-based analysis. In combination with Planck data, we measure 𝑤 =−1.3⁢4$^{+0.22}_{−0.15}$, or a 2.2⁢𝜎 difference with a cosmological constant. We use the cluster abundance to measure 𝜎8 in five redshift bins between 0.25 and 1.8, and we find the results to be consistent with structure growth as predicted by the Λ⁢ CDM model fit to Planck primary CMB data.

79 ASTRONOMY AND ASTROPHYSICS↗