Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Prediction of the Al-rich part of the Al-Cu phase diagram using cluster expansion and statistical mechanics

In this work, the thermodynamic properties of α-Al and other phases (GP zones, θ'', θ' and θ) in the Al-rich part of the Al-Cu system have been obtained by means of the cluster expansion formalism in combination with statistical mechanics. This information was used to build the Al-rich part of the Al-Cu phase-diagram taking into account vibrational entropic contributions for θ', as those of the other phases were negligible. The simulation predictions of the phase boundaries between α-Al and either θ'', θ' or θ phases as a function of temperature are in good agreement with experimental data and extend the phase boundaries to a wider temperature range. The DFT calculations reveal the presence of a number of metastable Guinier-Preston-zone type configurations that may coexist with α-Al and θ'' at low temperatures. They also demonstrate that θ' is the stable phase below 550K but it is replaced by θ above this temperature due to the vibrational entropic contribution to the Gibbs energy of θ'. This work shows how the combination of cluster expansion and statistical mechanics can be used to expand our knowledge of the phase diagram of metallic alloys and to provide Gibbs free energies of different phases that can be used as input in mesoscale simulations of precipitation.

36 MATERIALS SCIENCE↗

Nearest neighbour distributions: New statistical measures for cosmological clustering

ABSTRACT The use of summary statistics beyond the two-point correlation function to analyse the non-Gaussian clustering on small scales, and thereby, increasing the sensitivity to the underlying cosmological parameters, is an active field of research in cosmology. In this paper, we explore a set of new summary statistics – the k-Nearest Neighbour Cumulative Distribution Functions (kNN-CDF). This is the empirical cumulative distribution function of distances from a set of volume-filling, Poisson distributed random points to the k-nearest data points, and is sensitive to all connected N-point correlations in the data. The kNN-CDF can be used to measure counts in cell, void probability distributions, and higher N-point correlation functions, all using the same formalism exploiting fast searches with spatial tree data structures. We demonstrate how it can be computed efficiently from various data sets – both discrete points, and the generalization for continuous fields. We use data from a large suite of N-body simulations to explore the sensitivity of this new statistic to various cosmological parameters, compared to the two-point correlation function, while using the same range of scales. We demonstrate that the use of kNN-CDF improves the constraints on the cosmological parameters by more than a factor of 2 when applied to the clustering of dark matter in the range of scales between 10 and $40\, h^{-1}\, {\rm Mpc}$. We also show that relative improvement is even greater when applied on the same scales to the clustering of haloes in the simulations at a fixed number density, both in real space, as well as in redshift space. Since the kNN-CDF are sensitive to all higher order connected correlation functions in the data, the gains over traditional two-point analyses are expected to grow as progressively smaller scales are included in the analysis of cosmological data, provided the higher order correlation functions are sensitive to cosmology on the scales of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Detection of the significant impact of source clustering on higher order statistics with DES Year 3 weak gravitational lensing data

We measure the impact of source galaxy clustering on higher order summary statistics of weak gravitational lensing data. By comparing simulated data with galaxies that either trace or do not trace the underlying density field, we show that this effect can exceed measurement uncertainties for common higher order statistics for certain analysis choices. We evaluate the impact on different weak lensing observables, finding that third moments and wavelet phase harmonics are more affected than peak count statistics. Using Dark Energy Survey (DES) Year 3 (Y3) data, we construct null tests for the source-clustering-free case, finding a p-value of p = 4 × 10 −3 (2.6σ) using third-order map moments and p = 3 × 10 −11 (6.5σ) using wavelet phase harmonics. The impact of source clustering on cosmological inference can be either included in the model or minimized through ad hoc procedures (e.g. scale cuts). We verify that the procedures adopted in existing DES Y3 cosmological analyses were sufficient to render this effect negligible. Failing to account for source clustering can significantly impact cosmological inference from higher order gravitational lensing statistics, e.g. higher order N-point functions, wavelet-moment observables, and deep learning or field-level summary statistics of weak lensing maps.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward Accurate Modeling of Galaxy Clustering on Small Scales: Constraining the Galaxy-halo Connection with Optimal Statistics

Applying halo models to analyze the small-scale clustering of galaxies is a proven method for characterizing the connection between galaxies and their host halos. Such works are often plagued by systematic errors or limited to clustering statistics that can be predicted analytically. In this work, we employ a numerical mock-based modeling procedure to examine the clustering of Sloan Digital Sky Survey DR7 galaxies. We apply a standard halo occupation distribution (HOD) model to dark matter only simulations with a ΛCDM cosmology. To constrain the theoreStical models, we utilize a combination of galaxy number density and selected scales of the projected correlation function, redshift-space correlation function, group multiplicity function, average group velocity dispersion, mark correlation function, and counts-in-cells statistics. We design an algorithm to choose an optimal combination of measurements that yields tight and accurate constraints on our model parameters. Compared to previous work using fewer clustering statistics, we find a significant improvement in the constraints on all parameters of our halo model for two different luminosity-threshold galaxy samples. Most interestingly, we obtain unprecedented high-precision constraints on the scatter in the relationship between galaxy luminosity and halo mass. However, our best-fit model results in significant tension (>4σ) for both samples, indicating the need to add second-order features to the standard HOD model. To guarantee the robustness of these results, we perform an extensive analysis of the systematic and statistical errors in our modeling procedure, including a first of its kind study of the sensitivity of our constraints to changes in the halo mass function due to baryonic physics.

79 ASTRONOMY AND ASTROPHYSICS↗

2D k -th nearest neighbour statistics: a highly informative probe of galaxy clustering

ABSTRACT Beyond standard summary statistics are necessary to summarize the rich information on non-linear scales in the era of precision galaxy clustering measurements. For the first time, we introduce the 2D k-th nearest neighbour (kNN) statistics as a summary statistic for discrete galaxy fields. This is a direct generalization of the standard 1D kNN by disentangling the projected galaxy distribution from the redshift-space distortion signature along the line-of-sight. We further introduce two different flavours of 2D kNNs that trace different aspects of the galaxy field: the standard flavour which tabulates the distances between galaxies and random query points, and a ‘DD’ flavour that tabulates the distances between galaxies and galaxies. We showcase the 2D kNNs’ strong constraining power both through theoretical arguments and by testing on realistic galaxy mocks. Theoretically, we show that 2D kNNs are computationally efficient and directly generate other statistics such as the popular two-point correlation function (2PCF), voids probability function, and counts-in-cell statistics. In a more practical test, we apply the 2D kNN statistics to simulated galaxy mocks that fold in a large range of observational realism and recover parameters of the underlying extended halo occupation distribution (HOD) model that includes velocity bias and galaxy assembly bias. We find unbiased and significantly tighter constraints on all aspects of the HOD model with the 2D kNNs, both compared to the standard 1D kNN, and the classical redshift-space 2PCF.

79 ASTRONOMY AND ASTROPHYSICS↗

Hydrated Anions: From Clusters to Bulk Solution with Quasi-Chemical Theory

The interactions of hydrated ions with molecular and macromolecular solution and interface partners are strong on a chemical energy scale. Here we recount the foremost ab initio theory for the evaluation of the hydration free energies of ions, namely, quasi-chemical theory (QCT). We focus on anions, particularly halides but also the hydroxide anion, because they have been outstanding challenges for all theories. For example, this work supports understanding the high selectivity for F – over Cl – in fluoride-selective ion channels despite the identical charge and the size similarity of these ions. QCT is built by the identification of inner-shell clusters, separate treatment of those clusters, and then the integration of those results into the broader-scale solution environment. Recent work has focused on a close comparison with mass-spectrometric measurements of ion-hydration equilibria. We delineate how ab initio molecular dynamics (AIMD) calculations on ion-hydration clusters, elementary statistical thermodynamics, and electronic structure calculations on cluster structures sampled from the AIMD calculations obtain just the free energies extracted from the cluster experiments. That theory–experiment comparison has not been attempted before the work discussed here, but the agreement is excellent with moderate computational effort. This agreement reinforces both theory and experiment and provides a numerically accurate inner-shell contribution to QCT. The inner-shell complexes involving heavier halides display strikingly asymmetric hydration clusters. Asymmetric hydration structures can be problematic for the evaluation of the QCT outer-shell contribution with the polarizable continuum model (PCM). Nevertheless, QCT provides a favorable setting for the exploitation of PCM when the inner-shell material shields the ion from the outer solution environment. For the more asymmetrically hydrated, and thus less effectively shielded, heavier halide ions clustered with waters, the PCM is less satisfactory. We therefore investigate an inverse procedure in which the inner-shell structures are sampled from readily available AIMD calculations on the bulk solutions. This inverse procedure is a remarkable improvement; our final results are in close agreement with a standard tabulation of hydration free energies, and the final composite results are independent of the coordination number on the chemical energy scale of relevance, as they should be. Finally, a comparison of anion hydration structure in clusters and bulk solutions from AIMD simulations emphasize some differences: the asymmetries of bulk solution inner-shell structures are moderated compared with clusters but are still present, and inner hydration shells fill to slightly higher average coordination numbers in bulk solution than in clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient computation of N -point correlation functions in D dimensions

We present efficient algorithms for computing the N-point correlation functions (NPCFs) of random fields in arbitrary D-dimensional homogeneous and isotropic spaces. Such statistics appear throughout the physical sciences and provide a natural tool to describe stochastic processes. Typically, algorithms for computing the NPCF components have $\mathscr O$(n N ) complexity (for a dataset containing n particles); their application is thus computationally infeasible unless N is small. By projecting the statistic onto a suitably defined angular basis, we show that the estimators can be written in a separable form, with complexity $\mathscr O$(n 2 ) or $\mathscr O$(n g log n g ) if evaluated using a Fast Fourier Transform on a grid of size n g . Our decomposition is built upon the D-dimensional hyperspherical harmonics; these form a complete basis on the (D – 1) sphere and are intrinsically related to angular momentum operators. Concatenation of (N – 1) such harmonics gives states of definite combined angular momentum, forming a natural separable basis for the NPCF. As N and D grow, the number of basis components quickly becomes large, providing a practical limitation to this (and all other) approaches: However, the dimensionality is greatly reduced in the presence of symmetries; for example, isotropic correlation functions require only states of zero combined angular momentum. We provide a Julia package implementing our estimators and show how they can be applied to a variety of scenarios within cosmology and fluid dynamics. The efficiency of such estimators will allow higher-order correlators to become a standard tool in the analysis of random fields.

97 MATHEMATICS AND COMPUTING↗

Network Modeling of Complex Data Sets

We demonstrate a selection of network and machine learning techniques useful in the analysis of complex datasets, including 2-way similarity networks, Markov clustering, enrichment statistical networks, FCROS differential analysis, and random forests. We demonstrate each of these techniques on the Populus trichocarpa gene expression atlas.

Jones, Piet C.↗

Operation and performance of VRF systems: Mining a large-scale dataset

The energy consumption of air-conditioning systems has gained increasing attention as it contributes significantly to the global building energy use. The variable refrigerant flow (VRF) system is a common air-conditioning system applied widely in residential and office buildings in China. Understanding the actual operation and performance of VRF systems is fundamental for the energy-efficient design and operation of VRF systems. Previous research on VRF system operation used either limited field data covering certain building types and climate zones or used a questionnaire to obtain a larger dataset. However, they did not capture the wide applications of VRF systems quantitatively across all building types, climate zones, and operating conditions. To fill this gap, statistical and clustering analysis was conducted on the newly proposed key performance indicators of approximately 287,000 VRF systems for residential and commercial buildings in all five climate zones in China. In this work, the main findings are: (1) VRF systems are mainly used for cooling in all climate zones in China; (2) among all building types, the duration of use is lowest in residential buildings and highest in hotels and medical buildings; (3) the distribution of the ideal VRF cooling coefficient of performance (COP) is similar across all climate zones and building types; whereas the COPs of ideal VRF heating in the Severe Cold region and Cold regions are lower than those in other climate zones; and (4) partial load operations for VRF systems are common in residential buildings and office buildings due to the part-time-part-space operation mode. These findings can inform the actual application of VRF systems in China, supporting the design, operation, industry standard development, and performance optimization of VRF systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The cosmic web of X-ray active galactic nuclei seen through the eROSITA Final Equatorial Depth Survey (eFEDS)

Which galaxies in the general population turn into active galactic nuclei (AGNs) is a keystone of galaxy formation and evolution. Thanks to SRG/eROSITA’s contiguous 140 square degree pilot survey field, we constructed a large, complete, and unbiased soft X-ray flux-limited (FX > 6.5 × 10 -15 erg s -1 cm -2 ) AGN sample at low redshift, 0.05 < z < 0.55. Two summary statistics, the clustering using spectra from SDSS-V and galaxy-galaxy lensing with imaging from HSC, are measured and interpreted with halo occupation distribution and abundance matching models. Both models successfully account for the observations. We obtain an exceptionally complete view of the AGN halo occupation distribution. The population of AGNs is broadly distributed among halos with a mean mass of 3.9 -2.4 +2.0 × 10 12 M ⊙ . This corresponds to a large-scale halo bias of b(z = 0.34) = 0.99 -0.10 +0.08 . The central occupation has a large transition parameter, σ log 10 (M) = 1.28 ± 0.2. The satellite occupation distribution is characterized by a shallow slope, α sat = 0.73 ± 0.38. We find that AGNs in satellites are rare, with f sat < 20%. Most soft X-ray-selected AGNs are hosted by central galaxies in their dark matter halo. A weak correlation between soft X-ray luminosity and large-scale halo bias is confirmed (3.3σ). We discuss the implications of environmental-dependent AGN triggering. This study paves the way toward fully charting, in the coming decade, the coevolution of X-ray AGNs, their host galaxies, and dark matter halos by combining eROSITA with SDSS-V, 4MOST, DESI, LSST, and Euclid data.

79 ASTRONOMY AND ASTROPHYSICS↗

Galaxy cluster profiles: a Gaussian mixture model approach to halo miscentering

Measurements of the galaxy density and weak-lensing profiles of galaxy clusters typically rely on an assumed cluster center, which is taken to be the brightest cluster galaxy or other proxies for the true halo center defined as the minimum in the potential well. Departure of the assumed cluster center from the true halo center bias the resultant profile measurements, an effect known as miscentering bias. Currently, miscentering is typically modeled in stacked profiles of clusters with a two parameter model. We use an alternate approach in which the profiles of individual clusters are used with the corresponding likelihood computed using a Gaussian mixture model. We test the approach using halos and the corresponding subhalo profiles from the IllustrisTNG hydrodynamic simulations. We obtain significantly improved estimates of the miscentering parameters for both 3D and projected 2D profiles relevant for imaging surveys. We discuss applications to upcoming cosmological surveys. Our Python package for the Gaussian mixture model is publicly available at https://github.com/KyleMiller1/Halo-Miscentering-Mixture-Model.

Bayesian reasoning↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

State-level suicide mortality insights: a comparative study of VHA veterans and the whole US population

Background: Suicide is a leading cause of death in the US Comparative State-level spatial analysis between Veterans Health Administration (VHA veterans) and the whole US population can reveal differences in conditions for targeted interventions and intricate geographical patterns. Methods: The study population contains 2018 and 2019 suicide deaths of VHA veterans and the whole US population. They were used to calculate state-level rates. States were classified by whether their VHA veteran and whole US population rates were above or below respective mean rates. Local Moran’s I was leveraged to examine spatial autocorrelation. Results: State-level suicide mortality rates and disparities among states were generally higher for VHA veterans (2018: 37.3 ± 7.2; 2019: 46.8 ± 8.3) than for the whole US population (2018: 16.6 ± 4.3; 2019: 16.4 ± 4.4). For both populations, there were statistically significant clusters with high suicide rates. Over one-fourth of states demonstrated inverse relationships, with rates above mean for one group but below for other. VHA veterans are at higher risk with over one-third of states had greater than average veteran suicide risk ratio. Conclusions: VHA veterans are at higher risk than the whole population across all states. Mortality disparities among states and clusters of states with high and low rates suggest targeted interventions and cooperative health strategies may help address these differences.

60 APPLIED LIFE SCIENCES↗

LBNL CRADA (FP00009949) with the American Public Power Association: Electricity Reliability Metrics, Analysis, and Planning (Final Technical Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Electricity Reliability Metrics, Analysis, and Planning (CRADA Final Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method. The team developed suggestions on how comparisons across groupings of feeders and review of the relative contributions of the constituents of SAIFI and SAIDI could be used to help prioritize utility actions to improve reliability. However, they also caution that their suggestions require further evaluation because they are based on only one year of information from a modest number of small utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Test Vector Development for Verification and Validation of Heavy-Duty Autonomous Vehicle Operations

The current focus in the ongoing development of autonomous driving systems (ADS) for heavy duty vehicles is that of vehicle operational safety. To this end, developers and researchers alike are working towards a complete understanding of the operating environments and conditions that autonomous vehicles are subject to during their mission. This understanding is critical to the testing and validation phases of the development of autonomous vehicles and allows for the identification of both the nominal and edge case scenarios encountered by these systems. Previous work by the authors saw the development of a comprehensive scenario generation framework to identify an operating domain specification (ODS), or external and internal conditions an autonomous driving system can expect to encounter on its mission to form critical scenario groups for autonomous vehicle testing and validating using statistical patterns, clustering, and correlation. Continuing this prior work, this paper focuses on the generation of test cases based on the critical scenarios identified that can be used to prioritize either the most common nominal driving scenarios or the least common severe driving scenarios. These test cases can then be used validate, through simulation or real-world testing, the operating design domain (ODD) for a generalized driving mission and built upon to identify spatial and temporal impacts on the driving mission of an autonomous vehicle.

Siekmann, Adam↗