Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Galaxy cluster profiles: a Gaussian mixture model approach to halo miscentering

Measurements of the galaxy density and weak-lensing profiles of galaxy clusters typically rely on an assumed cluster center, which is taken to be the brightest cluster galaxy or other proxies for the true halo center defined as the minimum in the potential well. Departure of the assumed cluster center from the true halo center bias the resultant profile measurements, an effect known as miscentering bias. Currently, miscentering is typically modeled in stacked profiles of clusters with a two parameter model. We use an alternate approach in which the profiles of individual clusters are used with the corresponding likelihood computed using a Gaussian mixture model. We test the approach using halos and the corresponding subhalo profiles from the IllustrisTNG hydrodynamic simulations. We obtain significantly improved estimates of the miscentering parameters for both 3D and projected 2D profiles relevant for imaging surveys. We discuss applications to upcoming cosmological surveys. Our Python package for the Gaussian mixture model is publicly available at https://github.com/KyleMiller1/Halo-Miscentering-Mixture-Model.

Bayesian reasoning↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

State-level suicide mortality insights: a comparative study of VHA veterans and the whole US population

Background: Suicide is a leading cause of death in the US Comparative State-level spatial analysis between Veterans Health Administration (VHA veterans) and the whole US population can reveal differences in conditions for targeted interventions and intricate geographical patterns. Methods: The study population contains 2018 and 2019 suicide deaths of VHA veterans and the whole US population. They were used to calculate state-level rates. States were classified by whether their VHA veteran and whole US population rates were above or below respective mean rates. Local Moran’s I was leveraged to examine spatial autocorrelation. Results: State-level suicide mortality rates and disparities among states were generally higher for VHA veterans (2018: 37.3 ± 7.2; 2019: 46.8 ± 8.3) than for the whole US population (2018: 16.6 ± 4.3; 2019: 16.4 ± 4.4). For both populations, there were statistically significant clusters with high suicide rates. Over one-fourth of states demonstrated inverse relationships, with rates above mean for one group but below for other. VHA veterans are at higher risk with over one-third of states had greater than average veteran suicide risk ratio. Conclusions: VHA veterans are at higher risk than the whole population across all states. Mortality disparities among states and clusters of states with high and low rates suggest targeted interventions and cooperative health strategies may help address these differences.

60 APPLIED LIFE SCIENCES↗

LBNL CRADA (FP00009949) with the American Public Power Association: Electricity Reliability Metrics, Analysis, and Planning (Final Technical Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Electricity Reliability Metrics, Analysis, and Planning (CRADA Final Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method. The team developed suggestions on how comparisons across groupings of feeders and review of the relative contributions of the constituents of SAIFI and SAIDI could be used to help prioritize utility actions to improve reliability. However, they also caution that their suggestions require further evaluation because they are based on only one year of information from a modest number of small utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Test Vector Development for Verification and Validation of Heavy-Duty Autonomous Vehicle Operations

The current focus in the ongoing development of autonomous driving systems (ADS) for heavy duty vehicles is that of vehicle operational safety. To this end, developers and researchers alike are working towards a complete understanding of the operating environments and conditions that autonomous vehicles are subject to during their mission. This understanding is critical to the testing and validation phases of the development of autonomous vehicles and allows for the identification of both the nominal and edge case scenarios encountered by these systems. Previous work by the authors saw the development of a comprehensive scenario generation framework to identify an operating domain specification (ODS), or external and internal conditions an autonomous driving system can expect to encounter on its mission to form critical scenario groups for autonomous vehicle testing and validating using statistical patterns, clustering, and correlation. Continuing this prior work, this paper focuses on the generation of test cases based on the critical scenarios identified that can be used to prioritize either the most common nominal driving scenarios or the least common severe driving scenarios. These test cases can then be used validate, through simulation or real-world testing, the operating design domain (ODD) for a generalized driving mission and built upon to identify spatial and temporal impacts on the driving mission of an autonomous vehicle.

Siekmann, Adam↗

New bound on neutrino dipole moments from globular-cluster stars

Neutrino dipole moments mu(nu) would increase the core mass of red giants at the helium flash by delta(Mc) = 0.015 solar mass x mu(nu)/10 to the -12th muB (where muB is the Bohr magneton) because of enhanced neutrino losses. Existing measurements of the bolometric magnitudes of the brightest red giants in 26 globular clusters, number counts of horizontal-branch stars and red giants in 15 globular clusters, and statistical parallax determinations of field RR Lyr luminosities yield delta(Mc) = 0.009 + or - 0.012 solar mass, so that conservatively mu(nu) is less than 3 x 10 to the -12th muB.

Raffelt, Georg G.↗

Estimation of the nucleation rate by differential scanning calorimetry

A realistic computer model is presented for calculating the time-dependent volume fraction transformed during the devitrification of glasses, assuming the classical theory of nucleation and continuous growth. Time- and cluster-dependent nucleation rates are calculated by modeling directly the evolving cluster distribution. Statistical overlap in the volume fraction transformed is taken into account using the standard Johnson-Mehl-Avrami formalism. Devitrification behavior under isothermal and nonisothermal conditions is described. The model is used to demonstrate that the recent suggestion by Ray and Day (1990) that nonisothermal DSC studies can be used to determine the temperature for the peak nucleation rate, is qualitatively correct for lithium disilicate, the glass investigated.

Kelton, Kenneth F.↗

A study of various methods for calculating locations of lightning events

This article reports on the results of numerical experiments on finding the location of lightning events using different numerical methods. The methods include linear least squares, nonlinear least squares, statistical estimations, cluster analysis and angular filters and combinations of such techniques. The experiments involved investigations of methods for excluding fake solutions which are solutions that appear to be reasonable but are in fact several kilometers distant from the actual location. Some of the conclusions derived from the study are that bad data produces fakes, that no fool-proof method of excluding fakes was found, that a short base-line interferometer under development at Kennedy Space Center to measure the direction cosines of an event shows promise as a filter for excluding fakes. The experiments generated a number of open questions, some of which are discussed at the end of the report.

Cannon, John R.↗

Nanophase and Composite Optical Materials

This talk will focus on accomplishments, current developments, and future directions of our work on composite optical materials for microgravity science and space exploration. This research spans the order parameter from quasi-fractal structures such as sol-gels and other aggregated or porous media, to statistically random cluster media such as metal colloids, to highly ordered materials such as layered media and photonic bandgap materials. The common focus is on flexible materials that can be used to produce composite or artificial materials with superior optical properties that could not be achieved with homogeneous materials. Applications of this work to NASA exploration goals such as terraforming, biosensors, solar sails, solar cells, and vehicle health monitoring, will be discussed.

Source record↗

Sulfur in Cometary Dust

The computer-intensive project consisted of the analysis and synthesis of existing data on composition of comet Halley dust particles. The main objective was to obtain a complete inventory of sulfur containing compounds in the comet Halley dust by building upon the existing classification of organic and inorganic compounds and applying a variety of statistical techniques for cluster and cross-correlational analyses. A student hired for this project wrote and tested the software to perform cluster analysis. The following tasks were carried out: (1) selecting the data from existing database for the proposed project; (2) finding access to a standard library of statistical routines for cluster analysis; (3) reformatting the data as necessary for input into the library routines; (4) performing cluster analysis and constructing hierarchical cluster trees using three methods to define the proximity of clusters; (5) presenting the output results in different formats to facilitate the interpretation of the obtained cluster trees; (6) selecting groups of data points common for all three trees as stable clusters. We have also considered the chemistry of sulfur in inorganic compounds.

Fomenkova, M. N.↗

Detection of spatial clustering in the 1000 richest SDSS DR8 redMaPPer clusters with nearest neighbor distributions

ABSTRACT Distances to the k-nearest-neighbor (kNN) data points from volume-filling query points are a sensitive probe of spatial clustering. Here, we present the first application of kNN summary statistics to observational clustering measurement, using the 1000 richest redMaPPer clusters (0.1 ≤ z ≤ 0.3) from the SDSS DR8 catalog. A clustering signal is defined as a difference in the cumulative distribution functions (CDFs) of kNN distances from fixed query points to the observed clusters versus a set of unclustered random points. We find that the k = 1, 2-NN CDFs of redMaPPer deviate significantly from the randoms’ across scales of 35 to 155 Mpc, which is a robust signature of clustering. In addition to kNN, we also measure the two-point correlation function for the same set of redMaPPer clusters versus random points, which shows a noisier and less significant clustering signal within the same radial scales. Quantitatively, the χ2 distribution for both the kNN-CDFs and the two-point correlation function measured on the randoms peak at χ2 ∼ 50 (null hypothesis), whereas the kNN-CDFs (χ2 ∼ 300, p = 1.54 × 10−36) pick up a much more significant clustering signal than the two-point function (χ2 ∼ 100, p = 1.16 × 10−6) when measured on redMaPPer. Finally, the measured 3NN and 4NN CDFs deviate from the predicted k = 3, 4-NN CDFs assuming an ideal Gaussian field, indicating a non-Gaussian clustering signal for redMaPPer clusters, although its origin might not be cosmological due to observational systematics. Therefore, kNN serves as a more sensitive probe of clustering complementary to the two point correlation function, providing a novel approach for constraining cosmology and galaxy–halo connection.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling neutrino-induced scale-dependent galaxy clustering for photometric galaxy surveys

Abstract The increasing statistical precision of photometric redshift surveys requires improved accuracy of theoretical predictions for large-scale structure observables to obtain unbiased cosmological constraints. In ΛCDM cosmologies, massive neutrinos stream freely at small cosmological scales, suppressing the small-scale power spectrum. In massive neutrino cosmologies, galaxy bias modeling needs to accurately relate the scale-dependent growth of the underlying matter field to observed galaxy clustering statistics. In this work, we implement a computationally efficient approximation of the neutrino-induced scale-dependent bias (NISDB). Through simulated likelihood analyses of Dark Energy Survey Year 3 (DESY3) and Legacy Survey of Space and Time Year 1 (LSSTY1) synthetic data that contain an appreciable NISDB, we examine the impact of linear galaxy bias and neutrino mass modeling choices on cosmological parameter inference. We find model misspecification of the NISDB approximation and neutrino mass models to decrease the constraining power of photometric galaxy surveys and cause parameter biases in the cosmological interpretation of future surveys. We quantify these biases and devise mitigation strategies.

Astronomy & Astrophysics↗

Deep, wide-field, multi-band imaging of z approximately equal to 0.4 clusters and their environs

The existence of an excess population of blue galaxies in the cores of distant, rich clusters of galaxies, commonly referred to as the 'Butcher-Oemler' effect is now well established. Spectroscopy of clusters at z = 0.2-0.4 has confirmed that the luminous blue populations comprise as much as 20 percent of these clusters. This fraction is much higher that the 2 percent blue fraction found for nearby rich clusters, such as Coma, indicating that rapid galaxy evolution has occurred on a relatively short time scale. Spectroscopy has also shown that the 'blue' galaxies can basically be divided into three classes: 'starburst' galaxies with large (O II) equivalent widths, 'post-starburst' E+A galaxies (i.e. galaxies with strong Balmer lines shortward of 4000A but elliptical-like colors, and normal spiral/irregulars. Unfortunately, it is difficult to obtain enough spectra of individual galaxies in these intermediate redshift clusters to say anything statistically meaningful. Thus, limited information is available about the relative numbers of these three classes of 'blue' galaxies and the associated E/SO population in these intermediate redshift clusters. More statistically meaningful results can be derived from deep imaging of these clusters. However, the best published data to date (e.g. MacLaren et al. 1988; Dressler & Gunn 1992) are limited to the cluster cores and do not sample the galaxy luminosity functions very deeply at the bluest wavelengths. Furthermore, only limited spectro-energy distribution data is available below 4000A in the observed cluster rest frame providing limited sensitivity to 'recent' star formation activity. To improve this situation, we are currently obtaining deep, wide-field UBRI images of all known rich clusters at z approx. equals 0.4. Our main objective is to obtain the necessary color information to distinguish between the E+SO, 'E+A', and spiral/irregular galaxy populations throughout the cluster/supercluster complex. At this redshift, UBRI correspond to rest-frame 2500A/UVR bandpasses. The rest-frame UVR system provides a powerful 'blue' galaxy discriminate given the expected color distribution. Moreover, since 'hot' stars peak near 2500A, that bandpass is a powerful probe of recent star formation activity in all classes of galaxies. In particular, it is sensitive to ellipticals with 'UV excess' populations (MacLaren et al. 1988).

Silva, David R.↗

ICAP: An Interactive Cluster Analysis Procedure for analyzing remotely sensed data

An Interactive Cluster Analysis Procedure (ICAP) was developed to derive classifier training statistics from remotely sensed data. The algorithm interfaces the rapid numerical processing capacity of a computer with the human ability to integrate qualitative information. Control of the clustering process alternates between the algorithm, which creates new centroids and forms clusters and the analyst, who evaluate and elect to modify the cluster structure. Clusters can be deleted or lumped pairwise, or new centroids can be added. A summary of the cluster statistics can be requested to facilitate cluster manipulation. The ICAP was implemented in APL (A Programming Language), an interactive computer language. The flexibility of the algorithm was evaluated using data from different LANDSAT scenes to simulate two situations: one in which the analyst is assumed to have no prior knowledge about the data and wishes to have the clusters formed more or less automatically; and the other in which the analyst is assumed to have some knowledge about the data structure and wishes to use that information to closely supervise the clustering process. For comparison, an existing clustering method was also applied to the two data sets.

Wharton, S. W.↗

Voids and constraints on nonlinear clustering of galaxies

Void statistics of the galaxy distribution in the Center for Astrophysics Redshift Survey provide strong constraints on galaxy clustering in the nonlinear regime, i.e., on scales R equal to or less than 10/h Mpc. Computation of high-order moments of the galaxy distribution requires a sample that (1) densely traces the large-scale structure and (2) covers sufficient volume to obtain good statistics. The CfA redshift survey densely samples structure on scales equal to or less than 10/h Mpc and has sufficient depth and angular coverage to approach a fair sample on these scales. In the nonlinear regime, the void probability function (VPF) for CfA samples exhibits apparent agreement with hierarchical scaling (such scaling implies that the N-point correlation functions for N greater than 2 depend only on pairwise products of the two-point function xi(r)) However, simulations of cosmological models show that this scaling in redshift space does not necessarily imply such scaling in real space, even in the nonlinear regime; peculiar velocities cause distortions which can yield erroneous agreement with hierarchical scaling. The underdensity probability measures the frequency of 'voids' with density rho less than 0.2 -/rho. This statistic reveals a paucity of very bright galaxies (L greater than L asterisk) in the 'voids.' Underdensities are equal to or greater than 2 sigma more frequent in bright galaxy samples than in samples that include fainter galaxies. Comparison of void statistics of CfA samples with simulations of a range of cosmological models favors models with Gaussian primordial fluctuations and Cold Dark Matter (CDM)-like initial power spectra. Biased models tend to produce voids that are too empty. We also compare these data with three specific models of the Cold Dark Matter cosmogony: an unbiased, open universe CDM model (omega = 0.4, h = 0.5) provides a good match to the VPF of the CfA samples. Biasing of the galaxy distribution in the 'standard' CDM model (omega = 1, b = 1.5; see below for definitions) and nonzero cosmological constant CDM model (omega = 0.4, h = 0.6 lambda(sub 0) = 0.6, b = 1.3) produce voids that are too empty. All three simulations match the observed VPF and underdensity probability for samples of very bright (M less than M asterisk = -19.2) galaxies, but produce voids that are too empty when compared with samples that include fainter galaxies.

Vogeley, Michael S.↗