Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical sampling techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

Galaxy blending effects in deep imaging cosmic shear probes of cosmology

ABSTRACT Upcoming deep imaging surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time will be confronted with challenges that come with increased depth. One of the leading systematic errors in deep surveys is the blending of objects due to higher surface density in the more crowded images; a considerable fraction of the galaxies which we hope to use for cosmology analyses will overlap each other on the observed sky. In order to investigate these challenges, we emulate blending in a mock catalogue consisting of galaxies at a depth equivalent to 1.3 yr of the full 10-yr Rubin Observatory that includes effects due to weak lensing, ground-based seeing, and the uncertainties due to extraction of catalogues from imaging data. The emulated catalogue indicates that approximately 12 per cent of the observed galaxies are ‘unrecognized’ blends that contain two or more objects but are detected as one. Using the positions and shears of half a billion distant galaxies, we compute shear–shear correlation functions after selecting tomographic samples in terms of both spectroscopic and photometric redshift bins. We examine the sensitivity of the cosmological parameter estimation to unrecognized blending employing both jackknife and analytical Gaussian covariance estimators. An ∼0.025 decrease in the derived structure growth parameter S8 = σ8(Ωm/0.3)0.5 is seen due to unrecognized blending in both tomographies with a slight additional bias for the photo-z-based tomography. This bias is greater than the 2σ statistical error in measuring S8.

79 ASTRONOMY AND ASTROPHYSICS↗

nautilus : boosting Bayesian importance nested sampling with deep learning

ABSTRACT We introduce a novel approach to boost the efficiency of the importance nested sampling (INS) technique for Bayesian posterior and evidence estimation using deep learning. Unlike rejection-based sampling methods such as vanilla nested sampling (NS) or Markov chain Monte Carlo (MCMC) algorithms, importance sampling techniques can use all likelihood evaluations for posterior and evidence estimation. However, for efficient importance sampling, one needs proposal distributions that closely mimic the posterior distributions. We show how to combine INS with deep learning via neural network regression to accomplish this task. We also introduce nautilus, a reference open-source python implementation of this technique for Bayesian posterior and evidence estimation. We compare nautilus against popular NS and MCMC packages, including emcee, dynesty, ultranest, and pocomc, on a variety of challenging synthetic problems and real-world applications in exoplanet detection, galaxy SED fitting and cosmology. In all applications, the sampling efficiency of nautilus is substantially higher than that of all other samplers, often by more than an order of magnitude. Simultaneously, nautilus delivers highly accurate results and needs fewer likelihood evaluations than all other samplers tested. We also show that nautilus has good scaling with the dimensionality of the likelihood and is easily parallelizable to many CPUs.

97 MATHEMATICS AND COMPUTING↗

To have value, comparisons of high-throughput phenotyping methods need statistical tests of bias and variance

The gap between genomics and phenomics is narrowing. The rate at which it is narrowing, however, is being slowed by improper statistical comparison of methods. Quantification using Pearson’s correlation coefficient ( r ) is commonly used to assess method quality, but it is an often misleading statistic for this purpose as it is unable to provide information about the relative quality of two methods. Using r can both erroneously discount methods that are inherently more precise and validate methods that are less accurate. These errors occur because of logical flaws inherent in the use of r when comparing methods, not as a problem of limited sample size or the unavoidable possibility of a type I error. A popular alternative to using r is to measure the limits of agreement (LOA). However both r and LOA fail to identify which instrument is more or less variable than the other and can lead to incorrect conclusions about method quality. An alternative approach, comparing variances of methods, requires repeated measurements of the same subject, but avoids incorrect conclusions. Variance comparison is arguably the most important component of method validation and, thus, when repeated measurements are possible, variance comparison provides considerable value to these studies. Statistical tests to compare variances presented here are well established, easy to interpret and ubiquitously available. The widespread use of r has potentially led to numerous incorrect conclusions about method quality, hampering development, and the approach described here would be useful to advance high throughput phenotyping methods but can also extend into any branch of science. The adoption of the statistical techniques outlined in this paper will help speed the adoption of new high throughput phenotyping techniques by indicating when one should reject a new method, outright replace an old method or conditionally use a new method.

59 BASIC BIOLOGICAL SCIENCES↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The development and application of the stirred‐reactor coupon analysis (SRCA) test method

A new technique, termed the stirred‐reactor coupon analysis (SRCA) method, has been developed to measure the rate of glass dissolution in forward‐rate conditions. Monolithic glass coupons are partially masked with an inert material before placement in a large volume of well‐mixed solution with known chemistry and temperature for a predetermined duration. After the test, the mask is removed, and the difference in step height between the protected area and the exposed corroded portions of the sample coupon is measured to determine the extent of glass dissolution. The step height is converted to a rate measurement using the test duration and glass density. Test parameters such as sample surface preparation and test duration were evaluated to determine their effects on the measured rates. Additionally, results from an interlaboratory study (ILS) consisting of 12 laboratories from 11 different institutions are presented, where each laboratory performed 12 independent tests. When removing experimental outlier data, the 95% reproducibility limits for the SRCA method has no statistical difference with previously published standardized test methods used to determine the forward rate of glass dissolution. Overall, this paper describes steps necessary to perform the test method and provides the statistical calculations to evaluate test accuracy.

chemical durability↗

The DESI One-Percent Survey: Modelling the clustering and halo occupation of all four DESI tracers with U CHUU

We present results from a set of mock lightcones for the DESI One-Percent Survey, created from the UCHUU simulation. This 8 h −3 Gpc 3 N-body simulation comprises 2.1 trillion particles and provides high-resolution dark matter (sub)haloes in the framework of the Planck-based ΛCDM cosmology. Employing the subhalo abundance matching (SHAM) technique, we populated the UCHUU (sub)haloes with all four DESI tracers – Bright Galaxy Survey (BGS), luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars (QSOs) – to z = 2.1. Our method accounts for redshift evolution as well as the clustering dependence on luminosity and stellar mass. The two-point clustering statistics of the DESI One-Percent Survey generally agree with predictions from UCHUU across scales ranging from 0.3 h −1 Mpc to 100 h −1 Mpc for the BGS and across scales ranging from 5 h −1 Mpc to 100 h −1 Mpc for the other tracers. We observed some differences in clustering statistics that can be attributed to incompleteness of the massive end of the stellar mass function of LRGs, our use of a simplified galaxy-halo connection model for ELGs and QSOs, and cosmic variance. We find that at the high precision of UCHUU, the shape of the halo occupation distribution (HOD) of the BGS and LRG samples is smaller bias values, likely due to cosmic variance. The bias dependence on absolute magnitude, stellar mass, and redshift aligns with that of previous surveys. These results provide DESI with tools to generate high-fidelity lightcones for the remainder of the survey and enhance our understanding of the galaxy-halo connection.

cosmology↗

Characterizing the Sample Selection for Supernova Cosmology

Type Ia supernovae (SNe Ia) are used as distance indicators to infer the cosmological parameters that specify the expansion history of the universe. Parameter inference depends on the criteria by which the analysis SN sample is selected. Only for the simplest selection criteria and population models can the likelihood be calculated analytically, otherwise it needs to be determined numerically, a process that inherently has error. Numerical errors in the likelihood lead to errors in parameter inference. This article presents toy examples where the distance modulus is inferred given a set of SNe at a single redshift. Parameter estimators and their uncertainties are calculated using Monte Carlo techniques. The relationship between the number of Monte Carlo realizations and numerical errors is presented. The procedure can be applied to more realistic models and used to determine the computational and data management requirements of the transient analysis pipeline.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic inference of the structure and orbit of Milky Way satellites with semi-analytic modelling

Semi-analytic modelling furnishes an efficient avenue for characterizing dark matter haloes associated with satellites of Milky Way-like systems, as it easily accounts for uncertainties arising from halo-to-halo variance, the orbital disruption of satellites, baryonic feedback, and the stellar-to-halo mass (SMHM) relation. We use the SatGen semi-analytic satellite generator, which incorporates both empirical models of the galaxy–halo connection as well as analytic prescriptions for the orbital evolution of these satellites after accretion onto a host to create large samples of Milky Way-like systems and their satellites. By selecting satellites in the sample that match observed properties of a particular dwarf galaxy, we can infer arbitrary properties of the satellite galaxy within the cold dark matter paradigm. For the Milky Way’s classical dwarfs, we provide inferred values (with associated uncertainties) for the maximum circular velocity v max and the radius r max at which it occurs, varying over two choices of baryonic feedback model and two prescriptions for the SMHM relation. While simple empirical scaling relations can recover the median inferred value for v max and r max , this approach provides realistic correlated uncertainties and aids interpretability. We also demonstrate how the internal properties of a satellite’s dark matter profile correlate with its orbit, and we show that it is difficult to reproduce observations of the Fornax dwarf without strong baryonic feedback. Furthermore, the technique developed in this work is flexible in its application of observational data and can leverage arbitrary information about the satellite galaxies to make inferences about their dark matter haloes and population statistics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

Sampling Size Optimization for Bioburden Density Estimation in Planetary Protection

Planetary protection (PP) is a discipline that focuses on minimizing the biological contamination of spacecraft to ensure compliance with international policy. Precise estimation of bioburden - the total number of microbes in or on spacecraft hardware – and the bioburden density are of utmost importance for PP. Such estimation is the way concordance with requirements is demonstrated, and it is critical for quantifying the potential risk of inadvertently contaminating other planetary bodies. Although a suite of molecular techniques have been used to thoroughly characterize and profile the microbiome of various cleanroom environments and spacecraft, the gold standard remains the physical enumeration of microbes via culturing of samples directly taken from spacecraft and associated surfaces. However, due to technical, budgetary, and programmatic constraints, only a manageable portion (around 10%) of the entire spacecraft surface is directly sampled with cotton swabs or wipes. To generate the bioburden current best estimate (CBE) for components not directly verifiable, the accepted approach is to apply a NASA-defined bioburden estimate based on the components’ manufacturing or assembly environment. This approach utilizes a prespecified bioburden density estimation that applies a maximum value across the total surface area of the specified component. For hardware components that underwent similar assembly processes, an implied bioburden is adopted for all components, based on a direct verification of a representative component within the same lot. Once all components have a CBE, the bioburden estimates are generated. In previous publication [ 1], we have shown that statistical risks quantifying the accuracy of the estimates for sampled, prespecified, and implied components can be derived and ranked. For mean squared error (MSE) function, the risks are available analytically and hence a cost function can be obtained to optimize the risks with respect to the sampling area and sampling cost. Since the sampling area and sampling cost are two complimentary variables, their sum will have a well-defined minimum. This paper presents the multivariate optimization of the integrated risk of an empirical Bayes estimator to determine the optimal sampling schedule for a given number of components. It is assumed that given a number of components, N, the bioburden density for each component can either be sampled, implied, or prespecified. The multivariate optimization searches through different options to sample, imply or prespecify the bioburden density for a component, and account for the component’s surface area and cost of sampling. The idea of the optimization is based on the observation that the statistical risk of using an estimator is a monotonically decreasing function of the sampled area. The larger the sampled area, the lower the risk of using the estimator as the estimator becomes more and more accurate as the sampling area increases. On the other hand, the cost of sampling is monotonically increasing as the sampled surface grows. This makes the risk and total cost of sampling complimentary variables which can be counterbalanced to achieve an optimal overall value with respect to the sampled surface. In this paper, the integrated risk has been used to quantify the accuracy of the estimator. This risk has been selected because it depends on neither the true value of the parameter nor on the collected data. The cost of each sample was also available to obtain the total cost of sampling of N components. The paper will present the results based on computer-simulated data as well as the data collected during the InSight mission. The computer-simulated data have N components with randomly generated total areas and each component assigned to one of the three categories according to the method of estimating of bioburden density: sampled, implied, or prespecified. The cost of sampling is also available. The cost of sampling is estimated based on a cost model provided by the planetary protection group at JPL. For this paper, the overall cost was assumed to be a linear function of exposure. The optimization process finds the allocation of the components to the three categories that minimizes the tradeoff between integrated risk and total cost. For the InSight data, a set of components is selected representing all three categories, and optimization is performed to determine if the performed allocation was optimal or if a better allocation could have been obtained. To the best of our knowledge, this work is the first attempt not only perform an accurate estimation of bioburden density but also do it in an optimal way.

97 - MATHEMATICS AND COMPUTING↗

Evaluating User Errors and Temporal Trends in Marine Fish Communities Using 360-Degree Underwater Photography

The use of environmental DNA (eDNA) sampling has been proposed as a complementary method to monitor fish species in marine environments, offering a non-invasive and potentially more efficient approach to marine species observations. eDNA monitoring could be especially useful in and around sites targeted for marine energy generation as these regions need regular monitoring that would be impractical with traditional techniques. Before we can fully rely upon eDNA, we must first verify its accuracy against other proven methods, such as the use of underwater photography. In this study, I deployed a 360-degree camera in the tidal channel of Sequim Bay once a month during several hours overlapping slack tide. I investigated how having multiple people identify and count fish on underwater images could affect the overall results. Using chi square tests in R, I compared my fish identifications and counts to those made by another intern on the same images recorded in August. I found significant differences in the number of species identified and the total individual counts between the two different datasets. I also tested the statistical differences in both Shannon diversity and Pielou evenness indices between the August, September, and November camera deployments using a Hutcheson t-test. Only one significant difference was found in the Shannon index comparisons, and none were found between the Pielou evenness comparisons. These findings show that if multiple identifiers are used to process underwater images, quality control checks must be made to reduce the potential for error. This also points toward the possibility to leverage more advanced image analysis processes, such as automated image analysis software. The findings from this study also show that the dynamics of marine fish communities can vary over a few months; however, further analysis is needed to determine the extent of the seasonal changes in Sequim Bay.

59 BASIC BIOLOGICAL SCIENCES↗

An In Situ , Automated High-Explosives Aging Method Utilizing Two-Dimensional Gas Chromatography–Mass Spectrometry

Understanding chemical changes that occur in high explosives as they age is of great importance to the safe employment and storage of these compounds. Traditional methods of aging high explosives even under accelerated aging conditions are time intensive with durations on the order of months to years. The nature of traditional aging analyses reduces each sample to a snapshot data point often separated widely in time, requiring many assumptions as to how the degradation products develop. Further complicating matters, several analytical techniques are typically employed for each sample analysis in order to ascertain an entire picture of the decomposition pathways. To address these shortcomings with existing methods, a new method of accelerated aging of high explosives utilizing comprehensive two-dimensional gas chromatography coupled to high-resolution mass spectrometry (GC × GC-HRMS) was developed using 2,4,6,8,10,12-hexanitro-2,4,6,8,10,12-hexaazaisowurtzitane (CL-20) as a model compound for method development. This in situ automated method reduces the time scale of aging to a matter of hours using the inlet of the GC × GC as the aging vessel. GC × GC in combination with HRMS allowed for the collection of both evolved gases and other decomposition products produced during the entire aging process in real time with HRMS providing far greater certainty in identification of explosives aging products. Additionally, this method allowed for a higher throughput of samples with greatly simplified sample preparation. Chemometric analysis of the GC × GC-HRMS data set via the alteration analysis (ALA) enabled discovery of statistically significant chemical changes providing insight into the variation of decomposition pathways with varying aging temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

How to avoid multiple scattering in strongly scattering SANS and USANS samples

Small Angle Neutron Scattering (SANS) and Ultra Small Angle Neutron Scattering (USANS) are the only available experimental techniques to provide seamless non-destructive measurements of the geometry of the accessible and inaccessible pore structure of rocks from sub-nanopore size to the scale of macropores. They have therefore become the measurement of choice for tight reservoir rocks such as organic rich shales. A simplifying assumption in the analysis is, however, that during the path of neutrons through the sample each neutron is only scattered once. Shales are samples with a high scattering power and Multiple Scattering (MS) may occur which requires special modelling for deconvolution of the results. The approach to avoid MS is to simply reduce the sample thickness to <0.15–0.5 mm. Here, in this work, we present a systematic method on wavelength selection and preparation of samples to optimise extraction of microstructural data and minimise parasitic errors. Experimentally measured SAS transmission (TSAS) values are used as a practical criterion for estimation of the extent of MS. Generous beamtime allocations allowed robust testing revealing that sample thicknesses can be twice as thick as predicted using the standard protocol. Analysing thicker samples is particularly beneficial for statistically relevant characterisation of heterogeneous samples making the new protocol the method of choice for such samples.

(U)SANS↗

On the dark matter haloes of optical and IR-selected AGNs in the local universe

ABSTRACT We use the technique of total satellite luminosity, Lsat, to probe the dark matter haloes around active galactic nuclei (AGNs) in the SDSS Main Galaxy Sample. Our results focus on galaxies and AGNs that are the central galaxy of their halo. Our two AGN samples are constructed from optical emission-line diagnostics and from Wide-Field Infrared Survey Explorer (WISE) infrared colours. Both optically selected and WISE-selected AGN have Lsat values twice as high as non-active galaxy samples when controlling for stellar mass and mean stellar age. This implies that the haloes are twice as massive, but we cannot rule out that the increase in Lsat is due to these AGNs residing in younger haloes at the same mass. When only controlling for host galaxy stellar mass, WISE-selected AGNs also have higher Lsat values than optical AGNs at the factor of two level, consistent with previous results comparing the clustering of obscured and unobscured AGNs. However, controlling for stellar age in the two populations of host galaxies removes half of this difference, attenuating the statistical significance of the difference. We perform permutation tests to quantify the difference in the halo populations of each sample. The difference in star formation properties does not fully explain the difference in the two AGN populations, however. Although AGN luminosity correlates with mean stellar age, the difference in stellar age between the WISE and optical samples cannot be fully explained by differences in their AGN luminosity distributions.

Alpaslan, Mehmet (ORCID:0000000303211033)↗

The persistence of large scale structures. Part I. Primordial non-Gaussianity

Abstract We develop an analysis pipeline for characterizing the topology of large scale structure and extracting cosmological constraints based on persistent homology . Persistent homology is a technique from topological data analysis that quantifies the multiscale topology of a data set, in our context unifying the contributions of clusters, filament loops, and cosmic voids to cosmological constraints. We describe how this method captures the imprint of primordial local non-Gaussianity on the late-time distribution of dark matter halos, using a set of N-body simulations as a proxy for real data analysis. For our best single statistic, running the pipeline on several cubic volumes of size 40 (Gpc/h) 3 , we detect f NL loc =10 at 97.5% confidence on ~ 85% of the volumes. Additionally wetest our ability to resolve degeneracies betweenthe topological signature of f NL loc and variation of σ 8 and argue that correctly identifying nonzero f NL loc in this case is possible via an optimal template method. Our method relies on information living at $\mathcal{O}$(10) Mpc/h, a complementary scale with respect to commonly used methods such as the scale-dependent bias in the halo/galaxy power spectrum. Therefore, while still requiring a large volume, our method does not require sampling long-wavelength modes to constrain primordial non-Gaussianity. Moreover, our statistics are interpretable: we are able to reproduce previous results in certain limits and we make new predictions for unexplored observables, such as filament loops formed by dark matter halos in a simulation box.

Astronomy & Astrophysics↗

Measurement Error and Resolution in Quantitative Stable Isotope Probing: Implications for Experimental Design

Quantitative stable isotope probing (qSIP) estimates isotope tracer incorporation into DNA of individual microbes and can link microbial biodiversity and biogeochemistry in complex communities. As with any quantitative estimation technique, qSIP involves measurement error, and a fuller understanding of error, precision, and statistical power benefits qSIP experimental design and data interpretation. We used several qSIP data sets—from soil and seawater microbiomes—to evaluate how variance in isotope incorporation estimates depends on organism abundance and resolution of the density fractionation scheme. We assessed statistical power for replicated qSIP studies, plus sensitivity and specificity for unreplicated designs. As a taxon’s abundance increases, the variance of its weighted mean density declines. Nine fractions appear to be a reasonable trade-off between cost and precision for most qSIP applications. Increasing the number of density fractions beyond that reduces variance, although the magnitude of this benefit declines with additional fractions. Our analysis suggests that, if a taxon has an isotope enrichment of 10 atom% excess, there is a 60% chance that this will be detected as significantly different from zero (with alpha 0.1). With five replicates, isotope enrichment of 5 atom% could be detected with power (0.6) and alpha (0.1). Finally, we illustrate the importance of internal standards, which can help to calibrate per sample conversions of %GC to mean weighted density. These results should benefit researchers designing future SIP experiments and provide a useful reference for metagenomic SIP applications where both financial and computational limitations constrain experimental scope.

59 BASIC BIOLOGICAL SCIENCES↗