Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

BEYONDPLANCK IV. On end-to-end simulations in CMB analysis — Bayesian versus frequentist statistics

End-to-end simulations play a key role in the analysis of any high-sensitivity cosmic microwave background (CMB) experiment, providing high-fidelity systematic error propagation capabilities that are unmatched by any other means. In this paper, we address an important issue regarding such simulations, namely, how to define the inputs in terms of sky model and instrument parameters. These may either be taken as a constrained realization derived from the data or as a random realization independent from the data. We refer to these as posterior and prior simulations, respectively. We show that the two options lead to significantly different correlation structures, as prior simulations (contrary to posterior simulations) effectively include cosmic variance, but they exclude realization-specific correlations from non-linear degeneracies. Consequently, they quantify fundamentally different types of uncertainties. We argue that as a result, they also have different and complementary scientific uses, even if this dichotomy is not absolute. In particular, posterior simulations are in general more convenient for parameter estimation studies, while prior simulations are generally more convenient for model testing. Before BEYONDPLANCK, most pipelines used a mix of constrained and random inputs and applied the same hybrid simulations for all applications, even though the statistical justification for this is not always evident. BEYONDPLANCK represents the first end-to-end CMB simulation framework that is able to generate both types of simulations and these new capabilities have brought this topic to the forefront. The BEYONDPLANCK posterior simulations and their uses are described extensively in a suite of companion papers. In this work, we consider one important applications of the corresponding prior simulations, namely, code validation. Specifically, we generated a set of one-year LFI 30 GHz prior simulations with known inputs and we used these to validate the core low-level BEYONDPLANCK algorithms dealing with gain estimation, correlated noise estimation, and mapmaking.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward Discovering the Structure and Dynamics of the Sun-Interstellar Medium System (Los Alamos LDRD Report)

We have taken major steps toward developing the tools necessary to explore with unprecedented resolution the plasma structure and dynamics of the Sun-interstellar medium (ISM) interaction region, known as the heliosheath. We have applied a rigorous imaging technique known as Generalized Ridge Regression (GRR) to construct statistically robust sky maps of hydrogen energetic neutral atoms (ENAs) emanating from this region and detected by the Los Alamos-led IBEX-Hi ENA imager [1] on NASA’s Interstellar Boundary Explorer (IBEX) mission [2]. Our methods go far beyond the map reconstruction process currently applied by the IBEX Science Operations Center (ISOC), opening up the possibility of new discovery science with the IBEX data set and positioning LANL for a lead science role for the upcoming Interstellar Mapping and Acceleration Probe (IMAP) mission [3] for which LANL is providing two key experiments. We have successfully demonstrated that the new maps are higher resolution and are capable of revealing structures that are not presently resolveable in standard ISOC maps

58 GEOSCIENCES↗

Grid-based minimization at scale: Feldman-Cousins corrections for light sterile neutrino search

High Energy Physics (HEP) experiments generally employ sophisticated statistical methods to present results in searches of new physics. In the problem of searching for sterile neutrinos, likelihood ratio tests are applied to short-baseline neutrino oscillation experiments to construct confidence intervals for the parameters of interest. The test statistics of the form Δχ2 is often used to form the confidence intervals, however, this approach can lead to statistical inaccuracies due to the small signal rate in the region-of-interest. In this paper, we present a computational model for the computationally expensive Feldman-Cousins corrections to construct a statistically accurate confidence interval for neutrino oscillation analysis. The program performs a grid-based minimization over oscillation parameters and is written in C++. Our algorithms make use of vectorization through Eigen3, yielding a single-core speed-up of 350 compared to the original implementation, and achieve MPI data parallelism by employing DIY. We demonstrate the strong scaling of the application at High-Performance Computing (HPC) sites. We utilize HDF5 along with HighFive to write the results of the calculation to file.

Wospakrik, Marianette↗

A road map to cosmological parameter analysis with third-order shear statistics: III. Efficient estimation of third-order shear correlation functions and an application to the KiDS-1000 data

Context. Third-order lensing statistics contain a wealth of cosmological information that is not captured by second-order statistics. However, the computational effort it takes to estimate such statistics in forthcoming stage IV surveys is prohibitively expensive. Aims. We derive and validate an efficient estimation procedure for the three-point correlation function (3PCF) of polar fields such as weak lensing shear. We then use our approach to measure the shear 3PCF and the third-order aperture mass statistics on the KiDS-1000 survey. Methods We constructed an efficient estimator for third-order shear statistics that builds on the multipole decomposition of the 3PCF. We then validated our estimator on mock ellipticity catalogs obtained from N -body simulations. Finally, we applied our estimator to the KiDS-1000 data and presented a measurement of the third-order aperture statistics in a tomographic setup. Results. Our estimator provides a speedup of a factor of ∼100–1000 compared to the state-of-the-art estimation procedures. It is also able to provide accurate measurements for squeezed and folded triangle configurations without additional computational effort. We report a significant detection of tomographic third-order aperture mass statistics in the KiDS-1000 data (S/N = 6.69). Conclusions. Our estimator will make it computationally feasible to measure third-order shear statistics in forthcoming stage IV surveys. Furthermore, it can be used to construct empirical covariance matrices for such statistics.

Astronomy & Astrophysics↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Statistical evaluation of microscale stress conditions leading to void nucleation in the weak shock regime

Here, we investigate the heterogeneity of the stress state driven by anisotropic deformation response at the single crystal level through five statistical volume element (SVE) calculations of polycrystalline BCC tantalum. This work focuses on grain boundaries as a prominent material defect type prone to void nucleation based upon experimental observations of predominantly intergranular void nucleation in this material. The SVEs are constructed to be statistically representative of larger volumes of material and are meshed such that mean and standard deviation of grain size and orientation information is reconstructed. The computational meshes feature hexahedral (brick) elements and smooth conformal grain boundaries where significant stress concentration is known to occur, a tail effect of interest in the extreme events process of dynamic ductile damage. An existing micromechanical crystallographic plasticity model shown to capture the single crystal behavior of BCC tantalum well is used to perform the polycrystal calculations. The model includes representation of the non-Schmid effect of non-planar screw dislocation kinetics in tantalum. A three-dimensional stress state time profile predicted by damage modeling of a flyer plate impact experiment is applied as boundary conditions to each SVE. Resulting grain boundary stress state statistics are strongly non-Gaussian. Significant structural evolution is observed within the compressive hold before unloading into tension in the stress profile. Strong angular dependence of grain boundary traction magnitude with shock direction is observed. Non-Schmid effects continue to suggest their influence on propensity of microstructural defect types to nucleate voids. A general void nucleation criterion is proposed using probability theory. The general framework is specified to polycrystalline BCC tantalum in the weak shock regime to include the SVE calculations and literature molecular dynamics calculations of grain boundary void nucleation strength. Probability density functions (PDFs) are used to describe the interaction between the local stress state heterogeneity and the distributed grain boundary void nucleation strength state. A causation entropy maximization procedure removes the requirement for ad hoc selection of a PDF functional form and provides a rigorous procedure for data-based PDF determination. The resulting physically informed PDF describes the spatial appearance frequency of nucleated voids as a function of applied macroscale pressure. Lower length scale physics are thus packaged in a precise and computationally efficient way to provide computational plasticity insight to macroscale dynamic ductile damage models.

36 MATERIALS SCIENCE↗

Optimal binning of correlated measurements

Experimental measurements are commonly represented on a discrete grid, requiring a balance between granularity and statistical noise. Two strategies have traditionally been used to improve such representations: selecting an appropriate bin width to control discretization error and applying kernel-based smoothing to suppress fluctuations. Despite their shared goal, these approaches have largely developed independently, without a unified statistical description of how discretization and correlation jointly determine measurement precision. Here, we extend the discussion of optimal interval averaging to a correlation-aware setting by Gaussian process regression, which explicitly accounts for correlations among neighboring bins. Starting from first principles, we derive the mean-squared error of discretized measurements and obtain closed-form asymptotic expressions for the optimal bin width and correlation length. When recast in reduced variables, the theory reveals distinct universal scaling laws governing the error in the correlation-free and correlation-controlled regimes. Characterized by intrinsically smooth intensity profiles and counting-based statistics, neutron scattering measurements are well suited for demonstrating the enhanced error contraction enabled by inter-bin correlations. We show that such improvement is achievable over the experimentally accessible Q-range and across multiple instruments and material systems. These results show that explicitly accounting for correlations systematically reshapes the limits of precision in discretized, noise-limited measurements. More broadly, the framework provides a transferable statistical foundation for optimizing data representation, inference, and experimental design across the physical and data sciences.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

Evidence of galaxy assembly bias in SDSS DR7 galaxy samples from count statistics

We present observational constraints on the galaxy–halo connection, focusing particularly on galaxy assembly bias from a novel combination of counts-in-cylinders statistics, P(N CIC ), with the standard measurements of the projected two-point correlation function w p (r p ), and number density n gal of galaxies. We measure n gal , w p (r p ), and P(N CIC ) for volume-limited, luminosity-threshold samples of galaxies selected from SDSS DR7, and use them to constrain halo occupation distribution (HOD) models, including a model in which galaxy occupation depends upon a secondary halo property, namely halo concentration. We detect significant positive central assembly bias for the M r < -20.0 and M r <-19.5 samples. Central galaxies preferentially reside within haloes of high concentration at fixed mass. Positive central assembly bias is also favoured in the M r < -20.5 and M r < -19.0 samples. We find no evidence of central assembly bias in the M r < -21.0 sample. We observe only a marginal preference for negative satellite assembly bias in the M r < -20.0 and M r < -19.0 samples, and non-zero satellite assembly bias is not indicated in other samples. Our findings underscore the necessity of accounting for galaxy assembly bias when interpreting galaxy survey data, and demonstrate the potential of count statistics in extracting information from the spatial distribution of galaxies, which could be applied to both galaxy–halo connection studies and cosmological analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Shared micromobility as a first- and last-mile transit solution? Spatiotemporal insights from a novel dataset

The first- and last-mile (FM/LM) problem is a major deterrent to public transit use. With the rise of shared micromobility options such as shared e-scooters in recent years, there is a growing interest in understanding their potential to serve as a last-mile transit solution. However, empirical data regarding the integrated use of shared micromobility and public transit have been limited so far. As a result, much is unknown regarding the spatiotemporal patterns and characteristics of shared micromobility trips serving as an FM/LM connection to transit. Here, this paper addresses these knowledge gaps by leveraging a novel dataset (i.e., the Spin post-ride survey dataset) that records thousands of transit-connecting shared e-scooter trips in Washington DC. Specifically, we used the dataset to reveal the spatiotemporal patterns of transit-connecting shared e-scooter trips in Washington DC, resulting in some major policy insights regarding the integral use of shared e-scooters and public transit. We further leveraged the dataset to validate if and to what extent a commonly applied buffer-zone approach can infer FM/LM micromobility trips accurately. Statistical tests showed that the actual FM/LM Spin e-scooter trips differ from inferred FM/LM Spin e-scooter trips in both spatial and temporal dimensions. This indicates that the common practice of inferring FM/LM micromobility trips with a buffer-zone approach can lead to inaccurate estimates of transit-connecting micromobility trips.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Subseasonal Clustering of Atmospheric Rivers Over the Western United States

Abstract The serial occurrence of atmospheric rivers (ARs) along the US West Coast can lead to prolonged and exacerbated hydrologic impacts, threatening flood‐control and water‐supply infrastructure due to soil saturation and diminished recovery time between storms. Here a statistical approach for quantifying subseasonal temporal clustering among extreme events is applied to a 41‐year (1979–2019) wintertime AR catalog across the western United States (US). Observed AR occurrence, compared against a randomly distributed AR timeseries with the same average event density, reveals temporal clustering at a greater‐than‐random rate across the western US with a distinct geographical pattern. Compared to the Pacific Northwest, significant AR clusters over the northern Coastal Range of California and Sierra Nevada are more frequent and occur over longer time periods. Clusters along the California Coastal Range typically persist for 2 weeks, are composed of 4–5 ARs per cluster, and account for over 85% of total AR occurrence. Across the northwest Coast‐Cascade Ranges, clusters account for ∼50% of total AR occurrence, typically last 8–10 days, and contain 3–4 individual AR events. Based on precipitation data from a high‐resolution dynamical downscaling of reanalysis, the fractions of total and extreme hourly precipitation attributable to AR clusters are largest along the northern California coast and in the Sierra Nevada. Interannual variability among clusters highlights their importance for determining whether a particular water year is anomalously wet or dry. The mechanisms behind this unusual clustering are unclear and require further research.

Meteorology & Atmospheric Sciences↗

Emerging opportunities for hybrid perovskite solar cells using machine learning

While there are several bottlenecks in hybrid organic–inorganic perovskite (HOIP) solar cell production steps, including composition screening, fabrication, material stability, and device performance, machine learning approaches have begun to tackle each of these issues in recent years. Different algorithms have successfully been adopted to solve the unique problems at each step of HOIP development. Specifically, high-throughput experimentation produces vast amount of training data required to effectively implement machine learning methods. Here, we present an overview of machine learning models, including linear regression, neural networks, deep learning, and statistical forecasting. Experimental examples from the literature, where machine learning is applied to HOIP composition screening, thin film fabrication, thin film characterization, and full device testing, are discussed. These paradigms give insights into the future of HOIP solar cell research. As databases expand and computational power improves, increasingly accurate predictions of the HOIP behavior are becoming possible.

Hering, Abigail R. (ORCID:0000000270806953)↗

A Data-Driven Method for Estimating Behind-the-Meter Photovoltaic Generation in Hawaii

Due to the increasing penetration of distributed behind-the-meter photovoltaic (PV) systems and the installed utility revenue metering limited to monitoring only the net power import/export of the household, it is increasingly challenging for utilities to effectively plan and operate the grid. This paper proposes a methodology that estimates behind-the-meter PV generation using a selected subset of monitored PV systems. It is a data-driven approach, and the PV output is estimated utilizing a statistic regression model. A Minimum Redundancy Maximum Relevance (MRMR) algorithm is applied to preselect the optimal subset of the monitored PV systems. The performance of this approach is compared with a spatial interpolation method and a model-based approach. The proposed method is validated using high-resolution meter data recorded from 18 residential rooftop PV systems located on the island of Maui, Hawaii.

Data-driven modeling↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part II: Statistical Downscaling

Abstract Capacity factors (CFs) derived from daily expected power at 22 operating wind farms in different regions of North America are used as predictands to train statistical downscaling algorithms using output from ERA5. The statistical downscaling models are then used to make CF projections for a suite of CMIP6 Earth System Models (ESMs). Downscaling is performed using a hybrid statistical approach that employs synoptic types derived using k -means clustering applied to sea level pressure fields with variance corrections applied as a function of the pressure gradient intensity. ESMs exhibit marked variability in terms of the skill with which the frequency of synoptic types and pressure gradients are reproduced relative to ERA5, and that differential skill is used to infer differential credibility in the associated CF projections. Projections of median annual mean CF [P50(CF)] in each 20-yr period from 1980 to 2099 show evidence of declines at most wind farms except in parts of the southern Great Plains, although the magnitude of the changes is strongly dependent on the ESM. For example, P50(CF) in 2080–99 deviate from those in 1980–99 by from −3.1 to +0.2 percentage points in the Northeast. The largest-magnitude declines in P50(CF) ranging from −3.9 to −2 percentage points are projected for the southern West Coast. CF trends exhibit marked seasonality and are strongly linked to changes in the relative intensity of future synoptic patterns, with much less impact from shifts in the occurrence of synoptic types over time. Internal climate modes continue to play a significant role in inducing interannual variability in wind power production, even under high radiative forcing scenarios. Significance Statement We describe how future climate changes may affect wind resources and wind power generation. Near-term changes in projected wind power electricity generation potential at operating wind farms over North America are small, but by the end of the current century electricity production is projected to decrease in many areas but may increase in parts of the southern Great Plains. The amount of change in projected wind power production is a strong function of the Earth system model that is downscaled and also depends on the continued presence of internally forced climate variability. An additional dependence on the amount of greenhouse gas–induced global warming indicates the transition of the energy sector to low-carbon sources may assist in maintaining the abundant U.S. wind resource.

Meteorology & Atmospheric Sciences↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Advanced Laboratory and Field Arrays: Debris Modeling, Detection,& Mitigation (Task 1)

The statement of project objectives for this task was: develop tools, methods and models to assess, and mitigate the risk of damage to MHK infrastructure from woody debris. Develop the capability to detect woody debris using sonar and/or physical methods for purposes of characterizing debris statistics in river (at UAF’s Tanana River Test Site) and near-shore wave (at Yakutat, AK) environments and to activate debris mitigation measures. Develop debris impact risk maps and tables using statistics on debris size, geometry, type, prevalence, mobility and location. Improve and apply the COUPi discrete element method (DEM) to develop models of debris movement and impact on MHK infrastructure to evaluate risk of damage, and interference, to operations from debris. The proposed final deliverable for the task was a set of tools or techniques for providing estimates of the probability of debris impact, and resulting impact forces, on MHK infrastructure as a function of debris size, type, wave regime, and current velocity. Such estimates are required to assess damage risks to operational MHK infrastructure.

13 HYDRO ENERGY↗

Large‐Scale Statistically Meaningful Patterns (LSMPs) Associated With Precipitation Extremes Over Northern California

Abstract We analyze large‐scale statistically meaningful patterns (LSMPs) that precede extreme precipitation (PEx) events over Northern California (NorCal). We find LSMPs by applying k‐means clustering to the two leading principal components of daily 500 hPa geopotential height anomalies two days before the onset, from October to March during 1948–2015. Statistical significance testing based on Monte Carlo simulations suggests a minimum of four statistically distinguished LSMP clusters. The four LSMP clusters are characterized as Northwest continental negative height anomaly, Eastward positive “Pacific‐North American Pattern (PNA),” Westward negative “PNA,” and Prominent Alaskan ridge. These four clusters, shown in multiple variables, evolve very differently and have differing links to the Arctic and tropical Pacific regions. Using binary forecast skill measures and a new copula‐based framework for predicting PEx events, we find LSMP indices that are useful predictors of NorCal PEx events, with moisture‐based variables being the best predictors of PEx events at least 6 days before the onset, and the lower atmospheric variables being better than their upper atmospheric counterparts any day in advance tested. To ensure statistical rigor, the LSMPs analyzed here (with the modified acronym) include local tests of both significance and consistency, which are not always featured in the literature on large‐scale meteorological patterns.

54 ENVIRONMENTAL SCIENCES↗

Full forward model of galaxy clustering statistics with AbacusSummit light cones

ABSTRACT Novel summary statistics beyond the standard 2-point correlation function (2PCF) are necessary to capture the full astrophysical and cosmological information from the small-scale (r < 30h−1Mpc) galaxy clustering. However, the analysis of beyond-2PCF statistics on small scales is challenging because we lack the appropriate treatment of observational systematics for arbitrary summary statistics of the galaxy field. In this paper, we develop a full forward modelling pipeline for a wide range of summary statistics using the large high-fidelity AbacusSummit light cones that account for many systematic effects as well as remain flexible and computationally efficient to enable posterior sampling. We apply our forward model approach to a fully realistic mock galaxy catalog and demonstrate that we can recover unbiased constraints on the underlying galaxy–halo connection model using two separate summary statistics: the standard 2PCF and the novel k-th nearest neighbour (kNN) statistics, which are sensitive to correlation functions of all orders. We will demonstrate its strong constraining power on extended galaxy–halo connection models and cosmology in follow up papers. We expect this to become a powerful approach when applying to upcoming surveys such as DESI where we can leverage a multitude of summary statistics across a wide redshift range to maximally extract information from the non-linear scales.

79 ASTRONOMY AND ASTROPHYSICS↗