Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Cosmology with second- and third-order shear statistics for the Dark Energy Survey: Methods and simulated analysis

We present a new pipeline designed for the robust inference of cosmological parameters using both second- and third-order shear statistics. We build a theoretical model for rapid evaluation of three-point correlations using our fastnc code and integrate it into the cosmosis framework. We measure the two-point functions 𝜉 ± and the full configuration-dependent three-point shear correlation functions across all auto- and cross-redshift bins. We compress the three-point functions into the mass aperture statistic ⟨ℳ$^{3}_{ap}$⟩ for a set of 796 simulated shear maps designed to model the Dark Energy Survey Year 3 data. We estimate from it the full covariance matrix and model the effects of intrinsic alignments, shear calibration biases and photometric redshift uncertainties. We apply scale cuts to minimize the contamination from the baryonic signal as modeled through hydrodynamical simulations. We find a significant improvement of 83% on the figure of merit in the Ω m − 𝑆 8 plane when we add the ⟨ℳ$^{3}_{ap}$⟩ data to 𝜉 ± . Here, we present our findings for all relevant cosmological and systematic uncertainty parameters and discuss the complementarity of third-order and second-order statistics.

79 ASTRONOMY AND ASTROPHYSICS↗

Improving statistical precision in Monte Carlo samples with negative weights via reweighting and uncertainty quantification

High statistical precision is critical for Monte Carlo (MC) samples in high energy physics and is degraded by negatively weighted events. This paper investigates a procedure to learn the relationship between the negative and positive weight distributions of any sample, allowing the reduction of statistical uncertainty by reweighting kinematically equivalent events with the same sign. A robust uncertainty quantification method is required for the practical application of such method. Two methods for the estimation of the reweighting uncertainty are developed: one at the event and another one at the final observable level. The latter method is strongly favored. The gains in statistical precision are then quantified. The method is demonstrated on Sherpa vector boson plus jets samples when using all generated events and when restricted to the signal region of a mock analysis. It is demonstrated to significantly reduce stochastic behavior in sparse MC samples while decreasing the overall uncertainty with a sufficiently well-known reweighting function.

Monte Carlo methods↗

Statistical uncertainties of the v n { 2 k } harmonics from Q cumulants

Analytic formulas to calculate statistical uncertainties of v n { 2 k } harmonics extracted from the Q cumulants are presented. The Q cumulants are multivariate polynomial functions of the weighted means of 2m-particle azimuthal correlations, $\langle\langle 2 m \rangle\rangle$. Variances and covariances of the $\langle\langle 2 m \rangle\rangle$ are included in the analytic formulas of the uncertainties that can be calculated simultaneously with the calculations of the v n { 2 k } harmonics. The calculations are performed using a simple toy model, which roughly simulates elliptic flow azimuthal anisotropy with magnitudes around 0.05. The results are compared with the results obtained by the many data subsets, and by the bootstrapping method. The first one is a common way of estimation of the statistical uncertainties of the v n { 2 k } harmonics in a real experiment. In order to increase precision in the measurement of the v n { 2 k } harmonics, a large number of 15 000 subsets and the same number of the resampling in the bootstrap method is used. Unlike the other ways of the analytic calculation of the v n { 2 k } statistical uncertainties, our proposal that includes the use of squared weights in the calculation of both the variances and covariances, gives the best agreement with the results obtained from the subsets and bootstrap method. Additionally, a recurrence equation between Q cumulants of any order is also presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Using kernel-based statistical distance to study the dynamics of charged particle beams in particle-based simulation codes

Measures of discrepancy between probability distributions (statistical distance) are widely used in the fields of artificial intelligence and machine learning. We describe how certain measures of statistical distance can be implemented as numerical diagnostics for simulations involving charged-particle beams. Related measures of statistical dependence are also described. The resulting diagnostics provide sensitive measures of dynamical processes important for beams in nonlinear or high-intensity systems, which are otherwise difficult to characterize. Here, the focus is on kernel-based methods such as maximum mean discrepancy, which have a well-developed mathematical foundation and reasonable computational complexity. Several benchmark problems and examples involving intense beams are discussed. While the focus is on charged-particle beams, these methods may also be applied to other many-body systems such as plasmas or gravitational systems.

47 OTHER INSTRUMENTATION↗

Statistical upscaling of ecosystem CO 2 fluxes across the terrestrial tundra and boreal domain: Regional patterns and uncertainties

Abstract The regional variability in tundra and boreal carbon dioxide (CO 2 ) fluxes can be high, complicating efforts to quantify sink‐source patterns across the entire region. Statistical models are increasingly used to predict (i.e., upscale) CO 2 fluxes across large spatial domains, but the reliability of different modeling techniques, each with different specifications and assumptions, has not been assessed in detail. Here, we compile eddy covariance and chamber measurements of annual and growing season CO 2 fluxes of gross primary productivity (GPP), ecosystem respiration (ER), and net ecosystem exchange (NEE) during 1990–2015 from 148 terrestrial high‐latitude (i.e., tundra and boreal) sites to analyze the spatial patterns and drivers of CO 2 fluxes and test the accuracy and uncertainty of different statistical models. CO 2 fluxes were upscaled at relatively high spatial resolution (1 km 2 ) across the high‐latitude region using five commonly used statistical models and their ensemble, that is, the median of all five models, using climatic, vegetation, and soil predictors. We found the performance of machine learning and ensemble predictions to outperform traditional regression methods. We also found the predictive performance of NEE‐focused models to be low, relative to models predicting GPP and ER. Our data compilation and ensemble predictions showed that CO 2 sink strength was larger in the boreal biome (observed and predicted average annual NEE −46 and −29 g C m −2 yr −1 , respectively) compared to tundra (average annual NEE +10 and −2 g C m −2 yr −1 ). This pattern was associated with large spatial variability, reflecting local heterogeneity in soil organic carbon stocks, climate, and vegetation productivity. The terrestrial ecosystem CO 2 budget, estimated using the annual NEE ensemble prediction, suggests the high‐latitude region was on average an annual CO 2 sink during 1990–2015, although uncertainty remains high.

Virkkala, Anna‐Maria↗

Divide and conquer: using RhizoVision Explorer to aggregate data from multiple root scans using image concatenation and statistical methods

Roots are important in agricultural and natural systems for determining plant productivity and soil carbon inputs. Sometimes, the amount of roots in a sample is too much to fit into a single scanned image, so the sample is divided among several scans, and there is no standard method to aggregate the data. Here, we describe and validate two methods for standardizing measurements across multiple scans: image concatenation and statistical aggregation. We developed a Python script that identifies which images belong to the same sample and returns a single, larger concatenated image. These concatenated images and the original images were processed with RhizoVision Explorer, a free and open-source software. An R script was developed, which identifies rows of data belonging to the same sample and applies correct statistical methods to return a single data row for each sample. These two methods were compared using example images from switchgrass, poplar, and various tree and ericaceous shrub species from a northern peatland and the Arctic. Most root measurements were nearly identical between the two methods except median diameter, which cannot be accurately computed by statistical aggregation. We believe the availability of these methods will be useful to the root biology community.

59 BASIC BIOLOGICAL SCIENCES↗

A Statistical Evaluation of Combining Human Productivity Metrics in the Indoor Environment

The potential of improving human productivity by providing healthy indoor environments has been a consistent interest in the building field for decades. This research field's long-standing challenge is to measure human productivity given the complex nature of office work. Previous studies have diversified productivity metrics, allowing greater flexibility in collecting human data; however, this diversity complicates the ability to combine productivity metrics from disparate studies within a meta-analysis. This study aims to categorize existing productivity metrics and statistically assess which categories show similar behavior when used to measure the impacts of indoor environmental quality. The 106 productivity metrics compiled were grouped into six productivity metric categories: neurobehavioral speed, accuracy, neurobehavioral response time, call handling time, self-reported productivity, and performance score. Then, this study set neurobehavioral speed as the baseline category given its fitness to the efficiency-based definition of productivity (i.e., output versus input) and conducted three statistical analyses with the other categories to evaluate their similarity. As results, the categories of neurobehavioral response time, self-reported productivity, and call handling time were found to have statistical similarity with neurobehavioral speed. This study contributes to creating a constructive research environment for future meta-analyses to understand which human productivity metrics can be combined with each other.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Perspective on Tsallis statistics for nuclear and particle physics

This is a concise introduction to the topic of nonextensive Tsallis statistics meant especially for those interested in its relation to high-energy proton–proton, proton–nucleus and nucleus–nucleus collisions. The three types of Tsallis statistics are reviewed. Only one of them is consistent with the fundamental hypothesis of equilibrium statistical mechanics. The single-particle distributions associated with it, namely Boltzmann, Fermi–Dirac and Bose–Einstein, are derived. These are not equilibrium solutions to the conventional Boltzmann transport equation which must be modified in a rather nonintuitive manner for them to be so. Nevertheless, the Boltzmann limit of the Tsallis distribution is extremely efficient in representing a wide variety of single-particle distributions in high-energy proton–proton, proton–nucleus and nucleus–nucleus collisions with only three parameters, one of them being the so-called nonextensitivity parameter [Formula: see text]. This distribution interpolates between an exponential at low transverse energy, reflecting thermal equilibrium, to a power law at high transverse energy, reflecting the asymptotic freedom of Quantum Chromodynamics (QCD). It should not be viewed as a fundamental new parameter representing nonextensive behavior in these collisions.

Physics↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part II: Statistical Downscaling

Abstract Capacity factors (CFs) derived from daily expected power at 22 operating wind farms in different regions of North America are used as predictands to train statistical downscaling algorithms using output from ERA5. The statistical downscaling models are then used to make CF projections for a suite of CMIP6 Earth System Models (ESMs). Downscaling is performed using a hybrid statistical approach that employs synoptic types derived using k -means clustering applied to sea level pressure fields with variance corrections applied as a function of the pressure gradient intensity. ESMs exhibit marked variability in terms of the skill with which the frequency of synoptic types and pressure gradients are reproduced relative to ERA5, and that differential skill is used to infer differential credibility in the associated CF projections. Projections of median annual mean CF [P50(CF)] in each 20-yr period from 1980 to 2099 show evidence of declines at most wind farms except in parts of the southern Great Plains, although the magnitude of the changes is strongly dependent on the ESM. For example, P50(CF) in 2080–99 deviate from those in 1980–99 by from −3.1 to +0.2 percentage points in the Northeast. The largest-magnitude declines in P50(CF) ranging from −3.9 to −2 percentage points are projected for the southern West Coast. CF trends exhibit marked seasonality and are strongly linked to changes in the relative intensity of future synoptic patterns, with much less impact from shifts in the occurrence of synoptic types over time. Internal climate modes continue to play a significant role in inducing interannual variability in wind power production, even under high radiative forcing scenarios. Significance Statement We describe how future climate changes may affect wind resources and wind power generation. Near-term changes in projected wind power electricity generation potential at operating wind farms over North America are small, but by the end of the current century electricity production is projected to decrease in many areas but may increase in parts of the southern Great Plains. The amount of change in projected wind power production is a strong function of the Earth system model that is downscaled and also depends on the continued presence of internally forced climate variability. An additional dependence on the amount of greenhouse gas–induced global warming indicates the transition of the energy sector to low-carbon sources may assist in maintaining the abundant U.S. wind resource.

Meteorology & Atmospheric Sciences↗

Overestimated prediction using polygenic prediction derived from summary statistics

When polygenic risk score (PRS) is derived from summary statistics, independence between discovery and test sets cannot be monitored. We compared two types of PRS studies derived from raw genetic data (denoted as rPRS) and the summary statistics for IGAP (sPRS). Two variables with the high heritability in UK Biobank, hypertension, and height, are used to derive an exemplary scale effect of PRS. sPRS without APOE is derived from International Genomics of Alzheimer’s Project (IGAP), which records ΔAUC and ΔR 2 of 0.051 ± 0.013 and 0.063 ± 0.015 for Alzheimer’s Disease Sequencing Project (ADSP) and 0.060 and 0.086 for Accelerating Medicine Partnership - Alzheimer’s Disease (AMP-AD). On UK Biobank, rPRS performances for hypertension assuming a similar size of discovery and test sets are 0.0036 ± 0.0027 (ΔAUC) and 0.0032 ± 0.0028 (ΔR 2 ). For height, ΔR 2 is 0.029 ± 0.0037. Considering the high heritability of hypertension and height of UK Biobank and sample size of UK Biobank, sPRS results from AD databases are inflated. Independence between discovery and test sets is a well-known basic requirement for PRS studies. However, a lot of PRS studies cannot follow such requirements because of impossible direct comparisons when using summary statistics. Thus, for sPRS, potential duplications should be carefully considered within the same ethnic group.

97 MATHEMATICS AND COMPUTING↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

Publishing statistical models: Getting the most out of particle physics experiments

The statistical models used to derive the results of experimental analyses are of incredible scientific value and are essential information for analysis preservation and reuse. In this paper, we make the scientific case for systematically publishing the full statistical models and discuss the technical developments that make this practical. By means of a variety of physics cases - including parton distribution functions, Higgs boson measurements, effective field theory interpretations, direct searches for new physics, heavy flavor physics, direct dark matter detection, world averages, and beyond the Standard Model global fits - we illustrate how detailed information on the statistical modelling can enhance the short- and long-term impact of experimental results.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES↗

A Stellar Activity F-statistic for Exoplanet Surveys (SAFE)

In the search for planets orbiting distant stars, the presence of stellar activity in the atmospheres of observed stars can obscure the radial velocity signal used to detect such planets. Furthermore, this stellar activity contamination is set by the star itself and cannot simply be avoided with better instrumentation. Various stellar activity indicators have been developed that may correlate with this contamination. We introduce a new stellar activity indicator called the Stellar Activity F-statistic for Exoplanet surveys (SAFE) that has higher statistical power (i.e., probability of detecting a true stellar activity signal) than many traditional stellar activity indicators in a simulation study of an active region on a Sun-like star with a moderate-to-high signal-to-noise ratio. Also through simulation, the SAFE is demonstrated to be associated with the projected area on the visible side of the star covered by active regions. We also demonstrate that the SAFE detects statistically significant stellar activity in most of the spectra for HD 22049, a star known to have high stellar variability. Additionally, the SAFE is calculated for recent observations of the three low-variability stars HD 34411, HD 10700, and HD 3651, the latter of which is known to have a planetary companion. As expected, the SAFE for these three only occasionally detects activity. Furthermore, initial exploration appears to indicate that the SAFE may be useful for disentangling stellar activity signals from planet-induced Doppler shifts.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Diagnostics of Magnetohydrodynamic Modes in the Interstellar Medium through Synchrotron Polarization Statistics

One of the biggest challenges in understanding magnetohydrodynamic (MHD) turbulence is identifying the plasma mode components from observational data. Previous studies on synchrotron polarization from the interstellar medium (ISM) suggest that the dominant MHD modes can be identified via statistics of Stokes parameters, which would be crucial for studying various ISM processes such as the scattering and acceleration of cosmic rays, star formation, and dynamo. In this paper, we present a numerical study of the synchrotron polarization analysis (SPA) method through systematic investigation of the statistical properties of the Stokes parameters. We derive the theoretical basis for our method from the fundamental statistics of MHD turbulence, recognizing that the projection of the MHD modes allows us to identify the modes dominating the energy fraction from synchrotron observations. Based on the discovery, we revise the SPA method using synthetic synchrotron polarization observations obtained from 3D ideal MHD simulations with a wide range of plasma parameters and driving mechanisms, and present a modified recipe for mode identification. We propose a classification criterion based on a new SPA+ fitting procedure, which allows us to distinguish between Alfvén mode and compressible/slow mode dominated turbulence. We further propose a new method to identify fast modes by analyzing the asymmetry of the SPA+ signature and establish a new asymmetry parameter to detect the presence of fast mode turbulence. Additionally, we confirm through numerical tests that the identification of the compressible and fast modes is not affected by Faraday rotation in both the emitting plasma and the foreground.

97 MATHEMATICS AND COMPUTING↗

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science↗

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]↗