Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

On the statistical theory of self-gravitating collisionless dark matter flow: High order kinematic and dynamic relations

Dark matter, if it exists, accounts for five times as much as ordinary baryonic matter. To better understand the self-gravitating collisionless dark matter flow on different scales, a statistical theory involving kinematic and dynamic relations must be developed for different types of flow, e.g., incompressible, constant divergence, and irrotational flow. This is mathematically challenging because of the intrinsic complexity of dark matter flow and the lack of a self-closed description of flow velocity. Here, this paper extends our previous work on second-order statistics Xu to kinematic relations of any order for any type of flow. Dynamic relations were also developed to relate statistical measures of different orders. The results were validated by N-body simulations. On large scales, we found that (i) third-order velocity correlations can be related to density correlation or pairwise velocity; (ii) the pth-order velocity correlations follow ∝ a (p+2)/2 for odd p and ∝ a p/2 for even p, where a is the scale factor; (iii) the overdensity δ is proportional to density correlation on the same scale, $\langle$δ$\rangle$∝$\langle$δδ'$\rangle$; (iv) velocity dispersion on a given scale r is proportional to the overdensity on the same scale. On small scales, (i) a self-closed velocity evolution is developed by decomposing the velocity into motion in haloes and motion of haloes; (ii) the evolution of vorticity and enstrophy are derived from the evolution of velocity; (iii) dynamic relations are derived to relate second- and third-order correlations; (iv) while the first moment of pairwise velocity follows $\langle$Δu L $\rangle$=-Har (H is the Hubble parameter), the third moment follows $\langle$(Δu L ) 3 $\rangle$ ∝ ε u ar that can be directly compared with simulations and observations, where ε u ≈ 10 -7 m 2 /s 3 is the constant rate for energy cascade; (v) the pth order velocity correlations follow ∝ a (3p-5)/4 for odd p and ∝ a 3p/4 for even p. Finally, the combined kinematic and dynamic relations lead to exponential and one-fourth power-law velocity correlations on large and small scales, respectively.

79 ASTRONOMY AND ASTROPHYSICS

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE

Map-level baryonification: unified treatment of weak lensing two-point and higher-order statistics

Precision cosmology benefits from extracting maximal information from cosmic structures, motivating the use of higher-order statistics (HOS) at small spatial scales. However, predicting how baryonic processes modify matter statistics at these scales has been challenging. The baryonic correction model (BCM) addresses this by modifying dark-matter-only simulations to mimic baryonic effects, providing a flexible, simulation-based framework for predicting both two-point and HOS. We show that a 3-parameter version of the BCM can jointly fit weak lensing maps' two-point statistics, wavelet phase harmonics coefficients, scattering coefficients, and the third and fourth moments to within 2% accuracy across all scales ℓ < 2000 and tomographic bins for a DES-Y3-like redshift distribution ( z ≲ 2), using the FLAMINGO simulations. These results demonstrate the viability of BCM-assisted, simulation-based weak lensing inference of two-point and HOS, paving the way for robust cosmological constraints that fully exploit non-Gaussian information on small spatial scales.

79 ASTRONOMY AND ASTROPHYSICS

Dark energy survey year 3 results: likelihood-free, simulation-based w CDM inference with neural compression of weak-lensing map statistics

We present simulation-based cosmological wcold dark matter (wCDM) inference using dark energy survey year 3 weak-lensing maps, via neural data compression of weak-lensing map summary statistics: power spectra, peak counts, and direct map-level compression/inference with convolutional neural networks (CNN). Using simulation-based inference, also known as likelihood-free or implicit inference, we use forward-modelled mock data to estimate posterior probability distributions of unknown parameters. This approach allows all statistical assumptions and uncertainties to be propagated through the forward-modelled mock data; these include sky masks, non-Gaussian shape noise, shape measurement bias, source galaxy clustering, photometric redshift uncertainty, intrinsic galaxy alignments, non-Gaussian density fields, neutrinos, and non-linear summary statistics. We include a series of tests to validate our inference results. This paper also describes the Gower Street simulation suite: 791 full-sky pkdgrav3 dark matter simulations, with cosmological model parameters sampled with a mixed active-learning strategy, from which we construct over 3000 mock dark energy survey lensing data sets. For wCDM inference, for which we allow –1 < w < –$\frac{1}{3}$⁠, our most constraining result uses power spectra combined with map-level (CNN) inference. Using gravitational lensing data only, this map-level combination gives Ω m = 0.283$^{+0.020}_{–0.027}$⁠, S 8 = 0.804$^{+0.025}_{–0.017⁠}$, and w < –0.80 (with a 68 per cent credible interval); compared to the power spectrum inference, this is more than a factor of two improvement in dark energy parameter (Ω⁠ DE , w⁠) precision.

79 ASTRONOMY AND ASTROPHYSICS

Cosmology with second- and third-order shear statistics for the Dark Energy Survey: Methods and simulated analysis

We present a new pipeline designed for the robust inference of cosmological parameters using both second- and third-order shear statistics. We build a theoretical model for rapid evaluation of three-point correlations using our fastnc code and integrate it into the cosmosis framework. We measure the two-point functions 𝜉 ± and the full configuration-dependent three-point shear correlation functions across all auto- and cross-redshift bins. We compress the three-point functions into the mass aperture statistic ⟨ℳ$^{3}_{ap}$⟩ for a set of 796 simulated shear maps designed to model the Dark Energy Survey Year 3 data. We estimate from it the full covariance matrix and model the effects of intrinsic alignments, shear calibration biases and photometric redshift uncertainties. We apply scale cuts to minimize the contamination from the baryonic signal as modeled through hydrodynamical simulations. We find a significant improvement of 83% on the figure of merit in the Ω m − 𝑆 8 plane when we add the ⟨ℳ$^{3}_{ap}$⟩ data to 𝜉 ± . Here, we present our findings for all relevant cosmological and systematic uncertainty parameters and discuss the complementarity of third-order and second-order statistics.

79 ASTRONOMY AND ASTROPHYSICS

Improving statistical precision in Monte Carlo samples with negative weights via reweighting and uncertainty quantification

High statistical precision is critical for Monte Carlo (MC) samples in high energy physics and is degraded by negatively weighted events. This paper investigates a procedure to learn the relationship between the negative and positive weight distributions of any sample, allowing the reduction of statistical uncertainty by reweighting kinematically equivalent events with the same sign. A robust uncertainty quantification method is required for the practical application of such method. Two methods for the estimation of the reweighting uncertainty are developed: one at the event and another one at the final observable level. The latter method is strongly favored. The gains in statistical precision are then quantified. The method is demonstrated on Sherpa vector boson plus jets samples when using all generated events and when restricted to the signal region of a mock analysis. It is demonstrated to significantly reduce stochastic behavior in sparse MC samples while decreasing the overall uncertainty with a sufficiently well-known reweighting function.

Monte Carlo methods

Divide and conquer: using RhizoVision Explorer to aggregate data from multiple root scans using image concatenation and statistical methods

Roots are important in agricultural and natural systems for determining plant productivity and soil carbon inputs. Sometimes, the amount of roots in a sample is too much to fit into a single scanned image, so the sample is divided among several scans, and there is no standard method to aggregate the data. Here, we describe and validate two methods for standardizing measurements across multiple scans: image concatenation and statistical aggregation. We developed a Python script that identifies which images belong to the same sample and returns a single, larger concatenated image. These concatenated images and the original images were processed with RhizoVision Explorer, a free and open-source software. An R script was developed, which identifies rows of data belonging to the same sample and applies correct statistical methods to return a single data row for each sample. These two methods were compared using example images from switchgrass, poplar, and various tree and ericaceous shrub species from a northern peatland and the Arctic. Most root measurements were nearly identical between the two methods except median diameter, which cannot be accurately computed by statistical aggregation. We believe the availability of these methods will be useful to the root biology community.

59 BASIC BIOLOGICAL SCIENCES

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES

Diagnostics of Magnetohydrodynamic Modes in the Interstellar Medium through Synchrotron Polarization Statistics

One of the biggest challenges in understanding magnetohydrodynamic (MHD) turbulence is identifying the plasma mode components from observational data. Previous studies on synchrotron polarization from the interstellar medium (ISM) suggest that the dominant MHD modes can be identified via statistics of Stokes parameters, which would be crucial for studying various ISM processes such as the scattering and acceleration of cosmic rays, star formation, and dynamo. In this paper, we present a numerical study of the synchrotron polarization analysis (SPA) method through systematic investigation of the statistical properties of the Stokes parameters. We derive the theoretical basis for our method from the fundamental statistics of MHD turbulence, recognizing that the projection of the MHD modes allows us to identify the modes dominating the energy fraction from synchrotron observations. Based on the discovery, we revise the SPA method using synthetic synchrotron polarization observations obtained from 3D ideal MHD simulations with a wide range of plasma parameters and driving mechanisms, and present a modified recipe for mode identification. We propose a classification criterion based on a new SPA+ fitting procedure, which allows us to distinguish between Alfvén mode and compressible/slow mode dominated turbulence. We further propose a new method to identify fast modes by analyzing the asymmetry of the SPA+ signature and establish a new asymmetry parameter to detect the presence of fast mode turbulence. Additionally, we confirm through numerical tests that the identification of the compressible and fast modes is not affected by Faraday rotation in both the emitting plasma and the foreground.

97 MATHEMATICS AND COMPUTING

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science

Recent statistical methods for orientation data

The application of statistical methods for determining the areas of animal orientation and navigation are discussed. The method employed is limited to the two-dimensional case. Various tests for determining the validity of the statistical analysis are presented. Mathematical models are included to support the theoretical considerations and tables of data are developed to show the value of information obtained by statistical analysis.

Batschelet, E.

Evaluation of two statistical models using the shock structure problem.

The accuracy of two statistical models for the collision integral of the Boltzmann equation has been evaluated by applying the models to the solution of the problem of shock structure in a monatomic gas and then comparing the theoretical results with available ex perimental data. The two models considered here are the Bhatnagar-Gross-Krook and ellipsoidal statistical models. The Mach number range covered is 1.59-10.7 and profiles for density and, where available, temperature are compared. The method of numerical solution is the discrete ordinate technique which looks quite promising for application to more complicated models. The results indicate that the ellipsoidal statistical model, which gives a correct value for the Prandtl number, gives accurate results for a low Mach number shock. However, the accuracy degenerates as the Mach number increases. The Bhatnagar-Gross-Krook model gives poorer agreement with experimental data in all cases examined.

Giddens, D. P.

Site preferences of Ni/2+/ and Co/2+/ in clinopyroxene and olivine - Limitations of the statistical approach.

Criticism of a statistical approach used by Dasgupta (1972) in analyzing Snyder's (1959) chemical data for minerals from the Duluth Complex in Minnesota. Apart from obvious mathematical objections to citing correlation coefficients to four significant figures from Snyder's relatively inaccurate analytical data, several more fundamental criticisms are leveled at the statistical approach of Dasgupta. These relate to compositional zoning and disequilibrium in the minerals, inhomogeneities of the samples caused by inclusions and exsolved phases, measured site population data for the major cations in olivines and pyroxenes, and the importance of coupled substitutions in the crystal structures. It is concluded that the crystal field predictions of relative enrichments of Ni(2+) and Co(2+) ions in olivine and pyroxene structures have not been disproved by Dasgupta's statistical approach.

Burns, R. G.

Statistical analysis of close pairs of QSOs

The observation of close pairs of QSOs with very different redshifts has been suggested by some as evidence in support of the noncosmological redshift hypothesis. A method is described for determining the statistical significance of such pairs. As an example, it is shown that the statistical significance of the pair 1548+115a,b is not well defined and ranges from approximately 99% confidence to about 60%. If statistical methods are to be used in such cases, they must not be argued a posteriori.

Burbidge, E. M.

Statistical separability of spectral classes of blighted corn

A study was conducted to determine the statistical separability of multispectral measurements from corn having varying levels of southern corn leaf blight severity. Multispectral scanner data in twelve spectral channels in the wavelength range 0.4 to 11.7 microns were analyzed for ten selected flightlines of the 1971 Corn Blight Watch Experiment. A total of 168 corn fields having 18,804 sample points were analyzed. The blight rating information for these fields was available from ground observations. Maximum average transformed divergence between spectral classes of all possible pairs of blight levels, maximized over a subset of channels, was computed in each of one, two, three, and four spectral channels for each of ten flightlines. From the statistical analysis of the values of average transformed divergence, it was concluded that the greater the difference between the blight levels, the more statistically separable they are.

Kumar, R.

Statistical separability of agricultural cover types in subsets of one to twelve spectral channels

The purpose of this study was to determine the statistical separability of multispectral measurements from agricultural cover types: corn, soybeans, green forage (hay and pasture) and forest, in one to twelve spectral channels. Multispectral scanner data in twelve spectral channels in the wavelength range 0.4 to 11.7 microns, acquired for three flightlines were analysed by applying automatic pattern recognition techniques. The same analysis was performed for the data acquired a month later over the same three flightlines to investigate the effect of time on statistical separability of agricultural cover types. In the subsets of one to six spectral channels, the combination of wavelength regions (where V, N, M and T denote the visible, near infrared, middle infrared and thermal infrared wavelength regions, respectively): V, V M, V N M, V N M T, V V N M T, V V N M M T, respectively, were found to be the best choices for getting good overall statistical separability of the agricultural cover types for the data acquired.

Kumar, R.