Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Streaming Statistics

In the context of a larger effort for in situ data analytics, there is a need to calculate basic statistics metrics (e.g., count, mean, median) online as new data points become available. Originally, the code for such online, or in other words streaming, statistics was part of the TALASS (Topological Analysis of Large- Scale Simulations) library. We isolated the relevant code and created a standalone library from it called Streaming Statistics. We also added an ability to serialize and deserialize the statistics objects so that the library can be used in distributed, task-based processing. To use the Streaming Statistics library, the user chooses a statistic, constructs an object for it, and then "adds" values to it, which means the statistic is augmented.

Shudler, Sergei↗

Hierarchical reconstruction of 3D well-connected porous media from 2D exemplars using statistics-informed neural network

The relationships between porous microstructures and transport properties are of fundamental importance in various scientific and engineering applications. Due to the intricacy, stochasticity and heterogeneity of porous media, reliable characterization and modeling of transport properties often require a complete dataset of internal microstructure samples. However, it is often an unbearable cost to acquire sufficient 3D digital microstructures by purely using microscopic imaging systems. Herein this paper presents a machine learning-based technique to hierarchically reconstruct 3D well-connected porous microstructures from one isotropic or several anisotropic low-cost 2D exemplar(s). To compactly characterize the large-scale microstructural features, a Gaussian image pyramid is built for each 2D exemplar. Local morphology patterns are collected from the Gaussian image pyramids, and then they serve as the training data to embed the 2D morphological statistics into feed-forward neural networks at multiple length levels. By using a specially-developed morphology integration scheme, the 3D morphological statistics at different levels can be inferred from the statistics-informed neural networks. Gibbs sampling is adopted to hierarchically reconstruct 3D microstructures by using multi-level 3D morphological statistics, where the large-scale, regional and local morphological patterns are statistically generated and successively added to the same 3D random field. The proposed method is tested on a series of porous media with distinct morphologies, and the statistical equivalence between the reconstructed and the real microstructures is systematically evaluated by comparing morphological descriptors and transport properties. The results demonstrate that the proposed 2D-to-3D microstructure reconstruction method is a universal and efficient approach to generating morphologically and physically realistic samples of porous media.

42 ENGINEERING↗

2D k -th nearest neighbour statistics: a highly informative probe of galaxy clustering

ABSTRACT Beyond standard summary statistics are necessary to summarize the rich information on non-linear scales in the era of precision galaxy clustering measurements. For the first time, we introduce the 2D k-th nearest neighbour (kNN) statistics as a summary statistic for discrete galaxy fields. This is a direct generalization of the standard 1D kNN by disentangling the projected galaxy distribution from the redshift-space distortion signature along the line-of-sight. We further introduce two different flavours of 2D kNNs that trace different aspects of the galaxy field: the standard flavour which tabulates the distances between galaxies and random query points, and a ‘DD’ flavour that tabulates the distances between galaxies and galaxies. We showcase the 2D kNNs’ strong constraining power both through theoretical arguments and by testing on realistic galaxy mocks. Theoretically, we show that 2D kNNs are computationally efficient and directly generate other statistics such as the popular two-point correlation function (2PCF), voids probability function, and counts-in-cell statistics. In a more practical test, we apply the 2D kNN statistics to simulated galaxy mocks that fold in a large range of observational realism and recover parameters of the underlying extended halo occupation distribution (HOD) model that includes velocity bias and galaxy assembly bias. We find unbiased and significantly tighter constraints on all aspects of the HOD model with the 2D kNNs, both compared to the standard 1D kNN, and the classical redshift-space 2PCF.

79 ASTRONOMY AND ASTROPHYSICS↗

Statistical Significance Testing for Mixed Priors: A Combined Bayesian and Frequentist Analysis

In many hypothesis testing applications, we have mixed priors, with well-motivated informative priors for some parameters but not for others. The Bayesian methodology uses the Bayes factor and is helpful for the informative priors, as it incorporates Occam’s razor via the multiplicity or trials factor in the look-elsewhere effect. However, if the prior is not known completely, the frequentist hypothesis test via the false-positive rate is a better approach, as it is less sensitive to the prior choice. We argue that when only partial prior information is available, it is best to combine the two methodologies by using the Bayes factor as a test statistic in the frequentist analysis. We show that the standard frequentist maximum likelihood-ratio test statistic corresponds to the Bayes factor with a non-informative Jeffrey’s prior. We also show that mixed priors increase the statistical power in frequentist analyses over the maximum likelihood test statistic. We develop an analytic formalism that does not require expensive simulations and generalize Wilks’ theorem beyond its usual regime of validity. In specific limits, the formalism reproduces existing expressions, such as the p-value of linear models and periodograms. We apply the formalism to an example of exoplanet transits, where multiplicity can be more than 10 7 . We show that our analytic expressions reproduce the $p$-values derived from numerical simulations. We offer an interpretation of our formalism based on the statistical mechanics. We introduce the counting of states in a continuous parameter space using the uncertainty volume as the quantum of the state. We show that both the $p$-value and Bayes factor can be expressed as an energy versus entropy competition.

97 MATHEMATICS AND COMPUTING↗

Estimating Uncertainty in Simulated ENSO Statistics

Abstract Large ensembles of model simulations are frequently used to reduce the impact of internal variability when evaluating climate models and assessing climate change induced trends. However, the optimal number of ensemble members required to distinguish model biases and climate change signals from internal variability varies across models and metrics. Here we analyze the mean, variance and skewness of precipitation and sea surface temperature in the eastern equatorial Pacific region often used to describe the El Niño–Southern Oscillation (ENSO), obtained from large ensembles of Coupled model intercomparison project phase 6 climate simulations. Leveraging established statistical theory, we develop and assess equations to estimate, a priori, the ensemble size or simulation length required to limit sampling‐based uncertainties in ENSO statistics to within a desired tolerance. Our results confirm that the uncertainty of these statistics decreases with the square root of the time series length and/or ensemble size. Moreover, we demonstrate that uncertainties of these statistics are generally comparable when computed using either pre‐industrial control or historical runs. This suggests that pre‐industrial runs can sometimes be used to estimate the expected uncertainty of statistics computed from an existing historical member or ensemble, and the number of simulation years (run duration and/or ensemble size) required to adequately characterize the statistic. This advance allows us to use existing simulations (e.g., control runs that are performed during model development) to design ensembles that can sufficiently limit diagnostic uncertainties arising from simulated internal variability. These results may well be applicable to variables and regions beyond ENSO.

54 ENVIRONMENTAL SCIENCES↗

Towards testing the theory of gravity with DESI: summary statistics, model predictions and future simulation requirements

Shortly after its discovery, General Relativity (GR) was applied to predict the behavior of our Universe on the largest scales, and later became the foundation of modern cosmology. Its validity has been verified on a range of scales and environments from the Solar system to merging black holes. However, experimental confirmations of GR on cosmological scales have so far lacked the accuracy one would hope for — its applications on those scales being largely based on extrapolation and its validity there sometimes questioned in the shadow of the discovery of the unexpected cosmic acceleration. Future astronomical instruments surveying the distribution and evolution of galaxies over substantial portions of the observable Universe, such as the Dark Energy Spectroscopic Instrument (DESI), will be able to measure the fingerprints of gravity and their statistical power will allow strong constraints on alternatives to GR. In this paper, based on a set of N-body simulations and mock galaxy catalogs, we study the predictions of a number of traditional and novel summary statistics beyond linear redshift distortions in two well-studied modified gravity models — chameleon f(R) gravity and a braneworld model — and the potential of testing these deviations from GR using DESI. These summary statistics employ a wide array of statistical properties of the galaxy and the underlying dark matter field, including two-point and higher-order statistics, environmental dependence, redshift space distortions and weak lensing. We find that they hold promising power for testing GR to unprecedented precision. The major future challenge is to make realistic, simulation-based mock galaxy catalogs for both GR and alternative models to fully exploit the statistic power of the DESI survey (by matching the volumes and galaxy number densities of the mocks to those in the real survey) and to better understand the impact of key systematic effects. Using these, we identify future simulation and analysis needs for gravity tests using DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Engineering Super–Poissonian Photon Statistics of Spatial Light Modes

The nature of light sources is defined by the statistical fluctuations of the electromagnetic field. As such, the photon statistics of light sources are typically associated with distinct emitters. Here, the possibility of producing light beams with various photon statistics through the spatial modulation of coherent light is demonstrated. This is achieved by the sequential encoding of controllable Kolmogorov phase screens in a digital micromirror device. Interestingly, the flexibility of this scheme allows for the shaping of spatial light modes with engineered photon statistics at different spatial positions. The performance of this scheme is assessed through the photon-number-resolving characterization of different families of spatial light modes with engineered photon statistics. Furthermore, it is believed that the possibility of controlling the photon fluctuations of the light field at arbitrary spatial locations has important implications for quantum spectroscopy, sensing, and imaging.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Statistical Downscaling of Climate Models for Solar Resource Assessment

This study presents the development of statistical models to efficiently downscale future projections of solar irradiance for solar energy applications. A climate data set simulated from a Regional Climate Model (RCM) obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) is selected as input to the statistical models to create high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). Our approach builds statistical downscaling models that (1) regrid RCM data (0.22 degree and daily spatiotemporal resolution), (2) correct bias of GHI projections, (3) downscale the future GHI project from daily-scale to hourly-scale, and (4) spatially downscale to generate GHI at 8-km resolution. To calibrate and validate the statistical models, we adapt and use the National Solar Radiation Database (NSRDB). Preliminary results show that the statistical downscaling approach downscales future projections of GHI under two climate scenarios (RCP4.5 and RCP8.5) with a nBIAS of 3%, nMAE of 34% and nRMSE of 46% estimated against NSRDB for the contiguous United State. This presentation will summarize the implemented methodology and validation results as well as future extension of this research.

climate data↗

First-passage time statistics on surfaces of general shape: Surface PDE solvers using Generalized Moving Least Squares (GMLS)

Here, we develop numerical methods for computing statistics of stochastic processes on surfaces of general shape with drift-diffusion dynamics d X t = a (X t ) dt + b(X t ) d W t . We formulate descriptions of Brownian motion and general drift-diffusion processes on surfaces. We consider statistics of the form u (x) = E x [$∫^{τ}_{0}$ g (X t ) dt ] + E x [ f (X τ )] for a domain Ω and the exit stopping time τ = inf t { t >0 | X i Ω}, where f , g are general smooth functions. For computing these statistics, we develop high-order Generalized Moving Least Squares (GMLS) solvers for associated surface PDE boundary-value problems based on Backward- Kolmogorov equations. We focus particularly on the mean First Passage Times (FPTs) given by the case f = 0, g = 1 where u (x) = E x [τ]. We perform studies for a variety of shapes showing our methods converge with high-order accuracy both in capturing the geometry and the surface PDE solutions. We then perform studies showing how statistics are influenced by the surface geometry, drift dynamics, and spatially dependent diffusivities.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

Computational assessment of smooth and rough parameter dependence of statistics in chaotic dynamical systems

An assumption of smooth response to small parameter changes, of statistics or long-time averages of a chaotic system, is generally made in the field of sensitivity analysis, and the parametric derivatives of statistical quantities are critically used in science and engineering. In this paper, we propose a numerical procedure to assess the differentiability of statistics with respect to parameters in chaotic systems. Here, we numerically show that the existence of the derivative depends on the Lebesgue integrability of a certain density gradient function, which we define as the derivative of logarithmic SRB density along the unstable manifold. We develop a recursive formula for the density gradient that can be efficiently computed along trajectories, and demonstrate its use in determining the differentiability of statistics. Our numerical procedure is illustrated on low-dimensional chaotic systems whose statistics exhibit both smooth and rough regions in parameter space.

42 ENGINEERING↗

Efficient high-fidelity TRISO statistical failure analysis using Bison: Applications to AGR-2 irradiation testing

The ability of tri-structural isotropic (TRISO) fuel to contain fission products is largely dictated by the quality of the manufacturing process, since most of the fission product release is expected to occur due to coating layer failure in a small number of particles containing defects. The Bison fuel performance code has capabilities to predict failure in individual particles, accounting for the presence of defects, and to apply statistical analysis methods to compute the probability of failure in a set of fuel particles. Bison has recently undergone significant development both to improve its physical representations of fuel particle behavior and to improve the efficiency of its statistical failure calculations. Physical model improvements include new capabilities to account for the pressure generated by fission gases on inner pyrolytic carbon (IPyC) crack surfaces and to use local material coordinate orientation to accurately incorporate the anisotropy in the material properties in aspherical particles. To improve statistical modeling efficiency, a direct integration approach which involves directly integrating the failure probability function associated with statistically varying parameters has been developed. The direct integration approach is much more efficient than the Monte Carlo (MC) schemes commonly employed, and allows Bison to directly run high-dimensional fuel performance models, which improves the accuracy of failure probability calculations. Finally, a set of benchmark problems is considered here to compare the MC and direct integration approaches, and a statistical failure analysis of compacts in the Advanced Gas Reactor (AGR)-2 experiments is performed using the direct integration approach.

36 MATERIALS SCIENCE↗

Accelerated statistical failure analysis of multifidelity TRISO fuel models

Statistical nuclear fuel failure analysis is critical for the design and development of advanced reactor technologies. Although Monte Carlo Sampling (MCS) is a standard method of statistical failure analysis for fuels, the low failure probabilities of some advanced fuel forms and the correspondingly large number of required model evaluations limit its application to low-fidelity (e.g., 1-D) fuel models. In this paper, we present four other statistical methods for fuel failure analysis in Bison, considering tri-structural isotropic (TRISO)-coated particle fuel as a case study. The statistical methods considered are Latin hypercube sampling (LHS), adaptive importance sampling (AIS), subset simulation (SS), and the Weibull theory. Using these methods, we analyzed both 1-D and 2-D representations of TRISO models to compute failure probabilities and the distributions of fuel properties that result in failures. The results of these methods compare well across all TRISO models considered. Overall, SS and the Weibull theory were deemed the most efficient, and can be applied to both 1-D and 2-D TRISO models to compute failure probabilities. Moreover, since SS also characterizes the distribution of parameters that cause TRISO failures, and can consider failure modes not described by the Weibull criterion, it may be preferred over the other methods. Finally, a discussion on the efficacy of different statistical methods of assessing nuclear fuel safety is provided.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Spatiotemporal and Statistical Mapping of Transition Metal Equilibria in Alkaline Media

Transition metal dissolution and redeposition (D/R) kinetics in alkaline media play a critical role in various chemical and electrochemical processes. Competitive reaction kinetics between different transition metals can modulate individual metal behavior in these processes. To date, these phenomena have remained largely unmeasured, and even when captured, they are difficult to statistically characterize due to their dynamic nature, simultaneous occurrence, and spatially heterogeneous nature. Here, in this study, we develop a statistical analysis framework based on in situ and operando X-ray fluorescence microscopy (XFM) to investigate the relative D/R kinetics of multiple transition metals in alkaline media. By employing statistical analysis, we quantify the spatial distribution of D/R species and assess the rate at which the system reaches equilibrium under varying reaction conditions. We show that pH does not simply change the rate of dissolution and redeposition, but reorganizes the cross-element kinetic correlations among Ni, Fe, and Mn and accelerates the spatial equilibration of D/R events, as quantified through correlation analysis, reaction-rate estimation, probability function distributions, and texture-based monitoring statistics. Additionally, we demonstrate how modifying the solvent environment can influence D/R kinetics, providing a pathway for tuning materials synthesis and process optimization. Our study offers valuable insights into the complex interplay between different transition metals and provides a reliable statistical framework for spatial analysis of diverse imaging data sets, enabling deeper extraction of latent information across multiple modalities.

36 MATERIALS SCIENCE↗

Alternating Conditional Expectations: Introducing a Non‐Parametric Statistical Method to Interpret Long‐Term Greenhouse Gas Flux Measurements Over Semi‐Arid and Wetland Ecosystems

Abstract We explore the potential of using a non‐parametric statistical method called Alternating Conditional Expectations, ACE, to quantify functional relationships in biogeosciences. Here, ACE is used to quantify the non‐linear and multi‐faceted responses of greenhouse gas fluxes to a set of biophysical forcings, when the shapes of those response surfaces are unknown. We evaluated the statistical method over two contrasting ecosystems and two contrasting time steps. One case involved quantifying the biophysical controls of water vapor and carbon dioxide (CO 2 ) fluxes over a semi‐arid oak savanna using daily integrated fluxes. The other case evaluated the responses of CO 2 and methane (CH 4 ) flux measurements to a set of biophysical forcings at a restored tidal wetland using thirty‐minute averages. The statistical model, based on 4 independent variables, explained up over 90% of the variation in daily integrated flux densities of water vapor and net carbon dioxide exchange at the savanna site. This fit was defined by distinct non‐linear responses to such drivers as gross primary production, photosynthetically active radiation, air temperature, vapor pressure deficit and soil moisture. At the tidal wetland site, we evaluated net carbon dioxide and methane fluxes with short‐term measurements to capture the influence of rising and falling tides and seasonality in biological activity. The statistical model defined the shape of the forcing of fluxes due to the roles of carbon exudates, water table depth, oxygen level in the water column, temperature and vegetation status. The statistical fits of the greenhouse gas fluxes were less precise than the savanna case. The fetch varies on a run‐to‐run basis as it is comprised of a heterogeneous mosaic of open water and vegetation. Furthermore, it is difficult to monitor the environmental conditions of the archaea and bacteria in the sediments that produce methane and carbon dioxide.

Environmental Sciences & Ecology↗

Quantum statistical plasmonic metacrystals

Engineering materials that control quantum many-body dynamics remains challenging, as multiparticle interactions typically produce complex emergent behaviour that is difficult to predict. Here we introduce quantum statistical plasmonic metacrystals, structures in which the multiparticle dynamics mediated by optical near fields produce forbidden quantum statistical bands that enable selective transmission of different types of light. This functionality arises from a plasmonic structure composed of nanoantennas acting as meta-atoms. Multiphoton fields with statistics within the allowed bands propagate without distortion, whereas fields in forbidden bands are suppressed or driven towards the nearest accessible statistical state. We show that these bands are determined by the geometry and collective arrangement of the meta-atoms, providing a deterministic route to engineering quantum statistical transport. This platform establishes a room-temperature quantum material intrinsically sensitive to the quantum coherence of many-body photonic systems, enabling their robust manipulation and transport. Our results have implications for coherence-sensitive photonic materials for energy harvesting and scalable many-body quantum technologies.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗