Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Quantitative Predictive Theories through Integrating Quantum, Statistical, Equilibrium, and Nonequilibrium Thermodynamics

Today's thermodynamics is largely based on the combined law for equilibrium systems and statistical mechanics derived by Gibbs in 1873 and 1901, respectively, while irreversible thermodynamics for nonequilibrium systems resides essentially on the Onsager Theorem as a separate branch of thermodynamics developed in 1930s. Between them, quantum mechanics was invented and was quantitatively solved in terms of density functional theory (DFT) in 1960s. Furthermore, these three scientific domains operate based on different principles and are very much separated from each other. In analogy to the parable of the blind men and the elephant articulated by Perdew, they individually represent different portions of a complex system and thus are incomplete by themselves alone, resulting in the lack of quantitative agreement between their predictions and experimental observations. Over the last two decades, the author's group has developed a multiscale entropy approach (recently termed as zentropy theory) that integrates DFT-based quantum mechanics and Gibbs statistical mechanics and is capable of accurately predicting entropy and free energy of complex systems. Furthermore, in combination with the combined law for nonequilibrium systems developed by Hillert, the author developed the theory of cross phenomena beyond the phenomenological Onsager Theorem. The zentropy theory and theory of cross phenomena jointly provide quantitative predictive theories for systems from electronic to any observable scales as reviewed in the present work.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Evidence of galaxy assembly bias in SDSS DR7 galaxy samples from count statistics

We present observational constraints on the galaxy–halo connection, focusing particularly on galaxy assembly bias from a novel combination of counts-in-cylinders statistics, P(N CIC ), with the standard measurements of the projected two-point correlation function w p (r p ), and number density n gal of galaxies. We measure n gal , w p (r p ), and P(N CIC ) for volume-limited, luminosity-threshold samples of galaxies selected from SDSS DR7, and use them to constrain halo occupation distribution (HOD) models, including a model in which galaxy occupation depends upon a secondary halo property, namely halo concentration. We detect significant positive central assembly bias for the M r < -20.0 and M r <-19.5 samples. Central galaxies preferentially reside within haloes of high concentration at fixed mass. Positive central assembly bias is also favoured in the M r < -20.5 and M r < -19.0 samples. We find no evidence of central assembly bias in the M r < -21.0 sample. We observe only a marginal preference for negative satellite assembly bias in the M r < -20.0 and M r < -19.0 samples, and non-zero satellite assembly bias is not indicated in other samples. Our findings underscore the necessity of accounting for galaxy assembly bias when interpreting galaxy survey data, and demonstrate the potential of count statistics in extracting information from the spatial distribution of galaxies, which could be applied to both galaxy–halo connection studies and cosmological analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Robust cosmological inference from non-linear scales with k -th nearest neighbour statistics

ABSTRACT We present the methodology for deriving accurate and reliable cosmological constraints from non-linear scales ($\lt 50\, h^{-1}$ Mpc) with k-th nearest neighbour (kNN) statistics. We detail our methods for choosing robust minimum scale cuts and validating galaxy–halo connection models. Using cross-validation, we identify the galaxy–halo model that ensures both good fits and unbiased predictions across diverse summary statistics. We demonstrate that we can model kNNs effectively down to transverse scales of $r_{\rm p}\sim 3\, h^{-1}$ Mpc and achieve precise and unbiased constraints on the matter density and clustering amplitude, leading to a 2 per cent constraint on σ8. Our simulation-based model pipeline is resilient to varied model systematics, spanning simulation codes, halo finding, and cosmology priors. We demonstrate the effectiveness of this approach through an application to the Beyond-2p mock challenge. We propose further explorations to test more complex galaxy–halo connection models and tackle potential observational systematics.

79 ASTRONOMY AND ASTROPHYSICS↗

Distinct universality classes of diffusive transport from full counting statistics

The hydrodynamic transport of local conserved densities furnishes an effective coarse-grained description of the dynamics of a many-body quantum system. However, the full quantum dynamics contains much more structure beyond the simplified hydrodynamic description. Here we show that systems with the same hydrodynamics can nevertheless belong to distinct dynamical universality classes, as revealed by new classes of experimental observables accessible in synthetic quantum systems, which can, for instance, measure simultaneous site-resolved snapshots of all of the particles in a system. Specifically, we study the full counting statistics of spin transport, whose first moment is related to linear-response transport, but the higher moments go beyond. We present an analytic theory of the full counting statistics of spin transport in various integrable and nonintegrable anisotropic one-dimensional spin models, including the XXZ spin chain. We find that spin transport, while diffusive on average, is governed by a distinct non-Gaussian dynamical universality class in the models considered. We consider a setup in which the left and right half of the chain are initially created at different magnetization densities, and consider the probability distribution of the magnetization transferred between the two half-chains. We derive a closed-form expression for the probability distribution of the magnetization transfer, in terms of random walks on the half-line. We show that this distribution strongly violates the large-deviation form expected for diffusive chaotic systems, and explain the physical origin of this violation. Here, we discuss the crossovers that occur as the initial state is brought closer to global equilibrium. Our predictions can directly be tested in experiments using quantum gas microscopes or superconducting qubit arrays.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Full Counting Statistics of Charge in Chaotic Many-Body Quantum Systems

We investigate the full counting statistics of charge transport in U(1)-symmetric random unitary circuits. We consider an initial mixed state prepared with a chemical potential imbalance between the left and right halves of the system and study the fluctuations of the charge transferred across the central bond in typical circuits. Using an effective replica statistical mechanics model and a mapping onto an emergent classical stochastic process valid at large on-site Hilbert space dimension, we show that charge transfer fluctuations approach those of the symmetric exclusion process at long times, with subleading t –1/2 quantum corrections. Here, we discuss our results in the context of fluctuating hydrodynamics and macroscopic fluctuation theory of classical nonequilibrium systems and check our predictions against direct matrix-product state calculations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Extracting Resilience Metrics From Distribution Utility Data Using Outage and Restore Process Statistics

Resilience curves track the accumulation and restoration of outages during an event on an electric distribution grid. We show that a resilience curve generated from utility data can always be decomposed into an outage process and a restore process and that these processes generally overlap in time. We use many events in real utility data to characterize the statistics of these processes, and derive formulas based on these statistics for resilience metrics such as restore duration, customer hours not served, and outage and restore rates. The formulas express the mean value of these metrics as a function of the number of outages in the event. We also give a formula for the variability of restore duration, which allows us to predict a maximum restore duration with 95% confidence. Overall, we give a simple and general way to decompose resilience curves into outage and restore processes and then show how to use these processes to extract resilience metrics from standard distribution system data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Statistical Spread of Transmission Outages on a Fast Protection Time Scale Based on Utility Data

When there is a fault, the protection system automatically removes one or more transmission lines on a fast time scale of less than one minute. The outaged lines form a pattern in the transmission network. We extract these patterns from utility outage data, determine some key statistics of these patterns, and then show how to generate new patterns consistent with these statistics. The generated patterns provide a new and easily feasible way to model the overall effect of the protection system at the scale of a large transmission system. This new data-driven generative modeling of protection is expected to contribute to simulations of disturbances in large grids so that they can better quantify the risk of blackouts. Analysis of the pattern sizes suggests an index that describes how much outages spread in the transmission network at the fast timescale.

Transmission↗

Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning

Various modeling techniques are used to predict the capacity fade of Li-ion batteries. Algebraic reduced-order models, which are inherently interpretable and computationally fast, are ideal for use in battery controllers, technoeconomic models, and multi-objective optimizations. For Li-ion batteries with graphite anodes, solid-electrolyte-interphase (SEI) growth on the graphite surface dominates fade. This fade is often modeled using physically informed equations, such as square-root of time for predicting solvent-diffusion limited SEI growth, and Arrhenius and Tafel-like equations predicting the temperature and state-of-charge rate dependencies. In some cases, completely empirical relationships are proposed. However, statistical validation is rarely conducted to evaluate model optimality, and only a handful of possible models are usually investigated. This article demonstrates a novel procedure for automatically identifying reduced-order degradation models from millions of algorithmically generated equations via bi-level optimization and symbolic regression. Identified models are statistically validated using cross-validation, sensitivity analysis, and uncertainty quantification via bootstrapping. On a LiFePO 4 /Graphite cell calendar aging data set, automatically identified models utilizing square-root, power law, stretched exponential, and sigmoidal functions result in greater accuracy and lower uncertainty than models identified by human experts, and demonstrate that previously known physical relationships can be empirically "rediscovered" using machine learning.

25 ENERGY STORAGE↗

Object-Based Evaluation of Dynamical and Statistical Downscaled Precipitation Products over CONUS

High-resolution precipitation data, generated through dynamical downscaling (DD) or statistical downscaling (SD) of global climate model output, provide critical information for regional climate assessment and adaptation planning. Most downscaling development and validation have focused on accurate gridscale precipitation construction and ignored the spatial structure of precipitation across model grids and at the event scale. However, many applications, e.g., hydrologic modeling and the analysis using the downscaled precipitation, require a reasonable representation of the spatial structure of precipitation within watersheds. Therefore, a set of standard metrics to evaluate the representation of the spatial structure of individual storms across diverse downscaled precipitation products is desired. To address this need, we conducted an object-based evaluation of precipitation in decades-long DD and SD products over the contiguous United States (CONUS). Specifically, we evaluate their ability to reproduce various features of precipitation objects in the observations: total volume, precipitation area, peak intensity, and spatial structure. Multiple metrics (bias, Perkins score, and nonparametric statistical tests) are used to quantify model performance. Our evaluation reveals notable variations in performance among individual products across different climate zones and seasons, as well as between extreme and nonextreme events. In general, most DD products exhibit balanced performance across the four precipitation object features, while SD products vary more significantly in their performance across products. Based on this comprehensive evaluation, we provide guidance on choosing downscaled products for specific regions, seasons, and precipitation object features. These findings and recommendations can inform precipitation-relevant modeling and analysis over CONUS, guide future downscaling technique developments, and provide actionable information for climate impact assessment and adaptation.

Downscaling↗

Differential credibility assessment for statistical downscaling

Climate science is increasingly using (i) ensembles of climate projections from multiple models derived using different assumptions and/or scenarios and (ii) process-oriented diagnostics of model fidelity. Efforts to assign differential credibility to projections and/or models are also rapidly advancing. A framework to quantify and depict the credibility of statistically downscaled model output is presented and demonstrated. Here, the approach employs transfer functions in the form of robust and resilient generalized linear models applied to downscale daily minimum and maximum temperature anomalies at 10 locations using predictors drawn from ERA-Interim reanalysis and two global climate models (GCM; GFDL-ESM2M and MPI-ESM-LR). The downscaled time series are used to derive several impact relevant CLIMDEX temperature indices that are assigned credibility based on (1) the reproduction of relevant large-scale predictors by the GCMs (i.e. fraction of regression beta-weights derived from predictors that are well-reproduced) and (2) the degree of variance in the observations reproduced in the downscaled series following application of a new variance inflation technique. Credibility of the downscaled predictands varies across locations, between the two GCM and is generally higher for minimum temperature than maximum temperature. The differential credibility assessment framework demonstrated here is easy to use and flexible. It can be applied as is to inform decision makers regarding projection confidence, and/or extended to include other components of the transfer functions, and/or used to weight members of a statistically downscaled ensemble.

54 ENVIRONMENTAL SCIENCES↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

MIDAS: Modeling Individual Differences using Advanced Statistics

This research explores novel methods for extracting relevant information from EEG data to characterize individual differences in cognitive processing. Our approach combines expertise in machine learning, statistics, and cognitive science, advancing the state-of-the art in all three domains. Specifically, by using cognitive science expertise to interpret results and inform algorithm development, we have developed a generalizable and interpretable machine learning method that can accurately predict individual differences in cognition. The output of the machine learning method revealed surprising features of the EEG data that, when interpreted by the cognitive science experts, provided novel insights to the underlying cognitive task. Additionally, the outputs of the statistical methods show promise as a principled approach to quickly find regions within the EEG data where individual differences lie, thereby supporting cognitive science analysis and informing machine learning models. This work lays methodological ground work for applying the large body of cognitive science literature on individual differences to high consequence mission applications.

97 MATHEMATICS AND COMPUTING↗

Advances in statistical methods for cancer surveillance research: an age-period-cohort perspective

Background: Analysis of Lexis diagrams (population-based cancer incidence and mortality rates indexed by age group and calendar period) requires specialized statistical methods. However, existing methods have limitations that can now be overcome using new approaches. Methods: We assembled a “toolbox” of novel methods to identify trends and patterns by age group, calendar period, and birth cohort. We evaluated operating characteristics across 152 cancer incidence Lexis diagrams compiled from United States (US) Surveillance, Epidemiology and End Results Program data for 21 leading cancers in men and women in four race and ethnicity groups (the “cancer incidence panel”). Results: Nonparametric singular values adaptive kernel filtration (SIFT) decreased the estimated root mean squared error by 90% across the cancer incidence panel. A novel method for semi-parametric age-period-cohort analysis (SAGE) provided optimally smoothed estimates of age-period-cohort (APC) estimable functions and stabilized estimates of lack-of-fit (LOF). SAGE identified statistically significant birth cohort effects across the entire cancer panel; LOF had little impact. As illustrated for colon cancer, newly developed methods for comparative age-period-cohort analysis can elucidate cancer heterogeneity that would otherwise be difficult or impossible to discern using standard methods. Conclusions: Cancer surveillance researchers can now identify fine-scale temporal signals with unprecedented accuracy and elucidate cancer heterogeneity with unprecedented specificity. Birth cohort effects are ubiquitous modulators of cancer incidence in the US. The novel methods described here can advance cancer surveillance research.

60 APPLIED LIFE SCIENCES↗

Collaborative Exploration of Scientific Datasets Using Immersive and Statistical Visualization: Preprint

We discuss the value of collaborative, immersive visualization for the exploration of scientific datasets and review techniques and tools that have been developed and deployed at the National Renewable Energy Laboratory (NREL). We believe that collaborative visualizations linking statistical interfaces and graphics on laptops and high-performance computing (HPC) with 3D visualizations on immersive displays (head-mounted displays and large-scale immersive environments) enable scientific workflows that further rapid exploration of large, high-dimensional datasets by teams of analysts. We present a framework, PlottyVR, that blends statistical tools, general-purpose programming environments, and simulation with 3D visualizations. To contextualize this framework, we propose a categorization and loose taxonomy of collaborative visualization and analysis techniques. Finally, we describe how scientists and engineers have adopted this framework to investigate large, complex datasets.

collaborative visualization↗

Collaborative Exploration of Scientific Datasets Using Immersive and Statistical Visualization

We discuss the value of collaborative, immersive visualization for the exploration of scientific datasets and review techniques and tools that have been developed and deployed at the National Renewable Energy Laboratory (NREL). We believe that collaborative visualizations linking statistical interfaces and graphics on laptops and high-performance computing (HPC) with 3D visualizations on immersive displays (head-mounted displays and large-scale immersive environments) enable scientific workflows that further rapid exploration of large, high-dimensional datasets by teams of analysts. We present a framework, PlottyVR, that blends statistical tools, general-purpose programming environments, and simulation with 3D visualizations. To contextualize this framework, we propose a categorization and loose taxonomy of collaborative visualization and analysis techniques. Finally, we describe how scientists and engineers have adopted this framework to investigate large, complex datasets.

collaborative visualization↗

Wormholes from heavy operator statistics in AdS/CFT

We construct higher dimensional Euclidean AdS wormhole solutions that reproduce the statistical description of the correlation functions of an ensemble of heavy CFT operators. We consider an operator which effectively backreacts on the geometry in the form of a thin shell of dust particles. Assuming dynamical chaos in the form of the ETH ansatz, we demonstrate that the semiclassical path integral provides an effective statistical description of the microscopic features of the thin shell operator in the CFT. The Euclidean wormhole solutions provide microcanonical saddlepoint contributions to the cumulants of the correlation functions over the ensemble of operators. We finally elaborate on the role of these wormholes in the context of non-perturbative violations of bulk global symmetries in AdS/CFT.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of precipitation indices in suites of dynamically and statistically downscaled regional climate models over Florida

Abstract The present work evaluates historical precipitation and its indices defined by the Expert Team on Climate Change Detection and Indices (ETCCDI) in suites of dynamically and statistically downscaled regional climate models (RCMs) against NOAA’s Global Historical Climatology Network Daily (GHCN-Daily) dataset over Florida. The models examined here are: (1) nested RCMs involved in the North American CORDEX (NA-CORDEX) program, (2) variable resolution Community Earth System Models (VR-CESM), (3) Coupled Model Intercomparison Project phase 5 (CMIP5) models statistically downscaled using localized constructed analogs (LOCA) technique. To quantify observational uncertainty, three in situ-based (PRISM, Livneh, CPC) and three reanalysis (ERA5, MERRA2, NARR) datasets are also evaluated against the station data. The reanalyses and dynamically downscaled RCMs generally underestimate the magnitude of the monthly precipitation and the frequency of the extreme rainfall in summer. The models forced with CanESM2 miss the phase of the seasonality of extreme precipitation. All models and reanalyses severely underestimate both the mean and interannual variability of mean wet-day precipitation (SDII), consecutive dry days (CDD), and overestimate consecutive wet days (CWD). Metric analysis suggests large uncertainty across NA-CORDEX models. Both the LOCA and VR-CESM models perform better than the majority of models. Overall, RegCM4 and WRF models perform poorer than the median model performance. The performance uncertainty across models is comparable to that in the reanalyses. Specifically, NARR performs poorer than the median model performance in simulating the mean indices and MERRA2 performs worse than the majority of models in capturing the interannual variability of the indices.

54 ENVIRONMENTAL SCIENCES↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗