Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Cellular Statistical Models of Broken Cloud Fields. Part IV: Effects of Pixel Size on Idealized Satellite Observations

In the fourth part of our “Cellular Statistical Models of Broken Cloud Fields” series we use the binary Markov processes framework for quantitative investigation of the effects of low resolution of idealized satellite observations on the statistics of the retrieved cloud masks. We assume that the cloud fields are Markovian and are characterized by the “actual” cloud fraction (CF) and scale length. We use two different models of observations: a simple discrete-point sampling and a more realistic “pixel” protocol. The latter is characterized by a state attribution function (SAF) which has the meaning of the probability that the pixel with a certain CF is declared cloudy in the observed cloud mask. The stochasticity of the SAF means that the cloud/clear attribution is not ideal and can be affected by external or unknown factors. We show that the observed cloud masks can be accurately described as Markov chains of pixels and use the master-matrix formalism (introduced in Part III of the series) for analytical computation of their parameters: the “observed” CF and scale length. This procedure allows us to establish a quantitative relationship (which is pixel-size dependent) between the actual and the observed cloud-field statistics. The feasibility of restoring the former from the latter is considered. The adequacy of our analytical approach to idealized observations is evaluated using numerical simulations. Comparison of the observed parameters of the simulated datasets with their theoretical expectations showed an agreement within 0.005 for the CF, while for the scale length it is within 1% in the sampling case and within 4% in the pixel case.

Mikhail D. Alexandrov↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

Statistical Hypothesis Testing in Wavelet Analysis: Theoretical Developments and Applications to Indian Rainfall

Statistical hypothesis tests in wavelet analysis are methods that assess the degree to which a wavelet quantity(e.g., power and coherence) exceeds background noise. Commonly, a point-wise approach is adopted in which a wavelet quantity at every point in a wavelet spectrum is individually compared to the critical level of the point-wise test. However, because adjacent wavelet coefficients are correlated and wavelet spectra often contain many wavelet quantities, the point-wise test can produce many false positive results that occur in clusters or patches. To circumvent the point-wise test drawbacks, it is necessary to implement the recently developed area-wise, geometric, cumulative area-wise, and topological significance tests, which are reviewed and developed in this paper. To improve the computational efficiency of the cumulative area-wise test, a simplified version of the testing procedure is created based on the idea that its output is the mean of individual estimates of statistical significance calculated from the geometric test applied at a set of point-wise significance levels. Ideal examples are used to show that the geometric and cumulative area-wise tests are unable to differentiate wavelet spectral features arising from singularity-like structures from those associated with periodicities. A cumulative arc-wise test is therefore developed to strictly test for periodicities by using normalized arc length, which is defined as the number of points composing a cross section of a patch divided by the wavelet scale in question. A previously proposed topological significance test is formalized using persistent homology profiles (PHPs) measuring the number of patches and holes corresponding to the set of all point-wise significance values. Ideal examples show that the PHPs can be used to distinguish time series containing signal components from those that are purely noise. To demonstrate the practical uses of the existing and newly developed statistical methodologies, a first comprehensive wavelet analysis of Indian rainfall is also provided. An R software package has been written by the author to implement the various testing procedures.

Wind speed↗

A Global Model for Estimating Atmospheric Phase Scintillation Statistics

Since 2007, the National Aeronautics and Space Administration (NASA) has been collecting atmospheric phase turbulence data from various NASA ground stations throughout the world. The goal of these measurement campaigns has been to generate statistics to characterize the local site turbulence conditions and their impact on widely distributed ground based antenna arrays. This is of critical importance for the situation of uplink arraying, in which a priori knowledge of the fast varying turbulent conditions of water vapor in the troposphere may not be known, and will impact the power combining efficiency of ground based transmitting arrays. Therefore, the design of these type of systems will be dependent on the local climatology of the particular ground station site. Based on the 30+ station years of data collected characterizing atmospheric phase scintillation statistics at various sites, a global model is presented which attempts to predict the average phase statistics of a generic site based on local surface weather data, such as surface pressure, temperature, relative humidity, wind speed, and median wind direction. A model is proposed and based on a standard log power distribution similar to amplitude scintillation models trained on the existing data sets and shows reasonable accuracy against existing data sets.

propagation↗

Statistical learning framework for safety and failure analysis of a DNN-based autonomous aircraft system

Deep Neural Networks (DNNs) and Machine Learning technology is increasingly used for safety-critical applications in the Aerospace domain. To ensure safe operations, the DNN and the system must undergo rigorous verification and validation, including advanced statistical analyses. Performance and safety of the DNN and system behavior must not only be analyzed for the nominal case, but under numerous off-nominal and failure cases. In this paper we will describe how our statistical learning framework SYSAI can efficiently perform such analyses using the tool’s unique combination of advanced learning modeling and statistical analysis techniques. SYSAI can effectively explore the high-dimensional state and failure space of the system under test; geometrical shape detection of safety regions and boundaries support explainability of the results to the designer. In this paper, we report experiments and results obtained with a vision-based DNN control system (ACT) that is capable of autonomously steering an aircraft down a runway.

Yuning He↗

Truncated ARQ Statistical Link Analysis for Dynamic Links

The future deep space links are migrating towards higher frequency bands such as Ka band and optical. These links are susceptible to non Gaussian and non linear effects such as atmospheric turbulence, scintillation, antenna mis-pointing, jitter, etc. These dynamic links thus will experience various degrees of fading loss, and some of these link disruptions cannot be effectively mitigated by forward error correction coding and/or interleaving. One effective way to ensure reliable communication is by using Automatic Repeat Request (ARQ) protocol, where the receiver acknowledges to the transmitter whether or not a data unit is successfully received. If a data unit is not successfully received (such as after a pre-set time-out), the transmitter would then re-transmit the lost data unit to the receiver. In a previous paper, we derived a statistical link analysis method of finding the optimal operating Signal-to-Noise Ratio (SNR) and estimating the latency of an ARQ scheme. In a more recent paper, we demonstrated the above method using the SNR distribution constructed from the Ka-band (32 GHz) flight data. To simplify the discussion, we considered the academic approach that the ARQ scheme allows for an infinite number of retransmissions. In this paper, we consider the more practical case of a truncated ARQ scheme, where there is a limit on the number of retransmissions. We derive the error probability, the optimal SNR setting, and the latency statistics of the correctly received frames of the truncated ARQ schemes. We first discuss the truncated ARQ link analysis principles using the Gaussian assumption for SNR distribution with a large variance. Next, we demonstrate the statistical truncated ARQ link analysis using the SNR distribution constructed from the Ka-band flight data. The results in this paper can be applied in the design of reliable communication systems such as the Consultative Committee for Space Data System (CCSDS) File Transfer Protocol (CFTP) and the Delay Tolerant Network (DTN).

Morabito, David↗

Markovian Statistical Model of Cloud Optical Thickness. Part I: Theory and Examples

We present a generalization of the binary-value Markovian model previously used for statistical characterization of cloud masks to a continuous-value model describing 1D fields of cloud optical thickness (COT). This model has simple functional expressions and is specified by four parameters: the cloud fraction, the autocorrelation (scale) length, and the two parameters of the normalized probability density function of (non-zero) COT values (this PDF is assumed to have gamma-distribution form). Cloud masks derived from this model by separation between the values above and below some threshold in COT appear to have the same statistical properties as in binary-value model described in our previous publications. We demonstrate the ability of our model to generate examples of various cloud-field types by using it to statistically imitate actual cloud observations made by the Research Scanning Polarimeter (RSP) during two field experiments.

binary-value Markovian model↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]↗

On the statistical theory of self-gravitating collisionless dark matter flow: Scale and redshift variation of velocity and density distributions

The statistics of velocity and density fields are crucial for cosmic structure formation and evolution. Here, this paper extends our previous work on the two-point second-order statistics for the velocity field [Phys. Fluids 35, 077105 (2023)] to one-point probability distributions for both density and velocity fields. The scale and redshift variation of density and velocity distributions are studied by a halo-based non-projection approach. First, all particles are divided into halo and out-of-halo particles so that the redshift variation can be studied via generalized kurtosis of distributions for halo and out-of-halo particles, respectively. Second, without projecting particle fields onto a structured grid, the scale variation is analyzed by identifying all particle pairs on different scales $r$. We demonstrate that: (i) Delaunay tessellation can be used to reconstruct the density field. The density correlation, spectrum, and dispersion functions were obtained, modeled, and compared with the N-body simulation; (ii) the velocity distributions are symmetric on both small and large scales and are non-symmetric with a negative skewness on intermediate scales due to the inverse energy cascade on small scales with a constant rate $\varepsilon_u$; (iii) On small scales, the even order moments of pairwise velocity $\Delta u_L$ follow a two-thirds law $\propto{(-\varepsilon_ur)}^{2/3}$, while the odd order moments follow a linear scaling $\langle(\Delta u_L)^{2n+1}\rangle=(2n+1)\langle(\Delta u_L)^{2n}\rangle\langle\Delta u_L\rangle\propto{r}$; (iv) The scale variation of the velocity distributions was studied for longitudinal velocities $u_L$ or $u_L^{'}$, pairwise velocity (velocity difference) $\Delta u_L$=$u_L^{'}$-$u_L$ and velocity sum $\Sigma u_L$=$u^{'}_L$+$u_L$. Fully developed velocity fields are never Gaussian on any scale, despite that they can initially be Gaussian; (v) On small scales, $u_L$ and $\Sigma u_L$ can be modeled by a $X$ distribution to maximize the entropy of the system. The distribution of $\Delta u_L$ can be different; (vi) On large scales, $\Delta u_L$ and $\Sigma u_L$ can be modeled by a logistic or a $X$ distribution, while $u_L$ has a different distribution; (vii) the redshift variation of the velocity distributions follows the evolution of the $X$ distribution involving a shape parameter $\alpha(z)$ decreasing with time.

79 ASTRONOMY AND ASTROPHYSICS↗

Highly accelerated life testing (HALT): A review from a statistical perspective

Despite its use in one form or another for at least four decades, HALT and related techniques [e.g., highly accelerated-stress screening (HASS) and stress audits (HASA)] are not well understood within the statistical community and remain controversial. This largely reflects a conflict in motivation between engineers, testing under harsh conditions to discover and eliminate failure modes, and statisticians, taking a more cautious approach to develop quantitative estimates of parameters such as mean time between failures (MTBF). Here, this review article will clarify HALT concepts and methods and explain where it fits within the universe of methods that involve the application of accelerating factors to compress the time required to evaluate or enhance product reliability. A major distinction is between methods such as HALT, a high-stress test-analyze-fix-test iterative process directed at improving reliability by discovering and fixing weak points in a design, and quantitative accelerated life testing (QALT), whose goal is the estimation of product life for a fixed design. We discuss methods such as physics of failure that offer some hope of bridging the gap between the qualitative nature of HALT, and purely quantitative statistical methods. We present a variety of engineering applications of HALT including metal fatigue, piping and pressure vessels, structural damage, radiation damage, and rotating machinery. We also discuss potential synergies between HALT and QALT, such as rapid identification, through HALT, of failure modes requiring quantitative analysis. For further study, extensive references to the applicable literature are provided as well as an appendix that describes related methods.

97 MATHEMATICS AND COMPUTING↗

The CMS Statistical Analysis and Combination Tool: Combine

This paper describes the Combine software package used for statistical analyses by the CMS Collaboration. The package, originally designed to perform searches for a Higgs boson and the combined analysis of those searches, has evolved to become the statistical analysis tool presently used in the majority of measurements and searches performed by the CMS Collaboration. It is not specific to the CMS experiment, and this paper is intended to serve as a reference for users outside of the CMS Collaboration, providing an outline of the most salient features and capabilities. Readers are provided with the possibility to run Combine and reproduce examples provided in this paper using a publicly available container image. Since the package is constantly evolving to meet the demands of ever-increasing data sets and analysis sophistication, this paper cannot cover all details of Combine. However, the online documentation referenced within this paper provides an up-to-date and complete user guide.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Probing the limits of statistical neutron capture for the r process: Experimental constraints on 141 Cs nuclear level densities

The r-process abundance peaks, particularly near mass number A ∼ 130, reflect underlying nuclear structure effects such as closed neutron shells, yet modeling the nucleosynthesis in this region remains hindered by uncertain neutron-capture rates. These rates are especially sensitive to nuclear level densities (NLDs) and γ-ray strength functions of neutron-rich nuclei, where experimental data are scarce. We present the first experimental constraint on the NLD of 141 Cs using the β-Oslo method, extending sensitivity to the neutron-rich regime near the N = 82 closed shell. Our data allow for critical calibration of microscopic NLD models and reveal that 141 Cs lies near the limit of statistical model applicability. Using this experimental input, we evaluate radiative neutron-capture rates across neighboring isotones using both Hauser–Feshbach (HF) and High Fidelity Resonance (HFR) models. Our results show order-of-magnitude rate increases for nuclei along the N = 86 line, signaling a transition to resonance-dominated capture in this region. These findings underscore the importance of constraining NLDs to improve r-process reaction network predictions, particularly in environments where the validity of statistical models breaks down.

Nuclear level density↗

Size-Resolved Shape Evolution in Inorganic Nanocrystals Captured via High-Throughput Deep Learning-Driven Statistical Characterization

Precise size and shape control in nanocrystal synthesis is essential for utilizing nanocrystals in various industrial applications, such as catalysis, sensing, and energy conversion. However, traditional ensemble measurements often overlook the subtle size and shape distributions of individual nanocrystals, hindering the establishment of robust structure–property relationships. In this study, we uncover intricate shape evolutions and growth mechanisms in Co 3 O 4 nanocrystal synthesis at a subnanometer scale, enabled by deep-learning-assisted statistical characterization. By first controlling synthetic parameters such as cobalt precursor concentration and water amount then using high resolution electron microscopy imaging to identify the geometric features of individual nanocrystals, this study provides insights into the interplay between synthesis conditions and the sizedependent shape evolution in colloidal nanocrystals. Utilizing population-wide imaging data encompassing over 441,067 nanocrystals, we analyze their characteristics and elucidate previously unobserved size-resolved shape evolution. This high-throughput statistical analysis is essential for representing the entire population accurately and enables the study of the size dependency of growth regimes in shaping nanocrystals. Our findings provide experimental quantification of the growth regime transition based on the size of the crystals, specifically (i) for faceting and (ii) from thermodynamic to kinetic, as evidenced by transitions from convex to concave polyhedral crystals. Additionally, we introduce the concept of an “onset radius,” which describes the critical size thresholds at which these transitions occur. This discovery has implications beyond achieving nanocrystals with desired morphology; it enables finely tuned correlation between geometry and material properties, advancing the field of colloidal nanocrystal synthesis and its applications.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Statistical Spread of Transmission Outages on a Fast Protection Time Scale Based on Utility Data

When there is a fault, the protection system automatically removes one or more transmission lines on a fast time scale of less than one minute. The outaged lines form a pattern in the transmission network. We extract these patterns from utility outage data, determine some key statistics of these patterns, and then show how to generate new patterns consistent with these statistics. The generated patterns provide a new and easily feasible way to model the overall effect of the protection system at the scale of a large transmission system. This new data-driven generative modeling of protection is expected to contribute to simulations of disturbances in large grids so that they can better quantify the risk of blackouts. Analysis of the pattern sizes suggests an index that describes how much outages spread in the transmission network at the fast timescale.

Transmission↗

Object-Based Evaluation of Dynamical and Statistical Downscaled Precipitation Products over CONUS

High-resolution precipitation data, generated through dynamical downscaling (DD) or statistical downscaling (SD) of global climate model output, provide critical information for regional climate assessment and adaptation planning. Most downscaling development and validation have focused on accurate gridscale precipitation construction and ignored the spatial structure of precipitation across model grids and at the event scale. However, many applications, e.g., hydrologic modeling and the analysis using the downscaled precipitation, require a reasonable representation of the spatial structure of precipitation within watersheds. Therefore, a set of standard metrics to evaluate the representation of the spatial structure of individual storms across diverse downscaled precipitation products is desired. To address this need, we conducted an object-based evaluation of precipitation in decades-long DD and SD products over the contiguous United States (CONUS). Specifically, we evaluate their ability to reproduce various features of precipitation objects in the observations: total volume, precipitation area, peak intensity, and spatial structure. Multiple metrics (bias, Perkins score, and nonparametric statistical tests) are used to quantify model performance. Our evaluation reveals notable variations in performance among individual products across different climate zones and seasons, as well as between extreme and nonextreme events. In general, most DD products exhibit balanced performance across the four precipitation object features, while SD products vary more significantly in their performance across products. Based on this comprehensive evaluation, we provide guidance on choosing downscaled products for specific regions, seasons, and precipitation object features. These findings and recommendations can inform precipitation-relevant modeling and analysis over CONUS, guide future downscaling technique developments, and provide actionable information for climate impact assessment and adaptation.

Downscaling↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗