Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

The look-elsewhere effect from a unified Bayesian and frequentist perspective

When searching over a large parameter space for anomalies such as events, peaks, objects, or particles, there is a large probability that spurious signals with seemingly high signi ficance will be found. This is known as the look-elsewhere effect and is prevalent throughout cosmology, (astro)particle physics, and beyond. To avoid making false claims of detection, one must account for this effect when assigning the statistical significance of an anomaly. This is typically accomplished by considering the trials factor, which is generally computed numerically via potentially expensive simulations. In this paper we develop a continuous generalization of the Bonferroni and Sidak corrections by applying the Laplace approximation to evaluate the Bayes factor, and in turn relating the trials factor to the prior-to-posterior volume ratio. Here, we use this to define a test statistic whose frequentist properties have a simple interpretation in terms of the global p-value, or statistical significance. We apply this method to various physics-based examples and show it to work well for the full range of p-values, i.e. in both the asymptotic and non-asymptotic regimes. We also show that this method naturally accounts for other model complexities such as additional degrees of freedom, generalizing Wilks' theorem. This provides a fast way to quantify statistical significance in light of the look-elsewhere effect, without resorting to expensive simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Maximum entropy distributions of dark matter in ΛCDM cosmology

Context. Small-scale challenges to ΛCDM cosmology require a deeper understanding of dark matter physics. Aims. This paper aims to develop the maximum entropy distributions for dark matter particle velocity (denoted by X ), speed (denoted by Z ), and energy (denoted by E ) that are especially relevant on small scales where system approaches full virialization. Methods. For systems involving long-range interactions, a spectrum of halos of different sizes is required to form to maximize system entropy. While the velocity in halos can be Gaussian, the velocity distribution throughout the entire system, involving all halos of different sizes, is non-Gaussian. With the virial theorem for mechanical equilibrium, we applied the maximum entropy principle to the statistical equilibrium of entire system, such that the maximum entropy distribution of velocity (the X distribution) could be analytically derived. The halo mass function was not required in this formulation, but it did indeed result from the maximum entropy. Results. The predicted X distribution involves a shape parameter α and a velocity scale, v 0 . The shape parameter α reflects the nature of force ( α → 0 for long-range force or α → ∞ for short-range force). Therefore, the distribution approaches Laplacian with α → 0 and Gaussian with α → ∞. For an intermediate value of α , the distribution naturally exhibits a Gaussian core for v ≪ v 0 and exponential wings for v ≫ v 0 , as confirmed by N -body simulations. From this distribution, the mean particle energy of all dark matter particles with a given speed, v , follows a parabolic scaling for low speeds (∝ v 2 for v ≪ v 0 in halo core region, i.e., “Newtonian”) and a linear scaling for high speeds (∝ v for v ≫ v 0 in halo outskirt, i.e., exhibiting “non-Newtonian” behavior due to long-range gravity). We compared our results against N -body simulations and found a good agreement.

79 ASTRONOMY AND ASTROPHYSICS↗

Informed total-error-minimizing priors: Interpretable cosmological parameter constraints despite complex nuisance effects

While Bayesian inference techniques are standard in cosmological analyses, it is common to interpret resulting parameter constraints with a frequentist intuition. This intuition can fail, for example, when marginalizing high-dimensional parameter spaces onto subsets of parameters, because of what has come to be known as projection effects or prior volume effects. We present the method of informed total-error-minimizing (ITEM) priors to address this problem. An ITEM prior is a prior distribution on a set of nuisance parameters, such as those describing astrophysical or calibration systematics, intended to enforce the validity of a frequentist interpretation of the posterior constraints derived for a set of target parameters (e.g., cosmological parameters). Our method works as follows. For a set of plausible nuisance realizations, we generate target parameter posteriors using several different candidate priors for the nuisance parameters. We reject candidate priors that do not accomplish the minimum requirements of bias (of point estimates) and coverage (of confidence regions among a set of noisy realizations of the data) for the target parameters on one or more of the plausible nuisance realizations. Of the priors that survive this cut, we select the ITEM prior as the one that minimizes the total error of the marginalized posteriors of the target parameters. As a proof of concept, we applied our method to the density split statistics measured in Dark Energy Survey Year 1 data. We demonstrate that the ITEM priors substantially reduce prior volume effects that otherwise arise and that they allow for sharpened yet robust constraints on the parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying uncertainty in analysis of shockless dynamic compression experiments on platinum. II. Bayesian model calibration

Dynamic shockless compression experiments provide the ability to explore material behavior at extreme pressures but relatively low temperatures. Typically, the data from these types of experiments are interpreted through an analytic method called Lagrangian analysis. Here, in this work, alternative analysis methods are explored using modern statistical methods. Specifically, Bayesian model calibration is applied to a new set of platinum data shocklessly compressed to 570 GPa. Several platinum equation-of-state models are evaluated, including traditional parametric forms as well as a novel non-parametric model concept. The results are compared to those in Paper I obtained by inverse Lagrangian analysis. The comparisons suggest that Bayesian calibration is not only a viable framework for precise quantification of the compression path, but also reveals insights pertaining to trade-offs surrounding model form selection, sensitivities of the relevant experimental uncertainties, and assumptions and limitations within Lagrangian analysis. The non-parametric model method, in particular, is found to give precise unbiased results and is expected to be useful over a wide range of applications. The calibration results in estimates of the platinum principal isentrope over the full range of experimental pressures to a standard error of 1.6%, which extends the results from Paper I while maintaining the high precision required for the platinum pressure standard.

Brown, Justin Lee↗

Progress Toward Interpretable Machine Learning–Based Disruption Predictors Across Tokamaks

Here in this paper we lay the groundwork for a robust cross-device comparison of data-driven disruption prediction algorithms on DIII-D and JET tokamaks. In order to consistently carry on a comparative analysis, we define physics-based indicators of disruption precursors based on temperature, density, and radiation profiles that are currently not used in many other machine learning predictors for DIII-D data. These profile-based indicators are shown to well-describe impurity accumulation events in both DIII-D and JET discharges that eventually disrupt. The univariate analysis of the features used as input signals in the data-driven algorithms applied on the data of both tokamaks statistically highlights the differences in the dominant disruption precursors. JET with its ITER-like wall is more prone to impurity accumulation events, while DIII-D is more subject to edge-cooling mechanisms that destabilize dangerous magnetohydrodynamic modes. Even though the analyzed data sets are characterized by such intrinsic differences, we show through a few examples that the inclusion of physics-based disruption markers in data-driven algorithms is a promising path toward the realization of a uniform framework to predict and interpret disruptive scenarios across different tokamaks. As long as the destabilizing precursors are diagnosed in a device-independent way, the knowledge that data-driven algorithms learn on one device can be re-used to explain a disruptive behavior on another device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A new method for measuring the 3D turbulent velocity dispersion of molecular clouds

ABSTRACT The structure and star formation activity of a molecular cloud are fundamentally linked to its internal turbulence. However, accurately measuring the turbulent velocity dispersion is challenging due to projection effects and observational limitations, such as telescope resolution, particularly for clouds that include non-turbulent motions, such as large-scale rotation. Here, we develop a new method to recover the 3D turbulent velocity dispersion (σv,3D) from position–position–velocity (PPV) data. We simulate a rotating, turbulent, collapsing molecular cloud, and compare its intrinsic σv,3D with three different measures of the velocity dispersion accessible in PPV space: (1) the spatial mean of the 2nd-moment map, σi, (2) the standard deviation of the gradient/rotation-corrected 1st-moment map, σ(c − grad), and (3) a combination of (1) and (2), called the ‘gradient-corrected parent velocity dispersion’, $\sigma _{\mathrm{(p}-\mathrm{grad)}}=(\sigma _{\mathrm{i}}^2+\sigma _{(\mathrm{c}-\mathrm{grad)}}^2)^{1/2}$. We show that the gradient correction is crucial in order to recover purely turbulent motions of the cloud, independent of the orientation of the cloud with respect to the line of sight. We find that with a suitable correction factor and appropriate filters applied to the moment maps, all three statistics can be used to recover σv,3D, with method 3 being the most robust and reliable. We determine the correction factor as a function of the telescope beam size for different levels of cloud rotation, and find that for a beam full width at half-maximum f and cloud radius R, the 3D turbulent velocity dispersion can best be recovered from the gradient-corrected parent velocity dispersion via $\sigma _{v,\mathrm{3D}}= \left[(-0.29\pm 0.26)\, f/R + 1.93 \pm 0.15\right] \sigma _{\mathrm{(p}-\mathrm{grad)}}$ for f/R < 1, independent of the level of cloud rotation or LOS orientation.

Stewart, Madeleine↗

Systematics in asteroseismic modelling: application of a correlated noise model for oscillation frequencies

ABSTRACT The detailed modelling of stellar oscillations is a powerful approach to characterizing stars. However, poor treatment of systematics in theoretical models leads to misinterpretations of stars. Here, we propose a more principled statistical treatment for the systematics to be applied to fitting individual mode frequencies with a typical stellar model grid. We introduce a correlated noise model based on a Gaussian process (GP) kernel to describe the systematics given that mode frequency systematics are expected to be highly correlated. We show that tuning the GP kernel can reproduce general features of frequency variations for changing model input physics and fundamental parameters. Fits with the correlated noise model better recover stellar parameters than traditional methods that either ignore the systematics or treat them as uncorrelated noise.

Li, Tanda (ORCID:0000000163962563)↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

Developing and applying quantifiable metrics for diagnostic and experiment design on Z

This project applies methods in Bayesian inference and modern statistical methods to quantify the value of new experimental data, in the form of new or modified diagnostic configurations and/or experiment designs. We demonstrate experiment design methods that can be used to identify the highest priority diagnostic improvements or experimental data to obtain in order to reduce uncertainties on critical inferred experimental quantities and select the best course of action to distinguish between competing physical models. Bayesian statistics and information theory provide the foundation for developing the necessary metrics, using two high impact experimental platforms on Z as exemplars to develop and illustrate the technique. We emphasize that the general methodology is extensible to new diagnostics (provided synthetic models are available), as well as additional platforms. We also discuss initial scoping of additional applications that began development in the last year of this LDRD.

97 MATHEMATICS AND COMPUTING↗

Myriad World Baseline: Global Geodemographic Estimates

The LandScan Myriad World Baseline (MWB) method produces global, residential (nighttime/home-location) gridded geodemographic estimates based on 5-year age/gender cohorts—at 30-arcsecond (≈1 km) resolution. MWB is designed to fill gaps where detailed, georeferenced survey data (e.g., Demographic and Health Surveys (DHS)) are missing or outdated, and to provide a baseline that can support human security analysis, including consequence assessment, “patterns of life” modeling, and scenario-based population futures. MWB’s workflow spatializes household-level age/gender characteristics from the GLOPOP-S dataset by conflating household and gridded expected relative wealth adapted from Global Gridded Relative Deprivation Index (GRDI), then adjusts them to a target year of interest. Age/gender estimates are then applied to harmonize lowest-administrative-level statistics with LandScan residential counts, yielding final geodemographic estimates. Two validation case studies are presented: Ghana (2021) and Tokyo/Kanagawa, Japan (2020), illustrating spatial variability in demographic cohorts and comparing MWB outputs to official gridded statistics. Results show close overall alignment relative to validation criteria including population pyramids and age-dependency ratios.

Tuccillo, Joe [ORNL] (ORCID:0000000259300943)↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Uncertainties in turbulent statistics and fluxes of CO2 associated with density effect corrections

Density effect corrections (DEC) are applied to adjust raw CO2 fluxes measured by eddy covariance (EC) systems with open-path gas analyzers. DEC is also required for adjusting the measured CO2 concentration fluctuations to obtain the adjusted CO2 for analyzing turbulent statistics or quantifying fluxes. However, our data show that the power spectra of the DEC-adjusted CO2 are distorted in the high frequency range, as compared with the corresponding spectra of temperature and water vapor density. This contradicts the similarity behavior of scalars, as suggested by Monin-Obukhov similarity theory. It is demonstrated that such a distortion is caused by the DEC-induced spikes in the DEC-adjusted CO2, altering turbulent statistics of CO2 and scalar similarity between CO2 and other scalars. Our results suggest that CO2 fluxes are overestimated by applying DEC especially under high Bowen ratio conditions, potentially leading to substantial uncertainties in long-term ecosystem carbon exchange in dry regions.

Gao, Zhongming↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Changes in When and Where People are Spending Time in Response to COVID-19

The COVID-19 pandemic has resulted in a significant change in driving behavior as people respond to the new environment. However, existing methods for analyzing driver behavior such as travel surveys and travel demand models are not suited for incorporating abrupt environmental disruptions. To address this, we analyze a set of high-resolution trip data and introduce two new metrics for quantifying driving behavioral shifts as a function of time, allowing us to compare the time periods before and after pandemic began. We apply these metrics to the Denver, Colorado metropolitan statistical area (MSA) to demonstrate the utility of the metrics. Then, we present a case study for comparing two distinct MSAs, Louisville, Kentucky; and Des Moines, Iowa which exhibit significant differences in the makeup of their labor markets. The results indicate that although the regions of study exhibit certain unique driving behavioral shifts, emerging trends can be seen when comparing between seemingly distinct regions. For instance, drivers in all three MSAs are generally shown to have spent more time at residential locations and less time in workplaces in the time period after the pandemic started. In addition, workplaces that may be incompatible with remote working, such as hospitals and certain retail locations, generally retained much of their pre-pandemic travel activity.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗

Applied Risk Analysis for Guiding Homeland Security Policy

Risk analysis methods may be qualitative, semi-quantitative, or quantitative; adopt probabilistic and statistical theories; and implement concepts from core disciplines including operations research, reliability engineering, systems engineering, and applied mathematics. These methods continue to develop and evolve and have successfully been applied to address various homeland security mission challenges in recent years. The objective of this book is to: 1) highlight the role of risk analysis for informing homeland security policy decisions, and 2) describe case studies from academia, government, and industry that apply risk analysis methods for addressing challenges within each of the DHS missions.

national security, risk assessment, risk managemen↗

Concentration fluctuations and flammability of cryo-compressed hydrogen and methane jets

Compressed hydrogen stored at cryogenic temperatures has a much higher density than room-temperature storage, which enables large-scale hydrogen storage and transport. An understanding of the release of cryogenic hydrogen from pressurized vessels is needed to evaluate the risk and safety concerns with the use of this fuel. Here, the present work extends the analysis of previous experimental studies that measured the gas concentrations of cryo-compressed hydrogen jets and methane jets using a laser Raman scattering diagnostic system. Since the Raman signals are very small, a denoising algorithm was applied to significantly reduce the noise to enable statistical analysis of the data. The transient features of the turbulent jets were characterized by their concentration intermittencies and probability density functions (PDFs). A two-part PDF was developed to predict the bimodal features of the jet concentration distributions. Then, the flammability factors of the cryogenic jets were calculated based on the intermittency and the PDF.

33 ADVANCED PROPULSION SYSTEMS↗

Hyperspectral Detection of the Fluorescence Shift between Chirality-Sorted Empty and Water-Filled Single-Wall Carbon Nanotube Enantiomers

Single-wall carbon nanotubes (SWCNTs) have extraordinary electronic and optical properties that depend strongly on their exact chiral structure and their interaction with their inner and outer environment. The fluorescence (PL) of semiconducting SWCNTs, for instance, will shift depending on the molecules with which the SWCNT’s hollow core is filled. These interaction-induced shifts are challenging to resolve on the ensemble level in samples containing a mixture of different filling contents due to the relatively large inhomogeneous line width of the ensemble SWCNT PL compared to the size of these shifts. To circumvent this inhomogeneous broadening, single-tube spectroscopy and hyperspectral imaging are often applied, which until now required time-consuming statistical studies. Here, we present hyperspectral PL microscopy combined with automated SWCNT segmenting based on either principal component analysis or a convolutional neural network, capable of both spatially and spectrally resolving the PL along the length of many individual SWCNTs at the same time and automatically fitting peak positions and line widths of individual SWCNTs. The methodology is demonstrated by accurately determining the emission shifts and line widths of thousands of left- and right-handed empty and water-filled SWCNTs coated with a chiral surfactant, resulting in four statistical distributions which cannot be resolved in ensemble spectroscopy of unsorted samples. The results demonstrate a robust method to quickly probe ensemble properties with single-enantiomer spectral resolution. Moreover, it promises to be an absolute quantitative method to characterize the relative abundances of SWCNTs with different handedness or filling content in macroscopic samples, simply by counting individual species.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗