Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Bayesian characterization of uncertainties surrounding fluvial flood hazard estimates

Fluvial floods drive severe risk to riverine communities. There is strong evidence of increasing flood hazards in many regions around the world. The choice of methods and assumptions used in flood hazard estimates can impact the design of risk management strategies. In this study, we characterize the expected flood hazards conditioned on the uncertain model structures, model parameters, and prior distributions of the parameters. We construct a Bayesian framework for river stage return level estimation using a nonstationary statistical model that relies exclusively on the Indian Ocean Dipole Index. We show that ignoring uncertainties can lead to biased estimation of expected flood hazards. We find that the considered model parametric uncertainty is more influential than model structures and model priors. Our results highlight the importance of incorporating uncertainty in extreme flood stage estimates, and are of practical use for informing water infrastructure designs in a changing climate.

54 ENVIRONMENTAL SCIENCES↗

Neutron Induced Nuclear Reaction Cross Sections for Radiochemistry in the Region of Thallium, Lead, and Bismuth

We have developed a set of modeled nuclear reaction cross sections for use in radiochemical detector diagnostics. Systematics for the input parameters required by the Hauser-Feshbach statistical model developed in the TALYS code system are used to calculate neutron induced nuclear reaction cross sections for targets ranging from Thallium (Z = 81) to Bismuth (Z = 83).

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗

Understanding and Modeling Pooled Rideshare Acceptance: Influential Factors, Preferred User Experiences, and Implications

Ridesharing allows people to share a vehicle with others traveling in the same direction, which can reduce costs and traffic congestion. Pooled rideshare (PR) services, such as UberX Share and Lyft Shared, offer an economical and environmentally friendly alternative by matching passengers traveling similar routes. However, despite these benefits, PR adoption remains low due to concerns about safety, privacy, and convenience. This research explores the factors influencing PR adoption and provides recommendations to improve user acceptance. A nationwide survey of 5,385 participants across the U.S. was conducted to understand why people choose or avoid PR. The study identified five key factors influencing PR consideration: safety, service experience, privacy, traffic/environment, and time/cost. Additional research examined ways to optimize PR experiences by identifying four critical factors: comfort/ease of use, convenience, vehicle technology/accessibility, and passenger safety. To measure the impact of these factors, a statistical model called the Pooled Rideshare Acceptance Model (PRAM) was developed, providing insights into how each element influences PR adoption. Further analysis using the Pooled Rideshare Acceptance Model Multigroup Analyses (PRAMMA) revealed how demographic characteristics such as age, gender, income, and past rideshare experience shape PR perceptions. Some key findings from the multigroup analyses showed that younger users valued technological features and environmental benefits, while older users prioritized reliability and service transparency. Additionally, privacy concerns were more significant for female users, while convenience was critical for higher-income groups. These results emphasize that a 'onesize-fits-all' approach to PR service design is not effective, highlighting the need for tailored strategies to address different user segments. Further, workshops were conducted with researchers and students to translate the findings into real-world solutions. These workshops and 3 all the statistical analyses led to the development of 95 actionable recommendations. The recommendations focus on key areas such as safety, service reliability, user education, and accessibility, offering tangible improvements to PR services. The insights from this study provide valuable guidance for policymakers, transportation network companies (TNCs), and researchers aiming to make PR services safer, more accessible, and widely accepted. By addressing user concerns, PR can become a more viable transportation option, supporting sustainable urban mobility and reducing reliance on private vehicles. Additionally, these findings emphasize the importance of user-centric service design in encouraging broader PR adoption. Future research should explore evolving trends in PR preferences, technological advancements, and policy changes to ensure continued improvements. By implementing these recommendations, PR services can better align with user expectations, enhance trust in shared mobility, and contribute to a more efficient transportation ecosystem.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Large-scale tearing-mode hazard function analysis with standard matched equilibrium reconstructions

The association between features from standard tokamak equilibrium reconstructions and the onset of n = 1 tearing modes (TMs) is analyzed at scale. The TM onset rate is directly modeled with a ‘hazard’ function which gives the expected number of onsets (per unit time spent) in a given equilibrium parameter region. In particular the different statistical modeling performance achieved for magnetics-only reconstructions and motional Stark effect (MSE) enhanced reconstructions is studied. It is observed that a better hazard model for the TM onset rate can be built with the MSE-enhanced equilibria compared to the matched magnetics-only situation. This advantage disappears if internal profile details are withheld from the matched analysis. Plausibility of the hazard function is further demonstrated with visualizations of global trends in the operational space, and time-traces from specific tokamak discharges. As a result, TMs typically degrade tokamak plasma performance and may lead to plasma termination, motivating this statistical study.

equilibrium↗

Machine Learning for Daily Forecasts of Arctic Sea Ice Motion: An Attribution Assessment of Model Predictive Skill

Physics-based simulations of Arctic sea ice are highly complex, involving transport between different phases, length scales, and time scales. Resultantly, numerical simulations of sea ice dynamics have a high computational cost and model uncertainty. We employ data-driven machine learning (ML) to make predictions of sea ice motion. The ML models are built to predict present-day sea ice velocity given present-day wind velocity and previous-day sea ice concentration and velocity. Models are trained using reanalysis winds and satellite-derived sea ice properties. We compare the predictions of three different models: persistence (PS), linear regression (LR), and a convolutional neural network (CNN). We quantify the spatiotemporal variability of the correlation between observations and the statistical model predictions. Additionally, we analyze model performance in comparison to variability in properties related to ice motion (wind velocity, ice velocity, ice concentration, distance from coast, bathymetric depth) to understand the processes related to decreases in model performance. Results indicate that a CNN makes skillful predictions of daily sea ice velocity with a correlation up to 0.81 between predicted and observed sea ice velocity, while the LR and PS implementations exhibit correlations of 0.78 and 0.69, respectively. The correlation varies spatially and seasonally: lower values occur in shallow coastal regions and during times of minimum sea ice extent. LR parameter analysis indicates that wind velocity plays the largest role in predicting sea ice velocity on 1-day time scales, particularly in the central Arctic. Regions where wind velocity has the largest LR parameter are regions where the CNN has higher predictive skill than the LR.

54 ENVIRONMENTAL SCIENCES↗

Advanced Error Modeling for Prompt Yield Estimation [Slides]

Multi-Phenomenology Explosion Monitoring (MultiPEM) Toolbox provides a capability for post-detonation characterization focusing on yield estimation; Synthesize information from multiple sensor types; Leverage historical databases to inform statistical models; Allows for rapid or comprehensive analysis; Implemented in R code.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

EFFORT (EFFectiveness Of Rate sTructure for enabling demand response)

This tool helps utilities evaluate potential time of use price-plans and find effective hours for price plan on peak periods and the potential price-elastic demand response. An optimization model can be used to find tiered prices for the time of use tariffs and analyze the impact on load profiles, energy consumption, utility, and customer savings. The tool is equipped with an optimization model to determine the energy consumption of electric appliances that can leverage survey data containing information on the number of appliances owned by customers. Lastly, a statistical model can be used to predict system-level load based on exogenous parameters such as weather, time of day, month, type of day, and year. The tool is scripted in Python and uses the Pyomo python optimization language.

Duwadi, Kapil↗

A semiparametric latent factor model for large scale temporal data with heteroscedasticity

Large scale temporal data have flourished in a vast array of applications, and their sophisticated structures, especially the heteroscedasticity among subjects with inter- and intra-temporal dependence, have fueled a great demand for new statistical models. In this paper, with covariate information, we consider a flexible model for large scale temporal data with subject-specific heteroscedasticity. Formally, the model employs latent semiparametric factors to simultaneously account for the subject-specific heteroscedasticity and the contemporaneous and/or serial correlations. The subject-specific heteroscedasticity is modeled as the product of the unobserved factor process and subject’s covariate effect, which is further characterized via additive models. For estimation, we propose a two-step procedure. First, the latent factor process and nonparametric loading are recovered through projection-based methods, and following, we estimate the regression components by approaches motivated from the generalized least squares. By scrupulously examining the non-asymptotic rates for recovering the factor process and its loading, we show the consistency and efficiency of estimated regression coefficients in the absence of prior knowledge of latent factor process and subject’s covariate effect. Here, the statistical guarantees remain valid even for finite time points that makes our method particularly appealing when the subjects significantly outnumber the observation time points. Using comprehensive simulations, we demonstrate the finite sample performance of our method, which corroborates the theoretical findings. Finally, we apply our method to a data set of air quality and energy consumption collected at 129 monitoring sites in the United States in 2015.

97 MATHEMATICS AND COMPUTING↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

Structure and Reactivity of Binuclear Cu Active Sites in Cu-CHA Zeolites for Stoichiometric Partial Methane Oxidation to Methanol

Aluminosilicate zeolites exchanged with copper ions facilitate partial methane oxidation (PMO) to methanol in stoichiometric oxidation and reduction cycles, yet the identities of active Cu sites and details of the reaction mechanism remain debated. Here, we use the high-symmetry chabazite (CHA) zeolite framework as a model support to probe the relationship between bulk composition, Cu speciation, and response to various oxidizing and reducing treatments. Density functional theory and first-principles thermodynamics combined with statistical models reveal that Cu speciation and composition depend strongly on Al configuration and external gas conditions. Cu-CHA samples were synthesized to survey broad regions of Si/Al and Cu/Al composition space and framework Al proximity. Characterization by in situ X-ray absorption and UV–visible spectroscopy during exposure to different oxidation conditions reveal that the extent of Cu oxidation is sensitive to activation conditions and thus that both kinetic and thermodynamic factors influence Cu oxidizability in a given material. Similar characterizations during CO reduction reveal that CO titrates Cu 2+ in amounts suggesting the presence of both O- and O 2 -bridged species. In contrast, CH 4 and autoreduction (He) treatments reduce similar but smaller numbers of Cu sites than CO, implicating O 2 -bridged Cu dimers as a potential common intermediate in the former reduction pathways. Furthermore, a systematic increase in methanol yields (per Cu) in stoichiometric PMO cycles increase with the fraction of binuclear O x -bridged Cu sites suggests these species as active sites, as depicted in an updated PMO reaction mechanism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advanced data analysis in inertial confinement fusion and high energy density physics

Bayesian analysis enables flexible and rigorous definition of statistical model assumptions with well-characterized propagation of uncertainties and resulting inferences for single-shot, repeated, or even cross-platform data. This approach has a strong history of application to a variety of problems in physical sciences ranging from inference of particle mass from multi-source high-energy particle data to analysis of black-hole characteristics from gravitational wave observations. The recent adoption of Bayesian statistics for analysis and design of high-energy density physics (HEDP) and inertial confinement fusion (ICF) experiments has provided invaluable gains in expert understanding and experiment performance. In this Review, we discuss the basic theory and practical application of the Bayesian statistics framework. We highlight a variety of studies from the HEDP and ICF literature, demonstrating the power of this technique. Due to the computational complexity of multi-physics models needed to analyze HEDP and ICF experiments, Bayesian inference is often not computationally tractable. Two sections are devoted to a review of statistical approximations, efficient inference algorithms, and data-driven methods, such as deep-learning and dimensionality reduction, which play a significant role in enabling use of the Bayesian framework. We provide additional discussion of various applications of Bayesian and machine learning methods that appear to be sparse in the HEDP and ICF literature constituting possible next steps for the community. We conclude by highlighting community needs, the resolution of which will improve trust in data-driven methods that have proven critical for accelerating the design and discovery cycle in many application areas.

47 OTHER INSTRUMENTATION↗

HYBRD (High Resolution HYBrid Regional Downscaling) Model: Input data and Code

The HYBRD (HYBrid Regional Downscaling) model is a high-resolution urban land downscaling model that can be used to downscale intermediate urban land use and land cover (LULC) products into a high-resolution (30-meters). HYBRD uses a sequential hybrid process, combining statistical models with cellular-automata-based spatial algorithms. This repository contains all the necessary model code and inputs needed to successfully run HYBRD for Los Angeles, California. The repo also contains example outputs of each model step, except the final simulated raster outputs. Examples of simulated raster outputs for multiple scenarios for Los Angeles are available at DOI: 10.57931/2575233. Please refer to Related Works below.

Land↗

A proposed framework for the development and qualitative evaluation of West Nile virus models and their application to local public health decision-making

West Nile virus (WNV) is a globally distributed mosquito-borne virus of great public health concern. The number of WNV human cases and mosquito infection patterns vary in space and time. Many statistical models have been developed to understand and predict WNV geographic and temporal dynamics. However, these modeling efforts have been disjointed with little model comparison and inconsistent validation. In this paper, we describe a framework to unify and standardize WNV modeling efforts nationwide. WNV risk, detection, or warning models for this review were solicited from active research groups working in different regions of the United States. A total of 13 models were selected and described. The spatial and temporal scales of each model were compared to guide the timing and the locations for mosquito and virus surveillance, to support mosquito vector control decisions, and to assist in conducting public health outreach campaigns at multiple scales of decision-making. Our overarching goal is to bridge the existing gap between model development, which is usually conducted as an academic exercise, and practical model applications, which occur at state, tribal, local, or territorial public health and mosquito control agency levels. The proposed model assessment and comparison framework helps clarify the value of individual models for decision-making and identifies the appropriate temporal and spatial scope of each model. This qualitative evaluation clearly identifies gaps in linking models to applied decisions and sets the stage for a quantitative comparison of models. Specifically, whereas many coarse-grained models (county resolution or greater) have been developed, the greatest need is for fine-grained, short-term planning models (m–km, days–weeks) that remain scarce. We further recommend quantifying the value of information for each decision to identify decisions that would benefit most from model input.

60 APPLIED LIFE SCIENCES↗

Short-lead seasonal precipitation forecast in northeastern Brazil using an ensemble of artificial neural networks

This study assesses the deterministic and probabilistic forecasting skill of a 1-month-lead ensemble of Artificial Neural Networks (EANN) based on low-frequency climate oscillation indices. The predictand is the February-April (FMA) rainfall in the Brazilian state of Ceará, which is a prominent subject in climate forecasting studies due to its high seasonal predictability. Additionally, the study proposes combining the EANN with dynamical models into a hybrid multi-model ensemble (MME). The forecast verification is carried out through a leave-one-out cross-validation based on 40 years of data. The EANN forecasting skill is compared with traditional statistical models and the dynamical models that compose Ceará’s operational seasonal forecasting system. A spatial comparison showed that the EANN was among the models with the smallest Root Mean Squared Error (RMSE) and Ranked Probability Score (RPS) in most regions. Moreover, the analysis of the area-aggregated reliability showed that the EANN is better calibrated than the individual dynamical models and has better resolution than Multinomial Logistic Regression for above-normal (AN) and below-normal (BN) categories. It is also shown that combining the EANN and dynamical models into a hybrid MME reduces the overconfidence of the extreme categories observed in a dynamically-based MME, improving the reliability of the forecasting system.

54 ENVIRONMENTAL SCIENCES↗

sparse_bias

This is a python package used to fit an unknown function from data that potential contains systematic biases related to metadata. The model fits the unknown function and uses a Bayesian horseshoe prior model to impose sparsity on the bias terms. This code has been generalized from research code developed for AIACHNE into a package that should have more general application in a wider class of statistical models.

Walton, Noah↗

Predictive performance of multi-model ensemble forecasts of COVID-19 across European nations

Background: Short-term forecasts of infectious disease burden can contribute to situational awareness and aid capacity planning. Based on best practice in other fields and recent insights in infectious disease epidemiology, one can maximise the predictive performance of such forecasts if multiple models are combined into an ensemble. Here, we report on the performance of ensembles in predicting COVID-19 cases and deaths across Europe between 08 March 2021 and 07 March 2022. Methods: We used open-source tools to develop a public European COVID-19 Forecast Hub. We invited groups globally to contribute weekly forecasts for COVID-19 cases and deaths reported by a standardised source for 32 countries over the next 1–4 weeks. Teams submitted forecasts from March 2021 using standardised quantiles of the predictive distribution. Each week we created an ensemble forecast, where each predictive quantile was calculated as the equally-weighted average (initially the mean and then from 26th July the median) of all individual models’ predictive quantiles. We measured the performance of each model using the relative Weighted Interval Score (WIS), comparing models’ forecast accuracy relative to all other models. We retrospectively explored alternative methods for ensemble forecasts, including weighted averages based on models’ past predictive performance. Results: Over 52 weeks, we collected forecasts from 48 unique models. We evaluated 29 models’ forecast scores in comparison to the ensemble model. We found a weekly ensemble had a consistently strong performance across countries over time. Across all horizons and locations, the ensemble performed better on relative WIS than 83% of participating models’ forecasts of incident cases (with a total N=886 predictions from 23 unique models), and 91% of participating models’ forecasts of deaths (N=763 predictions from 20 models). Across a 1–4 week time horizon, ensemble performance declined with longer forecast periods when forecasting cases, but remained stable over 4 weeks for incident death forecasts. In every forecast across 32 countries, the ensemble outperformed most contributing models when forecasting either cases or deaths, frequently outperforming all of its individual component models. Among several choices of ensemble methods we found that the most influential and best choice was to use a median average of models instead of using the mean, regardless of methods of weighting component forecast models. Conclusions: Our results support the use of combining forecasts from individual models into an ensemble in order to improve predictive performance across epidemiological targets and populations during infectious disease epidemics. Our findings further suggest that median ensemble methods yield better predictive performance more than ones based on means. Our findings also highlight that forecast consumers should place more weight on incident death forecasts than incident case forecasts at forecast horizons greater than 2 weeks. Funding: AA, BH, BL, LWa, MMa, PP, SV funded by National Institutes of Health (NIH) Grant 1R01GM109718, NSF BIG DATA Grant IIS-1633028, NSF Grant No.: OAC-1916805, NSF Expeditions in Computing Grant CCF-1918656, CCF-1917819, NSF RAPID CNS-2028004, NSF RAPID OAC-2027541, US Centers for Disease Control and Prevention 75D30119C05935, a grant from Google, University of Virginia Strategic Investment Fund award number SIF160, Defense Threat Reduction Agency (DTRA) under Contract No. HDTRA1-19-D-0007, and respectively Virginia Dept of Health Grant VDH-21-501-0141, VDH-21-501-0143, VDH-21-501-0147, VDH-21-501-0145, VDH-21-501-0146, VDH-21-501-0142, VDH-21-501-0148. AF, AMa, GL funded by SMIGE - Modelli statistici inferenziali per governare l'epidemia, FISR 2020-Covid-19 I Fase, FISR2020IP-00156, Codice Progetto: PRJ-0695. AM, BK, FD, FR, JK, JN, JZ, KN, MG, MR, MS, RB funded by Ministry of Science and Higher Education of Poland with grant 28/WFSN/2021 to the University of Warsaw. BRe, CPe, JLAz funded by Ministerio de Sanidad/ISCIII. BT, PG funded by PERISCOPE European H2020 project, contract number 101016233. CP, DL, EA, MC, SA funded by European Commission - Directorate-General for Communications Networks, Content and Technology through the contract LC-01485746, and Ministerio de Ciencia, Innovacion y Universidades and FEDER, with the project PGC2018-095456-B-I00. DE., MGu funded by Spanish Ministry of Health / REACT-UE (FEDER). DO, GF, IMi, LC funded by Laboratory Directed Research and Development program of Los Alamos National Laboratory (LANL) under project number 20200700ER. DS, ELR, GG, NGR, NW, YW funded by National Institutes of General Medical Sciences (R35GM119582; the content is solely the responsibility of the authors and does not necessarily represent the official views of NIGMS or the National Institutes of Health). FB, FP funded by InPresa, Lombardy Region, Italy. HG, KS funded by European Centre for Disease Prevention and Control. IV funded by Agencia de Qualitat i Avaluacio Sanitaries de Catalunya (AQuAS) through contract 2021-021OE. JDe, SMo, VP funded by Netzwerk Universitatsmedizin (NUM) project egePan (01KX2021). JPB, SH, TH funded by Federal Ministry of Education and Research (BMBF; grant 05M18SIA). KH, MSc, YKh funded by Project SaxoCOV, funded by the German Free State of Saxony. Presentation of data, model results and simulations also funded by the NFDI4Health Task Force COVID-19 ( https://www.nfdi4health.de/task-force-covid-19-2 ) within the framework of a DFG-project (LO-342/17-1). LP, VE funded by Mathematical and Statistical modelling project (MUNI/A/1615/2020), Online platform for real-time monitoring, analysis and management of epidemic situations (MUNI/11/02202001/2020); VE also supported by RECETOX research infrastructure (Ministry of Education, Youth and Sports of the Czech Republic: LM2018121), the CETOCOEN EXCELLENCE (CZ.02.1.01/0.0/0.0/17-043/0009632), RECETOX RI project (CZ.02.1.01/0.0/0.0/16-013/0001761). NIB funded by Health Protection Research Unit (grant code NIHR200908). SAb, SF funded by Wellcome Trust (210758/Z/18/Z).

60 APPLIED LIFE SCIENCES↗

Denitrification and the challenge of scaling microsite knowledge to the globe

Here, our knowledge of microbial processes—who is responsible for what, the rates at which they occur, and the substrates consumed and products produced—is imperfect for many if not most taxa, but even less is known about how microsite processes scale to the ecosystem and thence the globe. In both natural and managed environments, scaling links fundamental knowledge to application and also allows for global assessments of the importance of microbial processes. But rarely is scaling straightforward: More often than not, process rates in situ are distributed in a highly skewed fashion, under the influence of multiple interacting controls, and thus often difficult to sample, quantify, and predict. To date, quantitative models of many important processes fail to capture daily, seasonal, and annual fluxes with the precision needed to effect meaningful management outcomes. Nitrogen cycle processes are a case in point, and denitrification is a prime example. Statistical models based on machine learning can improve predictability and identify the best environmental predictors but are—by themselves—insufficient for revealing process-level knowledge gaps or predicting outcomes under novel environmental conditions. Hybrid models that incorporate well-calibrated process models as predictors for machine learning algorithms can provide both improved understanding and more reliable forecasts under environmental conditions not yet experienced. Incorporating trait-based models into such efforts promises to improve predictions and understanding still further, but much more development is needed.

59 BASIC BIOLOGICAL SCIENCES↗