Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Advanced data analysis in inertial confinement fusion and high energy density physics

Bayesian analysis enables flexible and rigorous definition of statistical model assumptions with well-characterized propagation of uncertainties and resulting inferences for single-shot, repeated, or even cross-platform data. This approach has a strong history of application to a variety of problems in physical sciences ranging from inference of particle mass from multi-source high-energy particle data to analysis of black-hole characteristics from gravitational wave observations. The recent adoption of Bayesian statistics for analysis and design of high-energy density physics (HEDP) and inertial confinement fusion (ICF) experiments has provided invaluable gains in expert understanding and experiment performance. In this Review, we discuss the basic theory and practical application of the Bayesian statistics framework. We highlight a variety of studies from the HEDP and ICF literature, demonstrating the power of this technique. Due to the computational complexity of multi-physics models needed to analyze HEDP and ICF experiments, Bayesian inference is often not computationally tractable. Two sections are devoted to a review of statistical approximations, efficient inference algorithms, and data-driven methods, such as deep-learning and dimensionality reduction, which play a significant role in enabling use of the Bayesian framework. We provide additional discussion of various applications of Bayesian and machine learning methods that appear to be sparse in the HEDP and ICF literature constituting possible next steps for the community. We conclude by highlighting community needs, the resolution of which will improve trust in data-driven methods that have proven critical for accelerating the design and discovery cycle in many application areas.

47 OTHER INSTRUMENTATION↗

HYBRD (High Resolution HYBrid Regional Downscaling) Model: Input data and Code

The HYBRD (HYBrid Regional Downscaling) model is a high-resolution urban land downscaling model that can be used to downscale intermediate urban land use and land cover (LULC) products into a high-resolution (30-meters). HYBRD uses a sequential hybrid process, combining statistical models with cellular-automata-based spatial algorithms. This repository contains all the necessary model code and inputs needed to successfully run HYBRD for Los Angeles, California. The repo also contains example outputs of each model step, except the final simulated raster outputs. Examples of simulated raster outputs for multiple scenarios for Los Angeles are available at DOI: 10.57931/2575233. Please refer to Related Works below.

Land↗

A proposed framework for the development and qualitative evaluation of West Nile virus models and their application to local public health decision-making

West Nile virus (WNV) is a globally distributed mosquito-borne virus of great public health concern. The number of WNV human cases and mosquito infection patterns vary in space and time. Many statistical models have been developed to understand and predict WNV geographic and temporal dynamics. However, these modeling efforts have been disjointed with little model comparison and inconsistent validation. In this paper, we describe a framework to unify and standardize WNV modeling efforts nationwide. WNV risk, detection, or warning models for this review were solicited from active research groups working in different regions of the United States. A total of 13 models were selected and described. The spatial and temporal scales of each model were compared to guide the timing and the locations for mosquito and virus surveillance, to support mosquito vector control decisions, and to assist in conducting public health outreach campaigns at multiple scales of decision-making. Our overarching goal is to bridge the existing gap between model development, which is usually conducted as an academic exercise, and practical model applications, which occur at state, tribal, local, or territorial public health and mosquito control agency levels. The proposed model assessment and comparison framework helps clarify the value of individual models for decision-making and identifies the appropriate temporal and spatial scope of each model. This qualitative evaluation clearly identifies gaps in linking models to applied decisions and sets the stage for a quantitative comparison of models. Specifically, whereas many coarse-grained models (county resolution or greater) have been developed, the greatest need is for fine-grained, short-term planning models (m–km, days–weeks) that remain scarce. We further recommend quantifying the value of information for each decision to identify decisions that would benefit most from model input.

60 APPLIED LIFE SCIENCES↗

Short-lead seasonal precipitation forecast in northeastern Brazil using an ensemble of artificial neural networks

This study assesses the deterministic and probabilistic forecasting skill of a 1-month-lead ensemble of Artificial Neural Networks (EANN) based on low-frequency climate oscillation indices. The predictand is the February-April (FMA) rainfall in the Brazilian state of Ceará, which is a prominent subject in climate forecasting studies due to its high seasonal predictability. Additionally, the study proposes combining the EANN with dynamical models into a hybrid multi-model ensemble (MME). The forecast verification is carried out through a leave-one-out cross-validation based on 40 years of data. The EANN forecasting skill is compared with traditional statistical models and the dynamical models that compose Ceará’s operational seasonal forecasting system. A spatial comparison showed that the EANN was among the models with the smallest Root Mean Squared Error (RMSE) and Ranked Probability Score (RPS) in most regions. Moreover, the analysis of the area-aggregated reliability showed that the EANN is better calibrated than the individual dynamical models and has better resolution than Multinomial Logistic Regression for above-normal (AN) and below-normal (BN) categories. It is also shown that combining the EANN and dynamical models into a hybrid MME reduces the overconfidence of the extreme categories observed in a dynamically-based MME, improving the reliability of the forecasting system.

54 ENVIRONMENTAL SCIENCES↗

sparse_bias

This is a python package used to fit an unknown function from data that potential contains systematic biases related to metadata. The model fits the unknown function and uses a Bayesian horseshoe prior model to impose sparsity on the bias terms. This code has been generalized from research code developed for AIACHNE into a package that should have more general application in a wider class of statistical models.

Walton, Noah↗

Predictive performance of multi-model ensemble forecasts of COVID-19 across European nations

Background: Short-term forecasts of infectious disease burden can contribute to situational awareness and aid capacity planning. Based on best practice in other fields and recent insights in infectious disease epidemiology, one can maximise the predictive performance of such forecasts if multiple models are combined into an ensemble. Here, we report on the performance of ensembles in predicting COVID-19 cases and deaths across Europe between 08 March 2021 and 07 March 2022. Methods: We used open-source tools to develop a public European COVID-19 Forecast Hub. We invited groups globally to contribute weekly forecasts for COVID-19 cases and deaths reported by a standardised source for 32 countries over the next 1–4 weeks. Teams submitted forecasts from March 2021 using standardised quantiles of the predictive distribution. Each week we created an ensemble forecast, where each predictive quantile was calculated as the equally-weighted average (initially the mean and then from 26th July the median) of all individual models’ predictive quantiles. We measured the performance of each model using the relative Weighted Interval Score (WIS), comparing models’ forecast accuracy relative to all other models. We retrospectively explored alternative methods for ensemble forecasts, including weighted averages based on models’ past predictive performance. Results: Over 52 weeks, we collected forecasts from 48 unique models. We evaluated 29 models’ forecast scores in comparison to the ensemble model. We found a weekly ensemble had a consistently strong performance across countries over time. Across all horizons and locations, the ensemble performed better on relative WIS than 83% of participating models’ forecasts of incident cases (with a total N=886 predictions from 23 unique models), and 91% of participating models’ forecasts of deaths (N=763 predictions from 20 models). Across a 1–4 week time horizon, ensemble performance declined with longer forecast periods when forecasting cases, but remained stable over 4 weeks for incident death forecasts. In every forecast across 32 countries, the ensemble outperformed most contributing models when forecasting either cases or deaths, frequently outperforming all of its individual component models. Among several choices of ensemble methods we found that the most influential and best choice was to use a median average of models instead of using the mean, regardless of methods of weighting component forecast models. Conclusions: Our results support the use of combining forecasts from individual models into an ensemble in order to improve predictive performance across epidemiological targets and populations during infectious disease epidemics. Our findings further suggest that median ensemble methods yield better predictive performance more than ones based on means. Our findings also highlight that forecast consumers should place more weight on incident death forecasts than incident case forecasts at forecast horizons greater than 2 weeks. Funding: AA, BH, BL, LWa, MMa, PP, SV funded by National Institutes of Health (NIH) Grant 1R01GM109718, NSF BIG DATA Grant IIS-1633028, NSF Grant No.: OAC-1916805, NSF Expeditions in Computing Grant CCF-1918656, CCF-1917819, NSF RAPID CNS-2028004, NSF RAPID OAC-2027541, US Centers for Disease Control and Prevention 75D30119C05935, a grant from Google, University of Virginia Strategic Investment Fund award number SIF160, Defense Threat Reduction Agency (DTRA) under Contract No. HDTRA1-19-D-0007, and respectively Virginia Dept of Health Grant VDH-21-501-0141, VDH-21-501-0143, VDH-21-501-0147, VDH-21-501-0145, VDH-21-501-0146, VDH-21-501-0142, VDH-21-501-0148. AF, AMa, GL funded by SMIGE - Modelli statistici inferenziali per governare l'epidemia, FISR 2020-Covid-19 I Fase, FISR2020IP-00156, Codice Progetto: PRJ-0695. AM, BK, FD, FR, JK, JN, JZ, KN, MG, MR, MS, RB funded by Ministry of Science and Higher Education of Poland with grant 28/WFSN/2021 to the University of Warsaw. BRe, CPe, JLAz funded by Ministerio de Sanidad/ISCIII. BT, PG funded by PERISCOPE European H2020 project, contract number 101016233. CP, DL, EA, MC, SA funded by European Commission - Directorate-General for Communications Networks, Content and Technology through the contract LC-01485746, and Ministerio de Ciencia, Innovacion y Universidades and FEDER, with the project PGC2018-095456-B-I00. DE., MGu funded by Spanish Ministry of Health / REACT-UE (FEDER). DO, GF, IMi, LC funded by Laboratory Directed Research and Development program of Los Alamos National Laboratory (LANL) under project number 20200700ER. DS, ELR, GG, NGR, NW, YW funded by National Institutes of General Medical Sciences (R35GM119582; the content is solely the responsibility of the authors and does not necessarily represent the official views of NIGMS or the National Institutes of Health). FB, FP funded by InPresa, Lombardy Region, Italy. HG, KS funded by European Centre for Disease Prevention and Control. IV funded by Agencia de Qualitat i Avaluacio Sanitaries de Catalunya (AQuAS) through contract 2021-021OE. JDe, SMo, VP funded by Netzwerk Universitatsmedizin (NUM) project egePan (01KX2021). JPB, SH, TH funded by Federal Ministry of Education and Research (BMBF; grant 05M18SIA). KH, MSc, YKh funded by Project SaxoCOV, funded by the German Free State of Saxony. Presentation of data, model results and simulations also funded by the NFDI4Health Task Force COVID-19 ( https://www.nfdi4health.de/task-force-covid-19-2 ) within the framework of a DFG-project (LO-342/17-1). LP, VE funded by Mathematical and Statistical modelling project (MUNI/A/1615/2020), Online platform for real-time monitoring, analysis and management of epidemic situations (MUNI/11/02202001/2020); VE also supported by RECETOX research infrastructure (Ministry of Education, Youth and Sports of the Czech Republic: LM2018121), the CETOCOEN EXCELLENCE (CZ.02.1.01/0.0/0.0/17-043/0009632), RECETOX RI project (CZ.02.1.01/0.0/0.0/16-013/0001761). NIB funded by Health Protection Research Unit (grant code NIHR200908). SAb, SF funded by Wellcome Trust (210758/Z/18/Z).

60 APPLIED LIFE SCIENCES↗

Denitrification and the challenge of scaling microsite knowledge to the globe

Here, our knowledge of microbial processes—who is responsible for what, the rates at which they occur, and the substrates consumed and products produced—is imperfect for many if not most taxa, but even less is known about how microsite processes scale to the ecosystem and thence the globe. In both natural and managed environments, scaling links fundamental knowledge to application and also allows for global assessments of the importance of microbial processes. But rarely is scaling straightforward: More often than not, process rates in situ are distributed in a highly skewed fashion, under the influence of multiple interacting controls, and thus often difficult to sample, quantify, and predict. To date, quantitative models of many important processes fail to capture daily, seasonal, and annual fluxes with the precision needed to effect meaningful management outcomes. Nitrogen cycle processes are a case in point, and denitrification is a prime example. Statistical models based on machine learning can improve predictability and identify the best environmental predictors but are—by themselves—insufficient for revealing process-level knowledge gaps or predicting outcomes under novel environmental conditions. Hybrid models that incorporate well-calibrated process models as predictors for machine learning algorithms can provide both improved understanding and more reliable forecasts under environmental conditions not yet experienced. Incorporating trait-based models into such efforts promises to improve predictions and understanding still further, but much more development is needed.

59 BASIC BIOLOGICAL SCIENCES↗

Kronecker-structured covariance models for multiway data

Many applications produce multiway data of exceedingly high dimension. Modeling such multi-way data is important in multichannel signal and video processing where sensors produce multi-indexed data, e.g. over spatial, frequency, and temporal dimensions. We will address the challenges of covariance representation of multiway data and review some of the progress in statistical modeling of multiway covariance over the past two decades, focusing on tensor-valued covariance models and their inference. We will illustrate through a space weather application: predicting the evolution of solar active regions over time.

97 MATHEMATICS AND COMPUTING↗

Nepheline Crystallization Studies and the Structural Integrity of the Residual (SIR) - FY 2022 Study Glasses

In FY22 a set of 367 HLW glasses with compositions representative of the Waste Treatment and Immobilization Plant (WTP) processing region were aggregated and analyzed with the Structural Integrity of the Residual (SIR) model developed by the Savannah River National Laboratory (SRNL). An additional 12 glasses targeting 27 – 37 wt % alumina (in glass) and at waste loadings up to ~70 wt % were fabricated and added to the database. Collectively, this data set was analyzed with regression statistics and partitioning functions to reveal correlations between boron release and the chemical composition. FY22 results indicate that the SIR theoretical alkali-silicate precipitation term, Al 2 O 3 and B 2 O 3 concentration, and SIR computed non-bridging oxygen are significant predictors of chemical durability. A statistical model using the SIR parameters was shown to have greater overall accuracy than the nepheline discriminator model while maintaining false predictions of pour durability below 1%. In FY23, we will continue this work to focus on refining the SIR with glass compositions that challenge the SIR correlation terms, namely the theoretical crystalline precipitation and non-bridging oxygen terms.

36 MATERIALS SCIENCE↗

Sheaves as a Framework for Understanding and Interpreting Model Fit

As data grows in size and complexity, finding frameworks which aid in interpretation and analysis has become critical. This is particularly true when data comes from complex systems where extensive structure is available, but must be drawn from peripheral sources. In this paper we argue that in such situations, sheaves can provide a natural framework to analyze how well a statistical model fits at the local level (that is, on subsets of related datapoints) vs the global level (on all the data). The sheaf-based approach that we propose is suitably general enough to be useful in a range of applications, from analyzing sensor networks to understanding the feature space of a deep learning model.

Kvinge, Henry J.↗

In Situ Inference for Earth System Predictability

An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

Accurate prediction of oxygen vacancy concentration with disordered A-site cations in high-entropy perovskite oxides

Abstract Entropic stabilized ABO 3 perovskite oxides promise many applications, including the two-step solar thermochemical hydrogen (STCH) production. Using binary and quaternary A-site mixed {A}FeO 3 as a model system, we reveal that as more cation types, especially above four, are mixed on the A-site, the cell lattice becomes more cubic-like but the local Fe–O octahedrons are more distorted. By comparing four different Density Functional Theory-informed statistical models with experiments, we show that the oxygen vacancy formation energies ( $${E}_{V}^{f}$$ E V f ) distribution and the vacancy interactions must be considered to predict the oxygen non-stoichiometry ( δ ) accurately. For STCH applications, the $${E}_{V}^{f}$$ E V f distribution, including both the average and the spread, can be optimized jointly to improve Δ δ (difference of δ between the two-step conditions) in some hydrogen production levels. This model can be used to predict the range of water splitting that can be thermodynamically improved by mixing cations in {A}FeO 3 perovskites.

08 HYDROGEN↗

Evaluation of obstacle modelling approaches for resource assessment and small wind turbine siting: case study in the northern Netherlands

Abstract. Growth in adoption of distributed wind turbines for energy generation is significantly impacted by challenges associated with siting and accurate estimation of the wind resource. Small turbines, at hub heights of 40 m or less, are greatly impacted by terrestrial obstacles such as built structures and vegetation that can cause complex wake effects. While some progress in high-fidelity complex fluid dynamics (CFD) models has increased the potential accuracy for modelling the impacts of obstacles on turbulent wind flow, these models are too computationally expensive for practical siting and resource assessment applications. To understand the efficacy of available models in situ, this study evaluates classic and commonly used methods alongside new state-of-the-art lower-order models derived from CFD simulations and machine learning approaches. This evaluation is conducted using a subset of an extensive original dataset of measurements from more than 300 operational wind turbines in the northern Netherlands. The results show that data-driven methods (e.g. machine learning and statistical modelling) are most effective at predicting production at real sites with an average error in annual energy production of 2.5 %. When sufficient data may not be available de novo to support these data-driven approaches, models derived from high-fidelity simulations show promise and reliably outperform classic methods. On average these models have 6.3 %–11.5 % error compared with 26 % for classic methods and 27 % baseline error for reanalysis data without obstacle correction. While more performant on average, these methods are also sensitive to the quality of obstacle descriptions and reanalysis inputs.

17 WIND ENERGY↗

Classification of Photovoltaic Failures with Hidden Markov Modeling, an Unsupervised Statistical Approach

Failure detection methods are of significant interest for photovoltaic (PV) site operators to help reduce gaps between expected and observed energy generation. Current approaches for field-based fault detection, however, rely on multiple data inputs and can suffer from interpretability issues. In contrast, this work offers an unsupervised statistical approach that leverages hidden Markov models (HMM) to identify failures occurring at PV sites. Using performance index data from 104 sites across the United States, individual PV-HMM models are trained and evaluated for failure detection and transition probabilities. This analysis indicates that the trained PV-HMM models have the highest probability of remaining in their current state (87.1% to 93.5%), whereas the transition probability from normal to failure (6.5%) is lower than the transition from failure to normal (12.9%) states. A comparison of these patterns using both threshold levels and operations and maintenance (O&M) tickets indicate high precision rates of PV-HMMs (median = 82.4%) across all of the sites. Although additional work is needed to assess sensitivities, the PV-HMM methodology demonstrates significant potential for real-time failure detection as well as extensions into predictive maintenance capabilities for PV.

classification↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

Optimal experimental design: Formulations and computations

Questions of ‘how best to acquire data’ are essential to modelling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates computational methods to answer them. This article presents a systematic survey of modern OED, from its foundations in classical design theory to current research involving OED for complex models. We begin by reviewing criteria used to formulate an OED problem and thus to encode the goal of performing an experiment. We emphasize the flexibility of the Bayesian and decision-theoretic approach, which encompasses information-based criteria that are well-suited to nonlinear and non-Gaussian statistical models. We then discuss methods for estimating or bounding the values of these design criteria; this endeavour can be quite challenging due to strong nonlinearities, high parameter dimension, large per-sample costs, or settings where the model is implicit. A complementary set of computational issues involves optimization methods used to find a design; we discuss such methods in the discrete (combinatorial) setting of observation selection and in settings where an exact design can be continuously parametrized. Finally we present emerging methods for sequential OED that build non-myopic design policies, rather than explicit designs; these methods naturally adapt to the outcomes of past experiments in proposing new experiments, while seeking coordination among all experiments to be performed. Throughout, we highlight important open questions and challenges.

97 MATHEMATICS AND COMPUTING↗

Quantitative assessment of fitting errors associated with streak camera noise in Thomson scattering data analysis

Thomson scattering measurements in high energy density experiments are often recorded using optical streak cameras. In the low-signal regime, noise introduced by the streak camera can become an important and sometimes the dominant source of measurement uncertainty. In this paper, we present a formal method of accounting for the presence of streak camera noise in our measurements. We present a phenomenological description of the noise generation mechanisms and present a statistical model that may be used to construct the covariance matrix associated with a given measurement. This model is benchmarked against simulations of streak camera images. We demonstrate how this covariance may then be used to weight fitting of the data and provide quantitative assessments of the uncertainty in the fitting parameters determined by the best fit to the data and build confidence in the ability to make statistically significant measurements in the low-signal regime, where spatial correlations in the noise become apparent. These methods will have general applicability to other measurements made using optical streak cameras.

47 OTHER INSTRUMENTATION↗