Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The WRF-Solar Ensemble Prediction System to Provide Solar Irradiance Probabilistic Forecasts

In this study, we introduce the recently developed WRF-solar ensemble prediction system and a calibration method. The performances of forecast models are evaluated using the National Solar Radiation Database observational analysis for day-ahead solar irradiance predictions. The results demonstrate that the ensemble forecast improves the quality of the forecasts by considering the uncertainty of each ensemble member. The analog ensemble calibration contributed to the reduction of positive bias and an overall improvement in the probabilistic attributes, such as reliability and statistical consistency.

14 SOLAR ENERGY↗

A method to fuse multiphysics waveforms and improve predictive explosion detection: theory, experiment and performance

Natural and human-made sources of transient energy often emit multiple geophysical signatures that include mechanical and electromagnetic waveforms. We present a constructive method to fuse and evaluate statistics that we derive from such multiphysics waveforms that improves our capability to detect small, near-ground explosions over similar methods that consume single signature waveforms. Our method advances Fisher's Combined Probability Test (Fisher's Method) to operate under both hypotheses of a binary test on noisy data and provide researchers with the density functions required to forecast the ability of Fisher's Method to screen fused explosion signatures from noise. We apply this method against 12 d, multisignature explosion and noise records to show (1) that a fused multiphysics waveform statistic that combines radio, acoustic and seismic waveform data can identify explosions roughly 0.8 magnitude units lower than an acoustic emission, STA/LTA detector for the same detection probability and (2) that we can quantitatively predict how this fused, multiphysics statistic performs with Fisher's Method. Our work thereby offers a baseline method for predictive waveform fusion that supports multiphenomenological explosion monitoring (multiPEM) and is applicable to any binary testing problem in observational geophysics.

58 GEOSCIENCES↗

Explainable tokamak-agnostic forecasting of fusion plasma instability via megahertz turbulent fluctuations

Scientific applications of artificial intelligence (AI) often remain limited by device-specific training and unexplained “black-box” approaches, creating fundamental barriers to cross-system generalization. This challenge is critical for nuclear fusion, where future reactors will have limited operational data for AI training. Here, we demonstrate that our neural network, trained solely on megahertz-scale turbulence measurements from one machine (DIII-D), forecasts Type-I edge localized mode (ELM) onsets in a different tokamak (KSTAR) through zero-shot weight transfer following physics-consistent preprocessing without device-specific retraining. Through an explainable AI framework combining gradient-weighted class activation mapping with physics validation, we reveal that our network can internalize physics relationships governing the ELM instabilities rather than memorizing device-specific patterns. The network perceives spatiotemporal features that correlate consistently with independently calculated instability growth rates, magnetohydrodynamic stability limits, and pedestal structure dynamics. Statistical analyses of dimensionally-reduced saliency features reveal the identical triangular features between the saliency representations, instability growth rates, and prediction probability across tokamaks, providing evidence that our forecasting system can show tokamak-agnostic generalization. This work contributes to a foundation for explainable scientific AI systems, where cross-system developments are essential for transcending traditional domain-specific constraints.

AI↗

Capturing non-Gaussianity of the large-scale structure with weighted skew-spectra

The forthcoming generation of wide-field galaxy surveys will probe larger volumes and galaxy densities, thus allowing for a much larger signal-to-noise ratio for higher-order clustering statistics, in particular the galaxy bispectrum. Extracting this information, however, is more challenging than using the power spectrum due to more complex theoretical modeling, as well as significant computational cost of evaluating the bispectrum signal and the error budget. To overcome these challenges, several proxy statistics have been proposed in the literature, which partially or fully capture the information in the bispectrum, while being computationally less expensive than the bispectrum. One such statistics are weighted skew-spectra, which are cross-spectra of the density field and appropriately weighted quadratic fields. Using Fisher forecasts, we show that the information in these skew-spectra is equivalent to that in the bispectrum for parameters that appear as amplitudes in the bispectrum model, such as galaxy bias parameters or the amplitude of primordial non-Gaussianity. We consider three shapes of the primordial bispectrum: local, equilateral and that due to massive particles with spin two during inflation. In conclusion, to obtain constraints that match those from a measurement of the full bispectrum, we find that it is crucial to account for the full covariance matrix of the skew-spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

Predicting September Arctic Sea Ice: A Multimodel Seasonal Skill Comparison

This study quantifies the state of the art in the rapidly growing field of seasonal Arctic sea ice prediction. A novel multimodel dataset of retrospective seasonal predictions of September Arctic sea ice is created and analyzed, consisting of community contributions from 17 statistical models and 17 dynamical models. Prediction skill is compared over the period 2001–20 for predictions of pan-Arctic sea ice extent (SIE), regional SIE, and local sea ice concentration (SIC) initialized on 1 June, 1 July, 1 August, and 1 September. This diverse set of statistical and dynamical models can individually predict linearly detrended pan-Arctic SIE anomalies with skill, and a multimodel median prediction has correlation coefficients of 0.79, 0.86, 0.92, and 0.99 at these respective initialization times. Regional SIE predictions have similar skill to pan-Arctic predictions in the Alaskan and Siberian regions, whereas regional skill is lower in the Canadian, Atlantic, and central Arctic sectors. The skill of dynamical and statistical models is generally comparable for pan-Arctic SIE, whereas dynamical models outperform their statistical counterparts for regional and local predictions. The prediction systems are found to provide the most value added relative to basic reference forecasts in the extreme SIE years of 1996, 2007, and 2012. SIE prediction errors do not show clear trends over time, suggesting that there has been minimal change in inherent sea ice predictability over the satellite era. Overall, this study demonstrates that there are bright prospects for skillful operational predictions of September sea ice at least 3 months in advance.

54 ENVIRONMENTAL SCIENCES↗

Assessing the Impact of a Forest Canopy on Near-Surface Wind Statistics

Representing the forest canopy in atmospheric numerical models should improve simulated winds within and above the canopy up to a few hundred meters above the ground. Here, in this study, we implement a forest canopy parameterization into the Weather Research and Forecasting (WRF) Model in a large-eddy simulation (LES) mode by applying drag forces across multiple layers within the canopy height. We use unique observations from the Lidar Experiments for Assessing Flow over Forests (LEAFF) field campaign at the Wind River Experimental Forest (WREF) in the U.S. Pacific Northwest to evaluate model performance. In a 2-day case study, the canopy parameterization improved wind predictions both within and above the canopy, particularly during the daytime and at finer grid resolution. Without it, winds were frequently overpredicted above the canopy. Similarly, derived quantities such as the wind shear index also yielded estimates closer to observations with the canopy parameterization implemented. These findings suggest that representing the canopy using drag forces alone can improve simulated mean winds up to 200 m above the surface. Furthermore, second-order statistical moments of wind were more sensitive to canopy density than first-order moments, especially during the daytime. This increased sensitivity and the improved daytime performance in wind speed—evidenced by the lowest bias from observations (3% compared to 20% over diurnal cycle)—imply that winds above the canopy layer are strongly influenced by how well turbulence above the canopy is modeled. The results of this study can serve as a foundation for parameterizing forest canopy effects in coarser weather forecast models.

Energy - Wind↗

Precision redshift-space galaxy power spectra using Zel'dovich control variates

Numerical simulations in cosmology require trade-offs between volume, resolution and run-time that limit the volume of the Universe that can be simulated, leading to sample variance in predictions of ensemble-average quantities such as the power spectrum or correlation function(s). Sample variance is particularly acute at large scales, which is also where analytic techniques can be highly reliable. This provides an opportunity to combine analytic and numerical techniques in a principled way to improve the dynamic range and reliability of predictions for clustering statistics. In this paper we extend the technique of Zel'dovich control variates, previously demonstrated for 2-point functions in real space, to reduce the sample variance in measurements of 2-point statistics of biased tracers in redshift space. We demonstrate that with this technique, we can reduce the sample variance of these statistics down to their shot-noise limit out to k ~ 0.2 h Mpc -1 . This allows a better matching with perturbative models and improved predictions for the clustering of e.g. quasars, galaxies and neutral Hydrogen measured in spectroscopic redshift surveys at very modest computational expense. We discuss the implementation of ZCV, give some examples and provide forecasts for the efficacy of the method under various conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Tropical Interbasin Interaction as Effective Predictors of Late-Spring Precipitation Variability in the Southern Great Plains

Abstract The southern Great Plains experience fluctuating precipitation extremes that significantly impact agriculture and water management. Despite ongoing efforts to enhance forecast accuracy, the underlying causes of these climatic phenomena remain inadequately understood. This study elucidates the relative influence of the tropical Pacific and Atlantic basins on April–May–June precipitation variability in this region. Our partial ocean assimilation experiments using the Community Earth System Model unveil the prominent role of interbasin interaction, with the Pacific and Atlantic contributing approximately 70% and 30%, respectively, to these interbasin contrasts. Our statistical analyses suggest that these tropical interbasin contrasts could serve as a more reliable indicator for late-spring precipitation anomalies than El Niño–Southern Oscillation. The conclusions are reinforced by analyses of seven climate forecasting systems within the North American Multi-Model Ensemble, offering an optimistic outlook for enhancing real-time forecasting of late-spring precipitation in the southern plains. However, the current predictive skills of the interbasin contrasts across the prediction systems are hindered by the lower predictability of the tropical Atlantic Ocean, pointing to the need for future research to refine climate prediction models further. Significance Statement Agriculture and infrastructure in the southern plains face challenges from severe late-spring precipitation extremes. Traditional predictors like El Niño–Southern Oscillation (ENSO) lose effectiveness during the critical spring-to-summer transition, creating a forecasting gap. This study introduces the concept of tropical interbasin interactions, known to enhance seasonal predictability for late-spring precipitation in the southern plains. Novel climate model experiments highlight contributions from the tropical Pacific and Atlantic, offering a promising predictability that potentially surpasses the limitations of ENSO-based predictions. These outcomes hold the potential for developing operational forecasts of late-spring precipitation anomalies in the southern plains, enabling proactive risk management.

Chikamoto, Yoshimitsu↗

Evaluating downscaled products with expected hydroclimatic co-variances

Abstract. There has been widespread adoption of downscaled products amongst practitioners and stakeholders to ascertain risk from climate hazards at the local scale (e.g., ∼ 5 km resolution). Such products must nevertheless be consistent with physical laws to be credible and of value to users. Here we evaluate statistically and dynamically downscaled products by examining local co-evolution of downscaled temperature and precipitation during convective and frontal precipitation events (two mechanisms testable with just temperature and precipitation). We find that two widely used statistical downscaling techniques (Localized Constructed Analogs version 2, LOCA2, and Seasonal Trends and Analysis of Residuals Empirical Statistical Downscaling Model, STAR-ESDM) generally preserve expected co-variances during convective precipitation events over the historical and future projected intervals as compared to European Centre for Medium-Range Weather Forecasts Reanalysis v5 (ERA5) and two observation-based data products (Livneh and nClimGrid-Daily). However, both techniques dampen future intensification of frontal precipitation that is otherwise robustly captured in global climate models (i.e., prior to downscaling) and with process-based dynamical downscaling across five different regional climate models. In the case of LOCA2, this leads to appreciable underestimation of future frontal precipitation event intensity. This study is one of the first to quantify a likely ramification of the stationarity assumption underlying statistical downscaling methods and identify a phenomenon where projections of future change diverge depending on data production method employed. Finally, our work proposes expected co-variances during convective and frontal precipitation as useful evaluation diagnostics that can be universally applied to a wide range of statistically downscaled products.

54 ENVIRONMENTAL SCIENCES↗

Forecasting generative amplification

Generative networks are perfect tools to enhance the speed and precision of LHC simulations. Especially when generating events beyond the size of the training dataset, it is important to understand their statistical precision. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is already possible in specific regions of phase space.

Bahl, Henning [Heidelberg Univ. (Germany)] (ORCID:↗

Stochastic Analysis for Long Term Capital Structures, Systems, and Components Refurbishment and Replacement

As commercial Nuclear Power Plants (NPPs) pursue extended plant operation in the form of Second License Renewal (SLR), opportunities exist for these plants to provide capital investments to ensure long-term safe and economic performance. At the current time, several utilities have announced an intention to pursue extended operation for one or more of their NPPs via SLR . The goal of this research is to develop a risk-informed approach to evaluate and prioritize plant capital investments made in preparation for, and during the period of, extended plant operations to support decisions for NPP operations. Since the capital investments are influenced by various factors, such as markets, safety and regulatory, the decision-making process of NPP operations should take into account relevant factors for balancing risks, costs and profits. The traditional method of capital budgeting is based on the priority list of candidate projects using economic measures such as benefit-investment ratio, net present value (NPV) and internal rate of return. In the literatures, the problem of capital budgeting or the variant can be represented by an appropriate knapsack problem. The knapsack approach to capital budgeting takes as input as investment, along with the cost and profit of each project. The objective of capital budgeting is to find the combination of the binary decisions for every investment such that the overall profit is as large as possible. The output is a collection of projects to be carried out, and we refer this selected collection of projects as a project portfolio. One limitation of traditional optimization models for capital budgeting is that they do not account for risk/uncertainty in profit and cost streams associated with individual projects, they do not account for risk in resource availability in future years [1,2,3]. Projects can incur cost over-runs, especially when projects are large, performed infrequently, and when there is risk regarding technical viability, external contractors, and/or suppliers of requisite parts and materials. Occasionally, projects are performed ahead of schedule and with cost savings. Planned budgets for capital improvements can be cut and key personnel may be lost. Or, there may be surprise windfalls in budgets for maintenance activities due to decreased costs for “unplanned” maintenance. In these cases, how should we resolve capital budgeting when we have risk forecasts for costs, profits and budgets? One approach we proposed in this summary is to re-solve the optimization models based on assumed statistical distributions of given parameters. If these distributions were not available, a two-stage stochastic optimization approach can be used to provide priority lists to decision-makers to support better risk-informed decisions [4, 5]. In this summary, we will only focus on the first approach.

42 ENGINEERING↗

Integrating Deep Learning and Hydrodynamic Modeling to Improve the Great Lakes Forecast

The Laurentian Great Lakes, one of the world’s largest surface freshwater systems, pose a modeling challenge in seasonal forecast and climate projection. While physics-based hydrodynamic modeling is a fundamental approach, improving the forecast accuracy remains critical. In recent years, machine learning (ML) has quickly emerged in geoscience applications, but its application to the Great Lakes hydrodynamic prediction is still in its early stages. This work is the first one to explore a deep learning approach to predicting spatiotemporal distributions of the lake surface temperature (LST) in the Great Lakes. Our study shows that the Long Short-Term Memory (LSTM) neural network, trained with the limited data from hypothetical monitoring networks, can provide consistent and robust performance. The LSTM prediction captured the LST spatiotemporal variabilities across the five Great Lakes well, suggesting an effective and efficient way for monitoring network design in assisting the ML-based forecast. Furthermore, we employed an explainable artificial intelligence (XAI) technique named SHapley Additive exPlanations (SHAP) to uncover how the features impact the LSTM prediction. Our XAI analysis shows air temperature is the most influential feature for predicting LST in the trained LSTM. The relatively large bias in the LSTM prediction during the spring and fall was associated with substantial heterogeneity of air temperature during the two seasons. In contrast, the physics-based hydrodynamic model performed better in spring and fall yet exhibited relatively large biases during the summer stratification period. Finally, we developed a statistical integration of the hydrodynamic modeling and deep learning results based on the Best Linear Unbiased Estimator (BLUE). The integration further enhanced prediction accuracy, suggesting its potential for next-generation Great Lakes forecast systems.

Xue, Pengfei (ORCID:000000025702421X)↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗

Short-term nodal load forecasting based on machine learning techniques

This paper introduces an advanced Short-term Nodal Load Forecasting (STNLF) method that forecasts nodal load profiles for the next day in power systems, based on the combined use of three machine learning techniques. Least Absolute Shrinkage and Selection Operator (LASSO) is employed to reduce the number of features for a single nodal load forecasting. Principal Component Analysis (PCA) is used to capture the features of historical loads in low-dimensional space compared to the original high-dimensional load space where features are barely possible to depict. Additionally, Bayesian Ridge Regression (BRR) is utilized to decide the parameters of the prediction model from a statistics perspective. Tests based on modified PJM load data demonstrate the effectiveness of the proposed STNLF method compared to the state-of-the-art General Regression Neural Network (GRNN) method. Moreover, the reliability of the day-ahead Unit Commitment (UC) solution is shown to have been improved, based on the forecasted load data using the proposed STNLF method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Tilted lidar profiling: Development and testing of a novel scanning strategy for inhomogeneous flows

The most common profiling techniques for the atmospheric boundary layer based on a monostatic Doppler wind lidar rely on the assumption of horizontal homogeneity of the flow. This assumption breaks down in the presence of either natural or human-made obstructions that can generate significant flow distortions. The need to deploy ground-based lidars near operating wind turbines for the American WAKE experimeNt (AWAKEN) spurred a search for novel profiling techniques that could avoid the influence of the flow modifications caused by the wind farms. With this goal in mind, two well-established profiling scanning strategies have been retrofitted to scan in a tilted fashion and steer the beams away from the more severely inhomogeneous region of the flow. Results from a field test at the National Renewable Energy Laboratory's 135-m meteorological tower show that the accuracy of the horizontal mean flow reconstruction is insensitive to the tilt of the scan, although higher-order wind statistics are severely deteriorated at extreme tilts mainly due to geometrical error amplification. A numerical study of the AWAKEN domain based on the Weather Research and Forecasting Model and large-eddy simulation are also conducted to test the effectiveness of tilted profiling. It is shown that a threefold reduction of the error on inflow mean wind speed can be achieved for a lidar placed at the base of the turbine using tilted profiling.

17 WIND ENERGY↗

When to vaccinate for seasonal influenza: check the peak forecast

Background Seasonal influenza infects 5-20% of people every year in the United States, resulting in hospitalizations, deaths, and adverse economic impacts. To mitigate these impacts, influenza vaccines are developed and distributed annually; however, growing evidence suggests that vaccine effectiveness (VE) wanes over the course of a flu season. Delaying influenza vaccination for older adults has attracted attention as a potential public health strategy. However, given the uncertainties in seasonal peak, vaccine effectiveness, and waning rates, postponing vaccination could also lead to increased morbidity, motivating an evaluation of a range of potential scenarios. The aim of this study was to investigate favorable age group-specific vaccination schedules that could lead to the greatest disease burden reduction. Methods We systematically investigated a broad range of vaccination start times for five age groups under six combinations of initial effectiveness and waning rates, based on influenza cases and vaccine uptake data from 10 influenza seasons. We defined the most favorable vaccination schedule as the one that resulted in the greatest reduction in disease burden. Results In scenarios with fast waning, all age groups benefit from delaying vaccination regardless of initial VE and peak timing. In scenarios with slower waning, results are mixed. For the ≥65 group, high initial VE and slow waning suggests that in early-peaking seasons, early vaccination most effectively reduces disease burden, while in late-peaking seasons delaying vaccination is most effective. For the ≥65 group in medium and low initial VE, and slow waning scenarios, delaying vaccination appears to prevent the greatest number of cases, regardless of whether the season peaks early or late. Conclusion The most favorable vaccination schedule is sensitive to changes in initial VE, waning rate, and peak timing. Given estimates of these quantities from statistical and immunological models and observations, our methods can inform vaccination recommendations in order to most effectively reduce the annual disease burden caused by seasonal influenza. Specifically, accurate peak timing forecasts for the upcoming season have the potential to guide decisions on when to vaccinate.

59 BASIC BIOLOGICAL SCIENCES↗

Can Simple Metrics Identify the Process(es) Driving Extreme Precipitation?

This work seeks an automatic algorithm to determine the primary meteorological cause(s) of individual extreme precipitation events. Such determinations have been made before, but required a by-hand analysis of each separate event. This is very time-consuming and the field would benefit from an automatic process. This is especially relevant when comparing different datasets to determine which ones most closely hew towards reality. This paper tests three simple metrics over the continental United States using the European Center for Medium-Range Weather Forecasting’s (ECMWF) atmospheric reanalysis (ERA5). The metrics tested measure and compare the strength of three meteorological processes associated with extreme precipitation: fronts, convection, and cyclones. A multivariate statistical technique as well as individual case studies show evidence that the three meteorological processes of interest cannot be isolated from one another using these simple physical metrics. This shows the difficulty in finding “pure” cases of these precipitation-generating processes and suggests approaching these processes with an eye toward mixed-type events.

Swenson, Leif M. (ORCID:0000000199708735)↗