Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Updraft and Downdraft Core Size and Intensity as Revealed by Radar Wind Profilers: MCS Observations and Idealized Model Comparisons

This study explores the updraft and downdraft properties of mature stage mesoscale convective systems (MCSs) in terms of draft core width, shape, intensity, and mass flux characteristics. The observations use extended radar wind profiler (RWP) and surveillance radar datasets from the U.S. Department of Energy Atmospheric Radiation Measurement program for midlatitude (Oklahoma, USA) and tropical (Amazon, Brazil) sites. MCS drafts behave qualitatively similar to previous aircraft and RWP cloud summaries. The Oklahoma MCSs indicate larger and more intense convective up- and downdraft cores, and greater mass flux than Amazon MCS counterparts. However, similar size-intensity relationships and draft vertical profile behaviors are observed for both regions. Additional similarities include weak positive correlations between core intensity and core width (correlation coefficient r ~ 0.5), and increases in draft intensity with altitude. A model-observational intercomparison for draft properties (core width, intensity, mass flux) is also performed to illustrate the potential usefulness of statistical observed draft characterizations. Idealized simulations with the Weather Research and Forecasting model aligned with midlatitude MCS conditions are performed at model grid spacings (?x) that range from 4 km to 250 m. It is shown that the simulations performed at ?x = 250 m at similar mature MCS lifecycle stages are those that exhibit draft intensity, width, mass flux, and shape parameter performances best matching with observed properties.

radar, arm, wind profiler, goamazon↗

Deep Learning-Based Weather-Related Power Outage Prediction with Socio-Economic and Power Infrastructure Data

This paper presents a deep learning-based approach for hourly power outage probability prediction within census tracts encompassing a utility company's service territory. Two distinct deep learning models, conditional Multi-Layer Perceptron (MLP) and unconditional MLP, were developed to forecast power outage probabilities, leveraging a rich array of input features gathered from publicly available sources including weather data, weather station locations, power infrastructure maps, socio-economic and demographic statistics, and power outage records. Given a one-hour-ahead weather forecast, the models predict the power outage probability for each census tract, taking into account both the weather prediction and the location's characteristics. The deep learning models employed different loss functions to optimize prediction performance. Our experimental results underscore the significance of socio-economic factors in enhancing the accuracy of power outage predictions at the census tract level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Mapping Incidence and Prevalence Peak Data for SIR Modeling Applications

Infectious disease modeling and forecasting have played a key role in helping assess and respond to epidemics and pandemics. Recent work has leveraged data on disease peak infection and peak hospital incidence to fit compartmental models for the purpose of forecasting and describing the dynamics of a disease outbreak. Incorporating these data can greatly stabilize a compartmental model fit on early observations, where slight perturbations in the data may lead to model fits that forecast wildly unrealistic peak infection. We introduce a new method for incorporating historic data on the value and time of peak incidence of hospitalization into the fit for a Susceptible-Infectious-Recovered (SIR) model by formulating the relationship between an SIR model’s starting parameters and peak incidence as a system of two equations that can be solved computationally. We demonstrate how to calculate SIR parameter estimates – which describe disease dynamics such as transmission and recovery rates – using this method, and determine that there is a noticeable loss in accuracy whenever prevalence data is misspecified as incidence data. To exhibit the modeling potential, we update the Dirichlet-Beta State Space modeling framework to use hospital incidence data, as this framework was previously formulated to incorporate only data on total infections. This approach is assessed for practicality in terms of accuracy and speed of computation via simulation.

97 MATHEMATICS AND COMPUTING↗

From Optimization to Sampling Through Gradient Flows

Optimization and sampling algorithms play a central role in science and engineering as they enable finding optimal predictions, policies, and recommendations, as well as expected and equilibrium states of complex systems. The notion of “optimality” is formalized by the choice of an objective function, while the notion of an “expected” state is specified by a probabilistic model for the distribution of states. Optimizing rugged objective functions and sampling multimodal distributions is computationally challenging, especially in high-dimensional problems. Here, for this reason, many optimization and sampling methods have been developed by researchers working in disparate fields such as Bayesian statistics, molecular dynamics, genetics, quantum chemistry, machine learning, weather forecasting, econometrics, and medical imaging.

Trillos, N. García↗

Model orthogonalization and Bayesian forecast mixing via principal component analysis

One can improve predictability in the unknown domain by combining forecasts of imperfect complex computational models using a Bayesian statistical machine learning framework. In many cases, however, the models used in the mixing process are similar. In addition to contaminating the model space, the existence of such similar, or even redundant, models during the multimodeling process can result in misinterpretation of results and deterioration of predictive performance. In this paper we describe a method based on the principal component analysis that eliminates model redundancy. We show that by adding model orthogonalization to the proposed Bayesian model combination framework, one can arrive at better prediction accuracy and reach excellent uncertainty quantification performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

Data-driven occupant-behavior analytics for residential buildings

Many advances have been made in building technology to help save energy, but influencing the behavior of the occupants is still necessary to achieve low-energy use targets. One of the most practical ways to influence and change occupant behaviors is through incentives. Developing incentives for energy-saving and quantifying the impact of occupant behaviors are both active areas of research. Here, we propose a data analytics framework for detecting changes in occupant behaviors, which will help build an analytics feedback loop from behavior impact to incentive design. The framework has two major parts. The first forecasts energy consumption for each occupant, while the second determines a probability distribution for changes in energy consumption. The parts are interchangeable with other existing machine learning and statistical methods. A specific instantiation of the framework, using kernel ridge-regression for forecasting and k-means to find an empirical behavior distribution, is described in detail. An HVAC use-case with 5 different incentivized behaviors is used as an example to show that the framework can detect behavior changes induced by incentives. Furthermore, we show that some simpler behavior-change detection methods do not work, further justifying the use of advanced analytics.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Radiosonde Observations of Environments Supporting Deep Moist Convection Initiation during RELAMPAGO-CACTI

The Remote Sensing of Electrification, Lightning, and Mesoscale/Microscale Processes with Adaptive Ground Observations (RELAMPAGO) and Cloud, Aerosol, and Complex Terrain Interactions (CACTI) projects deployed a high-spatiotemporal-resolution radiosonde network to examine environments supporting deep convection in the complex terrain of central Argentina. This study aims to characterize atmospheric profiles most representative of the near-cloud environment (in time and space) to identify the mesoscale ingredients affecting storm initiation and growth. Spatiotemporal autocorrelation analysis of the soundings reveals that there is considerable environmental heterogeneity, with boundary layer thermodynamic and kinematic fields becoming statistically uncorrelated on scales of 1–2 h and 30 km. Using this as guidance, we examine a variety of environmental parameters derived from soundings collected within close proximity (30 km in space and 30 min in time) of 44 events over 9 days where the atmosphere either: 1) supported the initiation of sustained precipitating convection, 2) yielded weak and short-lived precipitating convection, or 3) produced no precipitating convection in disagreement with numerical forecasts from convection-allowing models (i.e., Null events). There are large statistical differences between the Null event environments and those supporting any convective precipitation. Null event profiles contained larger convective available potential energy, but had low free-tropospheric relative humidity, higher freezing levels, and evidence of limited horizontal convergence near the terrain at low levels that likely suppressed deep convective growth. Here, we also present evidence from the radiosonde and satellite measurements that flow–terrain interactions may yield gravity wave activity that affects CI outcome.

54 ENVIRONMENTAL SCIENCES↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

The winter central Arctic surface energy budget: A model evaluation using observations from the MOSAiC campaign

This study evaluates the simulation of wintertime (15 October, 2019, to 15 March, 2020) statistics of the central Arctic near-surface atmosphere and surface energy budget observed during the MOSAiC campaign with short-term forecasts from 7 state-of-the-art operational and experimental forecast systems. Five of these systems are fully coupled ocean-sea ice-atmosphere models. Forecast systems need to simultaneously simulate the impact of radiative effects, turbulence, and precipitation processes on the surface energy budget and near-surface atmospheric conditions in order to produce useful forecasts of the Arctic system. This study focuses on processes unique to the Arctic, such as, the representation of liquid-bearing clouds at cold temperatures and the representation of a persistent stable boundary layer. It is found that contemporary models still struggle to maintain liquid water in clouds at cold temperatures. Given the simple balance between net longwave radiation, sensible heat flux, and conductive ground flux in the wintertime Arctic surface energy balance, a bias in one of these components manifests as a compensating bias in other terms. This study highlights the different manifestations of model bias and the potential implications on other terms. Three general types of challenges are found within the models evaluated: representing the radiative impact of clouds, representing the interaction of atmospheric heat fluxes with sub-surface fluxes (i.e., snow and ice properties), and representing the relationship between stability and turbulent heat fluxes.

54 ENVIRONMENTAL SCIENCES↗

Daily Forecasting of Regional Epidemics of Coronavirus Disease with Bayesian Uncertainty Quantification, United States

To increase situational awareness and support evidence-based policymaking, we formulated a mathematical model for coronavirus disease transmission within a regional population. This compartmental model accounts for quarantine, self-isolation, social distancing, a nonexponentially distributed incubation period, asymptomatic persons, and mild and severe forms of symptomatic disease. We used Bayesian inference to calibrate region-specific models for consistency with daily reports of confirmed cases in the 15 most populous metropolitan statistical areas in the United States. We also quantified uncertainty in parameter estimates and forecasts. This online learning approach enables early identification of new trends despite considerable variability in case reporting.

59 BASIC BIOLOGICAL SCIENCES↗

The Potential Benefits of Handling Mixture Statistics via a Bi-Gaussian EnKF: Tests With All-Sky Satellite Infrared Radiances

The meteorological characteristics of cloudy atmospheric columns can be very different from their clear counterparts. Thus, when a forecast ensemble is uncertain about the presence/absence of clouds at a specific atmospheric column (i.e., some members are clear while others are cloudy), that column's ensemble statistics will contain a mixture of clear and cloudy statistics. Such mixtures are inconsistent with the ensemble data assimilation algorithms currently used in numerical weather prediction. Hence, ensemble data assimilation algorithms that can handle such mixtures can potentially outperform currently used algorithms. In this study, we demonstrate the potential benefits of addressing such mixtures through a bi-Gaussian extension of the ensemble Kalman filter (BGEnKF). The BGEnKF is compared against the commonly used ensemble Kalman filter (EnKF) using perfect model observing system simulated experiments (OSSEs) with a realistic weather model (the Weather Research and Forecast model). Synthetic all-sky infrared radiance observations are assimilated in this study. In these OSSEs, the BGEnKF outperforms the EnKF in terms of the horizontal wind components, temperature, specific humidity, and simulated upper tropospheric water vapor channel infrared brightness temperatures. This study is one of the first to demonstrate the potential of a Gaussian mixture model EnKF with a realistic weather model. Our results thus motivate future research toward improving numerical Earth system predictions though explicitly handling mixture statistics.

54 ENVIRONMENTAL SCIENCES↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Intercomparison of Dynamically and Statistically Downscaled Climate Change Projections over the Midwest and Great Lakes Region

Downscaling of global climate model (GCMs) simulations is a key element of regional-to-local-scale climate change projections that can inform impact assessments, long-term planning, and resource management in different sectors. Here, we conduct an intercomparison between statistically and dynamically downscaled GCMs simulations using the hybrid delta (HD) and the Weather Research and Forecast (WRF) Model, respectively, over the Midwest and Great Lakes region to 1) validate their performance in reproducing extreme daily precipitation (P) and daily maximum temperature (T max ) for summer and winter and 2) evaluate projections of extremes in the future. Our results show the HD statistical downscaling approach, which includes large-scale bias correction of GCM inputs, can reproduce observed extreme P and T max reasonably well for both summer and winter. However, raw historical WRF simulations show significant bias in both extreme P and T max for both seasons. Interestingly, the convection-permitting WRF simulation at 4-km grid spacing does not produce better results for seasonal extremes than the WRF simulation at 12 km using a parameterized convection scheme. Despite a broad similarity for winter extreme P projections, the projected changes in the future summer storms are quite different between downscaling methods; WRF simulations show substantial increases in summer extreme precipitation, while the changes projected by the HD approach exhibit moderate decreases overall. The WRF simulations at 4 km also show a pronounced decoupling effect between seasonal totals and extreme daily P for summer, which suggests that there could be more intense summer extremes at two different time scales, with more severe individual convective storms combined with longer summer droughts at the end of the twenty-first century.

54 ENVIRONMENTAL SCIENCES↗

Photometric redshift uncertainties in weak gravitational lensing shear analysis: models and marginalization

ABSTRACT Recovering credible cosmological parameter constraints in a weak lensing shear analysis requires an accurate model that can be used to marginalize over nuisance parameters describing potential sources of systematic uncertainty, such as the uncertainties on the sample redshift distribution n(z). Due to the challenge of running Markov chain Monte Carlo (MCMC) in the high-dimensional parameter spaces in which the n(z) uncertainties may be parametrized, it is common practice to simplify the n(z) parametrization or combine MCMC chains that each have a fixed n(z) resampled from the n(z) uncertainties. In this work, we propose a statistically principled Bayesian resampling approach for marginalizing over the n(z) uncertainty using multiple MCMC chains. We self-consistently compare the new method to existing ones from the literature in the context of a forecasted cosmic shear analysis for the HSC three-year shape catalogue, and find that these methods recover statistically consistent error bars for the cosmological parameter constraints for predicted HSC three-year analysis, implying that using the most computationally efficient of the approaches is appropriate. However, we find that for data sets with the constraining power of the full HSC survey data set (and, by implication, those upcoming surveys with even tighter constraints), the choice of method for marginalizing over n(z) uncertainty among the several methods from the literature may modify the 1σ uncertainties on Ωm–S8 constraints by ∼4 per cent, and a careful model selection is needed to ensure credible parameter intervals.

Zhang, Tianqing (ORCID:000000025596198X)↗

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

Coordinated Ramping Product and Regulation Reserve Procurements in CAISO and MISO using Multi-Scale Probabilistic Solar Power Forecasts (Pro2R)

How can probabilistic solar forecasts lower costs and improve reliability for independent system operator (ISO) markets? We tackle this question in three steps. First, we enhance an existing solar forecasting system to provide well-calibrated hours-ahead probabilistic forecasts. We then relate the degree of uncertainty in those forecasts to error distributions for net load ramps for the California ISO (CAISO) using statistical and machine learning methods. Projected net load errors conditioned on solar uncertainty are translated into flexible ramp requirements that therefore reflect real-time meteorological and solar conditions, improving on typical ISO procedures. Finally, a multi-period look-ahead production cost model quantifies how conditional ramp requirements can a) decrease operating costs by lowering requirements compared to often conservative unconditional methods, and b) reduce generation scarcity events and consequently improve reliability by increasing flexibility requirements at times when unconditional forecast-based requirements understate actual ramp uncertainty. In addition to the products just described (quantification of solar uncertainty, its translation into requirements for ramp capability product, and quantification of the benefits of more accurate ramp requirements), this project also developed a visualization system that alerts system operators of ramp and uncertainty conditions within the network based on solar forecasts. The system is called Resource Forecast and Ramp Visualization for Situational Awareness (RaVIS). These four products represent significant advances in the state-of-the-art of probabilistic solar forecasting, development of weather-informed reserve requirements, production costing methods for estimating the benefits of more accurate reserve requirements, and visualization of system status, respectively. Yet the products are also practical and can be immediately implemented, potentially enabling system operators to save millions of dollars in ramp product procurement costs per year.

14 SOLAR ENERGY↗

Cross-Market Price Difference Forecast Using Deep Learning for Electricity Markets

Price forecasting is in the center of decision making in electricity markets. Many researches have been done in forecasting energy prices while little research has been reported on forecasting price difference between day-ahead and realtime markets due to its high volatility, which however plays a critical role in virtual trading. To this end, this paper takes the first attempt to employ novel deep learning architecture with Bidirectional Long-Short Term Memory (LSTM) units to forecast the price difference between day-ahead and real-time markets for the same node. The raw data is collected from PJM market, processed and fed into the proposed network. The Root Mean Squared Error (RMSE) and customized performance metric are used to evaluate the performance of the proposed method. Case studies show that it outperforms the traditional statistical models like ARIMA, and machine learning models like XGBoost and SVR methods in both RMSE and the capability of forecasting the sign of price difference. Additionally, to cross-market price difference forecast, the proposed approach has the potential to be applied to solve other forecasting problems such as price spread forecast in DA market for Financial Transmission Right (FTR) trading purpose.

DA/RT price difference↗