Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Integrating Survival Analysis with Bayesian Statistics to Forecast the Remaining Useful Life of a Centrifugal Pump Conditional to Multiple Fault Types

To improve the viability of nuclear power plants, there is a need to reduce their operational costs. Operational costs account for a significant portion of a plant’s yearly budget, due to their scheduled-based maintenance approach. In order to reduce these costs, proactive methods are required that estimate and forecast the state of a machine in real time to optimize maintenance schedules. In this research, we use Bayesian networks to develop a framework that can forecast the remaining useful life of a centrifugal pump. To do so, we integrate survival analysis with Bayesian statistics to forecast the health of the pump conditional to its current state. We complete our research by successfully using the Bayesian network on a case study. This solution provides an informed probabilistic viewpoint of the pumping system for the purpose of predictive maintenance.

42 ENGINEERING↗

Improving Seasonal Forecast Using Probabilistic Deep Learning

The path toward realizing the potential of seasonal forecasting and its socioeconomic benefits relies on improving general circulation model (GCM) based dynamical forecast systems. To improve dynamical seasonal forecasts, it is crucial to set up forecast benchmarks, and clarify forecast limitations posed by model initialization errors, formulation deficiencies, and internal climate variability. With huge costs in generating large forecast ensembles, and limited observations for forecast verification, the seasonal forecast benchmarking and diagnosing task proves challenging. Here, we develop a probabilistic deep learning-based statistical forecast methodology, drawing on a wealth of climate simulations to enhance seasonal forecast capability and forecast diagnosis. By explicitly modeling the internal climate variability and GCM formulation differences, the proposed Conditional Generative Forecasting (CGF) methodology enables bypassing crucial barriers in dynamical forecast, and offers a top-down viewpoint to examine how complicated GCMs encode the seasonal predictability information. We apply the CGF methodology for global seasonal forecast of precipitation and 2 m air temperature, based on a unique data set consisting 52,201 years of climate simulation. Results show that the CGF methodology can faithfully represent the seasonal predictability information encoded in GCMs. We successfully apply this learned relationship in real-world seasonal forecast, achieving competitive performance compared to dynamical forecasts. Using this CGF as benchmark, we reveal the impact of insufficient forecast spread sampling that limits the skill of the considered dynamical forecast system. Finally, we introduce different strategies for composing ensembles using the CGF methodology, highlighting the potential for leveraging the strengths of multiple GCMs to achieve advantgeous seasonal forecast.

54 ENVIRONMENTAL SCIENCES↗

Power Electronics Materials and Bonded Interfaces - Reliability and Lifetime

High temperature operation of wide bandgap devices continue to be a challenge for the power electronics packages. Thermal performance and reliability are important factors that determine the viability of a bonded interface for operation at high temperatures. In this presentation, we present the technical approach and key results from the research on sintered silver, transient liquid phase alloy, and polymeric materials. A lifetime prediction model that incorporates the thermomechanical behavior of sintered silver at 200C was developed. The copper-aluminum transient alloy completed 350 thermal cycles from -40C to 200C and little increase in the defect level was observed. In addition to material research, we initiated a time-series analysis on the scanning acoustic microscope images of eutectic solder to explore statistical forecasting methods and machine learning techniques. Initial results that report the accuracy of a few different statistical models are presented.

ADVANCED PROPULSION SYSTEMS↗

Emerging opportunities for hybrid perovskite solar cells using machine learning

While there are several bottlenecks in hybrid organic–inorganic perovskite (HOIP) solar cell production steps, including composition screening, fabrication, material stability, and device performance, machine learning approaches have begun to tackle each of these issues in recent years. Different algorithms have successfully been adopted to solve the unique problems at each step of HOIP development. Specifically, high-throughput experimentation produces vast amount of training data required to effectively implement machine learning methods. Here, we present an overview of machine learning models, including linear regression, neural networks, deep learning, and statistical forecasting. Experimental examples from the literature, where machine learning is applied to HOIP composition screening, thin film fabrication, thin film characterization, and full device testing, are discussed. These paradigms give insights into the future of HOIP solar cell research. As databases expand and computational power improves, increasingly accurate predictions of the HOIP behavior are becoming possible.

Hering, Abigail R. (ORCID:0000000270806953)↗

Improving seasonal precipitation forecasts in the Western United States through statistical downscaling

Abstract Seasonal precipitation forecasts in the western United States are critical resources for water resource management, especially during winter. While current seasonal forecasting systems provide monthly precipitation forecasts operationally, their coarse resolution limits their effectiveness in capturing the localized precipitation patterns and snowpack conditions essential for water resource managers in the mountainous regions. Here, analog statistical downscaling is demonstrated as an effective approach to enhance the spatial resolution of operational seasonal forecasts provided by the North American Multi-Model Ensemble. Downscaling was performed by building an analog ‘library’, in which corresponding model forecasts and observed values during the training period were stored. In the testing period, unseen model forecasts referenced the closest historical forecast from the analog library and applied the corresponding observational value for each point. This analysis indicates that downscaled products can capture localized features more accurately than the original coarse resolution forecasts, reducing forecast error across the western United States. Moreover, downscaling individual ensemble members—rather than downscaling the ensemble mean—further reduces forecasting error for their multi-model ensemble mean products. The greatest error reductions in the downscaled product, measured by root mean squared error (RMSE), were observed at low to mid-elevations (500–2000 meters), with 50%–70% improvement relative to the original forecasts. In the higher elevations (2000 meters and above), changes in RMSE relative to the original forecast were limited to 10%–30% improvements. The improvement is more substantial for forecast systems with 10 ensemble members compared to that with 4 members, but this relationship does not hold for the system with 24 ensemble members. These findings show that analog statistical downscaling can effectively address the spatial limitations of seasonal precipitation forecasts with minimal computational cost, providing a valuable framework for enhancing coarse resolution forecasting products while providing insights into the timing of ensemble mean calculations during the downscaling process.

Vernon, B. (ORCID:0009000891670689)↗

Quantifying and simulating the weather forecast uncertainty for advanced building control

Weather forecast uncertainty is unavoidable despite technological advancements. Accurately quantifying and modelling this uncertainty is essential for developing and comparing advanced building controllers. In this study, we present a structured approach using a first-order autoregressive model (AR(1)) to model uncertainty in ambient temperature and global solar irradiation (GHI) forecasts. We analyzed weather data from four cities and employed Jensen–Shannon divergence (JSD) to evaluate the similarity between synthetic and actual forecast errors. The average JSD values for temperature are 0.027 (Berkeley), 0.021 (Leuven), 0.018 (Berlin), and 0.008 (Oslo), and for GHI, the average JSD values are 0.016 (Berkeley), 0.058 (Leuven), and 0.013 (Berlin). The low JSD values indicate a high similarity between the synthetic and real forecast error distributions. Further, our approach successfully generates synthetic weather forecasts that mirror the statistical properties of actual forecasts. The implementation of our method for uncertain forecast generation is being added to the BOPTEST framework.

54 ENVIRONMENTAL SCIENCES↗

Map-based cosmology inference with lognormal cosmic shear maps

ABSTRACT Most cosmic shear analyses to date have relied on summary statistics (e.g. ξ+ and ξ−). These types of analyses are necessarily suboptimal, as the use of summary statistics is lossy. In this paper, we forward-model the convergence field of the Universe as a lognormal random field conditioned on the observed shear data. This new map-based inference framework enables us to recover the joint posterior of the cosmological parameters and the convergence field of the Universe. Our analysis properly accounts for the covariance in the mass maps across tomographic bins, which significantly improves the fidelity of the maps relative to single-bin reconstructions. We verify that applying our inference pipeline to Gaussian random fields recovers posteriors that are in excellent agreement with their analytical counterparts. At the resolution of our maps – and to the extent that the convergence field can be described by the lognormal model – our map posteriors allow us to reconstruct all summary statistics (including non-Gaussian statistics). We forecast that a map-based inference analysis of LSST-Y10 data can improve cosmological constraints in the σ8–Ωm plane by $\approx\!{30}{{\ \rm per\ cent}}$ relative to the currently standard cosmic shear analysis. This improvement happens almost entirely along the $S_8=\sigma _8\Omega _{\rm m}^{1/2}$ directions, meaning map-based inference fails to significantly improve constraints on S8.

79 ASTRONOMY AND ASTROPHYSICS↗

Clustering with general photo- z uncertainties: application to Baryon Acoustic Oscillations

ABSTRACT Photometric data can be analysed using the 3D correlation function ξp to extract cosmological information via e.g. measurement of the Baryon Acoustic Oscillations (BAO). Previous studies modeled ξp assuming a Gaussian photo-z approximation. In this work we improve the modeling by incorporating realistic photo-z distribution. We show that the position of the BAO scale in ξp is determined by the photo-z distribution and the Jacobian of the transformation. The latter diverges at the transverse scale of the separation s⊥, and it explains why ξp traces the underlying correlation function at s⊥, rather than s, when the photo-z uncertainty σz/(1+ z) ≳ 0.02. We also obtain the Gaussian covariance for ξp. Due to photo-z mixing, the covariance of ξp shows strong off-diagonal elements. The high correlation of the data causes some issues to the data fitting. None the less, we find that either it can be solved by suppressing the largest eigenvalues of the covariance or it is not directly related to the BAO. We test our BAO fitting pipeline using a set of mock catalogs. The data set is dedicated for Dark Energy Survey Year 3 (DES Y3) BAO analyses and includes realistic photo-z distributions. The theory template is in good agreement with mock measurement. Based on the DES Y3 mocks, ξp statistic is forecast to constrain the BAO shift parameter α to be 1.001 ± 0.023, which is well consistent with the corresponding constraint derived from the angular correlation function measurements. Thus, ξp offers a competitive alternative for the photometric data analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Predictive Complexity of Quantum Subsystems

We define predictive states and predictive complexity for quantum systems composed of distinct subsystems. This complexity is a generalization of entanglement entropy. It is inspired by the statistical or forecasting complexity of predictive state analysis of stochastic and complex systems theory but is intrinsically quantum. Predictive states of a subsystem are formed by equivalence classes of state vectors in the exterior Hilbert space that effectively predict the same future behavior of that subsystem for some time. As an illustrative example, we present calculations in the dynamics of an isotropic Heisenberg model spin chain and show that, in comparison to the entanglement entropy, the predictive complexity better signifies dynamically important events, such as magnon collisions. It can also serve as a local order parameter that can distinguish long and short range entanglement.

Asplund, Curtis T. (ORCID:0000000305575850)↗

Comparison of Deterministic and Statistical Models for Water Quality Compliance Forecasting in the San Joaquin River Basin, California

Model selection for water quality forecasting depends on many factors including analyst expertise and cost, stakeholder involvement and expected performance. Water quality forecasting in arid river basins is especially challenging given the importance of protecting beneficial uses in these environments and the livelihood of agricultural communities. In the agriculture-dominated San Joaquin River Basin of California, real-time salinity management (RTSM) is a state-sanctioned program that helps to maximize allowable salt export while protecting existing basin beneficial uses of water supply. The RTSM strategy supplants the federal total maximum daily load (TMDL) approach that could impose fines associated with exceedances of monthly and annual salt load allocations of up to $1 million per year based on average year hydrology and salt load export limits. The essential components of the current program include the establishment of telemetered sensor networks, a web-based information system for sharing data, a basin-scale salt load assimilative capacity forecasting model and institutional entities tasked with performing weekly forecasts of river salt assimilative capacity and scheduling west-side drainage export of salt loads. Web-based information portals have been developed to share model input data and salt assimilative capacity forecasts together with increasing stakeholder awareness and involvement in water quality resource management activities in the river basin. Two modeling approaches have been developed simultaneously. The first relies on a statistical analysis of the relationship between flow and salt concentration at three compliance monitoring sites and the use of these regression relationships for forecasting. The second salt load forecasting approach is a customized application of the Watershed Analysis Risk Management Framework (WARMF), a watershed water quality simulation model that has been configured to estimate daily river salt assimilative capacity and to provide decision support for real-time salinity management at the watershed level. Analysis of the results from both model-based forecasting approaches over a period of five years shows that the regression-based forecasting model, run daily Monday to Friday each week, provided marginally better performance. However, the regression-based forecasting model assumes the same general relationship between flow and salinity which breaks down during extreme weather events such as droughts when water allocation cutbacks among stakeholders are not evenly distributed across the basin. A recent test case shows the utility of both models in dealing with an exceedance event at one compliance monitoring site recently introduced in 2020.

54 ENVIRONMENTAL SCIENCES↗

Temporal Forecasting of Distributed Temperature Sensing in a Thermal Hydraulic System With Machine Learning and Statistical Models

We benchmark performance of long-short term memory (LSTM) network machine learning model and autoregressive integrated moving average (ARIMA) statistical model in temporal forecasting of distributed temperature sensing (DTS). Data in this study consists of fluid temperature transient measured with two co-located Rayleigh scattering fiber optic sensors (FOS) in a forced convection mixing zone of a thermal tee. We treat each gauge of a FOS as an independent temperature sensor. We first study prediction of DTS time series using Vanilla LSTM and ARIMA models trained on prior history of the same FOS that is used for testing. The results yield maximum absolute percentage error (MaxAPE) and root mean squared percentage error (RMSPE) of 1.58% and 0.06% for ARIMA, and 3.14% and 0.44% for LSTM, respectively. Next, we investigate zero-shot forecasting (ZSF) with LSTM and ARIMA trained on history of the co-located FOS only, which is advantageous when limited training data is available. The ZSF MaxAPE and RMSPE values for ARIMA are comparable to those of the Vanilla use case, while the error values for LSTM increase. We show that in ZSF, performance of LSTM network can be improved by training on most correlated gauges between the two FOS, which are identified by calculating the Pearson correlation coefficient. The improved ZSF MaxAPE and RMSPE for LSTM are 4.4% and 0.33%, respectively. Performance of ZSF LSTM can be further enhanced through transfer learning (TL), where LSTM is re-trained on a subset of the FOS that is the target of forecasting. We show that LSTM pre-trained on correlated dataset and re-trained on 30% of testing target dataset achieves MaxAPE and RMSPE values of 2.32% and 0.28%, respectively.

ARIMA↗

A Statistical Interpolation Code for Ocean Analysis and Forecasting

Abstract We present a data assimilation package for use with ocean circulation models in analysis, forecasting, and system evaluation applications. The basic functionality of the package is centered on a multivariate linear statistical estimation for a given predicted/background ocean state, observations, and error statistics. Novel features of the package include support for multiple covariance models, and the solution of the least squares normal equations either using the covariance matrix or its inverse—the information matrix. The main focus of this paper, however, is on the solution of the analysis equations using the information matrix, which offers several advantages for solving large problems efficiently. Details of the parameterization of the inverse covariance using Markov random fields are provided and its relationship to finite-difference discretizations of diffusion equations are pointed out. The package can assimilate a variety of observation types from both remote sensing and in situ platforms. The performance of the data assimilation methodology implemented in the package is demonstrated with a yearlong global ocean hindcast with a 1/4° ocean model. The code is implemented in modern Fortran, supports distributed memory, shared memory, multicore architectures, and uses climate and forecasts compliant Network Common Data Form for input/output. The package is freely available with an open source license from www.tendral.com/tsis/ .

Srinivasan, Ashwanth↗

Revealing the Statistics of Extreme Events Hidden in Short Weather Forecast Data

Extreme weather events have significant consequences, dominating the impact of climate on society. While high-resolution weather models can forecast many types of extreme events on synoptic timescales, long-term climatological risk assessment is an altogether different problem. A once-in-a-century event takes, on average, 100 years of simulation time to appear just once, far beyond the typical integration length of a weather forecast model. Therefore, this task is left to cheaper, but less accurate, low-resolution or statistical models. But there is untapped potential in weather model output: despite being short in duration, weather forecast ensembles are produced multiple times a week. Integrations are launched with independent perturbations, causing them to spread apart over time and broadly sample phase space. Collectively, these integrations add up to thousands of years of data. We establish methods to extract climatological information from these short weather simulations. Using ensemble hindcasts by the European Center for Medium-range Weather Forecasting archived in the subseasonal-to-seasonal (S2S) database, we characterize sudden stratospheric warming (SSW) events with multi-centennial return times. Consistent results are found between alternative methods, including basic counting strategies and Markov state modeling. By carefully combining trajectories together, we obtain estimates of SSW frequencies and their seasonal distributions that are consistent with reanalysis-derived estimates for moderately rare events, but with much tighter uncertainty bounds, and which can be extended to events of unprecedented severity that have not yet been observed historically. These methods hold potential for assessing extreme events throughout the climate system, beyond this example of stratospheric extremes.

58 GEOSCIENCES↗

Optimizing Prediction Error for Time-dependent Solar Radiation Modeling

Numerical weather forecasting models and statistical methods have found wide use to help power companies estimate renewable output, but better methods are needed, particularly for extended forecasts. Machine learning approaches have been used here as well, but so far a major limitation is the ability to also predict the corresponding uncertainty in a forecast. Here we show that both can be done and demonstrate this using a long-term-short memory neural network where the difference between predicted and ground truth data are used to train a model for the corresponding forecast uncertainties.

97 MATHEMATICS AND COMPUTING↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

En echelon faults reactivated by wastewater disposal near Musreau Lake, Alberta

We use machine-learning and cross-correlation techniques to enhance earthquake detectability by two magnitude units for the earthquake sequence near Musreau Lake, Alberta, which is induced by wastewater disposal. This deep catalogue reveals a series of en echelon ~N–S oriented strike-slip faults that are favourably oriented for reactivation. These faults require only ~0.6 MPa overpressure for triggering to occur. Earthquake activity occurs in bursts, or episodes; episodes restricted to the largest fault tend to have earthquakes starting near the southern end (distant from injectors) and progressing northwards (towards the injectors). While most events are concentrated along these ~N–S oriented faults, we also delineate smaller faults. Together, these findings suggest pore pressure as the triggering mechanism, where a time-dependent increase in pore pressure likely caused these faults to progressively reawaken. Analysis of the ‘next record-breaking event’, a statistical model that forecasts the sequencing of earthquake magnitudes, suggests that the next largest event would be M L ~4.3. The seismically illuminated length of the largest fault indicates potential magnitudes as large as M w 5.3.

58 GEOSCIENCES↗