Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

125 records · Page 7

Identifiability and predictability of integer- and fractional-order epidemiological models using physics-informed neural networks

Here we analyze a plurality of epidemiological models through the lens of physics-informed neural networks (PINNs) that enable us to identify time-dependent parameters and data-driven fractional differential operators. In particular, we consider several variations of the classical susceptible-infectious-removed (SIR) model by introducing more compartments and fractional-order and time-delay models. We report the results for the spread of COVID-19 in New York City, Rhode Island and Michigan states and Italy, by simultaneously inferring the unknown parameters and the unobserved dynamics. For integer-order and time-delay models, we fit the available data by identifying time-dependent parameters, which are represented by neural networks. In contrast, for fractional differential models, we fit the data by determining different time-dependent derivative orders for each compartment, which we represent by neural networks. We investigate the structural and practical identifiability of these unknown functions for different datasets, and quantify the uncertainty associated with neural networks and with control measures in forecasting the pandemic.

60 APPLIED LIFE SCIENCES↗

Predicting resistive wall mode stability in NSTX through balanced random forests and counterfactual explanations

Abstract Recent progress in the disruption event characterization and forecasting framework has shown that machine learning guided by physics theory can be easily implemented as a supporting tool for fast computations of ideal stability properties of spherical tokamak plasmas. In order to extend that idea, a customized random forest (RF) classifier that takes into account imbalances in the training data is hereby employed to predict resistive wall mode (RWM) stability for a set of high beta discharges from the NSTX spherical tokamak. More specifically, with this approach each tree in the forest is trained on samples that are balanced via a user-defined over/under-sampler. The proposed approach outperforms classical cost-sensitive methods for the problem at hand, in particular when used in conjunction with a random under-sampler, while also resulting in a threefold reduction in the training time. In order to further understand the model’s decisions, a diverse set of counterfactual explanations based on determinantal point processes (DPP) is generated and evaluated. Via the use of DPP, the underlying RF model infers that the presence of hypothetical magnetohydrodynamic activity would have prevented the RWM from concurrently going unstable, which is a counterfactual that is indeed expected by prior physics knowledge. Given that this result emerges from the data-driven RF classifier and the use of counterfactuals without hand-crafted embedding of prior physics intuition, it motivates the usage of counterfactuals to simulate real-time control by generating the β N levels that would have kept the RWM stable for a set of unstable discharges.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dark Energy Survey Year 3 results: Simulation-based 𝑤CDM inference from weak lensing and galaxy clustering maps with deep learning: Analysis design

Data-driven approaches using deep learning are emerging as powerful techniques to extract non-Gaussian information from cosmological large-scale structure. Here, this work presents the first simulation-based inference (SBI) pipeline that combines weak lensing and galaxy clustering maps in a realistic Dark Energy Survey Year 3 (DES Y3) configuration and serves as preparation for a forthcoming analysis of the survey data. We develop a scalable forward model based on the CosmoGridV1 suite of N-body simulations to generate over one million self-consistent mock realizations of DES Y3 at the map level. Leveraging this large dataset, we train deep graph convolutional neural networks on the full survey footprint in spherical geometry to learn low-dimensional features that approximately maximize mutual information with target parameters. These learned compressions enable neural density estimation of the implicit likelihood via normalizing flows in a ten-dimensional parameter space spanning cosmological 𝑤CDM, intrinsic alignment, and linear galaxy bias parameters, while marginalizing over baryonic, photometric redshift, and shear bias nuisances. To ensure robustness, we extensively validate our inference pipeline using synthetic observations derived from both systematic contaminations in our forward model and independent Buzzard galaxy catalogs. Our forecasts yield significant improvements in cosmological parameter constraints, achieving 2−3× higher figures of merit in the 𝛺 𝑚 − 𝑆 8 plane relative to our implementation of baseline two-point statistics and effectively breaking parameter degeneracies through probe combination. These results demonstrate the potential of SBI analyses powered by deep learning for upcoming Stage-IV wide-field imaging surveys.

Thomsen, A. [Zurich, ETH] (ORCID:0000000203099021)↗

Probabilistic Voltage Sensitivity based Preemptive Voltage Monitoring in Unbalanced Distribution Networks

With increasing penetration of renewable energy and active consumers, control and management of power distribution networks has become challenging. Renewable energy sources can cause random voltage fluctuations as their output power depends on weather conditions. Conventional voltage control schemes such as tap changers and capacitor banks lack the foresight required to quickly alleviate voltage violations. Thus, there is an urgent need for effective approaches for predicting and mitigating voltage violations as a result of random fluctuations in power injections. This work proposes a novel voltage monitoring approach based on low-complexity, data-driven probabilistic voltage sensitivity analysis. The usefulness of this work is not only in predicting voltage violations in unbalanced distribution grids, but also in opening up the door for optimal voltage control. Using system data and forecasts, the proposed approach predicts the distribution of system node voltages which is then used to to identify nodes that may violate the nominal operational limits with high probability. The method is tested on the IEEE 37 node distribution system considering integrated distributed solar energy sources. The method is validated against the classic load flow based method and offers over 95% accuracy in predicting voltage violations.

Abujubbeh, Mohammad↗

Data-Driven Day-Ahead PV Estimation Using Autoencoder-LSTM and Persistence Model

Inherent variability in photovoltaic (PV) and associated impacts on power systems is a challenging problem for both the PV owners and the grid operators. Existing statistical and machine learning algorithms typically work well for weather conditions similar to historical data. Furthermore, uncertain weather conditions pose a great challenge to the estimation accuracy of the estimation models. With the enhanced integration of intelligent electronic devices and the realization of associated automation in the power grid, renewable energy data is becoming more accessible, which can be utilized by deep learning models and improve the PV power generation estimation accuracy. In this paper, a hybrid deep learning model driven by external weather data is proposed to do day-ahead PV output forecasting at 15-minute-interval. The proposed model is motivated by the recent advancement of Long-Short-Term-Memory (LSTM) networks and AutoEncoder (AE), which estimates uncertainties in sequence while making the prediction for complex weather conditions. Meanwhile, the persistence model (PM) is used to predict continuous sunny weather conditions. The forecasting result is validated with data from multiple locations

42 ENGINEERING↗

HPC Analytics of Fused Thermal Plants Data to Optimize Operating Envelope

In this project, ORNL extensively reviewed the ORAP RAM data, and it guided us to develop machine learning models that can predict time to next failures and forecast failure trends, which will be useful for optimizing power plant operation strategies. More specifically, we trained multiple random forest models and evaluated the model accuracy to validate with 10+ years of historical data. In addition, we implemented a web-based graphical user interface system for the models to show how our models can be used in more intuitive ways. This proof of concept allowed exploration of model use with power plant operators in mind. Developed machine learning models will be helpful for managing risks, planning maintenance and operation, ultimately reducing the down time and increasing the service hours. For future work, there are several interesting research topics including but not limited to model enhancement, creating synergy with traditional failure modeling approaches, and data-driven actionable recommendation and suggestions.

20 FOSSIL-FUELED POWER PLANTS↗

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Automated Cloud Based Long Short-Term Memory Neural Network Based SWE Prediction

Snow derived water is a critical component of the US water supply. Measurements of the Snow Water Equivalent (SWE) and associated predictions of peak SWE and snowmelt onset are essential inputs for water management efforts. This paper aims to develop an integrated framework for real-time data ingestion, estimation, prediction and visualization of SWE based on daily snow datasets. In particular, we develop a data-driven approach for estimating and predicting SWE dynamics using the Long Short-Term Memory neural network (LSTM) method. Our approach uses historical datasets (precipitation, air temperature, SWE, and snow thickness) collected at NRCS Snow Telemetry (SNOTEL) stations to train the LSTM network and current year data to predict SWE behavior. The performance of our prediction was compared for different prediction dates and prediction training datasets. Our results suggest that the proposed LSTM network can be an efficient tool for forecasting the SWE timeseries, as well as Peak SWE and snowmelt timing. Results showed that the window size impacts the model performance (where the Nash Sutcliffe efficiency (NSE) ranged from 0.96 to 0.85 and the Rooted Mean Square Error (RMSE) ranged from 0.038 to 0.07) with an optimum number that should be calibrated for different stations and climate conditions. In addition, by implementing the LSTM prediction capability in a cloud based site-monitoring platform, we automate model-data integration. By making the data accessible through a graphical web interface and an underlying API which exposes both training and prediction capabilities. The associated results can be made easily accessible to a broad range of stakeholders.

54 ENVIRONMENTAL SCIENCES↗

Monitoring of Liquid Metal Reactor Heater Zones with Recurrent Neural Network Learning of Temperature Time Series

Advanced high-temperature fluid reactors (ARs), such as sodium fast reactors (SFRs) and molten salt cooled reactors (MSCRs) utilize high-temperature fluids at ambient pressure. To melt the fluid during reactor startup and prevent fluid freezing during cooldown, the thermal–hydraulic systems of such ARs include heater zones consisting of specific heaters with controllers, temperature sensors, and thermal insulation. The failure of heater zones due to insulation material degradation or improper installation, resulting in parasitic heat losses, can lead to fluid freezing. The detection of faults using a heat-transfer model is difficult because of a lack of knowledge of the experimental details. Data-driven machine learning of heater zone temperature time series offers a viable alternative. In this study, we benchmarked the performance of recurrent neural networks (RNNs) in an analysis of heat-up transient temperature time series of heater zones installed on a liquid sodium vessel. The RNN models include long short-term memory (LSTM) and gated recurrent unit (GRU) networks, as well as their bi-directional variants, BiLSTM and BiGRU. Anomalous temperature points were designated using a percentile-based threshold applied to residual fluctuations in the detrended temperature time series. Additionally, the impact of the exponentially weighted moving average (EWMA) method on detection accuracy was examined. The RNN models’ performance was assessed using precision, recall, and F 1 score metrics. Results demonstrated that RNN models effectively detect anomalies in temperature time series with the best models for each heater zone achieving F 1 scores of over 93%. To explain the variations in RNN model performance across different heater zones, we used Kullback–Leibler (KL) divergence to quantify the relative entropy between training and testing data, and the Detrended Fluctuation Analysis (DFA) to assess long-range temporal correlations. For datasets with strong long-range correlations and minimal relative entropy between training and testing data, GRU is the best-performing model. When the data exhibits weaker long-term correlations and a significant relative entropy between training and testing distributions, BiGRU shows the best performance. For the data sets with intermediate values of both KL divergence and DFA, the best performance is obtained with LSTM and BiLSTM, respectively.

gated recurrent unit↗

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

Towards Trust-Augmented Visual Analytics for Data-Driven Energy Modeling

The promise of data-driven predictive modeling is being increasingly realized in various science and engineering disciplines, where experts are used to the more conventional, simulation-driven modeling practices. However, trust remains a bottleneck for greater adoption of machine learning-based models for domain experts, who might not be necessarily trained in data science. In this paper, we focus on the building energy domain, where physics-based simulations are being complemented or replaced by machine learning-based methods for forecasting energy supply and demand at various spatio-temporal scales. We study the trust problem in close collaboration with energy scientists and engineers and describe how visual analytics can be leveraged for alleviating this trust bottleneck for stakeholders with varying degrees of expertise and analytics goals in this domain.

Kandakatla, Akshith R.↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

Storm-DEPART (Damage Estimate Prediction and Recovery Tool)

Storm-DEPART (Damage Estimate Prediction and Restoration Tool): Each year hurricanes and tropical storms in the United States damage critical infrastructure assets, disrupt the services they provide, and cause millions to billions of dollars in economic impacts due to extended recovery times. The Storm-DEPART tool and analytical output enable more impactful data-driven decision-making capabilities and strengthen national-level disaster preparedness, response, and recovery. Storm-DEPART, built through multi-month collaboration between Entergy and INL, combines Entergy’s critical infrastructure inventory data with weather forecasts to predict damages to Electric utility’s assets due to natural disasters and the estimated recovery support needed, including time, materials, and resource allocation. In the event of an approaching hurricane, this innovative solution can assess potential damage to power generation capacity, transmission grids, distribution networks, and communications assets from wind bands, storm surge, and flooding. With more effective predictions, Entergy can more efficiently allocate resources to mitigate impacts and optimize recovery for customers. Storm-DEPART also allows Electric utilities the ability to apply a planning scenario and model expected damage to better inform infrastructure restoration needs leading to enhance system resiliency. The technology is fully transferrable to other electric utilities with the same damage estimating challenges. The INL team is working on the evolution of Storm-DEPART to include ice event damage prediction framework.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Storm-DEPART (Damage Estimate Prediction and Recovery Tool)

Storm-DEPART (Damage Estimate Prediction and Restoration Tool): Each year hurricanes and tropical storms in the United States damage critical infrastructure assets, disrupt the services they provide, and cause millions to billions of dollars in economic impacts due to extended recovery times. The Storm-DEPART tool and analytical output enable more impactful data-driven decision-making capabilities and strengthen national-level disaster preparedness, response, and recovery. Storm-DEPART, built through multi-month collaboration between Entergy and INL, combines Entergy’s critical infrastructure inventory data with weather forecasts to predict damages to Electric utility’s assets due to natural disasters and the estimated recovery support needed, including time, materials, and resource allocation. In the event of an approaching hurricane, this innovative solution can assess potential damage to power generation capacity, transmission grids, distribution networks, and communications assets from wind bands, storm surge, and flooding. With more effective predictions, Entergy can more efficiently allocate resources to mitigate impacts and optimize recovery for customers. Storm-DEPART also allows Electric utilities the ability to apply a planning scenario and model expected damage to better inform infrastructure restoration needs leading to enhance system resiliency. The technology is fully transferrable to other electric utilities with the same damage estimating challenges. The INL team is working on the evolution of Storm-DEPART to include ice event damage prediction framework.

24 POWER TRANSMISSION AND DISTRIBUTION↗