Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Solar Forecasting, Net Load Forecasting, and Data-Driven Distributed Solar Visibility Prizes (Final Technical Report)

The American-Made Solar Forecasting Prize, Net Load Forecasting Prize, and Data-Driven Distribution (3D) Solar Visibility Prize is a multimillion-dollar prize competition designed to energize U.S. solar innovation through a series of contests that accelerate the entrepreneurial process from years to months. The activities incentivized by these three prizes will support the governmentwide approach to increase American energy dominance by promoting innovation and early deployment of energy technologies, resulting in wider adoption, which is critical for secure, affordable, and reliable solar energy.

14 SOLAR ENERGY↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

APSO-enhanced algebraic derivative estimation approach for real-time traffic flow prediction on critical road sections during wildfire evacuation

In rapid-onset disaster scenarios such as wildfires, evacuation traffic often significantly deviates from historical patterns, rendering conventional data-driven forecasting methods less effective. To address this challenge, we propose an improved algebraic derivative estimation (ADE) incorporating particle swarm optimization (PSO) for real-time traffic flow prediction. Our approach dynamically adjusts the ADE prediction time window at each step by minimizing a cost function based on the mean and variance of accumulated forecasting errors within the window, thereby balancing bias and variability. We evaluate the method using traffic data from the January 2025 California wildfires, focusing on key road segments critical for large-scale evacuations. The results demonstrate that our approach surpasses established machine learning and deep learning models—XGBoost, LSTM, and GRU—in predictive accuracy and maintains high computational efficiency. Notably, the proposed method eliminates the need for offline model training. Moreover, rapid PSO-based tuning enables real-time deployment, which provides a crucial advantage in scenarios where evacuation timings and road closures change dynamically. In conclusion, these findings highlight the benefits of the PSO-enhanced ADE framework for emergency traffic management, where rapid, data-sparse forecasts are essential for effective evacuation planning.

Algebraic derivative estimation↗

Predictive Model for Starlink Maritime Performance Using Multi-Horizon RandomForest

Low Earth orbit (LEO) satellite systems have become a crucial enabler of broadband access for maritime industries, where traditional networks are unavailable. However, the high mobility of LEO constellations and constantly changing weather conditions result in unpredictable link fluctuations, limiting the ability of maritime platforms to plan bandwidth usage proactively. To the best of our knowledge, no prior work has developed a short-term predictive model for maritime LEO connectivity using real experimental field measurements. This paper proposes a data-driven forecasting model that predicts future downlink throughput using multi-horizon RandomForest regression. The model is trained using real experimental coastal measurement data incorporating recent throughput history, network-layer indicators, and environmental variables. The proposed approach reduces mean absolute error by approximately 31% compared to a persistence baseline for 15-minute horizons. It maintains a measurable improvement at 30 minutes, despite increased stochasticity. These findings confirm that proactive bandwidth awareness is feasible on maritime platforms and can effectively support operational decisions such as adaptive streaming, routing, and resource scheduling. The performance gap between forecasting horizons also highlights the need for expanded offshore datasets to improve prediction robustness under harsher maritime environments.

97 MATHEMATICS AND COMPUTING↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Enhancing Short-Range Weather Forecasts through Temporal Variation Encoding: A Multiperiod Embedding Approach

Machine learning (ML) techniques have emerged as promising approaches to improve regional weather forecast accuracy and reliability through data-driven methods. We propose a novel ML-based weather forecasting model, the Multiperiod Embed Net (MPENet). A key distinguishing feature of MPENet is its explicit utilization of the inherent cyclic nature in weather dynamics, unlike the autoregressive strategies commonly used in other ML weather forecasting approaches. Critical cyclic structures are identified via Fourier analyses of dynamic time series. Cyclicity in the convolutional representation is achieved by transforming one-dimensional time series of meteorological variables into two-dimensional tensors based on identified periods. This approach enables the model to leverage intrinsic weather patterns, enhancing regional forecast performance. To demonstrate the effectiveness of MPENet, we conduct a comparative analysis with Nvidia’s FourCastNet. Both models are trained on High-Resolution Rapid Refresh (HRRR) data from 2015 to 2022, over a 192 km × 192 km region in Tennessee. The comparisons are performed locally at two specific locations known to have different weather dynamics due to orographic effects: Crossville, on the relatively flat Cumberland Plateau with fewer topographic airflow disruptions, and Oak Ridge, in the ridge-and-valley region, where airflow is heavily influenced by surrounding valleys and mountains. Our results indicate that FourCastNet achieves strong accuracy at very short lead times, while MPENet maintains competitive skill and shows advantages in capturing temporal evolution over longer periods. Cross-correlation analyses of MPENet and FourCastNet predictions with the HRRR data suggest that encoding critical cyclicity into the network architecture leads to improvements in the forecasting skill.

Artificial intelligence↗

Feedforward-feedback ammonia control at a water resource recovery facility based on a digital twin with hybrid model

Ammonia-based aeration control (ABAC) at full-scale Water Resource Recovery Facilities (WRRFs) can be challenged by diurnal loading and transport delays. This work addressed these challenges using a hybrid feedforward–feedback controller built on Activated Sludge Model 1 (ASM1), marking the first full-scale deployment to pair a mechanistic feedforward core with data-driven corrections. The objectives were to improve ammonia setpoint tracking, assess performance of the mechanistic model when enhanced with data-driven corrections, and document full-scale operation. The hybrid model incorporates two data-driven components: (1) a Mechanistic Error Forecasting Engine (MEFE), consisting of a multivariate linear regressor and a long short-term memory (LSTM) ensemble. Defying expectations, low-parameter models outperformed more complex alternatives, reducing the mechanistic error by 71%. (2) A Residual Oscillation Forecasting Engine (ROFE), based on Fast Fourier Transform, reduced the remaining error by another 35%. Two proportional–integral (PI) feedback loops further (i) trim the feedforward output and (ii) eliminate residual controller error in the final aerobic zone. In full-scale operation, the controller reduced mean-squared error (MSE) by 94% over the baseline and produced more stable dissolved oxygen (DO) setpoints. Overall, it was proven that layering multi-timescale data-driven models on a mechanistic core can yield reliable ABAC performance at WRRFs.

54 ENVIRONMENTAL SCIENCES↗

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse↗

Digital Twin User Guide for Chelan County Public Utility District

This user manual offers a comprehensive guide for developing a Digital twin (DT) of a Kaplan turbine at Chelan County Public Utility District (Chelan PUD) using neural networks. As variable renewable generation expands, hydropower units must operate with optimal efficiency and stability. For Kaplan machines, this flexibility is achieved through coordinated control of guide vane (wicket gates) opening and runner blade pitch, which amplifies the plant’s inherent nonlinear behavior and challenges traditional physics-only modeling. The efficiency of the Kaplan turbine varies with different combinations of the guide vans (wicket gate) opening and the blade angle. Each guide van opening and blade angle has a corresponding highest efficiency point, forming a cam relationship that represents the optimal combination.The discharge of a hydraulic turbine is controlled by the opening angle of the guide vans. Therefore, for each value of head, there is a certain guide van opening and blade angle that corresponds to the highest efficiency. For a given head, different combinations of the guide van opening and blade angle have different efficiencies. Therefore, coordinate cam curves are used to describe the relationship between the wicket gate opening and blade angle with different water head. To address these challenges, the manual details a data-driven modeling and learning workflow centered on structured neural networks. The approach is designed to forecast critical operational variables—discharge flow, net head, penstock (or scroll-case) pressure, and generator electrical outputs—by leveraging real-time inputs such as the generator power control setpoint, exciter field current and field voltage, together with hydromechanical commands (e.g., gate position and, when available, runner blade-pitch angle). The neural models are trained and validated on operational data from a Kaplan unit operated by Chelan PUD, demonstrating that the structured NN architecture can learn the coupled gate–blade–electrical dynamics. The result is a robust DT that improves situational awareness and supports data-informed decision-making for Chelan PUD’s Kaplan turbine operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing shellfish water exposure to fecal bacteria pollution in Salish Sea: three-dimensional modeling and implications for monitoring

Fecal bacteria (FB) contamination poses significant risks to shellfish safety and management in coastal and estuarine waters. Despite extensive pollution identification and correction efforts, FB contamination in shellfish-growing areas persists in the Salish Sea, highlighting the need to identify overlooked sources and better understand FB transport from riverine and shoreline inputs to shellfish beds. To address this, a high-resolution three-dimensional hydrodynamic model coupled with FB kinetics was developed and applied to a case study site in Salish Sea—Portage Bay—to simulate freshwater plume circulation, flushing dynamics, and bacterial transport. Daily FB loading from the major freshwater inflow—Nooksack River was generated by both linear interpolation and integrating a machine learning approach (XGBoost), trained on historical hydrological and meteorological data. The model successfully reproduced both the magnitude and seasonal variation of FB concentrations in Portage Bay for the year of 2021, demonstrating that simplified FB kinetics with first-order decay due to mortality was effective in this dynamic coastal environment with short flushing time. Model results identified the Nooksack River as the dominant far-field FB source, while scenario simulations showed that near-field coastal stormwater outfalls elevated local FB levels following rainfall, particularly under low-flow conditions. The XGBoost prediction provided comparable or superior accuracy to linear interpolation, particularly during periods of missing observational data, by capturing short-term variability and event-driven loading more effectively. Integrating data-driven riverine FB inputs with mechanistic coastal numerical modeling provides a robust framework for operational forecasting of shellfish bed exposure risk and supports adaptive monitoring and management of shellfish growing areas in the Salish Sea and similar coastal systems.

Salish Sea↗

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

Latent Twins

Over the past decade, scientific machine learning has transformed the development of mathematical and computational frameworks for analyzing, modeling, and predicting complex systems. From inverse problems to numerical partial differential equations (PDEs), dynamical systems, and model reduction, these advances have pushed the boundaries of what can be simulated. Yet they have often progressed in parallel, with representation learning and algorithmic solution methods evolving largely as separate pipelines. With Latent Twins, we propose a unifying mathematical framework that creates a hidden surrogate in latent space for the underlying equations. Whereas digital twins mirror physical systems in the digital world, Latent Twins mirror mathematical systems in a learned latent space governed by operators. Through this lens, classical modeling, inversion, model reduction, and operator approximation all emerge as special cases of a single principle. We establish the fundamental approximation properties of Latent Twins for both ordinary differential equations (ODEs) and PDEs and demonstrate the framework across three representative settings: (i) canonical ODEs, capturing diverse dynamical regimes; (ii) a PDE benchmark using the shallow-water equations, contrasting Latent Twin simulations with deep operator network and forecasts with a four-dimensional variational method baseline; and (iii) a challenging real-data geopotential reanalysis dataset, reconstructing and forecasting from sparse, noisy observations. Latent Twins provide a compact, interpretable surrogate for solution operators that evaluate across arbitrary time gaps in a single-shot, while remaining compatible with scientific pipelines such as assimilation, control, and uncertainty quantification. Looking forward, this framework offers scalable, theory-grounded surrogates that bridge data-driven representation learning and classical scientific modeling across disciplines.

Latent Twins↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Physics-Informed Neural Network (PINN) Prediction of Mixed Mass-Heat-Crystallization Limited Methane Hydrate Formation and Dissociation in Micro-Confinement

The creation and use of Physics-Informed Neural Networks (PINNs) for simulating the dynamics of methane hydrate formation and dissociation will be presented. The PINN framework's main benefit is its capacity to impose physical consistency with only a partial comprehension of the governing equations. This makes the algorithm especially useful for systems with little experimental evidence or a lack of theoretical knowledge. A strong basis for forecasting methane hydrate behavior over the verified operating ranges of 30.0-80.9 bar pressure and 1.0-4.0 K sub-cooling conditions is provided by the combination of conductive heat transfer equations and mixed mass-transfer–crystallization kinetics. PINNs were more accurate at predicting the mixed mass-heat-crystallization limited kinetics than conventional Artificial Neural Networks (ANNs), demonstrating remarkable predictive accuracy for methane hydrate production over the ANN model. The efficiency of incorporating physical limitations from first principles into machine learning frameworks for methane hydrate crystallizations is reinforced by these findings. For hydrate-related applications in energy generation, carbon sequestration, and climate modelling, our study establishes PINNs as a computational tool that is both scalable and efficient. The proven capacity to close the gap between conventional physics-based simulations and solely data-driven models creates new opportunities for expedited hydrate research and practical applications.

Hartman, Ryan L [NYU Tandon School of Engineering]↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Forecasting Battery Electrode Performance via Electrochemical Fluorescence Microscopy and Machine-Learning

Predicting lithium-ion battery performance is hindered by microscale electrode heterogeneities invisible to conventional diagnostics. Here, we combine electrochemical fluorescence microscopy (EFM), which maps electronic connectivity by visualizing an electrofluorophore reaction distribution, with a multitask ElasticNet regression to forecast discharge capacity from spatial heterogeneity. Analyzing 196 images from six pilot-scale LiNi 0.5 Mn 0.3 Co 0.2 O 2 cathodes with varying carbon loadings, we extract 62 descriptors that capture morphology and texture. A compact five-feature model predicts capacity across eight discharge rates, achieving a per-target R 2 of up to 0.63 and an overall R 2 of 0.92, with a mean absolute percentage error of less than 2%. This performance rivals impedance-based approaches while avoiding their reliance on postformation data and incomplete electronic network information. Our facile and rapid, image-driven method may enable electrode quality control upstream of costly cell assembly to offer a transformative tool for data-driven battery research and manufacturing.

battery electrodes↗

Dark Energy Survey Year 3 results: Simulation-based 𝑤CDM inference from weak lensing and galaxy clustering maps with deep learning: Analysis design

Data-driven approaches using deep learning are emerging as powerful techniques to extract non-Gaussian information from cosmological large-scale structure. Here, this work presents the first simulation-based inference (SBI) pipeline that combines weak lensing and galaxy clustering maps in a realistic Dark Energy Survey Year 3 (DES Y3) configuration and serves as preparation for a forthcoming analysis of the survey data. We develop a scalable forward model based on the CosmoGridV1 suite of N-body simulations to generate over one million self-consistent mock realizations of DES Y3 at the map level. Leveraging this large dataset, we train deep graph convolutional neural networks on the full survey footprint in spherical geometry to learn low-dimensional features that approximately maximize mutual information with target parameters. These learned compressions enable neural density estimation of the implicit likelihood via normalizing flows in a ten-dimensional parameter space spanning cosmological 𝑤CDM, intrinsic alignment, and linear galaxy bias parameters, while marginalizing over baryonic, photometric redshift, and shear bias nuisances. To ensure robustness, we extensively validate our inference pipeline using synthetic observations derived from both systematic contaminations in our forward model and independent Buzzard galaxy catalogs. Our forecasts yield significant improvements in cosmological parameter constraints, achieving 2−3× higher figures of merit in the 𝛺 𝑚 − 𝑆 8 plane relative to our implementation of baseline two-point statistics and effectively breaking parameter degeneracies through probe combination. These results demonstrate the potential of SBI analyses powered by deep learning for upcoming Stage-IV wide-field imaging surveys.

Thomsen, A. [Zurich, ETH] (ORCID:0000000203099021)↗