Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Data-Driven Methodology for Contextual Unit Commitment Using Regression Residuals

Day after day, system operators are faced with the challenge of taking unit commitment (UC) decisions under uncertain net load conditions. The standard operating procedure for taking UC decisions begins by leveraging auxiliary data on covariates (such as the day of the week or latest weather information) to generate a point prediction for net load, which is used in solving a deterministic UC problem. Such an approach, however, is known to deliver a notoriously poor out-of-sample (OOS) performance, as it completely disregards the stochastic nature of net load. While stochastic programming models explicitly represent uncertainty, they mostly do so using a generic set of scenarios that neglect covariate observations, squandering useful auxiliary data that could be harnessed to glean insights into uncertainty. In this article, we discuss a contextual stochastic optimization approach to UC, which effectively exploits covariate observations while explicitly assessing uncertainty so as to boost the OOS performance of UC decisions. The key thrust of our approach is to leverage regression models, along with their empirical residuals, to set up and solve sample average approximation problems. Not only do we prove that our approach satisfies the requisite conditions for asymptotic optimality and consistency laid out in (Kannan et al., 2022), but we also assess its performance on several case studies conducted using real-world data collected in California ISO and New York ISO grids. In conclusion, results show that the proposed approach can significantly improve OOS performance compared to alternative methods proposed in the literature under varying dataset sizes.

Yurdakul, Ogun↗

SISSO (Sure-independence-screening sparsifying-operator regressor)

Implementation of the SISSO regression algorithm in MATLAB. The SISSO regression algorithm iteratively selects model features from candidates, converging even when the number of possible features is much greater than the number of available data points. Includes a test script that validates the algorithm by replicating the results using data and procedure from https://analytics-toolkit.nomad-coe.eu/hub/user-redirect/notebooks/tutorials/compressed_sensing.ipynb. Python code from the SISSO regressor in 'sisso.py', from the link above, was used as the basis for developing the MATLAB implementation. The SISSO regression algorithm is detailed by the original authors in R. Ouyang, S. Curtarolo, E. Ahmetcik et al., Phys. Rev. Mater. 2, 083802 (2018), R. Ouyang, E. Ahmetcik, C. Carbogno, M. Scheffler, and L. M. Ghiringhelli, J. Phys.: Mater. 2, 024002 (2019).

Gasper, Paul↗

Machine-learning-assisted high-temperature reservoir thermal energy storage optimization

High-temperature reservoir thermal energy storage (HT-RTES) has the potential to become an indispensable component in achieving the goal of the net-zero carbon economy, given its capability to balance the intermittent nature of renewable energy generation. In this study, a machine-learning-assisted computational framework is presented to co-optimize the performance metrics of HT-RTES by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Further, neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems.

15 GEOTHERMAL ENERGY↗

Operational-based annual energy production uncertainty: are its components actually uncorrelated?

Calculations of annual energy production (AEP) from a wind power plant – whether based on preconstruction or operational data – are critical for wind plant financial transactions. The uncertainty in the AEP calculation is especially important in quantifying risk and is a key factor in determining financing terms. A popular industry practice is to assume that different uncertainty components within an AEP calculation are uncorrelated and can therefore be combined as the sum of their squares. We assess the practical validity of this assumption for operational-based uncertainty by performing operational AEP estimates for more than 470 wind plants in the United States, mostly in simple terrain. We apply a Monte Carlo approach to quantify uncertainty in five categories: revenue meter data, wind speed data, regression relationship between density-corrected wind speed (from reanalysis data) and measured wind power, length of long-term-correction data set, and future interannual variability. We identify correlations between categories by comparing the results across all 470 wind plants. We observe a positive correlation between interannual variability and the linearized long-term correction; a negative correlation between wind resource interannual variability and linear regression; and a positive correlation between reference wind speed uncertainty and linear regression. Then, we contrast total operational AEP uncertainty values calculated by omitting and considering correlations between the uncertainty components. We quantify that ignoring these correlations leads to an underestimation of total AEP uncertainty of, on average, 0.1% and as large as 0.5% for specific sites. Although these are not large increases, these would still impact wind plant financing rates; further, we expect these values to increase for wind plants in complex terrain. Based on these results, we conclude that correlations between the identified uncertainty components should be considered when computing the total AEP uncertainty.

17 WIND ENERGY↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Evaluating Cell Temperature Models and the Effect of Wind Speed in PV System Capacity Testing: Preprint

Capacity testing is a routine procedure for assessing a photovoltaic system's performance relative to expectations. The most common test method involves fitting a regression model that predicts system output power using operating weather conditions including wind speed. Structural modifications to the regression model to incorporate wind in different ways improved the model's ability to fit measured system performance, but the observed improvements were small and unlikely to change the result of a capacity test. However, the results showed that the choice of reporting wind speed and inclusion or exclusion of wind speed in the performance model used as the test benchmark can significantly change the test result.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Optimizing bioenergy biofuel harvest: a comparative analysis of stepwise and integrated methods for economic and environmental sustainability

Switchgrass is a promising bioenergy feedstock due to its high biomass yield potential, adaptability to marginal lands, and low carbon intensity for feedstock production. However, accurate cost estimation and assessment of greenhouse gas (GHG) emissions for the energy-intensive harvesting process are essential for evaluating the sustainability of bioenergy. This study provides a comparative analysis of two harvesting methods: the Stepwise Method, which separates operations into multiple stages, and the Integrated Method, which combines mowing and raking into a single pass. The analysis was conducted under four scenarios based on field sizes and biomass yields. Using three years of field-scale switchgrass harvest data from 125 sites, GHG emissions, energy consumption, and harvesting costs were quantified using the GREET model and techno-economic analysis. Additionally, regression analysis identified key climate and operational factors affecting fuel consumption. The Stepwise method was the most cost-effective for large fields with high biomass yield, achieving the lowest harvesting costs ($37.70 per ton). In contrast, the Integrated Method performed better in small fields and low-yield conditions, reducing GHG emissions by 9 % and energy use by 5 %. Regression analysis confirmed that a larger field size reduced fuel consumption, while higher biomass yield and longer operational time increased fuel use. Maximum temperature also contributed to a slight increase in fuel consumption. Furthermore, these results provide actionable insights for optimizing harvesting strategies based on field-specific conditions and operational goals, contributing to the economic and environmental sustainability of bioenergy production.

60 APPLIED LIFE SCIENCES↗

A Deep Neural Network for Accurate and Robust Prediction of the Glass Transition Temperature of Polyhydroxyalkanoate Homo- and Copolymers

The purpose of this study was to develop a data-driven machine learning model to predict the performance properties of polyhydroxyalkanoates (PHAs), a group of biosourced polyesters featuring excellent performance, to guide future design and synthesis experiments. A deep neural network (DNN) machine learning model was built for predicting the glass transition temperature, Tg, of PHA homo- and copolymers. Molecular fingerprints were used to capture the structural and atomic information of PHA monomers. The other input variables included the molecular weight, the polydispersity index, and the percentage of each monomer in the homo- and copolymers. The results indicate that the DNN model achieves high accuracy in estimation of the glass transition temperature of PHAs. In addition, the symmetry of the DNN model is ensured by incorporating symmetry data in the training process. The DNN model achieved better performance than the support vector machine (SVD), a nonlinear ML model and least absolute shrinkage and selection operator (LASSO), a sparse linear regression model. The relative importance of factors affecting the DNN model prediction were analyzed. Sensitivity of the DNN model, including strategies to deal with missing data, were also investigated. Compared with commonly used machine learning models incorporating quantitative structure–property (QSPR) relationships, it does not require an explicit descriptor selection step but shows a comparable performance. The machine learning model framework can be readily extended to predict other properties.

quantitative structure–property relationship (QSPR↗

PSA 2025 Presentation: "Modeling and Sensitivity Analysis of a Generation IV Pebble Bed Reactor Using MELCOR 2.2"

Accompanying the advancement of reactor technologies is the need for computational modeling and simulation to predict their behavior under normal operating conditions and accident scenarios. New Generation IV reactor designs which employ non-conventional fuel have a particular need for modeling the behavior and release of radionuclides and other material from the fuel. In this work, MELCOR version 2.2, a system-level safety and accident scenario code developed by Sandia National Laboratories, was used to model a 200-MWth pebble bed modular reactor and calculate the inventories of circulating and deposited graphite, metal dust, and elemental components released from the fuel elements. A base case modeling the reactor under standard operating conditions was calculated using MELCOR and the inventories were extrapolated to 30 years of operation time using a logarithmic regression fit. A sensitivity analysis was also performed in which several key parameters for the base case model were modified to explore the effect of these changes on the inventories calculated by MELCOR. A set of transient scenario simulations for a depressurized loss of forced cooling (DLOFC) accident were also performed. The results of the sensitivity analysis and transient simulations are reported and discussed in relation to the modeling techniques used for this study.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Factors affecting powerhouse passage of spring migrant smolts at federally operated hydroelectric dams of the Snake and Columbia rivers

From 2008 to 2018, acoustic telemetry studies were conducted to evaluate dam passage survival of spring migrant Chinook salmon and steelhead smolts at seven of the eight federally operated dams on the lower Snake and Columbia rivers. Data from over 87 000 dam passage events were evaluated using regression modeling to identify the effect of spill operations, environmental conditions, and fish characteristics on powerhouse passage probability. In general, powerhouse passage was positively correlated with discharge, negatively correlated with forebay temperature and fish size, and higher for fish that passed the dam at night and for those that approached from the powerhouse side of the river, suggesting powerhouse passage is largely a function of smolt activity level and swimming ability. As such, spilling large volumes of water to reduce powerhouse passage is likely to be most effective during times of reduced activity and swimming ability (e.g., at night, high flows, and cold temperatures). This information can be used to develop dam- and time-specific spill operations that optimize smolt passage, power generation, and other competing demands, such as adult passage.

60 APPLIED LIFE SCIENCES↗

BOOSTR: A Dataset for Accelerator Control Systems

The Booster Operation Optimization Sequential Time-series for Regression (BOOSTR) dataset was created to provide a cycle-by-cycle time series of readings and settings from instruments and controllable devices of the Booster, Fermilab’s Rapid-Cycling Synchrotron (RCS) operating at 15 Hz. BOOSTR provides a time series from 55 device readings and settings that pertain most directly to the high-precision regulation of the Booster’s gradient magnet power supply (GMPS). To our knowledge, this is one of the first well-documented datasets of accelerator device parameters made publicly available. We are releasing it in the hopes that it can be used to demonstrate aspects of artificial intelligence for advanced control systems, such as reinforcement learning and autonomous anomaly detection.

Kafkes, Diana (ORCID:000000021716463X)↗

BOOSTR: A Dataset for Accelerator Control Systems

The Booster Operation Optimization Sequential Time-series for Regression (BOOSTR) dataset was created to provide a cycle-by-cycle time series of readings and settings from instruments and controllable devices of the Booster, Fermilab's Rapid-Cycling Synchrotron (RCS) operating at 15~Hz. BOOSTR provides a time series from 55 device readings and settings that pertain most directly to the high-precision regulation of the Booster's gradient magnet power supply (GMPS). To our knowledge, this is one of the first well-documented datasets of accelerator device parameters made publicly available. We are releasing it in the hopes that it can be used to demonstrate aspects of artificial intelligence for advanced control systems, such as reinforcement learning and autonomous anomaly detection.

43 PARTICLE ACCELERATORS↗

Data-Driven Mean-Corrected Recursive Estimation-Based Optimal DER Dispatch for Distribution System Voltage Control

Recent advances in smart inverters offer opportunities to mitigate adverse grid impacts caused by high penetrations of distributed photovoltaics (PV) in distribution grids, such as voltage violations. Here, this paper proposes a novel measurement-driven optimal power flow (OPF)-based distributed energy resource management system (DERMS) voltage regulation via recursive sensitivity estimation informed coordinated control of distributed PV inverters. The proposed approach leverages available grid and controllable DER measurements, eliminating reliance on system model information while being adaptive and robust to volatile operating conditions. A mean-corrected recursive ridge regression (MCRRR) algorithm is proposed for sensitivity estimation, continuously refining the sensitivity model through a closed-form solution. It effectively manages varying grid operating conditions, such as changes in power injections and topology reconfiguration, to facilitate a time-varying update of the Load Sensitivity Factors (LSF). The proposed approach is formulated as a linear programming (LP) problem and is thus scalable to larger-scale distribution systems. Its effectiveness and efficiency are demonstrated on a realistic distribution feeder with high PV penetrations in Southern California, USA.

14 SOLAR ENERGY↗

Multisublattice cluster expansion study of short-range ordering in iron-substituted strontium titanate

Owing to the challenges in obtaining realistic atomic configurations in large chemical phase spaces, it is not straightforward to describe structure–property relations in materials exhibiting configurational disorder. One example is iron-substituted strontium titanate (SrTi 1–x Fe x O 3–d , STF), a promising perovskite-derivative cathode material in solid oxide fuel cells that exhibits full solid solubility 0 ≤ x ≤ 1 and a tendency to exhibit short-range order. Here we demonstrate a multisublattice cluster expansion (CE) framework and apply it to STF across the full composition range. The CE approach is distinct from more traditional CE formulations in that clusters are defined explicitly by the chemical species distributed among multiple sublattices, rather than via cluster functions of occupation variables with decoration. The modified CE approach makes it easy to distinguish meaningful chemical interactions that are harder to extract from conventional CE, since for the latter chemical identity in a cluster is expressed as a product of site occupations. The least absolute shrinkage and selection operator (LASSO) is implemented as a regression analysis tool to select key clusters and avoid overfitting. We demonstrate this formulation on STF, and show that it can accurately predict configurational energies in comparison to conventional CE. From the key clusters, we identify that short-range ordering between substitutional Fe and oxygen vacancies (V O ) results in the formation of Fe–VO strings. In addition, we consider the stability of STF through CE-based Monte Carlo (MC) simulations and confirm the presence of superstructures that were previously observed in transmission electron microscopy. In this work, analysis of atomic configurations from MC samples reveals variations in the oxidation state of Fe atoms, which can be explained by the ordering tendency of Fe and V O . The cluster description and selection formalism described here may be applied to other disordered multisublattice systems for accurate and efficient material modeling.

36 MATERIALS SCIENCE↗

A DATA EFFICIENT SPARSE MODELING FRAMEWORK FOR POWER ESTIMATION IN WATER TREATMENT SENSING OPERATIONS

With increasing freshwater scarcity, advanced process design mechanisms such as Closed-Circuit Reverse Osmosis (CCRO) and Digital/Physical Twin systems are gaining traction in water treatment and reuse operations. While digital and physical twin models enable improved system insight and control, their development is often expensive and computationally intensive, requiring large volumes of synthetic or experimental data to characterize underlying process dynamics. This work introduces a sparse surrogate modeling framework to estimate power consumption from measured flow and pressure variables, along with their nonlinear polynomial and interaction expansions. To ensure model reliability and reduce overfitting, a two-stage pipeline is proposed. First, a dynamic data filtering algorithm is employed to remove uninformative observations and transient operational states. Second, a sparse penalized regression technique is applied to select a minimal set of parsimonious features. The proposed model achieves high sparsity, retaining only 7 out of 34 candidate features (≈79.41% sparsity) while delivering a root mean square error (RMSE) of 0.072 on the test dataset.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Non-intrusive data-driven model reduction for differential–algebraic equations derived from lifting transformations

In this paper we present a non-intrusive data-driven approach for model reduction of nonlinear systems. The approach considers the particular case of nonlinear partial differential equations (PDEs) that form systems of partial differential–algebraic equations (PDAEs) when lifted to polynomial form. Such systems arise, for example, when the governing equations include Arrhenius reaction terms (e.g., in reacting flow models) and thermodynamic terms (e.g., the Helmholtz free energy terms in a phase-field solidification model). Using the known structured form of the lifted algebraic equations, the approach computes the reduced operators for the algebraic equations explicitly, using straightforward linear algebra operations on the basis matrices. The reduced operators for the differential equations are inferred from lifted snapshot data using operator inference, which solves a linear least squares regression problem. The approach is illustrated for the nonlinear model of solidification of a pure material. The lifting transformations reformulate the solidification PDEs as a system of PDAEs that have cubic structure. The operators of the lifted system for this solidification example have affine dependence on key process parameters, permitting us to learn a parametric reduced model with operator inference. Numerical experiments show the effectiveness of the resulting reduced models in capturing key aspects of the solidification dynamics.

42 ENGINEERING↗

Intelligent Prediction of States in Multi-port Autonomous Reconfigurable Solar power plant (MARS)

In power electronics, prediction of states may be used for identification of faults, determination of aging of components, identification of bad data measurements, among others. Prediction of states in power electronics have broadly been based on: (a) physics-based models, (b) data-driven models, and (c) hybrid models. In this paper, data-driven approaches are presented for intelligent prediction of states in multi-port autonomous reconfigurable solar power plant (MARS) and compared. The data-set needed to train the data-driven models based on artificial intelligence (AI) algorithms has been identified and the trained models are evaluated under different extrapolated normal and abnormal operating conditions. The AI algorithms include nonlinear auto-regressive exogenous model (NARX), spiking neural networks (SNN), and decision tree. The models are compared and contrasted. The best model (NARX) is evaluated under different normal and abnormal operating conditions that have indicated accurate prediction.

Debnath, Suman↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗