Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Better Method to Calculate Fuel Burnup in Pebble Bed Reactors Using Machine Learning

Burnup measurement is an important step in material control and accountancy (MC&A) at nuclear reactors, and may be done by examining gamma spectra of fuel samples. Traditional approaches rely on known correlations to specific photopeaks (e.g. 137 Cs) and operate via a standard linear regression method. However, the quality of these regression methods is limited even in the best case, and is significantly poorer at short fuel cool-down times, due to the elevated radiation background by short life-time isotopes, and self-shielding effect of the fuel. For practical operation of pebble bed reactors (PBRs), quick measurements (in minutes) and short cooling times (in hours) are required from a safety and security perspective. We investigated the efficacy and performance of machine learning (ML) methods to predict the burnup of the pebble fuel from full gamma spectra (rather than specific discrete photopeaks) and found a full-spectrum ML approach to far outperform baseline regression predictions in all measurement and cooling conditions - including in operational-like measurement conditions. We also performed model and data ablation experiments to determine the relative performance impact of our ML methods' capacity to model data nonlinearities and the inherent additional information in full spectra. Applying our ML methods, we found a number of surprising results, including improved accuracy at shorter fuel cooling times (the opposite of the norm), remarkable robustness to spectrum compression (via rebinning), and competitive burnup predictions even when using background signal only (i.e. explicitly omitting known isotope photopeaks).

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

A Data-Driven Methodology for Contextual Unit Commitment Using Regression Residuals

Day after day, system operators are faced with the challenge of taking unit commitment (UC) decisions under uncertain net load conditions. The standard operating procedure for taking UC decisions begins by leveraging auxiliary data on covariates (such as the day of the week or latest weather information) to generate a point prediction for net load, which is used in solving a deterministic UC problem. Such an approach, however, is known to deliver a notoriously poor out-of-sample (OOS) performance, as it completely disregards the stochastic nature of net load. While stochastic programming models explicitly represent uncertainty, they mostly do so using a generic set of scenarios that neglect covariate observations, squandering useful auxiliary data that could be harnessed to glean insights into uncertainty. In this article, we discuss a contextual stochastic optimization approach to UC, which effectively exploits covariate observations while explicitly assessing uncertainty so as to boost the OOS performance of UC decisions. The key thrust of our approach is to leverage regression models, along with their empirical residuals, to set up and solve sample average approximation problems. Not only do we prove that our approach satisfies the requisite conditions for asymptotic optimality and consistency laid out in (Kannan et al., 2022), but we also assess its performance on several case studies conducted using real-world data collected in California ISO and New York ISO grids. In conclusion, results show that the proposed approach can significantly improve OOS performance compared to alternative methods proposed in the literature under varying dataset sizes.

Yurdakul, Ogun↗

Machine-learning-assisted high-temperature reservoir thermal energy storage optimization

High-temperature reservoir thermal energy storage (HT-RTES) has the potential to become an indispensable component in achieving the goal of the net-zero carbon economy, given its capability to balance the intermittent nature of renewable energy generation. In this study, a machine-learning-assisted computational framework is presented to co-optimize the performance metrics of HT-RTES by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Further, neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems.

15 GEOTHERMAL ENERGY↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimizing bioenergy biofuel harvest: a comparative analysis of stepwise and integrated methods for economic and environmental sustainability

Switchgrass is a promising bioenergy feedstock due to its high biomass yield potential, adaptability to marginal lands, and low carbon intensity for feedstock production. However, accurate cost estimation and assessment of greenhouse gas (GHG) emissions for the energy-intensive harvesting process are essential for evaluating the sustainability of bioenergy. This study provides a comparative analysis of two harvesting methods: the Stepwise Method, which separates operations into multiple stages, and the Integrated Method, which combines mowing and raking into a single pass. The analysis was conducted under four scenarios based on field sizes and biomass yields. Using three years of field-scale switchgrass harvest data from 125 sites, GHG emissions, energy consumption, and harvesting costs were quantified using the GREET model and techno-economic analysis. Additionally, regression analysis identified key climate and operational factors affecting fuel consumption. The Stepwise method was the most cost-effective for large fields with high biomass yield, achieving the lowest harvesting costs ($37.70 per ton). In contrast, the Integrated Method performed better in small fields and low-yield conditions, reducing GHG emissions by 9 % and energy use by 5 %. Regression analysis confirmed that a larger field size reduced fuel consumption, while higher biomass yield and longer operational time increased fuel use. Maximum temperature also contributed to a slight increase in fuel consumption. Furthermore, these results provide actionable insights for optimizing harvesting strategies based on field-specific conditions and operational goals, contributing to the economic and environmental sustainability of bioenergy production.

60 APPLIED LIFE SCIENCES↗

PSA 2025 Presentation: "Modeling and Sensitivity Analysis of a Generation IV Pebble Bed Reactor Using MELCOR 2.2"

Accompanying the advancement of reactor technologies is the need for computational modeling and simulation to predict their behavior under normal operating conditions and accident scenarios. New Generation IV reactor designs which employ non-conventional fuel have a particular need for modeling the behavior and release of radionuclides and other material from the fuel. In this work, MELCOR version 2.2, a system-level safety and accident scenario code developed by Sandia National Laboratories, was used to model a 200-MWth pebble bed modular reactor and calculate the inventories of circulating and deposited graphite, metal dust, and elemental components released from the fuel elements. A base case modeling the reactor under standard operating conditions was calculated using MELCOR and the inventories were extrapolated to 30 years of operation time using a logarithmic regression fit. A sensitivity analysis was also performed in which several key parameters for the base case model were modified to explore the effect of these changes on the inventories calculated by MELCOR. A set of transient scenario simulations for a depressurized loss of forced cooling (DLOFC) accident were also performed. The results of the sensitivity analysis and transient simulations are reported and discussed in relation to the modeling techniques used for this study.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Factors affecting powerhouse passage of spring migrant smolts at federally operated hydroelectric dams of the Snake and Columbia rivers

From 2008 to 2018, acoustic telemetry studies were conducted to evaluate dam passage survival of spring migrant Chinook salmon and steelhead smolts at seven of the eight federally operated dams on the lower Snake and Columbia rivers. Data from over 87 000 dam passage events were evaluated using regression modeling to identify the effect of spill operations, environmental conditions, and fish characteristics on powerhouse passage probability. In general, powerhouse passage was positively correlated with discharge, negatively correlated with forebay temperature and fish size, and higher for fish that passed the dam at night and for those that approached from the powerhouse side of the river, suggesting powerhouse passage is largely a function of smolt activity level and swimming ability. As such, spilling large volumes of water to reduce powerhouse passage is likely to be most effective during times of reduced activity and swimming ability (e.g., at night, high flows, and cold temperatures). This information can be used to develop dam- and time-specific spill operations that optimize smolt passage, power generation, and other competing demands, such as adult passage.

60 APPLIED LIFE SCIENCES↗

Data-Driven Mean-Corrected Recursive Estimation-Based Optimal DER Dispatch for Distribution System Voltage Control

Recent advances in smart inverters offer opportunities to mitigate adverse grid impacts caused by high penetrations of distributed photovoltaics (PV) in distribution grids, such as voltage violations. Here, this paper proposes a novel measurement-driven optimal power flow (OPF)-based distributed energy resource management system (DERMS) voltage regulation via recursive sensitivity estimation informed coordinated control of distributed PV inverters. The proposed approach leverages available grid and controllable DER measurements, eliminating reliance on system model information while being adaptive and robust to volatile operating conditions. A mean-corrected recursive ridge regression (MCRRR) algorithm is proposed for sensitivity estimation, continuously refining the sensitivity model through a closed-form solution. It effectively manages varying grid operating conditions, such as changes in power injections and topology reconfiguration, to facilitate a time-varying update of the Load Sensitivity Factors (LSF). The proposed approach is formulated as a linear programming (LP) problem and is thus scalable to larger-scale distribution systems. Its effectiveness and efficiency are demonstrated on a realistic distribution feeder with high PV penetrations in Southern California, USA.

14 SOLAR ENERGY↗

A DATA EFFICIENT SPARSE MODELING FRAMEWORK FOR POWER ESTIMATION IN WATER TREATMENT SENSING OPERATIONS

With increasing freshwater scarcity, advanced process design mechanisms such as Closed-Circuit Reverse Osmosis (CCRO) and Digital/Physical Twin systems are gaining traction in water treatment and reuse operations. While digital and physical twin models enable improved system insight and control, their development is often expensive and computationally intensive, requiring large volumes of synthetic or experimental data to characterize underlying process dynamics. This work introduces a sparse surrogate modeling framework to estimate power consumption from measured flow and pressure variables, along with their nonlinear polynomial and interaction expansions. To ensure model reliability and reduce overfitting, a two-stage pipeline is proposed. First, a dynamic data filtering algorithm is employed to remove uninformative observations and transient operational states. Second, a sparse penalized regression technique is applied to select a minimal set of parsimonious features. The proposed model achieves high sparsity, retaining only 7 out of 34 candidate features (≈79.41% sparsity) while delivering a root mean square error (RMSE) of 0.072 on the test dataset.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Hypothesis-Agnostic Network-Based Analysis of Real-World Data Suggests Ondansetron is Associated with Lower COVID-19 Any Cause Mortality

Background: The COVID-19 pandemic generated a massive amount of clinical data, which potentially hold yet undiscovered answers related to COVID-19 morbidity, mortality, long-term effects, and therapeutic solutions.Objectives: The objectives of this study were (1) to identify novel predictors of COVID-19 any cause mortality by employing artificial intelligence analytics on real-world data through a hypothesis-agnostic approach and (2) to determine if these effects are maintained after adjusting for potential confounders and to what degree they are moderated by other variables.Methods: A Bayesian statistics-based artificial intelligence data analytics tool (bAIcis®) within the Interrogative Biology® platform was used for Bayesian network learning and hypothesis generation to analyze 16,277 PCR+ patients from a database of 279,281 inpatients and outpatients tested for SARS-CoV-2 infection by antigen, antibody, or PCR methods during the first pandemic year in Central Florida. This approach generated Bayesian networks that enabled unbiased identification of significant predictors of any cause mortality for specific COVID-19 patient populations. These findings were further analyzed by logistic regression, regression by least absolute shrinkage and selection operator, and bootstrapping.Results: We found that in the COVID-19 PCR+ patient cohort, early use of the antiemetic agent ondansetron was associated with decreased any cause mortality 30 days post-PCR+ testing in mechanically ventilated patients.Conclusions: The results demonstrate how a real-world COVID-19-focused data analysis using artificial intelligence can generate unexpected yet valid insights that could possibly support clinical decision making and minimize the future loss of lives and resources.

60 APPLIED LIFE SCIENCES↗

Interpretable and flexible non-intrusive reduced-order models using reproducing kernel Hilbert spaces

This paper develops an interpretable, non-intrusive reduced-order modeling technique using regularized kernel interpolation. Existing non-intrusive approaches approximate the dynamics of a reduced-order model (ROM) by solving a data-driven least-squares regression problem for low-dimensional matrix operators. Our approach instead leverages regularized kernel interpolation, which yields an optimal approximation of the ROM dynamics from a user-defined reproducing kernel Hilbert space. We show that our kernel-based approach can produce interpretable ROMs whose structure mirrors full-order model structure by embedding judiciously chosen feature maps into the kernel. The approach is flexible and allows a combination of informed structure through feature maps and closure terms via more general nonlinear terms in the kernel. We also derive a computable a posteriori error bound that combines standard error estimates for intrusive projection-based ROMs and kernel interpolants. In conclusion, the approach is demonstrated in several numerical experiments that include comparisons to operator inference using both proper orthogonal decomposition and quadratic manifold dimension reduction.

Data-driven model reduction↗

A Data-Driven Exploration of the Impact of Renewable Energy on Inter-Area Oscillations in the U.S. Eastern Interconnection

As increasing amounts of renewable energy (RE) resources are incorporated into the bulk-power grid, power system oscillations are expected to change. This work investigates how RE generation impacts the frequency and damping ratio (DR) of two dominant inter-area modes in the U.S. Eastern Interconnection (EI) using regularly updated estimates collected over a 12-month period. Quantile regression is used to derive the correlation between operating conditions and mode properties, and a bootstrap method is used to quantify the uncertainty associated with the correlation estimates. Results show that with an increase in system load, the frequency of a mode decreases and DR increases. Evidence that increasing RE generation results in an increase in frequency and decline in DR was found for one of the two modes studied. This work shows that increasing RE levels will impact the properties of inter-area oscillations in the EI, but it does not indicate the presence of immediate threats to grid stability. The outlined approach can be used to periodically assess changing mode properties as RE levels continue to grow and flag stability concerns before they become serious reliability threats.

Inter-area oscillation, mode meters, quantile regr↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage↗

Mass Spectrometer Transient Analysis

This software implements a complete preprocessing pipeline for transient mass spectrometry (MS) data collected during TAP (Temporal Analysis of Products) experiments. It is designed to extract chemically meaningful fluxes from overlapping ion signals by applying a calibrated defragmentation matrix and solving the resulting linear system using non-negative least squares (NNLS) regression. The core script, preprocess_mass_spec.py, performs the following operations: Gain correction: Applies amplifier gain scalars derived from inert-packed calibration pulses to normalize signal intensities across AMUs and acquisition settings. Background subtraction: Removes experiment baselines using user-defined time windows, ensuring compatibility with slow-diffusing species and preventing negative values that would interfere with NNLS. Options to subtract before and after defragmentation. Defragmentation: Constructs a fragmentation matrix A from zeroth moments of calibration pulses (equal molar gas:inert mixtures) and solves Ax=b at each time point, where b is the raw MS signal and x is the estimated species flux. The matrix is normalized to inert signals and accounts for instrument-specific fragmentation behavior. Pulse-mode handling: Supports both averaged and individual pulse modes, enabling statistical treatment of fluxes and calculation of standard deviations. Integration and output: Computes zeroth moments (integrated fluxes) and exports time-resolved and integrated data in CSV format, suitable for downstream kinetic modeling. The software is validated using both virtual TAP simulations (VTAP) and experimental data from propane dehydrogenation (PDH) on CrOx/Al2O3 catalysts. It preserves temporal resolution by applying NNLS point-by-point across the pulse duration (typically 6,000+ time slices per pulse), leveraging the linear superposition principle to reconstruct full flux profiles. The defragmented outputs are compatible with kinetic extraction methods such as the G and Y procedures, which are used to derive rate–concentration relationships from TAP data. The details of these validations are discussed in detail in the supporting manuscript and supporting information. Example data and output files are also included. The methodology is robust to experimental noise and drift, with calibration protocols that account for pulse size effects, MS aging, and inert gas normalization. The software is modular, reproducible, and tailored for high-throughput TAP-MS workflows in catalysis research.

Kristy, Stephen [Idaho National Laboratory (INL), ↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Benefit Analysis of CO 2 Delivery Options for Offshore Storage or Enhanced Oil Recovery

The analysis presented in this report evaluates the benefits of CO₂ offshore transport via pipeline or ship within the GOM. It takes a top-down framework to estimate the costs. First, this analysis designed a reduced-order model (ROM) based on the cash flows in the FECM/NETL CO₂ Transport Cost Model (also known as CO2_T_COM). The ROM takes capital expenses (CAPEX) and operating expenses (OPEX) to calculate the CO₂ breakeven price based on the cash flows. Second, this analysis developed regression models utilizing published data from other analyses to estimate CAPEX and OPEX. Since the ROM is a simplified cash flow calculation, it is easy to exchange the core regression models to estimate various costs. The ROM and regression models provided a framework that can be easily used by other researchers, decision-makers, operators, and regulators. The objective of this analysis is to assess the CO₂ breakeven cost range for pipeline and ship transport of captured CO₂ given the CO₂ source and storage reservoir located in the GOM.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗