Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Measuring Success: A Refined Methodology for Estimating Long-term Continuous Improvement

Successful resource management systems require current and detailed feedback on operational and corporate-level performance. As corporate accountability concerns intensify, the precision and reliability of these performance metrics have become crucial. Traditional savings estimation methods can be difficult to understand, particularly linear regression, and can provide varying results. This paper reviews common efficiency metrics and highlights underlying mathematical inconsistencies when estimating total and percent savings with current methods. A refined approach to calculating long-term utility savings is proposed that simplifies current methodologies utilizing ratios to define an adjusted baseline, allowing for consistent and fair aggregation of results across multiple scales from resources to corporate performance. A simplified example demonstrates how the proposed methodology improves upon existing methods, especially in intermediate years. This paper’s major contributions are the simplified approach for converting modeled utility usage into estimated savings and the consistent roll-up methodology enabling more comparable, aggregable, and actionable results across scales.

Price, Chris [ORNL] (ORCID:0000000202007906)↗

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

Smart Methane Emission Detection System Development (Final Report)

Working with the Department of Energy’s National Energy Technology Laboratory, Southwest Research Institute® (SwRI®) developed a system to identify methane leaks reliably, accurately, and autonomously at critical midstream sections of the natural gas distribution network in real- time for the purpose of mitigating methane emissions using Optical Gas Imaging (OGI) cameras. SwRI’s Smart Leak Detection – Methane (SLED/M) adds a high degree of automation to the process of methane leak detection to minimize sources of human error, minimize response time to a leak event, and maximize midstream visibility. Furthermore, SwRI has been working towards integrating Quantitative OGI (QOGI) capabilities into this existing technology. By leveraging Deep Learning, SwRI now has the capability to estimate fugitive emission leak rates quickly and reliably, which allows operators to detect emissions, quantify leak rate, prioritize repairs, and validate the repairs in a single instrument. The next generation QOGI technology leverages the same cameras used in Leak Detection and Repair (LDAR) programs, with improvements in safety and speed for traditional quantification-based repairs, ultimately leading to less overhead cost for the operators. The goals for this research were to develop two types of models with the following goals: 1. Run in real-time on the edge (≥ 12 Hz) 2. Classification: Achieve less than 5% false positive detection 3. Classification: Achieve ≥ 95% methane plume detection rate 4. Regression: achieve ≤ 10 standard cubic feet per hour (scfh) prediction > 70% of the time In order to achieve these results, multiple infrared (IR) and other sensors were investigated in tandem with the midwave IR (MWIR) OGI to provide additional information to train the underlying models. Information on atmospheric conditions including humidity, temperature, pressure, and solar radiation was provided by a weather station. Several machine learning and deep learning architectures and methods, including looking at quantized classification networks and regressions networks, were explored. As further data was collected, curated, and labeled, it allowed for more refined regressive networks to be adequately trained, leading to better insight into the true flow rates being observed. An important valuable deliverable of this research effort was the development of an advanced network which underwent multiple iterations capable of giving a continuous output. The current network has a predicted mean average percentage error (MAPE) of 12.3% just outside our target goal of 10.00%, but an accuracy of 97.78% at ±50 scfh, well within the overall goal for the Department of Energy (DOE) program. Upon closer inspection, it was observed that more than 10% of datapoints contributing to the MAPE predictions were the result of low flow rate predictions and are beyond the sensitivity of instrument measurement as a result of normal operational variation and noise.

03 NATURAL GAS↗

mvBayesR

SAND2025-11559O The mvBayesR tool performs multivariate Bayesian analysis on generic data. It includes tools for regression modeling, diagnosis, basis decomposition, sensitivity analysis, and visualization. The tool compiles state-of-the-art methodology into one easy-to-use package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James [Sandia National Lab. (SNL-CA), Live↗

Loss Factors for Small Distributed Wind Turbines Based on Field Data in the United States

While wind energy production loss due to unavailability, environmental impacts, curtailment, and other causes has been studied and characterized at the utility-scale wind farm level, observation-based characterization of project loss is lacking for distributed wind energy, particularly for projects involving small wind turbines. Contemporary tools and research that support pre-construction distributed wind energy characterization present a wide range of default loss factors to convert gross energy estimates to net: 7-18%. We hypothesize that we can use generation observations from operational distributed wind projects to develop more accurate representations of loss. Using a density-based filtering technique on distributed wind power generation timeseries, we determine periods of typical performance and use them with regression algorithms in a measure-correlate-predict fashion to simulate what the generation would have been during periods of atypical or unreported performance. From there, the actual versus predicted generation leads to the establishment of observation-informed loss factors (median = 17%) for small, single turbine installation distributed wind projects.

17 WIND ENERGY↗

EVALUATION OF HRA METHODOLOGIES FOR APPLICATION IN SDP WORK

This study critically evaluates human reliability analysis (HRA) methodologies applicable to regulatory probabilistic safety assessment (PSA) model, with a particular focus on their role in supporting the significance determination process (SDP) in nuclear safety assessment. Firstly, three widely utilized HRA methods – IDHEAS-ECA, SPAR-H, and ASEP/THERP – were qualitatively and quantitatively assessed. Qualitative assessments were conducted using attributes from the NEA/CSNI/R(2015)1 report, while quantitative evaluations employed regression and correlation analyses to compare predicted human error probabilities (HEPs) against empirical data. Results reveal distinct strengths, for example, IDHEAS-ECA’s robust predictive accuracy and K-HRA’s alignment with operational practices. In addition, dependency analysis and recovery analysis were critically evaluated. For dependency analysis, the methods’ handling of inter-task dependencies and their impact on HEPs were examined, while recovery analysis highlighted strategies for mitigating failure events. Furthermore, strategies were proposed to evaluate performance-shaping factors under conditions of reduced human performance, such as stress, fatigue, or cognitive overload, addressing specific challenges faced in SDP evaluations. Human errors from KINS’s operational performance information system event reports were evaluated as a case study. This study identifies gaps and provides actionable insights to ensure their validity and applicability in SDP HRA applications. This paper is a part of research conducted by KINS, and it should be noted that this result does not represent the regulatory position of KINS.

99 - GENERAL AND MISCELLANEOUS↗

Frequency Nadir Constrained Unit Commitment for High Renewable Penetration Island Power Systems

The process of energy decarbonization in island power systems is accelerated due to the swift integration of inverter-based renewable energy resources (IBRs). The unique features of such systems, including rapid frequency changes resulting from potential generation outages or imbalances due to the unpredictability of renewable power, pose a significant challenge in maintaining the frequency nadir without external support. This paper presents a unit commitment (UC) model with data-driven frequency nadir constraints, including either frequency nadir or minimum inertia requirements, helping to limit frequency deviations after significant generator outages. The constraints are formulated using a linear regression model that takes advantage of real-world, year-long generation scheduling and dynamic simulation data. The efficacy of the proposed UC model is verified through a year-long simulation in an actual island power system using historical weather data. The alternative minimum inertia constraint, derived from actual system operation assumptions, is also evaluated. Findings demonstrate that the proposed frequency nadir constraint notably improves the system's frequency nadir under high photovoltaic (PV) penetration levels, albeit with a slight increase in generation costs, when compared to the alternative minimum inertia constraint.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Convolutional neural networks for intra-hour solar forecasting based on sky image sequences

Accurate and timely solar forecasts play an increasingly critical role in power systems. Compared to longer forecasting timescales, very short-term solar forecasting has lagged behind in both research and practice. In this paper, we propose deep convolutional neural networks (CNNs) to provide operational intra-hour (10-minute-ahead to 60-minute-ahead) solar forecasts. We develop two CNN structures inspired by a widely-used CNN architecture. The CNNs are tailored to our solar forecasting regression tasks and rely solely on sky image sequences. Case studies based on six years of data (over 150,000 data points) demonstrate that the best CNN model has forecast skill scores of 20%-39% over the naive persistence of cloudiness benchmark, even at these very short timescales. The CNNs also have consistently superior performance when compared to shallow machine learning models with meteorological predictors, where the improvement averages around 7%. The sensitivity analyses show that the sky image length, resolution, and weather conditions have impacts on the deep learning model accuracy. In our intra-hour problem with specific setups, two sky images with a 10-minute 128 x 128 resolution yield the most accurate forecasts. Current limitations, future work, and deployment challenges and solutions are also discussed.

14 SOLAR ENERGY↗

Evaluating Recursive Blind Forecast Against API and Baseline: A Puerto Rican Case Study on Solar Irradiance for Normal and Extreme Weather

This paper leverages ongoing work in a community microgrid in Adjuntas, Puerto Rico to forecast global horizontal irradiance (GHI) and compare performance in normal and extreme weather. Given a positive correlation of 0.98 between GHI and PV power, forecasting GHI can be an effective, indirect forecast of photovoltaic (PV) power, especially in microgrids where the end-users, owners, operators, or other stakeholders are reluctant to share data for training or validation due to privacy and security concerns. A recursive one-shot (termed as "blind") forecast is, hence, formulated, wherein a gradient-boosted regression tree (GBR) is built to forecast GHI for a 7-day horizon in normal weather, and a 2-day horizon in extreme weather. To demonstrate its resilience, the architecture is trained on normal and hurricane weather GHI from 2002-2022. It is generalized on February 9-16, 2023, and on the landfall of Hurricane Nicole (Nov 4-5, 2022), respectively. Forecasts from GBR are compared against that from a satellite-based API resource and three baselines: persistence, averaging, and exponential smoothing. Results show GBR and persistence outperform sophisticated API in both types of weather for this case study.

Sundararajan, Aditya↗

Assessment of small mechanical wastewater treatment plants: Relative life cycle environmental impacts of construction and operations

Many slow growing and shrinking rural communities struggle with aging or inadequate wastewater treatment plants (WWTPs), and face challenges in constructing and operating such facilities. Although existing literature has provided insight into the environmental sustainability of large facilities, including both the construction and operational phases, these studies have not examined small, rural facilities treating less than 7,000 m 3 /d (1.8 MGD) of wastewater in adequate depth and breadth. In this study, a detailed inventory of the construction and operational data for 16 case studies of small WWTPs was developed to elucidate their environmental life cycle impacts. Conventional LCA framework was followed. The results show that the environmental impacts of both the construction and operational phases are considerable. Operational impacts are highly related to energy usage. Improving energy efficiency of a plant may reduce the environmental impacts related to operations. Construction impacts can vary considerably between facilities. Process-related factors (e.g., concrete and reinforcing steel used in basins) are typically sized using the design flow; thus much of the variability in construction impacts among plants stems from the non-process related infrastructure. Multiple regression analysis was used as an exploratory tool to identify which non-process related plant aspects contribute to the variable environmental impact of small WWTPs. These factors include aluminum, cast iron, and the capacity utilization ratio (defined as the ratio of average flow to design flow). Furthermore, industry practitioners should consider these factors when aiming to reduce the environmental impacts of a small WWTP related to construction. Scenario sensitivity analyses found that the environmental impact of construction became smaller with longer design life, and the end-of-life consideration does not heavily influence the environmental sustainability of a WWTP.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Analysis and prediction of intersection traffic violations using automated enforcement system data

We report that the automated enforcement system (AES) is an effective way of supplementing traditional traffic enforcement, and the traffic violation data from AES can also be effectively used for safety research. In this study, traffic violation data were used to analyze the influencing factors associated with traffic violations and to predict the probability of violations at intersections. The potential factors influencing violations include 24 independent factors related to time, space, traffic and weather. Results from a logistic model showed that the midday period, weekends, residential districts, collector roads, congested traffic conditions, high traffic flow, lower wind speed and low temperature would increase the probability of traffic violations. The probability of violations was predicted by the random forest algorithm, which was proven to be the best traffic violation prediction model among logistic regression, Gaussian naive Bayes, and support vector machine. Moreover, the proximity weighted synthetic oversampling technique (ProWSyn) method was applied to reduce the impact of the imbalance ratio (IR) and improve the model’s prediction performance. The receiver operating characteristics (ROC) curves and Precision-Recall (PR) curves illustrated that the random forest algorithm using oversampling data had the best classifier prediction performance than undersampling data. The area under curve (AUC) and out-of-bag (OOB) error with IR = 1 reached 0.914 and 0.0787, which showed the better performance of the random forest algorithm using ProWSyn in dealing with imbalanced traffic violation data.

42 ENGINEERING↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Heterogeneity or illusion? Track the carbon Kuznets curve of global residential building operations

Residential buildings, the “last mile” sector in the global decarbonization, have become the most significant uncertain factor hindering carbon neutrality with increasing household energy demand. To track the operational carbon in buildings, this study investigates the carbon Kuznets curve (CKC) and the corresponding decoupling status of residential building operations at four emission scales by using the data of 30 countries from 2000 to 2019. The results show that (1) the CKC model can fit more than half of the samples. Most curves have an inverted U-shape, with 76% of emission per household and 82% of total emissions. (2) In the presence of the CKC, over four-fifths of global residential buildings peak regardless of any emission scale. The analysis denotes that the carbon emissions of developed countries reach their peaks earlier. In the total emissions, the samples’ peaking proportion is 20% and 25% with income per capita < 20,000 United States dollars (USD) and 20,000–40,000 USD, respectively. (3) The Tapio decoupling analysis and the threshold regression effectively verify the robustness and the heterogeneity of CKCs, respectively. Strong decoupling effects of CKCs in most countries are demonstrated at the scales of emission per floor space and the total emissions, and the heterogeneity proves the classic inverted U-shaped relationship between economy and emissions doesn’t exist in all emitters. Overall, this study tracks the historical carbon emission trajectories of residential building operations at a global scale, providing reference for different economies to simulate the dynamic of building carbon emissions along with the economic booming.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Meta-analysis of biogas upgrading to renewable natural gas through biological CO 2 conversion

Biogas upgrading through CO 2 conversion by hydrogenotrophic methanogenesis is receiving an increasing attention worldwide because of the demand for renewable natural gas. Herein, a holistic and statistical study of the operation conditions, driving forces, performances, and potential implementation of biogas upgrading via biological CO 2 conversion was conducted. Based on a systematic review and meta-analysis of 46 existing publications that were selected from 1475 papers, we have compiled a global dataset of CO 2 bioconversion biogas upgrading, encompassing 308 study cases. Subsequently, we employed a rigorous analytical framework incorporating data processing and mixed effects linear regression analysis to examine the dataset. This analysis revealed a significant positive relationship between the H 2 :CO 2 ratio and the methane percentage in the upgraded biogas when using the study as a random effect. Furthermore, we performed meta-analysis on observations taken when the ratio was close to 4:1 and found that ex situ reactors (91.93% [88.11%, 95.75%]) can perform better than in situ reactors (84.74% [80.69%, 88.80%]). No evidence of differential performance was found based on the present dataset between different temperature regimes or operation modes. Furthermore, those findings establish a database that will contribute to a deeper understanding of the biogas upgrading via biological CO 2 conversion.

Biogas upgrading↗

A Data-Driven Method for Estimating Behind-the-Meter Photovoltaic Generation in Hawaii

Due to the increasing penetration of distributed behind-the-meter photovoltaic (PV) systems and the installed utility revenue metering limited to monitoring only the net power import/export of the household, it is increasingly challenging for utilities to effectively plan and operate the grid. This paper proposes a methodology that estimates behind-the-meter PV generation using a selected subset of monitored PV systems. It is a data-driven approach, and the PV output is estimated utilizing a statistic regression model. A Minimum Redundancy Maximum Relevance (MRMR) algorithm is applied to preselect the optimal subset of the monitored PV systems. The performance of this approach is compared with a spatial interpolation method and a model-based approach. The proposed method is validated using high-resolution meter data recorded from 18 residential rooftop PV systems located on the island of Maui, Hawaii.

Data-driven modeling↗

mvBayes

SAND2026-16980O mvBayes implements multivariate Bayesian regression using MATLAB and decomposes a multivariate or functional response into components based on a user-specified orthogonal basis. This allows for independent modeling of each component with any chosen univariate Bayesian regression model. This tool includes methods for prediction and visualization, facilitating the evaluation of Bayesian surrogate models through the application of Bayesian theory and Markov Chain Monte Carlo (MCMC) sampling techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗