Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore’s Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to MLexperts because of lack of training data and relevant, easy-to-use workflows. Wepresent ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state. The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res2) and code(https://leap-stc.github.io/ClimSim)arereleasedopenlytosupport the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.

artificial intelligence, machine learning↗

Evaluation of cooling setpoint setback savings in commercial buildings using electricity and exterior temperature time series data

Commercial buildings account for a significant amount of total energy produced in the US, and the Heating Ventilation and Cooling (HVAC) systems are one of the most significant components of their overall consumption. In this study, we proposed a new data-driven approach to evaluate HVAC cooling systems in commercial buildings and identify savings opportunities. The focus is an investigation of the impact of thermostat setpoint setback but using only whole building, electricity data taken at 15-min intervals for the analysis. We conducted a comparative study of setpoint setback characteristics on 432 commercial buildings with 5 building usage types across the United States. To accomplish this, both piecewise and Random Forest regression algorithms were employed using electricity and exterior temperature datasets to identify operational characteristics and the effective setpoints in the building to determine the corresponding savings opportunities. Both occupied and unoccupied time periods were studied across cooling degree days (CDD), when air conditioning is typically operational. Here the results show that in commercial buildings, on average, cooling systems account for 9.5% of total consumption. When a one degree setback during the cooling season is applied, an average of approximately 1.1% of annual consumption is achieved; retail and office buildings demonstrate the highest potential for savings. Additionally, we identified that the number of cooling degree days and base to peak ratio (BPR) are the most important variables for predicting the magnitude of the consumption of cooling systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Modeling atmospheric data and identifying dynamics Temporal data-driven modeling of air pollutants

Atmospheric modelling has recently experienced a surge with the advent of deep learning. Most of these models, however, predict concentrations of pollutants following a data-driven approach in which the physical laws that govern their behaviors and relationships remain hidden. Furthermore, with the aid of real-world air quality data collected hourly in different stations throughout Madrid, we present a case study using a series of data-driven techniques with the following goals: (1) Find systems of ordinary differential equations that model the concentration of pollutants and their changes over time; (2) assess the performance and limitations of our model using stability analysis; (3) reconstruct the time series of chemical pollutants not measured in certain stations using delay coordinate embedding results.

54 ENVIRONMENTAL SCIENCES↗

Using machine learning to improve efficiency and accuracy of burnup measurements at PBR reactors [Slides]

The outline of the slides include: Motivations of the work; Modeling and simulation; Machine learning model; Results and comparison study with linear regression; and Conclusions. This work was done to help PBR designers and operators understand the burnup measurement better. We look forward to discussing the results in detail with industrial collaborators.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Stochastic AC optimal power flow: A data-driven approach

There is an emerging need for efficient solutions to stochastic AC Optimal Power Flow (AC-OPF) to ensure optimal and reliable grid operations in the presence of increasing demand and generation uncertainty. Herein this paper presents a highly scalable data-driven algorithm for stochastic AC-OPF that has extremely low sample requirement. The novelty behind the algorithm’s performance involves an iterative scenario design approach that merges information regarding constraint violations in the system with data-driven sparse regression. Compared to conventional methods with random scenario sampling, our approach is able to provide feasible operating points for realistic systems with much lower sample requirements. Furthermore, multiple sub-tasks in our approach can be easily paralleled and based on historical data to enhance its performance and application. We demonstrate the computational improvements of our approach through simulations on different test cases in the IEEE PES PGLib-OPF benchmark library.

42 ENGINEERING↗

Landsat 8 monitoring of multi-depth suspended sediment concentrations in Lake Erie’s Maumee River using machine learning

Satellite remote sensing has been widely used to map suspended sediment concentration (SSC) in waterbodies. However, due to the complexity of sediment-water interactions, it has been difficult to derive linear and non-linear regression equations to reliably predict SSC, especially when trying to estimate depth of integrated sediment. Herein, this study uses Landsat 8 OLI (Operational Land Imager) sensor to map SSC within the Maumee River in Ohio, USA, at multiple depth intervals (15, 61, 91, and 182 cm). Simple linear least squares regression (LLSR), and three common machine learning models: random forest (RF), support vector regression (SVR), and model averaged neural network (MANN) were used to estimate SSC at the depth intervals. All machine learning models significantly outperformed LLSR while RF performed the best. In both RF and MANN, R2 (coefficient of determination) increases with depth with a maximum R2 of 0.89 and 0.83, respectively, at a depth of 0–182 cm. The results show that machine learning models can implement nonlinear relationships that produce better predictions than traditional linear regression methods in estimating depth integrated SSC, especially when samples are limited.

47 OTHER INSTRUMENTATION↗

Performance study of a novel dew point evaporative cooler in the climate of central Europe using building simulation tools

This paper presents the performance research of a hybrid air conditioning system (System II) equipped with novel dew point evaporative cooler (DPIEC) for a retail building located in temperate climatic conditions. The new cooler has a dedicated, adaptable structure that maximizes its performance in the climate conditions of central Europe. The unit can operate as a typical regenerative air cooler when outdoor air is drier than indoor air, it can operate as a counter-flow heat recovery unit when indoor conditions are more dry than outdoor and it can operate as a combination of both arrangements. The year-round, hourly-stepped building energy simulations were carried-out using a prognostic tool of energy consumption. The novel dew point evaporative cooler application potential was established using the black-box model based on regression equations was established based on experimental tests on a prototype of novel DPIEC. The hybrid system operation was compared to the typical air handling unit equipped with a standard energy wheel for heat recovery (System I). Presented research shows that DPIEC can cover about 95% of total cooling loads while the energy wheel covers only 9% of total cooling load. System II allows us to save 65% of seasonal electricity consumption comparing to System I. It comes from the fact that during most of the cooling season DPIEC is in operation and it covers most of the cooling loads without the need to switch the chilled water system.

Air conditioning↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Short-term nodal load forecasting based on machine learning techniques

This paper introduces an advanced Short-term Nodal Load Forecasting (STNLF) method that forecasts nodal load profiles for the next day in power systems, based on the combined use of three machine learning techniques. Least Absolute Shrinkage and Selection Operator (LASSO) is employed to reduce the number of features for a single nodal load forecasting. Principal Component Analysis (PCA) is used to capture the features of historical loads in low-dimensional space compared to the original high-dimensional load space where features are barely possible to depict. Additionally, Bayesian Ridge Regression (BRR) is utilized to decide the parameters of the prediction model from a statistics perspective. Tests based on modified PJM load data demonstrate the effectiveness of the proposed STNLF method compared to the state-of-the-art General Regression Neural Network (GRNN) method. Moreover, the reliability of the day-ahead Unit Commitment (UC) solution is shown to have been improved, based on the forecasted load data using the proposed STNLF method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

Revenue prediction for integrated renewable energy and energy storage system using machine learning techniques

Revenue estimation for integrated renewable energy and energy storage systems is important to support plant owners or operators’ decisions in battery sizing selection that leads to maximized financial performances. A common approach to optimizing revenues of a hybrid hydro and energy storage system is using mixed-integer linear programming (MILP). Although MILP models can provide accurate production cost estimations, they are typically very computationally expensive. To provide a fast yet accurate first-step information to hydropower plant owners or operators who consider integrating energy storage systems, we propose an innovative approach to predicting optimal revenues of an integrated energy generation and storage system. In this study, we examined the performance of two prediction techniques: Generalized Additive Models (GAMs) and machine learning (ML) models developed based on artificial neural networks (ANN). Predictive equations and models are generated based on optimized solutions from a market participation optimization model, the Conventional Hydropower Energy and Environmental Resource System (CHEERS) model. The two predicting techniques reduce the computational time to evaluate annual revenue for one set of battery configurations from 3 h to 1 to 4 min per run while also being implementable with significantly less data. The model validation prediction errors of developed GAMs and ML models are generally below 5%; for model testing predictions, the ML models consistently outperform the regression equations in terms of root mean square errors. This new approach allows plant owners, operators, or potential investors to quickly access multiple battery configurations under different energy generation and market scenarios. This new revenue prediction method will therefore help reduce the barriers, and thereby promoting the deployment of battery hybridization with existing renewable energy sources.

13 HYDRO ENERGY↗

Short-Term Load Forecasting Considering EV Charging Loads with Prediction Interval Evaluation

Short-term load forecasting plays a critical role in power system planning and operation. Along with the electrification of various loads, electricity demands are becoming increasingly hard to predict. Notably, the recent rise in electric vehicles (EVs) has further contributed to this unpredictability. To address this issue, this paper proposes a probabilistic load forecasting strategy utilizing Gaussian process regression, structured in a day-ahead manner. While many works focus on deterministic prediction, probabilistic forecasting offers additional insights into variability and uncertainty, enabling more flexible and reliable operation for power systems. To enhance the accuracy of the load forecasting model, the inputs include features related to EV charging habits as well as commonly used weather information. The load forecasting results are evaluated using various metrics, including conventional ones that assess the accuracy of point forecasts, as well as additional metrics that test the reliability of prediction intervals. The proposed load forecasting method is finally tested on real residential power consumption data and EV charging data sampled from real-world sources. The results prove that the new features can greatly improve the performance of the load forecasting method.

electrical vehicle↗

Discrete generative diffusion models without stochastic differential equations: A tensor network approach

Diffusion models (DMs) are a class of generative machine learning methods that sample a target distribution by transforming samples of a trivial (often Gaussian) distribution using a learned stochastic differential equation. In standard DMs, this is done by learning a “score function” that reverses the effect of adding diffusive noise to the distribution of interest. Here we consider the generalisation of DMs to lattice systems with discrete degrees of freedom, and where noise is added via Markov chain jump dynamics. We show how to use tensor networks (TNs) to efficiently define and sample such “discrete diffusion models” (DDMs) without explicitly having to solve a stochastic differential equation. We show the following: (i) by parametrising the data and evolution operators as TNs, the denoising dynamics can be represented exactly; (ii) the auto-regressive nature of TNs allows to generate samples efficiently and without bias; (iii) for sampling Boltzmann-like distributions, TNs allow to construct an efficient learning scheme that integrates well with Monte Carlo. We illustrate this approach to study the equilibrium of two models with non-trivial thermodynamics, the d = 1 constrained Fredkin chain and the d = 2 Ising model. Published by the American Physical Society 2025

Causer, Luke (ORCID:0000000194243473)↗

Control-Oriented Modeling of Cycle-to-Cycle Combustion Variability at the Misfire Limit in SI Engines

A control-oriented model is presented that can capture the prior-cycle correlation of combustion cycles during conditions with high levels of exhaust gas recirculation (EGR). Combustion events are modeled in discrete time and the dynamic evolution is captured by the residual air, fuel, and inert gas trapped in the combustion chamber. The mathematical formulation of the model is presented together with the calibration procedure to emulate a particular engine operating condition. A cycle-to-cycle system identification methodology is described which allows regressing model parameters from experimental data. Simulations are presented and compared to real engine measurements to show the modeling potential for analysis and control of combustion events.

Maldonado Puente, Bryan↗

Adaptive Online Multivariate Signal Extraction With Locally Weighted Robust Polynomial Regression

High-frequency, multivariate data collected in real-time and used to control or make decisions regarding a process’ operation often contain some noise and outliers. Thus, a method to extract the signal is needed in order to reduce the number and magnitude of control-based adjustments that are implemented. Such a method must be (i) online, depending only on past and current observations; (ii) fast, producing a smooth value more quickly than the measurement frequency; (iii) robust, ignoring brief bursts of erroneously measured values; (iv) multivariate, ignoring observations that are jointly unusual; (v) adaptive, adjusting to periods of rapid fluctuation in the signal versus periods of stability; and (vi) purely data-driven, not incorporating any information about the process from which the data are collected. Most existing methods are only able to address a subset of these six features. Furthermore, we also require the method to be nonlinear, providing a local nonlinear estimate of the signal. In this work, we propose a novel, real-time signal extraction method based on a local, robust polynomial fit. We demonstrate the performance of our method compared to a state-of-the-art competitor through simulation. For illustration, the methodology is applied to data collected from a reverse osmosis water treatment process.

97 MATHEMATICS AND COMPUTING↗

Gaussian process regression constrained by boundary value problems

We develop a framework for Gaussian processes regression constrained by boundary value problems. The framework may be applied to infer the solution of a well-posed boundary value problem with a known second-order differential operator and boundary conditions, but for which only scattered observations of the source term are available. Scattered observations of the solution may also be used in the regression. The framework combines co-kriging with the linear transformation of a Gaussian process together with the use of kernels given by spectral expansions in eigenfunctions of the boundary value problem. Furthermore, it benefits from a reduced-rank property of covariance matrices. We demonstrate that the resulting framework yields more accurate and stable solution inference as compared to physics-informed Gaussian process regression without boundary condition constraints.

42 ENGINEERING↗

Deterministic symbolic regression with derivative information: General methodology and application to equations of state

Symbolic regression methods simultaneously determine the model functional form and the regression parameter values by generating expression trees. Symbolic regression can capture the complexity of real–world phenomena but the use of deterministic optimization for symbolic regression has been limited due to the complexity of the search space of existing formulations. Herein we present a novel deterministic mixed–integer nonlinear programming formulation for symbolic regression that incorporates derivative constraints through auxiliary expression trees. By applying the chain rule to mathematical operations, binary expression trees are capable of representing the calculation of first and second derivatives. We apply this formulation to illustrative examples using derivative information to show increased model discrimination capability. In addition, we perform a case study of a thermodynamic equation of state to gain insight on valid functional forms with thermodynamics–based constraints on the first and second derivatives.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modern deep neural networks for Direct Normal Irradiance forecasting: A classification approach

The escalating energy demand and the adverse environmental impacts of fossil-fuel use necessitate a shift towards cleaner and renewable alternatives. Concentrated Solar Power (CSP) technology emerges as a promising solution, offering a carbon-free alternative for power generation. The efficiency and profitability of CSP depend on the Direct Normal Irradiance (DNI) component of solar radiation; hence, accurate DNI forecasting can help optimize CSP plants’ operations and performance. The unpredictable nature of weather phenomena, particularly cloud cover, introduces uncertainty into DNI projections. Existing DNI forecasting models use meteorological factors, which are both challenging to estimate numerically over short prediction windows and expensive to model through data at a sufficiently high spatial and temporal resolution. This research addresses the challenge by presenting a novel approach that formulates DNI prediction as a multi-class classification problem, departing from conventional regression-based methods. The primary objective of this classification framework is to identify optimal periods aligning with specific operational thresholds for CSP plants, contributing to enhanced dispatch optimization strategies. We model the DNI classification problem using four advanced deep neural networks – rectified linear unit (ReLU) networks, 1D residual networks (ResNets), bidirectional long short-term memory (BiLSTM) networks, and transformers – achieving accuracies up to 93.5% without requiring meteorological parameters.

14 SOLAR ENERGY↗