Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Novel method for accurately estimating membrane transport properties and mass transfer coefficients in reverse osmosis

Here, we present a simple and robust method to simultaneously characterize the water and salt permeability (A, B) of reverse osmosis (RO) membranes and mass transfer coefficient (k) in membrane modules. The proposed methodology comprises a set of RO experiments performed at different operating pressures or stages. The measured water and salt fluxes in each stage are simultaneously fitted to the RO transport equations by performing a non-linear regression, using A, B, and k as regression parameters. We first perform a systematic accuracy analysis of the proposed method across the full operational range of RO. The assessment shows that the method accuracy is substantially higher than current methods and increases with number of experimental stages and driving forces. This assessment is used to inform the design of an experimental protocol that minimizes errors in estimated A, B, and k. We then evaluate two commercial RO membranes following the new protocol. For both membranes, A and B parameters decrease by 17% and 15% from the dilute solution to seawater concentrations, whereas the k parameter remains constant. Our study demonstrates that the proposed method, informed by data-driven experimental designs, provides a new approach for accurately characterizing transport phenomenon in membrane processes with feeds of less than 100 g/L total dissolved solids.

36 MATERIALS SCIENCE↗

Scalable Hybrid Classification-Regression Solution for High-Frequency Nonintrusive Load Monitoring

Residential buildings with the ability to monitor and control their net-load (sum of load and generation) can provide valuable flexibility to power grid operators. We present a novel multiclass nonintrusive load monitoring (NILM) approach that enables effective net-load monitoring capabilities at high-frequency with minimal additional equipment and cost. The proposed machine learning based solution provides accurate multiclass state predictions while operating at a faster timescale (able to provide a prediction for each 60- Hz ac cycle used in US power grid) without relying on event-detection techniques. We also introduce an innovative hybrid classification-regression method that allows for the prediction of not only load on/off states but also individual load operating power levels. A test bed with eight residential appliances is used for validating the NILM approach. Results show that the overall method has high accuracy, good scaling and generalization properties.

feature extraction↗

Non-intrusive nonlinear model reduction via machine learning approximations to low-dimensional operators

Abstract Although projection-based reduced-order models (ROMs) for parameterized nonlinear dynamical systems have demonstrated exciting results across a range of applications, their broad adoption has been limited by their intrusivity: implementing such a reduced-order model typically requires significant modifications to the underlying simulation code. To address this, we propose a method that enables traditionally intrusive reduced-order models to be accurately approximated in a non-intrusive manner. Specifically, the approach approximates the low-dimensional operators associated with projection-based reduced-order models (ROMs) using modern machine-learning regression techniques. The only requirement of the simulation code is the ability to export the velocity given the state and parameters; this functionality is used to train the approximated low-dimensional operators. In addition to enabling nonintrusivity, we demonstrate that the approach also leads to very low computational complexity, achieving up to $$10^3{\times }$$ 10 3 × in run time. We demonstrate the effectiveness of the proposed technique on two types of PDEs. The domain of applications include both parabolic and hyperbolic PDEs, regardless of the dimension of full-order models (FOMs).

42 ENGINEERING↗

BASS

SAND2026-17001O BASS implements Bayesian Adaptive Spline Surfaces in MATLAB and serves as a surrogate model for regression applications. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

BayesPPR

SAND2026-17002O BayesPPR performs Bayesian Projection Pursuit Regression (PPR) using MATLAB. A surrogate model for calibration applications, it enables users to efficiently analyze complex datasets and extract meaningful patterns through regression techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore’s Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to MLexperts because of lack of training data and relevant, easy-to-use workflows. Wepresent ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state. The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res2) and code(https://leap-stc.github.io/ClimSim)arereleasedopenlytosupport the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.

artificial intelligence, machine learning↗

Modeling atmospheric data and identifying dynamics Temporal data-driven modeling of air pollutants

Atmospheric modelling has recently experienced a surge with the advent of deep learning. Most of these models, however, predict concentrations of pollutants following a data-driven approach in which the physical laws that govern their behaviors and relationships remain hidden. Furthermore, with the aid of real-world air quality data collected hourly in different stations throughout Madrid, we present a case study using a series of data-driven techniques with the following goals: (1) Find systems of ordinary differential equations that model the concentration of pollutants and their changes over time; (2) assess the performance and limitations of our model using stability analysis; (3) reconstruct the time series of chemical pollutants not measured in certain stations using delay coordinate embedding results.

54 ENVIRONMENTAL SCIENCES↗

Using machine learning to improve efficiency and accuracy of burnup measurements at PBR reactors [Slides]

The outline of the slides include: Motivations of the work; Modeling and simulation; Machine learning model; Results and comparison study with linear regression; and Conclusions. This work was done to help PBR designers and operators understand the burnup measurement better. We look forward to discussing the results in detail with industrial collaborators.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Short-term nodal load forecasting based on machine learning techniques

This paper introduces an advanced Short-term Nodal Load Forecasting (STNLF) method that forecasts nodal load profiles for the next day in power systems, based on the combined use of three machine learning techniques. Least Absolute Shrinkage and Selection Operator (LASSO) is employed to reduce the number of features for a single nodal load forecasting. Principal Component Analysis (PCA) is used to capture the features of historical loads in low-dimensional space compared to the original high-dimensional load space where features are barely possible to depict. Additionally, Bayesian Ridge Regression (BRR) is utilized to decide the parameters of the prediction model from a statistics perspective. Tests based on modified PJM load data demonstrate the effectiveness of the proposed STNLF method compared to the state-of-the-art General Regression Neural Network (GRNN) method. Moreover, the reliability of the day-ahead Unit Commitment (UC) solution is shown to have been improved, based on the forecasted load data using the proposed STNLF method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

Revenue prediction for integrated renewable energy and energy storage system using machine learning techniques

Revenue estimation for integrated renewable energy and energy storage systems is important to support plant owners or operators’ decisions in battery sizing selection that leads to maximized financial performances. A common approach to optimizing revenues of a hybrid hydro and energy storage system is using mixed-integer linear programming (MILP). Although MILP models can provide accurate production cost estimations, they are typically very computationally expensive. To provide a fast yet accurate first-step information to hydropower plant owners or operators who consider integrating energy storage systems, we propose an innovative approach to predicting optimal revenues of an integrated energy generation and storage system. In this study, we examined the performance of two prediction techniques: Generalized Additive Models (GAMs) and machine learning (ML) models developed based on artificial neural networks (ANN). Predictive equations and models are generated based on optimized solutions from a market participation optimization model, the Conventional Hydropower Energy and Environmental Resource System (CHEERS) model. The two predicting techniques reduce the computational time to evaluate annual revenue for one set of battery configurations from 3 h to 1 to 4 min per run while also being implementable with significantly less data. The model validation prediction errors of developed GAMs and ML models are generally below 5%; for model testing predictions, the ML models consistently outperform the regression equations in terms of root mean square errors. This new approach allows plant owners, operators, or potential investors to quickly access multiple battery configurations under different energy generation and market scenarios. This new revenue prediction method will therefore help reduce the barriers, and thereby promoting the deployment of battery hybridization with existing renewable energy sources.

13 HYDRO ENERGY↗

Short-Term Load Forecasting Considering EV Charging Loads with Prediction Interval Evaluation

Short-term load forecasting plays a critical role in power system planning and operation. Along with the electrification of various loads, electricity demands are becoming increasingly hard to predict. Notably, the recent rise in electric vehicles (EVs) has further contributed to this unpredictability. To address this issue, this paper proposes a probabilistic load forecasting strategy utilizing Gaussian process regression, structured in a day-ahead manner. While many works focus on deterministic prediction, probabilistic forecasting offers additional insights into variability and uncertainty, enabling more flexible and reliable operation for power systems. To enhance the accuracy of the load forecasting model, the inputs include features related to EV charging habits as well as commonly used weather information. The load forecasting results are evaluated using various metrics, including conventional ones that assess the accuracy of point forecasts, as well as additional metrics that test the reliability of prediction intervals. The proposed load forecasting method is finally tested on real residential power consumption data and EV charging data sampled from real-world sources. The results prove that the new features can greatly improve the performance of the load forecasting method.

electrical vehicle↗

Discrete generative diffusion models without stochastic differential equations: A tensor network approach

Diffusion models (DMs) are a class of generative machine learning methods that sample a target distribution by transforming samples of a trivial (often Gaussian) distribution using a learned stochastic differential equation. In standard DMs, this is done by learning a “score function” that reverses the effect of adding diffusive noise to the distribution of interest. Here we consider the generalisation of DMs to lattice systems with discrete degrees of freedom, and where noise is added via Markov chain jump dynamics. We show how to use tensor networks (TNs) to efficiently define and sample such “discrete diffusion models” (DDMs) without explicitly having to solve a stochastic differential equation. We show the following: (i) by parametrising the data and evolution operators as TNs, the denoising dynamics can be represented exactly; (ii) the auto-regressive nature of TNs allows to generate samples efficiently and without bias; (iii) for sampling Boltzmann-like distributions, TNs allow to construct an efficient learning scheme that integrates well with Monte Carlo. We illustrate this approach to study the equilibrium of two models with non-trivial thermodynamics, the d = 1 constrained Fredkin chain and the d = 2 Ising model. Published by the American Physical Society 2025

Causer, Luke (ORCID:0000000194243473)↗

Adaptive Online Multivariate Signal Extraction With Locally Weighted Robust Polynomial Regression

High-frequency, multivariate data collected in real-time and used to control or make decisions regarding a process’ operation often contain some noise and outliers. Thus, a method to extract the signal is needed in order to reduce the number and magnitude of control-based adjustments that are implemented. Such a method must be (i) online, depending only on past and current observations; (ii) fast, producing a smooth value more quickly than the measurement frequency; (iii) robust, ignoring brief bursts of erroneously measured values; (iv) multivariate, ignoring observations that are jointly unusual; (v) adaptive, adjusting to periods of rapid fluctuation in the signal versus periods of stability; and (vi) purely data-driven, not incorporating any information about the process from which the data are collected. Most existing methods are only able to address a subset of these six features. Furthermore, we also require the method to be nonlinear, providing a local nonlinear estimate of the signal. In this work, we propose a novel, real-time signal extraction method based on a local, robust polynomial fit. We demonstrate the performance of our method compared to a state-of-the-art competitor through simulation. For illustration, the methodology is applied to data collected from a reverse osmosis water treatment process.

97 MATHEMATICS AND COMPUTING↗

Gaussian process regression constrained by boundary value problems

We develop a framework for Gaussian processes regression constrained by boundary value problems. The framework may be applied to infer the solution of a well-posed boundary value problem with a known second-order differential operator and boundary conditions, but for which only scattered observations of the source term are available. Scattered observations of the solution may also be used in the regression. The framework combines co-kriging with the linear transformation of a Gaussian process together with the use of kernels given by spectral expansions in eigenfunctions of the boundary value problem. Furthermore, it benefits from a reduced-rank property of covariance matrices. We demonstrate that the resulting framework yields more accurate and stable solution inference as compared to physics-informed Gaussian process regression without boundary condition constraints.

42 ENGINEERING↗

Deterministic symbolic regression with derivative information: General methodology and application to equations of state

Symbolic regression methods simultaneously determine the model functional form and the regression parameter values by generating expression trees. Symbolic regression can capture the complexity of real–world phenomena but the use of deterministic optimization for symbolic regression has been limited due to the complexity of the search space of existing formulations. Herein we present a novel deterministic mixed–integer nonlinear programming formulation for symbolic regression that incorporates derivative constraints through auxiliary expression trees. By applying the chain rule to mathematical operations, binary expression trees are capable of representing the calculation of first and second derivatives. We apply this formulation to illustrative examples using derivative information to show increased model discrimination capability. In addition, we perform a case study of a thermodynamic equation of state to gain insight on valid functional forms with thermodynamics–based constraints on the first and second derivatives.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modern deep neural networks for Direct Normal Irradiance forecasting: A classification approach

The escalating energy demand and the adverse environmental impacts of fossil-fuel use necessitate a shift towards cleaner and renewable alternatives. Concentrated Solar Power (CSP) technology emerges as a promising solution, offering a carbon-free alternative for power generation. The efficiency and profitability of CSP depend on the Direct Normal Irradiance (DNI) component of solar radiation; hence, accurate DNI forecasting can help optimize CSP plants’ operations and performance. The unpredictable nature of weather phenomena, particularly cloud cover, introduces uncertainty into DNI projections. Existing DNI forecasting models use meteorological factors, which are both challenging to estimate numerically over short prediction windows and expensive to model through data at a sufficiently high spatial and temporal resolution. This research addresses the challenge by presenting a novel approach that formulates DNI prediction as a multi-class classification problem, departing from conventional regression-based methods. The primary objective of this classification framework is to identify optimal periods aligning with specific operational thresholds for CSP plants, contributing to enhanced dispatch optimization strategies. We model the DNI classification problem using four advanced deep neural networks – rectified linear unit (ReLU) networks, 1D residual networks (ResNets), bidirectional long short-term memory (BiLSTM) networks, and transformers – achieving accuracies up to 93.5% without requiring meteorological parameters.

14 SOLAR ENERGY↗