Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

"Hidden" hydrothermal technical potential & technoeconomics: Revealing permeability & fluids with more data

Historical hydrothermal estimates have largely relied on temperature or heat flow estimates ignoring the need for natural flowing fluids. More accurate hydrothermal estimates require some indication of permeability and fluids that naturally exist in the subsurface. This paper describes a novel approach that includes proxies of permeability and fluids in hydrothermal estimates by leveraging the relatively data-rich Great Basin. Specifically, nameplate capacities (megawatts) of operating geothermal plants, negative (0 megawatt) locations and 48 geophysical and geologic features are used to used in eXtreme Gradient Boosting (XGBoost) regression to make hydrothermal capacity predictions. Additionally, this work inputs the XGBoost-based hydrothermal predictions into the Renewable Energy Potential (reV) model to quantify technical capacity, its uncertainty and techno-economics. Compared to historical hydrothermal estimates, these predictions adhere to the 37 operating geothermal plants and negative locations. We present a method for subsampling the negative sites to bring the labels into balance that uses the geologic domain knowledge to proportionally represent negatives. Overall, the distributions of the hydrothermal technical capacity and the site levelized cost of energy are respectively much tighter, lower and more accurate than the previous estimates for the Great Basin, as they include geological and geophysical surrogates for permeability and fluids. Percentile (50th and 90th, median and high estimate, respectively) models provide bookends for these metrics.

13 HYDRO ENERGY↗

Development of a Supercharged Octane Number and a Supercharged Octane Index

Gasoline knock resistance is characterized by the Research and Motor Octane Number (RON and MON), which are rated on the CFR octane rating engine at naturally aspirated conditions. However, modern automotive downsized boosted spark ignition (SI) engines generally operate at higher cylinder pressures and lower temperatures relative to the RON and MON tests. Using the naturally aspirated RON and MON ratings, the octane index (OI) characterizes the knock resistance of gasolines under boosted operation by linearly extrapolating into boosted “beyond RON” conditions via RON, MON, and a linear regression K factor. Using OI solely based on naturally aspirated RON and MON tests to extrapolate into boosted conditions can lead to significant errors in predicting boosted knock resistance between gasolines due to non-linear changes in autoignition and knocking characteristics with increasing pressure conditions. Here, a new “Supercharged Octane Number” (SON) method was developed on the CFR engine at increased intake pressures, which improved the correlation to boosted knock-limited automotive SI engine data over RON for several surrogate fuels and gasolines, including five “Co-Optima” RON 98 fuels and an E10 regular grade gasoline. Furthermore, the conventional OI was extended to a newly introduced Supercharged Octane Index (OI S ) based on SON and RON, which significantly improved the correlation to fuel knock resistance measurements from modern boosted SI engine knock-limited spark advance tests. This demonstrated the first proof of concept of a SON and OI S to better characterize a fuel’s knock resistance in modern boosted SI engines.

42 ENGINEERING↗

Learning Symbolic Expressions: Mixed-Integer Formulations, Cuts, and Heuristics

Here, in this paper, we consider the problem of learning a regression function without assuming its functional form. This problem is referred to as symbolic regression. An expression tree is typically used to represent a solution function, which is determined by assigning operators and operands to the nodes. Cozad and Sahinidis propose a nonconvex mixed-integer nonlinear program (MINLP), in which binary variables are used to assign operators and nonlinear expressions are used to propagate data values through nonlinear operators, such as square, square root, and exponential. We extend this formulation by adding new cuts that improve the solution of this challenging MINLP. We also propose a heuristic that iteratively builds an expression tree by solving a restricted MINLP. We perform computational experiments and compare our approach with a mixed-integer program–based method and a neural network–based method from the literature.

97 MATHEMATICS AND COMPUTING↗

Correlating Time-Resolved Pressure Measurements With Rim Sealing Effectiveness for Real-Time Turbine Health Monitoring

Purge flow is bled from the upstream compressor and supplied to the under-platform region to prevent hot main gas path ingress that damages vulnerable under-platform hardware components. A majority of turbine rim seal research has sought to identify methods of improving sealing technologies and understanding the physical mechanisms that drive ingress. While these studies directly support the design and analysis of advanced rim seal geometries and purge flow systems, the studies are limited in their applicability to real-time monitoring required for condition-based operation and maintenance. As operational hours increase for in-service engines, this lack of rim seal performance feedback results in progressive degradation of sealing effectiveness, thereby leading to reduced hardware life. To address this need for rim seal performance monitoring, this study utilizes measurements from a one-stage turbine research facility operating with true-scale engine hardware at engine-relevant conditions. Time-resolved pressure measurements collected from the rim seal region are regressed with sealing effectiveness through the use of common machine learning techniques to provide real-time feedback of sealing effectiveness. Two modeling approaches are presented that use a single sensor to predict sealing effectiveness accurately over a range of two turbine operating conditions. Here, the results show that an initial purely data-driven model can be further improved using domain knowledge of relevant turbine operations, which yields sealing effectiveness predictions within 3% of measured values.

42 ENGINEERING↗

Correlating Time-Resolved Pressure Measurements With Rim Sealing Effectiveness for Real-Time Turbine Health Monitoring

Purge flow is bled from the upstream compressor and supplied to the under-platform region to prevent hot main gas path ingress that damages vulnerable under-platform hardware components. A majority of turbine rim seal research has sought to identify methods of improving sealing technologies and understanding the physical mechanisms that drive ingress. While these studies directly support the design and analysis of advanced rim seal geometries and purge flow systems, the studies are limited in their applicability to real-time monitoring required for condition-based operation and maintenance. As operational hours increase for in-service engines, this lack of rim seal performance feedback results in progressive degradation of sealing effectiveness, thereby leading to reduced hardware life. To address this need for rim seal performance monitoring, the present study utilizes measurements from a one-stage turbine research facility operating with true-scale engine hardware at engine-relevant conditions. Time-resolved pressure measurements collected from the rim seal region are regressed with sealing effectiveness through the use of common machine learning techniques to provide real-time feedback of sealing effectiveness. Two modelling approaches are presented that use a single sensor to predict sealing effectiveness accurately over a range of two turbine operating conditions. Results show that an initial purely data-driven model can be further improved using domain knowledge of relevant turbine operations, which yields sealing effectiveness predictions within three percent of measured values.

Compressors↗

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Meta‐Analysis and Regression Modeling of the Impacts of Four Indoor Environmental Quality Metrics on Office Performance

Awareness of how buildings interact with the occupant experience—especially human performance—is becoming more prevalent, as seen by increasing interest and investment in healthy built environments. However, there is a need to synthesize the wide array of existing indoor environmental assessment and performance research in a way that can translate directly to building design and operation. Existing research in this area typically focuses on a single isolated metric and has not focused on making the results utilizable by building practitioners. The aim of this research is to investigate existing office performance literature through meta‐analyses and produce regression models for four indoor environmental quality (IEQ) metrics to support critical decision‐making for building operation and renovation. To reach this aim, a literature review was conducted to identify studies that measure the impact of changing ventilation rate, temperature, horizontal illuminance, and noise level in offices on occupant task performance. This repository of field and laboratory studies was analyzed to visualize the trends between the selected IEQ metrics and task performance. The temperature, ventilation rate, and horizontal illuminance regression models showed clear improvement potential when modifying indoor conditions toward the defined high‐performance range, while the regression model for noise level was inconclusive. The discussion notes the importance of designing holistically for all components of these IEQ categories to utilize the results, for example, good filtration on outdoor air for quantifying ventilation impact and uniform overhead lighting with low contrast for quantifying horizontal illuminance impact. The novelty of this work is in considering multiple facets of the indoor environment under a single, unified analysis schema and producing IEQ‐based performance gains that can directly inform cost‐benefit analyses of building design and renovation.

60 APPLIED LIFE SCIENCES↗

Stochastic projective splitting

Here, we present a new, stochastic variant of the projective splitting (PS) family of algorithms for inclusion problems involving the sum of any finite number of maximal monotone operators. This new variant uses a stochastic oracle to evaluate one of the operators, which is assumed to be Lipschitz continuous, and (deterministic) resolvents to process the remaining operators. Our proposal is the first version of PS with such stochastic capabilities. We envision the primary application being machine learning (ML) problems, with the method’s stochastic features facilitating “mini-batch” sampling of datasets. Since it uses a monotone operator formulation, the method can handle not only Lipschitz-smooth loss minimization, but also min–max and noncooperative game formulations, with better convergence properties than the gradient descent-ascent methods commonly applied in such settings. The proposed method can handle any number of constraints and nonsmooth regularizers via projection and proximal operators. We prove almost-sure convergence of the iterates to a solution and a convergence rate result for the expected residual, and close with numerical experiments on a distributionally robust sparse logistic regression problem.

97 MATHEMATICS AND COMPUTING↗

Measuring environmental and cost benefits of riparian buffers for drinking water production in a Midwest watershed

Here, this study focuses on the economic value and carbon benefits of riparian buffers in urban drinking water production. The impact of riparian buffers on the waterworks operation in the Raccoon River watershed was quantified using the following metrics: nitrate concentration, days of nitrate removal operation, material and energy cost (based on 17 years of historical records with a watershed model), regression, and cost analysis. The findings indicate that the presence of riparian buffers in agricultural land can substantially decrease nitrate concentration in the water intake of the waterworks during the crop-growing season: 19% in April, 9% in May, and 11% in June. These reductions mean less nitrate treatment of the plant intake flow: 23% in April, 12% in May, and 3% in June. These changes lead to significant resource savings: 425 metric tons of sodium chloride (NaCl), 147 810 kWh of electricity, 253 metric tons of powdered activated carbon, and 20.8 million liters of fresh water. The buffers would also reduce greenhouse gas emissions by 86.9 metric tons (CO 2 equivalent) in 17 years. The total cost saving was estimated at $\$$327 326, with the highest potential savings in May ($\$$215 100), followed by April ($\$$65 465), and June ($\$$46 761). When factoring in buffer installation, cropland loss, nitrate removal, and cost associated with buffer harvest for biofuel in an established biomass market, the benefit to the entire watershed community would be $\$$2.63 million annually. The results underscore the significant cost benefits and environmental benefits associated with cropland riparian buffers in a watershed community. The approach employed in this study holds promise for assessing riparian buffer benefits in other watersheds, contributes to an understanding of sustainable water management practices, and provides a basis for decision making in a wide range of agricultural regions.

54 ENVIRONMENTAL SCIENCES↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

A Framework for Closed-Loop Optimization of an Automated Mechanical Serial-Sectioning System via Run-to-Run Control as Applied to a Robo-Met.3D

Optimization of automated data collection is gaining increased interest for the purposes of enabling closed-loop self-correcting systems that inherently maximize operational efficiencies and reduce waste. Many data collection systems have several variables which influence data accuracy or consistency and which can require frequent user interaction to be monitored and maintained. Operating upon a Robo-MET.3D™ automated mechanical serial-sectioning system, a run-to-run control algorithm has been developed to accelerate data collection and reduce data inconsistency. Here, using historical data amassed over a decade of experiments, a linear regression model of the deterministic system dynamics is created and used to employ a run-to-run control algorithm that optimizes selected system inputs to reduce operator intervention and increase efficacy while reducing variance of system output.

42 ENGINEERING↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different datasets, i.e., 1) the unit generation dataset and 2) the system total inertia and headroom dataset, for the prediction of the frequency nadir. The training and testing datasets are generated through extensive generation scheduling simulations using Multi-timescale Integrated Dynamic and Scheduling (MI-DAS) toolbox on the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. Numerical results show that all five machine learning methods perform well in predicting the nadir frequency of the system. Among them, the gradient boosting and the XGBoost are clear winners yielding the best prediction accuracy in terms of four evaluation metrics.

data driven↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems: Preprint

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different sets of preprocess data for the prediction of the frequency nadir in the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. The training and testing data sets are collected by extensive generation scheduling simulations on the Multi-timescale Integrated Dynamic and Scheduling (MIDAS) toolbox. Numerical results show that all five machine learning methods can achieve high performance accuracy for power system nadir frequency estimation. Among them, the gradient boosting and the XGBoost are clear winners by providing the best prediction accuracy.

data driven↗

Situational awareness-enhancing community-level load mapping with opportunistic machine learning

Motivated by present and forthcoming challenges in the adoption and integration of distributed renewable energy, we develop a machine learning (ML) approach that builds short-fuse mappings connecting the occasionally-unobservable true load in one target community with information-rich signals collected from relatively more instrumented reference communities. Our setting is inspired by and tailored to target communities with significant unobservable behind-the-meter solar generation, where true load (a relatively well-behaved quantity of interest to grid operators) is hard to discern during daytime due to insufficient instrumentation and/or privacy reasons, but that can be related to reference communities with low unobservable distributed variable generation or with sufficient instrumentation. The developed mapping, herein realized with Support Vector Machine regression, is built using nighttime data from all communities, when their distributed generation is low or zero. Our ML algorithm opportunistically learns to correlate signals of interest and then is operationally used the next day to shed light into target community load evolution. The mapping is subsequently rebuilt, rolling its short-fuse scope perpetually forward in time. Here, we demonstrate the efficacy of our approach on nine synthetically generated topologies and associated timeseries stemming from real-world data, on which we observe cumulative error performance that yields lower than 10% and 15% daily-averaged mean absolute percentage errors in target community load estimation on more than about 75% and 90% of days, respectively, in multiple yearly evaluations that shed light on long-term performance also under seasonal and one-off effects. The proposed ML-powered methodology can offer grid operators much-improved visibility into a previously obscure space and can also serve as an additional source of information in broader, multi-modal solar disaggregation solutions.

14 SOLAR ENERGY↗

Congenital malformation and hemoglobin A1c in the first trimester among Japanese women with pregestational diabetes

Abstract Aim To investigate the incidence of major congenital malformations in Japanese women with pregestational diabetes, and to determine the cutoff value of hemoglobin A1c (HbA1c) in the first trimester associated with congenital malformations. Methods This retrospective cohort study included singleton pregnancies in Japanese women with pregestational diabetes, including type 1 and type 2 diabetes, and specific types of diabetes due to other causes. The primary outcome was the incidence of major congenital malformations. The secondary outcome was the incidence of all congenital malformations. The cutoff value of HbA1c for congenital malformations was calculated using receiver operating characteristic curve analysis. The adjusted odds ratios (aOR) of major congenital malformations were calculated using multiple logistic regression analyses. Results This study enrolled 292 patients, including 132 (45.2%) with type 1 diabetes, 156 (53.4%) with type 2 diabetes, and 4 (1.4%) with other specific types. The incidence rates of major congenital malformations and all congenital malformations were 7.2% (21/292) and 12.7% (37/292), respectively. The cutoff value of HbA1c in the first trimester for major malformations and for all congenital malformations was 6.5%. HbA1c ≥ 6.5% was significantly associated with major malformations (aOR 3.5; 95% confidence interval: 1.2–12.6; p = 0.018). Conclusion The incidence of major congenital malformations significantly increased in pregnant Japanese women with HbA1c values of 6.5% or higher. The recommended HbA1c value during the first trimester used in other countries can be applied to pregnant Japanese women.

Nakanishi, Kentaro↗