Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Predictive Model for Starlink Maritime Performance Using Multi-Horizon RandomForest

Low Earth orbit (LEO) satellite systems have become a crucial enabler of broadband access for maritime industries, where traditional networks are unavailable. However, the high mobility of LEO constellations and constantly changing weather conditions result in unpredictable link fluctuations, limiting the ability of maritime platforms to plan bandwidth usage proactively. To the best of our knowledge, no prior work has developed a short-term predictive model for maritime LEO connectivity using real experimental field measurements. This paper proposes a data-driven forecasting model that predicts future downlink throughput using multi-horizon RandomForest regression. The model is trained using real experimental coastal measurement data incorporating recent throughput history, network-layer indicators, and environmental variables. The proposed approach reduces mean absolute error by approximately 31% compared to a persistence baseline for 15-minute horizons. It maintains a measurable improvement at 30 minutes, despite increased stochasticity. These findings confirm that proactive bandwidth awareness is feasible on maritime platforms and can effectively support operational decisions such as adaptive streaming, routing, and resource scheduling. The performance gap between forecasting horizons also highlights the need for expanded offshore datasets to improve prediction robustness under harsher maritime environments.

97 MATHEMATICS AND COMPUTING↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Predictive Modeling and Operational Monitoring of Starlink Leo Satellite Network Performance in Maritime Environments

This thesis investigates short-term performance prediction for maritime LEO operational planning and application performance cannot anticipate available bandwidth or communication delay, which complicates variation alter signal conditions over short time scales. As a result, maritime users often fluctuations in link quality. In addition, Frequent satellite handoffs and Doppler-induced rapid satellite movement and changing atmospheric conditions cause substantial unavailable. Despite their availability, reliable operation at sea remains difficult because broadband connectivity in maritime regions where terrestrial infrastructure in

97 MATHEMATICS AND COMPUTING↗

Validation and Verification for INL Modelica-based TEDS models Via Experimental Results

This report provides an overview on the verification and validation (V&V) of the Thermal Energy Distribution System (TEDS) model developed in the Modelica process modeling ecosystem using experimental data. Model development has led to the creation of a dynamic process model of the experimental TEDS facility housed within the Energy Systems Laboratory (ESL) at Idaho National Laboratory (INL). The model was then used during the preconstruction phase of the experimental effort to inform experimental design (e.g., insulation requirements, bypass line placement, expected performance of components) and to test innovative control schemes prior to the initial operation. The TEDS model developed in Modelica includes the primary components of the TEDS experimental unit: a 200kW Chromalox heater; a single-tank packed-bed thermal energy storage system filled with 0.125-inch alumina (Al2O3) beads; an ethylene-glycol-to-Therminol-66 heat exchanger; system piping; five control valves; and all associated temperature, pressure, and volumetric flow sensors. Using the Institute of Electrical and Electronics Engineers (IEEE) V&V methodologies, considered the gold standard in the engineering field, the model was verified using a combination of static analysis, spatial convergence, and regression tests. Then using dynamic time warping (DTW) initial runs to validate and tune the TEDS model versus the experiment were conducted. This tuning method was accomplished using the INL Risk Analysis Virtual ENvironment (RAVEN) software package. Tuning is required to account for physical phenomena that are less understood within the empirical heat transfer correlations. Through the commencement of this work, a systems-level model of TEDS with associated control systems, sensors, piping diameters, and component capabilities has been created. This model was utilized in the pre-experimental phase to inform system design, insulation thicknesses, and potential control schemes to operate the system effectively and safely. Then, initial experimental startup and operational data were used to demonstrate the validation and tuning methodology. This process demonstrates the classical two-step approach of a model informing experimental design followed by the experiment validation and tuning the model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Image Analysis for Rapid Assessment and Quality-Based Sorting of Corn Stover

Imaging in the visible spectrum is a low-cost tool that can be readily deployed for in-field or over-belt monitoring of biomass quality for bio-refining operations. Rapid image analysis coupled with innovative preprocessing may reduce the impacts of feedstock variability through identification of contaminants or other material attributes to guide selective sorting and quality management. Image analysis was employed to evaluate the quality of corn stover in red-green-blue (RGB) chromatic space. This study used controlled, bench-scale imaging as a proof-of-concept for rapid quality assessment of corn stover based on variations in material attributes, including chemical and physical attributes, that relate to biological degradation and soil contamination. Additionally, logistic regression-based classification algorithms were used to develop a method for biomass screening as a function of biological degradation or soil contamination. This study demonstrated the use of image analysis to extract features from RGB color space to investigate variations in critical material attributes from chemical composition of corn stover. Fourier transform infrared (FT-IR) suggested a correlation between red band intensity and biological degradation, while detailed surface texture analysis was found to distinguish among variations in ash. These insights offer promise for development of a rapid screening tool that could be deployed by farmers for in-field assessment of biomass quality or biorefinery operators for in-line sorting and process optimization.

09 BIOMASS FUELS↗

The Application of Machine Learning Techniques to Meteorological Forecasting

Fog and inland-penetrating sea-breezes occur often at SRS and have a strong impact on site operations. Site personnel therefore require accurate forecasts of these events, but both are difficult to forecast using traditional techniques. Our goal is to apply machine learning (ML) techniques to the problem of forecasting fog and the sea breeze at the Savannah River Site. We apply several such techniques - decision trees, regression, and a series of classification/regression techniques – and train them using the large datasets collected by our group at SRS and from external organizations that maintain databases of regional meteorological variables.

54 ENVIRONMENTAL SCIENCES↗

Differential Property Prediction: A Machine Learning Approach to Experimental Design in Advanced Manufacturing

Advanced manufacturing techniques have enabled the production of materials with state-of-the-art properties. In many cases however, the development of physics-based models of these techniques lags behind their development in the lab. This means that material and process development proceeds largely via trial and error. This is sub-optimal since experiments are cost-, time-, and labor-intensive. In this work we propose a machine learning framework, differential property classification (DPC), which enables an experimenter to leverage machine learning's unparalleled pattern matching capability to pursue data-driven experimental design. DPC takes two possible experiment parameter sets and outputs a prediction of which will produce a material with a more desirable property specified by the operator. We demonstrate the success of DPC on AA7075 tube manufacturing process and mechanical property data using shear assisted processing and extrusion (ShAPE), an emerging solid phase processing technology. We show that by focusing on the experimenter's need to choose between multiple candidate experimental parameters, we can reframe the challenging regression task of predicting material properties from processing parameters, into a classification task on which machine learning models can achieve good performance.

advanced manufacturing, machine learning, ShAPE↗

What to expect when you're expecting engagement: Delivering procedural justice in large-scale solar energy deployment

Community engagement in the planning process to build large-scale solar (LSS) projects can win local support and advance procedural justice. However, an understanding of community engagement in current LSS development is lacking. Using responses from a U.S. nationwide survey (n = 979) of residential neighbors living within 3 miles (4.8 km) of completed LSS projects (i.e. “solar neighbors”) and project details from the U.S. Large-Scale Solar Photovoltaic Database (USPVDB), this study seeks to answer the following questions: How are solar neighbors' perceptions of community engagement associated with their attitudes toward their LSS projects? How do solar neighbors' perceptions of community engagement compare to their expectations? And, how do neighbors explain what they perceived about the planning process? We answer these questions using mixed methods, including regression modeling, a new gap analysis technique, and qualitative coding. We find that higher perceived engagement is associated with more positive attitudes toward the project, even when controlling for respondents who acted in opposition. Supporters and opponents alike expect more engagement than they perceived and information about projects both before construction and after operation is lacking. Awareness and engagement expectations increase at certain project size and proximity thresholds. However, most neighbors expect the public to offer input during engagement, but not make decisions. We contextualize these findings with explanatory comments from respondents.

14 SOLAR ENERGY↗

Generalized Tensor-on-Tensor Regression (GToTR)

SAND2026-23069O Generalized Tensor-on-Tensor Regression (GToTR) is a Python-based tool for conducting generalized tensor-on-tensor regression. It provides Canonical Polyadic (CP)-based generalized tensor regression models, support for generalized linear model-like families and links, alternating-optimization model fitting methods, and a standard statistics software interface. The tool supports tensor-valued responses and covariates using the open-source Python Tensor Toolbox (pyttb) software package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dunlavy, Daniel [Sandia National Lab. (SNL-CA), Li↗

Stand validation of lidar forest inventory modeling for a managed southern pine forest

We evaluated area-based approaches (ABAs) to light detection and ranging (lidar) predictions of plot- and stand-level forest attributes (tree count, height, basal area, volume, aboveground biomass, broadleaf/conifer, and diameter at breast height — “diameter”). ABA methods included post-stratification (PS), ordinary least squares (OLSs) regression, k nearest neighbors ( kNN), and random forest (RF). This study was conducted on the Savannah River Site in South Carolina, USA. Plot- and stand-level predictions were validated against fixed-radius 0.04 ha (0.1 acre) plots in 49 ≈2.0 ha (5 acre) stands. Our findings demonstrate that lidar can be incorporated operationally into forest inventory systems to provide stand-level inferences for a wide range of forest attributes. Volume predictions for specific diameter classes, however, often fared poorly (root mean squared error (RMSE) > 100%) for the methods we explored, especially for larger (less common) diameter trees. Stand-level results were consistently better than pixel-level results (10–200+ percentage points). kNN and RF performed similarly and better than OLS and PS, but RF was the most robust to model configurations, while kNN has practical advantages such as simultaneous predictions of many attributes.

Forestry↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

Plant-level performance and degradation of 31 GW DC of utility-scale PV in the United States

In this updated study, which samples 50% more capacity than the original and adds two additional years of operating history, we assess the performance of a fleet of 631 utility-scale PV plants totaling 31.0 GW DC (23.6 GW AC ) of capacity that achieved commercial operations in the United States from 2007-2018 and that have operated for at least two full calendar years. We use detailed information on individual plant characteristics, in conjunction with modeled irradiance data, to model expected or “ideal” capacity factors in each full calendar year of each plant’s operating history. A comparison of ideal versus actual first-year capacity factors finds that this fleet has modestly underperformed initial expectations (as modeled) on average, though perhaps due as much to modeling issues as to actual underperformance. We then analyze fleet-wide performance degradation in subsequent years by employing a “fixed effects” regression model to statistically isolate the impact of age on plant performance. The resulting average fleet-wide degradation rate of -1.2%/year (±0.1%) represents a slight improvement (seemingly driven by the oldest plants in our sample) over the -1.3%/year (±0.2%) found in our original study, yet is still of greater magnitude than is commonly found. We emphasize, however, that these fleet-wide estimates reflect both recoverable and unrecoverable degradation across the entire plant, and so will naturally be of greater magnitude than module- or cell-level studies, and/or studies that focus only on unrecoverable degradation. Moreover, when focusing on a sub-sample of newer and larger plants with higher DC:AC ratios—i.e., plants that more-closely resemble what is being built today—we find a more moderate sample-wide average performance decline of -0.7%/year (±0.4%), which is more in line with other estimates from the recent literature.

14 SOLAR ENERGY↗

Machine-Learning-Driven, Site-Specific Weather Forecasting for Grid-Interactive Efficient Buildings: Preprint

Emerging grid-interactive efficient buildings (GEBs) have great potential to provide much-needed demand flexibility to electric grids while fulfilling their own control targets by co-optimizing smart appliances, solar photovoltaics, electric vehicles, and energy storage at buildings. To enable the optimal operation of GEBs, site-specific weather information—such as temperature, solar irradiance, relative humidity, and wind speed—is crucial; however, this information is generally unavailable or expensive to obtain. This paper develops advanced machine learning methods to provide precise weather forecasts for individual building sites using readily available weather station data. Support vector regression and artificial neural networks have been employed to learn the spatiotemporal correlations between the weather conditions at nearby weather stations and the individual building site. The proposed site-specific weather forecasting methods have been validated using 1-year actual weather measurement data collected in the Denver metro area. Results show that the developed machine-learning-driven methods can accurately forecast the temperature at the target building site 1 hour ahead with mean absolute error less than 0.72°C and a 48% improvement over the persistence method. Site-specific weather forecasts will improve the understanding of the microclimate effect and its impact on building energy consumption. This information will drive efficiency upgrades and adjustments of building control strategies to improve energy savings and increase flexibility in building loads.

30 DIRECT ENERGY CONVERSION↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Hybrid Energy System Workflow for Energy Portfolio Optimization

This manuscript develops a workflow, driven by data analytics algorithms, to support the optimization of the economic performance of an Integrated Energy System. The goal is to determine the optimum mix of capacities from a set of different energy producers (e.g., nuclear, gas, wind and solar). A stochastic-based optimizer is employed, based on Gaussian Process Modeling, which requires numerous samples for its training. Each sample represents a time series describing the demand, load, or other operational and economic profiles for various types of energy producers. These samples are synthetically generated using a reduced order modeling algorithm that reads a limited set of historical data, such as demand and load data from past years. Numerous data analysis methods are employed to construct the reduced order models, including, for example, the Auto Regressive Moving Average, Fourier series decomposition, and the peak detection algorithm. All these algorithms are designed to detrend the data and extract features that can be employed to generate synthetic time histories that preserve the statistical properties of the original limited historical data. The optimization cost function is based on an economic model that assesses the effective cost of energy based on two figures of merit: the specific cash flow stream for each energy producer and the total Net Present Value. An initial guess for the optimal capacities is obtained using the screening curve method. The results of the Gaussian Process model-based optimization are assessed using an exhaustive Monte Carlo search, with the results indicating reasonable optimization results. The workflow has been implemented inside the Idaho National Laboratory’s Risk Analysis and Virtual Environment (RAVEN) framework. The main contribution of this study addresses several challenges in the current optimization methods of the energy portfolios in IES: First, the feasibility of generating the synthetic time series of the periodic peak data; Second, the computational burden of the conventional stochastic optimization of the energy portfolio, associated with the need for repeated executions of system models; Third, the inadequacies of previous studies in terms of the comparisons of the impact of the economic parameters. The proposed workflow can provide a scientifically defendable strategy to support decision-making in the electricity market and to help energy distributors develop a better understanding of the performance of integrated energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Open Call LDRD: Physically Informed Autoencoders for Galactic Redshift Regression

Physical constraints have been suggested to make neural network models more generalizable, act scientifically plausible, and be more data-efficient over unconstrained baselines. In this report, we present preliminary work on evaluating the effects of adding soft physical constraints to computer vision neural networks trained to estimate the conditional density of redshift on input galaxy images for the Sloan Digital Sky Survey. We introduce physically motivated soft constraint terms that are not implemented with differential or integral operators. We frame this work as a simple ablation study where the effect of including soft physical constraints is compared to an unconstrained baseline. We compare networks using standard point estimate metrics for photometric redshift estimation, as well as metrics to evaluate how faithful our conditional density estimate represents the probability over the ensemble of our test dataset. We find no evidence that the implemented soft physical constraints are more effective regularizers than augmentation.

97 MATHEMATICS AND COMPUTING↗

Decarbonizing or illusion? How carbon emissions of commercial building operations change worldwide

To lead the low-carbon transition in global buildings, this study is the first to use the generalized Divisia index method (GDIM) to identify the factors driving carbon emissions and assess the decarbonization performance in commercial building operations (CBOs) of sixteen countries during 2000–2019. Results show that (1) while the global carbon emissions from CBOs have increased at a modest rate of 0.9%/yr, this trend runs counter to the declining emissions in the U.S. (-1.1%/yr) and the significant growth in China (14.4%/yr), which can be attributed to the impact of economy-related factors. (2) The U.S. and China, as the largest emitters, contributed 66.8% of the samples’ decarbonization of CBOs (919.1 million tons of carbon dioxide). (3) Most countries’ CBOs had a decarbonization efficiency level of less than 10%, except for Spain (27.8%). Spain excelled in both areas of per capita and per floor area with an average efficiency of nearly 30%. Moreover, ridge regression successfully confirms the stability of GDIM results and it should be noted GDIM has limitations in characterizing the end-use activity. Overall, this study tracks the historical decarbonization of global CBOs and offers benchmarks for different emitters to forecast the dynamic of building emissions along with economic booms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Toward Discovering the Structure and Dynamics of the Sun-Interstellar Medium System (Los Alamos LDRD Report)

We have taken major steps toward developing the tools necessary to explore with unprecedented resolution the plasma structure and dynamics of the Sun-interstellar medium (ISM) interaction region, known as the heliosheath. We have applied a rigorous imaging technique known as Generalized Ridge Regression (GRR) to construct statistically robust sky maps of hydrogen energetic neutral atoms (ENAs) emanating from this region and detected by the Los Alamos-led IBEX-Hi ENA imager [1] on NASA’s Interstellar Boundary Explorer (IBEX) mission [2]. Our methods go far beyond the map reconstruction process currently applied by the IBEX Science Operations Center (ISOC), opening up the possibility of new discovery science with the IBEX data set and positioning LANL for a lead science role for the upcoming Interstellar Mapping and Acceleration Probe (IMAP) mission [3] for which LANL is providing two key experiments. We have successfully demonstrated that the new maps are higher resolution and are capable of revealing structures that are not presently resolveable in standard ISOC maps

58 GEOSCIENCES↗