Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Dynamic Ensemble Prediction of Cognitive Performance in Space

Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.

Danni Tu↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗

Urine monitoring system failure analysis and operational verification test report

Failure analysis and testing of a prototype urine monitoring system (UMS) are reported. System performance was characterized by a regression formula developed from volume measurement test data. When the volume measurement test data. When the volume measurement data was imputted to the formula, the standard error of the estimate calculated using the regression formula was found to be within 1.524% of the mean of the mass of the input. System repeatability was found to be somewhat dependent upon the residual volume of the system and the evaporation of fluid from the separator. The evaporation rate was determined to be approximately 1cc/minute. The residual volume in the UMS was determined by measuring the concentration of LiCl in the flush water. Observed results indicated residual levels in the range of 9-10ml, however, results obtained during the flushing efficiency test indicated a residual level of approximately 20ml. It is recommended that the phase separator pumpout time be extended or the design modified to minimize the residual level.

Glanfield, E. J.↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

A Framework for Closed-Loop Optimization of an Automated Mechanical Serial-Sectioning System via Run-to-Run Control as Applied to a Robo-Met.3D

Optimization of automated data collection is gaining increased interest for the purposes of enabling closed-loop self-correcting systems that inherently maximize operational efficiencies and reduce waste. Many data collection systems have several variables which influence data accuracy or consistency and which can require frequent user interaction to be monitored and maintained. Operating upon a Robo-MET.3D™ automated mechanical serial-sectioning system, a run-to-run control algorithm has been developed to accelerate data collection and reduce data inconsistency. Here, using historical data amassed over a decade of experiments, a linear regression model of the deterministic system dynamics is created and used to employ a run-to-run control algorithm that optimizes selected system inputs to reduce operator intervention and increase efficacy while reducing variance of system output.

42 ENGINEERING↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different datasets, i.e., 1) the unit generation dataset and 2) the system total inertia and headroom dataset, for the prediction of the frequency nadir. The training and testing datasets are generated through extensive generation scheduling simulations using Multi-timescale Integrated Dynamic and Scheduling (MI-DAS) toolbox on the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. Numerical results show that all five machine learning methods perform well in predicting the nadir frequency of the system. Among them, the gradient boosting and the XGBoost are clear winners yielding the best prediction accuracy in terms of four evaluation metrics.

data driven↗

Health Monitoring System for the SSME-fault detection algorithms

A Health Monitoring System (HMS) Framework for the Space Shuttle Main Engine (SSME) has been developed by United Technologies Corporation (UTC) for the NASA Lewis Research Center. As part of this effort, fault detection algorithms have been developed to detect the SSME faults with sufficient time to shutdown the engine. These algorithms have been designed to provide monitoring coverage during the startup, mainstage and shutdown phases of the SSME operation. The algorithms have the capability to detect multiple SSME faults, and are based on time series, regression and clustering techniques. This paper presents a discussion of candidate algorithms suitable for fault detection followed by a description of the algorithms selected for implementation in the HMS and the results of testing these algorithms with the SSME test stand data.

Tulpule, S.↗

Interactive Software Fault Analysis Tool for Operational Anomaly Resolution

Resolving software operational anomalies frequently requires a significant amount of resources for software troubleshooting activities. The time required to identify a root cause of the anomaly in the software may lead to significant timeline impacts and in some cases, may extend to compromise of mission and safety objectives. An integrated tool that supports software fault analysis based on the observed operational effects of an anomaly could significantly reduce the time required to resolve operational anomalies; increase confidence for the proposed solution; identify software paths to be re-verified during regression testing; and, as a secondary product of the analysis, identify safety critical software paths.

Chen, Ken↗

Productivity Analysis of Public and Private Airports: A Causal Investigation

Around the world, airports are being viewed as enterprises, rather than public services, which are expected to be managed efficiently and provide passengers with courteous customer services. Governments are, increasingly, turning to the private sectors for their efficiency in managing the operation, financing, and development, as well as providing security for airports. Operational and financial performance evaluation has become increasingly important to airport operators due to recent trends in airport privatization. Assessing performance allows the airport operators to plan for human resources and capital investment as efficiently as possible. Productivity measurements may be used as comparisons and guidelines in strategic planning, in the internal analysis of operational efficiency and effectiveness, and in assessing the competitive position of an airport in transportation industry. The primary purpose of this paper is to investigate the operational and financial efficiencies of 22 major airports in the United States and Europe. These airports are divided into three groups based on private ownership (7 British Airport Authority airports), public ownership (8 major United States airports), and a mix of private and public ownership (7 major European Union airports. The detail ownership structures of these airports are presented in Appendix A. Total factor productivity (TFP) model was utilized to measure airport performance in terms of financial and operational efficiencies and to develop a benchmarking tool to identify the areas of strength and weakness. A regression model was then employed to measure the relationship between TFP and ownership structure. Finally a Granger causality test was performed to determine whether ownership structure is a Granger cause of TFP. The results of the analysis presented in this paper demonstrate that there is not a significant relationship between airport TFP and ownership structure. Airport productivity and efficiency is, however dependent upon the level of competition, choice of the market, and regulatory control.

Vasigh, Bijan↗

Estimation of Stability and Control Derivatives of an F-15

A technique for real-time estimation of stability and control derivatives (derivatives of moment coefficients with respect to control-surface deflection angles) was used to support a flight demonstration of a concept of an indirect-adaptive intelligent flight control system (IFCS). Traditionally, parameter identification, including estimation of stability and control derivatives, is done post-flight. However, for the indirect-adaptive IFCS concept, parameter identification is required during flight so that the system can modify control laws for a damaged aircraft. The flight demonstration was carried out on a highly modified F-15 airplane (see Figure 1). The main objective was to estimate the stability and control derivatives of the airplane in nearly real time. A secondary goal was to develop a system to automatically assess the quality of the results, so as to be able to tell a learning neural network which data to use. Parameter estimation was performed by use of Fourier-transform regression (FTR) a technique developed at NASA Langley Research Center. FTR is an equation- error technique that operates in the frequency domain. Data are put into the frequency domain by use of a recursive Fourier transform for a discrete frequency set. This calculation simplifies many subsequent calculations, removes biases, and automatically filters out data beyond the chosen frequency range. FTR as applied here was tailored to work with pilot inputs, which produce correlated surface positions that prevent accurate parameter estimates, by replacing half the derivatives with predicted values. FTR was also set up to work only on a recent window of data, to accommodate changes in flight condition. A system of confidence measures was developed to identify quality-parameter estimates that a learning neural network could use. This system judged the estimates primarily on the basis of their estimated variances and of the level of aircraft response. The resulting FTR system was implemented in the Simulink software system and auto-coded in the C programming language for use on the Airborne Research Test System (ARTS II) computer installed in the F-15 airplane. The Simulink model was also used in a control room that utilizes the Ring Buffered Network Bus hardware and software, making it possible to evaluate test points during flights. In-flight parameter estimation was done for piloted and automated maneuvers, primarily at three test conditions. Figure 2 shows results for pitching moment due to symmetric stabilator actuations for a series of three pitch doublet maneuvers (in a doublet maneuver, a command to change attitude in a given direction by a given amount is followed immediately by a command to change attitude in the opposite direction by the same amount). A time window of 5 seconds was used. The portions of the curves shown in red are those that passed the confidence tests. The technique showed good convergence for most derivatives for both kinds of maneuvers - typically within a few seconds. The confidence tests were marginally successful, and it would be necessary to refine them for use in an IFCS.

Smith, Mark↗

Utilizing Airborne and Space-Based Remote Sensing Imagery to Implement the Unvegetated-Vegetated Ratio to Assess Salt Marsh Vulnerability in South Carolina

Among the most productive ecosystems on earth, salt marshes provide crucial ecosystem services including water filtration, shoreline protection, storm surge buffering, and flood mitigation. Marshes are largely dependent on their sediment budget which can significantly vary across a region and can be used to determine the life span of the marsh. Upstream land use change near Charleston, South Carolina, along with rising sea levels, are expected to alter sediment budgets and threaten marsh stability and long-term health. The unvegetated-vegetated ratio (UVVR), developed by researchers at USGS, is a scalable and efficient method to assess vulnerability. The NASA DEVELOP National Program collaborated with the South Carolina Department of Natural Resources, the South Carolina Department of Health and Environmental Control, and the United States Geological Survey Woods Hole Coastal and Marine Science Center to apply the UVVR method within Google Earth Engine. Marsh vulnerability was analyzed using UVVR derived from clustering and manual interpretation of National Agriculture Imagery Program (NAIP) high-resolution aerial imagery. NAIP derived UVVR was aggregated to Landsat 8 Operational Land Imager (OLI) and Landsat 7 Enhanced Thematic Mapper (ETM+) resolution and projection. A Random Forest Regression between Landsat derived data and UVVR was modeled to estimate a potential relationship. The estimation of this relationship was used to produce temporal change analysis maps of salt marsh vulnerability back to 1984. The NAIP imagery processed through Google Earth Engine allowed us to make detailed UVVR maps for 2009, 2015, 2017, and 2019 for decision making within South Carolina. Google Earth Engine scripting provided a novel approach to UVVR methodology that will allow decision makers to input new marsh regions and easily calculate marsh vulnerability without external data downloading. These results were used to understand what areas of the marsh need most resource allocation in the future.

NASA DEVELOP↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems: Preprint

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different sets of preprocess data for the prediction of the frequency nadir in the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. The training and testing data sets are collected by extensive generation scheduling simulations on the Multi-timescale Integrated Dynamic and Scheduling (MIDAS) toolbox. Numerical results show that all five machine learning methods can achieve high performance accuracy for power system nadir frequency estimation. Among them, the gradient boosting and the XGBoost are clear winners by providing the best prediction accuracy.

data driven↗

Situational awareness-enhancing community-level load mapping with opportunistic machine learning

Motivated by present and forthcoming challenges in the adoption and integration of distributed renewable energy, we develop a machine learning (ML) approach that builds short-fuse mappings connecting the occasionally-unobservable true load in one target community with information-rich signals collected from relatively more instrumented reference communities. Our setting is inspired by and tailored to target communities with significant unobservable behind-the-meter solar generation, where true load (a relatively well-behaved quantity of interest to grid operators) is hard to discern during daytime due to insufficient instrumentation and/or privacy reasons, but that can be related to reference communities with low unobservable distributed variable generation or with sufficient instrumentation. The developed mapping, herein realized with Support Vector Machine regression, is built using nighttime data from all communities, when their distributed generation is low or zero. Our ML algorithm opportunistically learns to correlate signals of interest and then is operationally used the next day to shed light into target community load evolution. The mapping is subsequently rebuilt, rolling its short-fuse scope perpetually forward in time. Here, we demonstrate the efficacy of our approach on nine synthetically generated topologies and associated timeseries stemming from real-world data, on which we observe cumulative error performance that yields lower than 10% and 15% daily-averaged mean absolute percentage errors in target community load estimation on more than about 75% and 90% of days, respectively, in multiple yearly evaluations that shed light on long-term performance also under seasonal and one-off effects. The proposed ML-powered methodology can offer grid operators much-improved visibility into a previously obscure space and can also serve as an additional source of information in broader, multi-modal solar disaggregation solutions.

14 SOLAR ENERGY↗

Northern Rockies Ecological Conservation: Leveraging Earth Observations to Monitor and Predict Populations of Federally Threatened Whitebark Pine (Pinus albicaulis) across the Intermountain West

Whitebark pine (WBP; Pinus albicaulis) is an ecologically important species in North America. As a federally listed threatened species, an understanding of WBP habitat, distribution, and health is important for the natural resource managers of the National Park Service, United States Forest Service, Bureau of Land Management, Fish and Wildlife Service, and non-profit organizations such as the Whitebark Pine Ecosystem Foundation. Previous attempts to develop models of WBP habitat suitability and distribution lack confidence in their validity and integrity for these organizations. The updated models of habitat suitability and distribution developed by this study would provide managers with a capability to be employed in the conservation and future research direction for WBP. Thus, we developed a habitat suitability model of WBP at a high spatial resolution (Landsat 9 Operational Land Image-2, National Land Cover Database, NASA Shuttle Radar Topography Mission; 30m pixels) using a generalized logistic regression with an area under the curve value of 0.754. We extracted spectral reflectance signatures from overlapped ground sample points and Sentinel-2 Multispectral Instrument. The spectral signature analysis indicates WBP is separable from other tree species. We also utilized a visual validation approach and random forest (RF) modeling to separate WBP from limber pine. Through visual validation the RF classifier successfully identified 8out of 10 WBP trees gathered through ground truth points. Additionally, we achieved an overall accuracy of 91%in our confusion matrix for the distribution model using a dependent validation approach. The derived products from this study allow project partners to assess current suitable habitat and apparent health status in areas of identified WBP occurrence, providing data to aid future research regarding WBP health.

Sentinel-2↗

A Robust Segmented Mixed Effect Regression Model for Baseline Electricity Consumption Forecasting

Renewable energy production has been surging around the world in recent years. To mitigate the increasing uncertainty and intermittency of the renewable generation, proactive demand response algorithms and programs are proposed and developed to further improve the utilization of load flexibility and increase the efficiency of power system operation. One of the biggest challenges to efficient control and operation of demand response resources is how to forecast the baseline electricity consumption and estimate the load impact from demand response resources accurately. In this paper, we propose a mixed effect segmented regression model and a new robust estimate for forecasting the baseline electricity consumption in Southern California, USA, by combining the ideas of random effect regression model, segmented regression model, and the least trimmed squares estimate. Since the log-likelihood of the considered model is not differentiable at breakpoints, we propose a new backfitting algorithm to estimate the unknown parameters. The estimation performance of the new estimation procedure has been demonstrated with both simulation studies and the real data application for the electric load baseline forecasting in Southern California.

42 ENGINEERING↗

Congenital malformation and hemoglobin A1c in the first trimester among Japanese women with pregestational diabetes

Abstract Aim To investigate the incidence of major congenital malformations in Japanese women with pregestational diabetes, and to determine the cutoff value of hemoglobin A1c (HbA1c) in the first trimester associated with congenital malformations. Methods This retrospective cohort study included singleton pregnancies in Japanese women with pregestational diabetes, including type 1 and type 2 diabetes, and specific types of diabetes due to other causes. The primary outcome was the incidence of major congenital malformations. The secondary outcome was the incidence of all congenital malformations. The cutoff value of HbA1c for congenital malformations was calculated using receiver operating characteristic curve analysis. The adjusted odds ratios (aOR) of major congenital malformations were calculated using multiple logistic regression analyses. Results This study enrolled 292 patients, including 132 (45.2%) with type 1 diabetes, 156 (53.4%) with type 2 diabetes, and 4 (1.4%) with other specific types. The incidence rates of major congenital malformations and all congenital malformations were 7.2% (21/292) and 12.7% (37/292), respectively. The cutoff value of HbA1c in the first trimester for major malformations and for all congenital malformations was 6.5%. HbA1c ≥ 6.5% was significantly associated with major malformations (aOR 3.5; 95% confidence interval: 1.2–12.6; p = 0.018). Conclusion The incidence of major congenital malformations significantly increased in pregnant Japanese women with HbA1c values of 6.5% or higher. The recommended HbA1c value during the first trimester used in other countries can be applied to pregnant Japanese women.

Nakanishi, Kentaro↗