Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Hypothesis-Agnostic Network-Based Analysis of Real-World Data Suggests Ondansetron is Associated with Lower COVID-19 Any Cause Mortality

Background: The COVID-19 pandemic generated a massive amount of clinical data, which potentially hold yet undiscovered answers related to COVID-19 morbidity, mortality, long-term effects, and therapeutic solutions.Objectives: The objectives of this study were (1) to identify novel predictors of COVID-19 any cause mortality by employing artificial intelligence analytics on real-world data through a hypothesis-agnostic approach and (2) to determine if these effects are maintained after adjusting for potential confounders and to what degree they are moderated by other variables.Methods: A Bayesian statistics-based artificial intelligence data analytics tool (bAIcis®) within the Interrogative Biology® platform was used for Bayesian network learning and hypothesis generation to analyze 16,277 PCR+ patients from a database of 279,281 inpatients and outpatients tested for SARS-CoV-2 infection by antigen, antibody, or PCR methods during the first pandemic year in Central Florida. This approach generated Bayesian networks that enabled unbiased identification of significant predictors of any cause mortality for specific COVID-19 patient populations. These findings were further analyzed by logistic regression, regression by least absolute shrinkage and selection operator, and bootstrapping.Results: We found that in the COVID-19 PCR+ patient cohort, early use of the antiemetic agent ondansetron was associated with decreased any cause mortality 30 days post-PCR+ testing in mechanically ventilated patients.Conclusions: The results demonstrate how a real-world COVID-19-focused data analysis using artificial intelligence can generate unexpected yet valid insights that could possibly support clinical decision making and minimize the future loss of lives and resources.

60 APPLIED LIFE SCIENCES↗

Interpretable and flexible non-intrusive reduced-order models using reproducing kernel Hilbert spaces

This paper develops an interpretable, non-intrusive reduced-order modeling technique using regularized kernel interpolation. Existing non-intrusive approaches approximate the dynamics of a reduced-order model (ROM) by solving a data-driven least-squares regression problem for low-dimensional matrix operators. Our approach instead leverages regularized kernel interpolation, which yields an optimal approximation of the ROM dynamics from a user-defined reproducing kernel Hilbert space. We show that our kernel-based approach can produce interpretable ROMs whose structure mirrors full-order model structure by embedding judiciously chosen feature maps into the kernel. The approach is flexible and allows a combination of informed structure through feature maps and closure terms via more general nonlinear terms in the kernel. We also derive a computable a posteriori error bound that combines standard error estimates for intrusive projection-based ROMs and kernel interpolants. In conclusion, the approach is demonstrated in several numerical experiments that include comparisons to operator inference using both proper orthogonal decomposition and quadratic manifold dimension reduction.

Data-driven model reduction↗

A Data-Driven Exploration of the Impact of Renewable Energy on Inter-Area Oscillations in the U.S. Eastern Interconnection

As increasing amounts of renewable energy (RE) resources are incorporated into the bulk-power grid, power system oscillations are expected to change. This work investigates how RE generation impacts the frequency and damping ratio (DR) of two dominant inter-area modes in the U.S. Eastern Interconnection (EI) using regularly updated estimates collected over a 12-month period. Quantile regression is used to derive the correlation between operating conditions and mode properties, and a bootstrap method is used to quantify the uncertainty associated with the correlation estimates. Results show that with an increase in system load, the frequency of a mode decreases and DR increases. Evidence that increasing RE generation results in an increase in frequency and decline in DR was found for one of the two modes studied. This work shows that increasing RE levels will impact the properties of inter-area oscillations in the EI, but it does not indicate the presence of immediate threats to grid stability. The outlined approach can be used to periodically assess changing mode properties as RE levels continue to grow and flag stability concerns before they become serious reliability threats.

Inter-area oscillation, mode meters, quantile regr↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage↗

Mass Spectrometer Transient Analysis

This software implements a complete preprocessing pipeline for transient mass spectrometry (MS) data collected during TAP (Temporal Analysis of Products) experiments. It is designed to extract chemically meaningful fluxes from overlapping ion signals by applying a calibrated defragmentation matrix and solving the resulting linear system using non-negative least squares (NNLS) regression. The core script, preprocess_mass_spec.py, performs the following operations: Gain correction: Applies amplifier gain scalars derived from inert-packed calibration pulses to normalize signal intensities across AMUs and acquisition settings. Background subtraction: Removes experiment baselines using user-defined time windows, ensuring compatibility with slow-diffusing species and preventing negative values that would interfere with NNLS. Options to subtract before and after defragmentation. Defragmentation: Constructs a fragmentation matrix A from zeroth moments of calibration pulses (equal molar gas:inert mixtures) and solves Ax=b at each time point, where b is the raw MS signal and x is the estimated species flux. The matrix is normalized to inert signals and accounts for instrument-specific fragmentation behavior. Pulse-mode handling: Supports both averaged and individual pulse modes, enabling statistical treatment of fluxes and calculation of standard deviations. Integration and output: Computes zeroth moments (integrated fluxes) and exports time-resolved and integrated data in CSV format, suitable for downstream kinetic modeling. The software is validated using both virtual TAP simulations (VTAP) and experimental data from propane dehydrogenation (PDH) on CrOx/Al2O3 catalysts. It preserves temporal resolution by applying NNLS point-by-point across the pulse duration (typically 6,000+ time slices per pulse), leveraging the linear superposition principle to reconstruct full flux profiles. The defragmented outputs are compatible with kinetic extraction methods such as the G and Y procedures, which are used to derive rate–concentration relationships from TAP data. The details of these validations are discussed in detail in the supporting manuscript and supporting information. Example data and output files are also included. The methodology is robust to experimental noise and drift, with calibration protocols that account for pulse size effects, MS aging, and inert gas normalization. The software is modular, reproducible, and tailored for high-throughput TAP-MS workflows in catalysis research.

Kristy, Stephen [Idaho National Laboratory (INL), ↗

The Financial Performance of Family versus Non-Family Firms Operating in Nautical Tourism

This article analyses the financial performance of family versus non-family firms operating in nautical tourism, in 2015–2019. The sample of 39 Portuguese companies was collected from the SABI database. We use a regression of financial performance, measured by three alternative proxies: return on assets, return on equity and operating profit margin, on liquidity, leverage, turnover of assets, asset structure, company size and age. The regressions are performed across Nuts II regions on mainland and across types of firms (family and non-family). The results uncover several patterns. First, family firms are larger and older, make higher investments and therefore are less liquid. Second, liquidity, leverage and investment in tangible assets impact negatively and significantly the corporate financial performance, while the turnover of assets, size and age impacts positively and significantly. Third, the sign of the impacts depends on the measure of performance. Finally, firms in the Northern region show superior performance, which can be explained by the higher share of family firms. These findings can serve as a roadmap for managers when selecting strategies to improve performance. Additionally, they will contribute to the understanding of tourism destination dynamics and competitiveness.

Santos, Eleonora (ORCID:0000000346930804)↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Benefit Analysis of CO 2 Delivery Options for Offshore Storage or Enhanced Oil Recovery

The analysis presented in this report evaluates the benefits of CO₂ offshore transport via pipeline or ship within the GOM. It takes a top-down framework to estimate the costs. First, this analysis designed a reduced-order model (ROM) based on the cash flows in the FECM/NETL CO₂ Transport Cost Model (also known as CO2_T_COM). The ROM takes capital expenses (CAPEX) and operating expenses (OPEX) to calculate the CO₂ breakeven price based on the cash flows. Second, this analysis developed regression models utilizing published data from other analyses to estimate CAPEX and OPEX. Since the ROM is a simplified cash flow calculation, it is easy to exchange the core regression models to estimate various costs. The ROM and regression models provided a framework that can be easily used by other researchers, decision-makers, operators, and regulators. The objective of this analysis is to assess the CO₂ breakeven cost range for pipeline and ship transport of captured CO₂ given the CO₂ source and storage reservoir located in the GOM.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Deep transfer operator learning for partial differential equations under conditional shift

Transfer learning enables the transfer of knowledge gained while learning to perform one task (source) to a related but different task (target), hence addressing the expense of data acquisition and labelling, potential computational power limitations and dataset distribution mismatches. Here, we propose a new transfer learning framework for task-specific learning (functional regression in partial differential equations) under conditional shift based on the deep operator network (DeepONet). Task-specific operator learning is accomplished by fine-tuning task-specific layers of the target DeepONet using a hybrid loss function that allows for the matching of individual target samples while also preserving the global properties of the conditional distribution of the target data. Inspired by conditional embedding operator theory, we minimize the statistical distance between labelled target data and the surrogate prediction on unlabelled target data by embedding conditional distributions onto a reproducing kernel Hilbert space. We demonstrate the advantages of our approach for various transfer learning scenarios involving nonlinear partial differential equations under diverse conditions due to shifts in the geometric domain and model dynamics. Our transfer learning framework enables fast and efficient learning of heterogeneous tasks despite considerable differences between the source and target domains.

42 ENGINEERING↗

QRF4P-NRT: Probabilistic Post-Processing of Near-Real-Time Satellite Precipitation Estimates Using Quantile Regression Forests

Accurate and reliable near-real-time satellite precipitation estimation is of great importance for operational large-scale flood forecasting and drought monitoring. The state-of-the-art precipitation post-processing model is based on a deterministic approach to construct relationships between satellites estimates and ground observations. We propose a probabilistic postprocessor, the Probabilistic Post-Processing of Near-Real-Time Satellite Precipitation Estimates using Quantile Regression Forests (QRF4P-NRT), based on quantile modeling, yielding both deterministic and probabilistic predictions. The experimental design incorporates different solutions of near-real-time predictors to further improve the model performance. Using the Integrated Multi-satellitE Retrievals Early Run for Global Precipitation Measurement Mission (IMERG-E) product as an example, we illustrate that the proposed method significantly improves the overall quality of the raw IMERG-E and is also superior to the bias-corrected product (IMERG Final Run, IMERG-F) at daily scale in a complex mountain basin. Evaluations of the corrected IMERG-E, raw IMERG-E, and IMERG-F using ground observation show that the corrected IMERG-E improves correlation coefficients (0.7), mean error (-0.14 mm/day) and root mean square error (3.3 mm/day) relative to the raw IMERG-E (0.31, -0.72 and 5.5 mm/day) and IMERG-F (0.34, -0.09 and 6.0 mm/day). The error decomposition further confirms that the QRF4P-NRT improves on the various deficiencies of the raw IMERG-E product. The ensemble assessment also demonstrates that the quantile outputs provide reliable prediction spread and sharp prediction intervals. The promising results indicate the great potential of the proposed method for probabilistic post-processing for near-real-time satellite precipitation estimates, and for further applications such as hydrological ensemble forecasting.

54 ENVIRONMENTAL SCIENCES↗

Impacts of climate change on subannual hydropower generation: a multi-model assessment of the United States federal hydropower plant

Abstract Hydropower is a low-carbon emission renewable energy source that provides competitive and flexible electricity generation and is essential to the evolving power grid in the context of decarbonization. Assessing hydropower availability in a changing climate is technically challenging because there is a lack of consensus in the modeling representation of key dynamics across scales and processes. Focusing on 132 US federal hydropower plants, in this study we evaluate the compounded impact of climate and reservoir-hydropower models’ structural uncertainties on monthly hydropower projections. In particular, instead of relying on one single regression-based hydropower model, we introduce another conceptual reservoir operations-hydropower model in the assessment framework. This multi-model assessment approach allows us to partition uncertainties associated with both climate and hydropower models for better clarity. Results suggest that while at least 70% of the uncertainties at the annual scale and 50% at the seasonal scale can be attributed to the choice of climate models, up to 50% of seasonal variability can be attributed to the choice of hydropower models, particularly in regions over the western US where the reservoir storage is substantial. The analysis identifies regions where multi-model assessments are needed and presents a novel approach to partition uncertainties in hydropower projections. Another outcome includes an updated evaluation of Coupled Model Intercomparison Project Phase 5 (CMIP5)-based federal hydropower projection, at the monthly scale and with a larger ensemble, which can provide a baseline for understanding future assessments based on CMIP6 and beyond.

54 ENVIRONMENTAL SCIENCES↗

An Artificial Intelligence-Assisted Method for Dementia Detection Using Images from the Clock Drawing Test

Background: Widespread dementia detection could increase clinical trial candidates and enable appropriate interventions. Since the Clock Drawing Test (CDT) can be potentially used for diagnosing dementia-related disorders, it can be leveraged to develop a computer-aided screening tool. Objective: To evaluate if a machine learning model that uses images from the CDT can predict mild cognitive impairment or dementia. Methods: Images of an analog clock drawn by 3,263 cognitively intact and 160 impaired subjects were collected during in-person dementia evaluations by the Framingham Heart Study. We processed the CDT images, participant’s age, and education level using a deep learning algorithm to predict dementia status. Results: When only the CDT images were used, the deep learning model predicted dementia status with an area under the receiver operating characteristic curve (AUC) of 81.3% ± 4.3%. A composite logistic regression model using age, level of education, and the predictions from the CDT-only model, yielded an average AUC and average F1 score of 91.9% ±1.1% and 94.6% ±0.4%, respectively. Conclusion: Our modeling framework establishes a proof-of-principle that deep learning can be applied on images derived from the CDT to predict dementia status. When fully validated, this approach can offer a cost-effective and easily deployable mechanism for detecting cognitive impairment.

Neurosciences & Neurology↗

Novel method for accurately estimating membrane transport properties and mass transfer coefficients in reverse osmosis

Here, we present a simple and robust method to simultaneously characterize the water and salt permeability (A, B) of reverse osmosis (RO) membranes and mass transfer coefficient (k) in membrane modules. The proposed methodology comprises a set of RO experiments performed at different operating pressures or stages. The measured water and salt fluxes in each stage are simultaneously fitted to the RO transport equations by performing a non-linear regression, using A, B, and k as regression parameters. We first perform a systematic accuracy analysis of the proposed method across the full operational range of RO. The assessment shows that the method accuracy is substantially higher than current methods and increases with number of experimental stages and driving forces. This assessment is used to inform the design of an experimental protocol that minimizes errors in estimated A, B, and k. We then evaluate two commercial RO membranes following the new protocol. For both membranes, A and B parameters decrease by 17% and 15% from the dilute solution to seawater concentrations, whereas the k parameter remains constant. Our study demonstrates that the proposed method, informed by data-driven experimental designs, provides a new approach for accurately characterizing transport phenomenon in membrane processes with feeds of less than 100 g/L total dissolved solids.

36 MATERIALS SCIENCE↗

Scalable Hybrid Classification-Regression Solution for High-Frequency Nonintrusive Load Monitoring

Residential buildings with the ability to monitor and control their net-load (sum of load and generation) can provide valuable flexibility to power grid operators. We present a novel multiclass nonintrusive load monitoring (NILM) approach that enables effective net-load monitoring capabilities at high-frequency with minimal additional equipment and cost. The proposed machine learning based solution provides accurate multiclass state predictions while operating at a faster timescale (able to provide a prediction for each 60- Hz ac cycle used in US power grid) without relying on event-detection techniques. We also introduce an innovative hybrid classification-regression method that allows for the prediction of not only load on/off states but also individual load operating power levels. A test bed with eight residential appliances is used for validating the NILM approach. Results show that the overall method has high accuracy, good scaling and generalization properties.

feature extraction↗

Non-intrusive nonlinear model reduction via machine learning approximations to low-dimensional operators

Abstract Although projection-based reduced-order models (ROMs) for parameterized nonlinear dynamical systems have demonstrated exciting results across a range of applications, their broad adoption has been limited by their intrusivity: implementing such a reduced-order model typically requires significant modifications to the underlying simulation code. To address this, we propose a method that enables traditionally intrusive reduced-order models to be accurately approximated in a non-intrusive manner. Specifically, the approach approximates the low-dimensional operators associated with projection-based reduced-order models (ROMs) using modern machine-learning regression techniques. The only requirement of the simulation code is the ability to export the velocity given the state and parameters; this functionality is used to train the approximated low-dimensional operators. In addition to enabling nonintrusivity, we demonstrate that the approach also leads to very low computational complexity, achieving up to $$10^3{\times }$$ 10 3 × in run time. We demonstrate the effectiveness of the proposed technique on two types of PDEs. The domain of applications include both parabolic and hyperbolic PDEs, regardless of the dimension of full-order models (FOMs).

42 ENGINEERING↗

BASS

SAND2026-17001O BASS implements Bayesian Adaptive Spline Surfaces in MATLAB and serves as a surrogate model for regression applications. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

BayesPPR

SAND2026-17002O BayesPPR performs Bayesian Projection Pursuit Regression (PPR) using MATLAB. A surrogate model for calibration applications, it enables users to efficiently analyze complex datasets and extract meaningful patterns through regression techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗