Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Intermediate Molecular Phenotypes to Identify Genetic Markers of Anthracycline-Induced Cardiotoxicity Risk

Cardiotoxicity due to anthracyclines (CDA) affects cancer patients, but we cannot predict who may suffer from this complication. CDA is a complex trait with a polygenic component that is mainly unidentified. We propose that levels of intermediate molecular phenotypes (IMPs) in the myocardium associated with histopathological damage could explain CDA susceptibility, so variants of genes encoding these IMPs could identify patients susceptible to this complication. Thus, a genetically heterogeneous cohort of mice (n = 165) generated by backcrossing were treated with doxorubicin and docetaxel. We quantified heart fibrosis using an Ariol slide scanner and intramyocardial levels of IMPs using multiplex bead arrays and QPCR. We identified quantitative trait loci linked to IMPs (ipQTLs) and cdaQTLs via linkage analysis. In three cancer patient cohorts, CDA was quantified using echocardiography or Cardiac Magnetic Resonance. CDA behaves as a complex trait in the mouse cohort. IMP levels in the myocardium were associated with CDA. ipQTLs integrated into genetic models with cdaQTLs account for more CDA phenotypic variation than that explained by cda-QTLs alone. Allelic forms of genes encoding IMPs associated with CDA in mice, including AKT1, MAPK14, MAPK8, STAT3, CAS3, and TP53, are genetic determinants of CDA in patients. Two genetic risk scores for pediatric patients (n = 71) and women with breast cancer (n = 420) were generated using machine-learning Least Absolute Shrinkage and Selection Operator (LASSO) regression. Thus, IMPs associated with heart damage identify genetic markers of CDA risk, thereby allowing more personalized patient management.

60 APPLIED LIFE SCIENCES↗

Serum bile acid and unsaturated fatty acid profiles of non-alcoholic fatty liver disease in type 2 diabetic patients

The understanding of bile acid (BA) and unsaturated fatty acid (UFA) profiles, as well as their dysregulation, remains elusive in individuals with type 2 diabetes mellitus (T2DM) coexisting with non-alcoholic fatty liver disease (NAFLD). Investigating these metabolites could offer valuable insights into the pathophy-siology of NAFLD in T2DM. Our aim is to identify potential metabolite biomarkers capable of distinguishing between NAFLD and T2DM. A training model was developed involving 399 participants, comprising 113 healthy controls (HCs), 134 individuals with T2DM without NAFLD, and 152 individuals with T2DM and NAFLD. External validation encompassed 172 participants. NAFLD patients were divided based on liver fibrosis scores. The analytical approach employed univariate testing, orthogonal partial least squares-discriminant analysis, logistic regression, receiver operating characteristic curve analysis, and decision curve analysis to pinpoint and assess the diagnostic value of serum biomarkers. Compared to HCs, both T2DM and NAFLD groups exhibited diminished levels of specific BAs. In UFAs, particular acids exhibited a positive correlation with NAFLD risk in T2DM, while the ω-6:ω-3 UFA ratio demonstrated a negative correlation. Levels of α-linolenic acid and γ-linolenic acid were linked to significant liver fibrosis in NAFLD. The validation cohort substantiated the predictive efficacy of these biomarkers for assessing NAFLD risk in T2DM patients. This study underscores the connection between altered BA and UFA profiles and the presence of NAFLD in individuals with T2DM, proposing their potential as biomarkers in the pathogenesis of NAFLD.

60 APPLIED LIFE SCIENCES↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

NCAPH drives breast cancer progression and identifies a gene signature that predicts luminal a tumour recurrence

Luminal A tumours generally have a favourable prognosis but possess the highest 10-year recurrence risk among breast cancers. Additionally, a quarter of the recurrence cases occur within 5 years post-diagnosis. Identifying such patients is crucial as long-term relapsers could benefit from extended hormone therapy, while early relapsers might require more aggressive treatment. We conducted a study to explore non-structural chromosome maintenance condensin I complex subunit H’s (NCAPH) role in luminal A breast cancer pathogenesis, both in vitro and in vivo, aiming to identify an intratumoural gene expression signature, with a focus on elevated NCAPH levels, as a potential marker for unfavourable progression. Our analysis included transgenic mouse models overexpressing NCAPH and a genetically diverse mouse cohort generated by backcrossing. A least absolute shrinkage and selection operator (LASSO) multivariate regression analysis was performed on transcripts associated with elevated intratumoural NCAPH levels. We found that NCAPH contributes to adverse luminal A breast cancer progression. The intratumoural gene expression signature associated with elevated NCAPH levels emerged as a potential risk identifier. Transgenic mice overexpressing NCAPH developed breast tumours with extended latency, and in Mouse Mammary Tumor Virus (MMTV)-NCAPH ErbB2 double-transgenic mice, luminal tumours showed increased aggressiveness. High intratumoural Ncaph levels correlated with worse breast cancer outcome and subpar chemotherapy response. A 10-gene risk score, termed Gene Signature for Luminal A 10 (GSLA10), was derived from the LASSO analysis, correlating with adverse luminal A breast cancer progression. The GSLA10 signature outperformed the Oncotype DX signature in discerning tumours with unfavourable outcomes, previously categorised as luminal A by Prediction Analysis of Microarray 50 (PAM50) across three independent human cohorts. This new signature holds promise for identifying luminal A tumour patients with adverse prognosis, aiding in the development of personalised treatment strategies to significantly improve patient outcomes.

60 APPLIED LIFE SCIENCES↗

Topological network analysis of patient similarity for precision management of acute blood pressure in spinal cord injury

Background: Predicting neurological recovery after spinal cord injury (SCI) is challenging. Using topological data analysis, we have previously shown that mean arterial pressure (MAP) during SCI surgery predicts long-term functional recovery in rodent models, motivating the present multicenter study in patients. Methods: Intra-operative monitoring records and neurological outcome data were extracted (n = 118 patients). We built a similarity network of patients from a low-dimensional space embedded using a non-linear algorithm, Isomap, and ensured topological extraction using persistent homology metrics. Confirmatory analysis was conducted through regression methods. Results: Network analysis suggested that time outside of an optimum MAP range (hypotension or hypertension) during surgery was associated with lower likelihood of neurological recovery at hospital discharge. Logistic and LASSO (least absolute shrinkage and selection operator) regression confirmed these findings, revealing an optimal MAP range of 76–[104-117] mmHg associated with neurological recovery. Conclusions: We show that deviation from this optimal MAP range during SCI surgery predicts lower probability of neurological recovery and suggest new targets for therapeutic intervention. Funding: NIH/NINDS: R01NS088475 (ARF); R01NS122888 (ARF); UH3NS106899 (ARF); Department of Veterans Affairs: 1I01RX002245 (ARF), I01RX002787 (ARF); Wings for Life Foundation (ATE, ARF); Craig H. Neilsen Foundation (ARF); and DOD: SC150198 (MSB); SC190233 (MSB); DOE: DE-AC02-05CH11231 (DM).

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning-Assisted High-Temperature Reservoir Thermal Energy Storage Optimization: Numerical Modeling and Machine Learning Input and Output Files

This data set includes the numerical modeling input files and output files used to synthesize data, and the reduced-order machine learning models trained from the synthesized data for reservoir thermal energy storage site identification. In this study, a machine-learning-assisted computational framework is presented to identify High-Temperature Reservoir Thermal Energy Storage (HT-RTES) site with optimal performance metrics by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems. All the simulations and the neural network model were done by Idaho National Laboratory. A detailed description of the work was reported in publication linked below.

15 GEOTHERMAL ENERGY↗

Nonintrusive projection-based reduced order modeling using stable learned differential operators

Nonintrusive projection-based reduced order models (ROMs) are essential for dynamics prediction in multi-query applications where underlying governing equations are known but the access to the source of the underlying full order model (FOM) is unavailable; that is, FOM is a glass-box. This article proposes a learn-then-project approach for nonintrusive model reduction. In the first step of this approach, high-dimensional stable sparse learned differential operators (S-LDOs) are determined using the generated data. In the second step, the ordinary differential equations, comprising these S-LDOs, are used with suitable dimensionality reduction and low-dimensional subspace projection methods to provide equations for the evolution of reduced states. This approach allows easy integration into the existing intrusive ROM framework to enable nonintrusive model reduction while allowing the use of Petrov–Galerkin projections. The applicability of the proposed approach is demonstrated for Galerkin and LSPG projection-based ROMs through four numerical experiments: 1-D scalar advection, 1-D Burgers, 2-D scalar advection and 1-D scalar advection–diffusion–reaction equations. In conclusion, the results indicate that the proposed nonintrusive ROM strategy provides accurate and stable dynamics prediction.

42 ENGINEERING↗

A Better Method to Calculate Fuel Burnup in Pebble Bed Reactors Using Machine Learning

Burnup measurement is an important step in material control and accountancy (MC&A) at nuclear reactors, and may be done by examining gamma spectra of fuel samples. Traditional approaches rely on known correlations to specific photopeaks (e.g. 137 Cs) and operate via a standard linear regression method. However, the quality of these regression methods is limited even in the best case, and is significantly poorer at short fuel cool-down times, due to the elevated radiation background by short life-time isotopes, and self-shielding effect of the fuel. For practical operation of pebble bed reactors (PBRs), quick measurements (in minutes) and short cooling times (in hours) are required from a safety and security perspective. We investigated the efficacy and performance of machine learning (ML) methods to predict the burnup of the pebble fuel from full gamma spectra (rather than specific discrete photopeaks) and found a full-spectrum ML approach to far outperform baseline regression predictions in all measurement and cooling conditions - including in operational-like measurement conditions. We also performed model and data ablation experiments to determine the relative performance impact of our ML methods' capacity to model data nonlinearities and the inherent additional information in full spectra. Applying our ML methods, we found a number of surprising results, including improved accuracy at shorter fuel cooling times (the opposite of the norm), remarkable robustness to spectrum compression (via rebinning), and competitive burnup predictions even when using background signal only (i.e. explicitly omitting known isotope photopeaks).

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

A Data-Driven Methodology for Contextual Unit Commitment Using Regression Residuals

Day after day, system operators are faced with the challenge of taking unit commitment (UC) decisions under uncertain net load conditions. The standard operating procedure for taking UC decisions begins by leveraging auxiliary data on covariates (such as the day of the week or latest weather information) to generate a point prediction for net load, which is used in solving a deterministic UC problem. Such an approach, however, is known to deliver a notoriously poor out-of-sample (OOS) performance, as it completely disregards the stochastic nature of net load. While stochastic programming models explicitly represent uncertainty, they mostly do so using a generic set of scenarios that neglect covariate observations, squandering useful auxiliary data that could be harnessed to glean insights into uncertainty. In this article, we discuss a contextual stochastic optimization approach to UC, which effectively exploits covariate observations while explicitly assessing uncertainty so as to boost the OOS performance of UC decisions. The key thrust of our approach is to leverage regression models, along with their empirical residuals, to set up and solve sample average approximation problems. Not only do we prove that our approach satisfies the requisite conditions for asymptotic optimality and consistency laid out in (Kannan et al., 2022), but we also assess its performance on several case studies conducted using real-world data collected in California ISO and New York ISO grids. In conclusion, results show that the proposed approach can significantly improve OOS performance compared to alternative methods proposed in the literature under varying dataset sizes.

Yurdakul, Ogun↗

Machine-learning-assisted high-temperature reservoir thermal energy storage optimization

High-temperature reservoir thermal energy storage (HT-RTES) has the potential to become an indispensable component in achieving the goal of the net-zero carbon economy, given its capability to balance the intermittent nature of renewable energy generation. In this study, a machine-learning-assisted computational framework is presented to co-optimize the performance metrics of HT-RTES by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Further, neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems.

15 GEOTHERMAL ENERGY↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Evaluating Cell Temperature Models and the Effect of Wind Speed in PV System Capacity Testing: Preprint

Capacity testing is a routine procedure for assessing a photovoltaic system's performance relative to expectations. The most common test method involves fitting a regression model that predicts system output power using operating weather conditions including wind speed. Structural modifications to the regression model to incorporate wind in different ways improved the model's ability to fit measured system performance, but the observed improvements were small and unlikely to change the result of a capacity test. However, the results showed that the choice of reporting wind speed and inclusion or exclusion of wind speed in the performance model used as the test benchmark can significantly change the test result.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Optimizing bioenergy biofuel harvest: a comparative analysis of stepwise and integrated methods for economic and environmental sustainability

Switchgrass is a promising bioenergy feedstock due to its high biomass yield potential, adaptability to marginal lands, and low carbon intensity for feedstock production. However, accurate cost estimation and assessment of greenhouse gas (GHG) emissions for the energy-intensive harvesting process are essential for evaluating the sustainability of bioenergy. This study provides a comparative analysis of two harvesting methods: the Stepwise Method, which separates operations into multiple stages, and the Integrated Method, which combines mowing and raking into a single pass. The analysis was conducted under four scenarios based on field sizes and biomass yields. Using three years of field-scale switchgrass harvest data from 125 sites, GHG emissions, energy consumption, and harvesting costs were quantified using the GREET model and techno-economic analysis. Additionally, regression analysis identified key climate and operational factors affecting fuel consumption. The Stepwise method was the most cost-effective for large fields with high biomass yield, achieving the lowest harvesting costs ($37.70 per ton). In contrast, the Integrated Method performed better in small fields and low-yield conditions, reducing GHG emissions by 9 % and energy use by 5 %. Regression analysis confirmed that a larger field size reduced fuel consumption, while higher biomass yield and longer operational time increased fuel use. Maximum temperature also contributed to a slight increase in fuel consumption. Furthermore, these results provide actionable insights for optimizing harvesting strategies based on field-specific conditions and operational goals, contributing to the economic and environmental sustainability of bioenergy production.

60 APPLIED LIFE SCIENCES↗

PSA 2025 Presentation: "Modeling and Sensitivity Analysis of a Generation IV Pebble Bed Reactor Using MELCOR 2.2"

Accompanying the advancement of reactor technologies is the need for computational modeling and simulation to predict their behavior under normal operating conditions and accident scenarios. New Generation IV reactor designs which employ non-conventional fuel have a particular need for modeling the behavior and release of radionuclides and other material from the fuel. In this work, MELCOR version 2.2, a system-level safety and accident scenario code developed by Sandia National Laboratories, was used to model a 200-MWth pebble bed modular reactor and calculate the inventories of circulating and deposited graphite, metal dust, and elemental components released from the fuel elements. A base case modeling the reactor under standard operating conditions was calculated using MELCOR and the inventories were extrapolated to 30 years of operation time using a logarithmic regression fit. A sensitivity analysis was also performed in which several key parameters for the base case model were modified to explore the effect of these changes on the inventories calculated by MELCOR. A set of transient scenario simulations for a depressurized loss of forced cooling (DLOFC) accident were also performed. The results of the sensitivity analysis and transient simulations are reported and discussed in relation to the modeling techniques used for this study.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Factors affecting powerhouse passage of spring migrant smolts at federally operated hydroelectric dams of the Snake and Columbia rivers

From 2008 to 2018, acoustic telemetry studies were conducted to evaluate dam passage survival of spring migrant Chinook salmon and steelhead smolts at seven of the eight federally operated dams on the lower Snake and Columbia rivers. Data from over 87 000 dam passage events were evaluated using regression modeling to identify the effect of spill operations, environmental conditions, and fish characteristics on powerhouse passage probability. In general, powerhouse passage was positively correlated with discharge, negatively correlated with forebay temperature and fish size, and higher for fish that passed the dam at night and for those that approached from the powerhouse side of the river, suggesting powerhouse passage is largely a function of smolt activity level and swimming ability. As such, spilling large volumes of water to reduce powerhouse passage is likely to be most effective during times of reduced activity and swimming ability (e.g., at night, high flows, and cold temperatures). This information can be used to develop dam- and time-specific spill operations that optimize smolt passage, power generation, and other competing demands, such as adult passage.

60 APPLIED LIFE SCIENCES↗

Data-Driven Mean-Corrected Recursive Estimation-Based Optimal DER Dispatch for Distribution System Voltage Control

Recent advances in smart inverters offer opportunities to mitigate adverse grid impacts caused by high penetrations of distributed photovoltaics (PV) in distribution grids, such as voltage violations. Here, this paper proposes a novel measurement-driven optimal power flow (OPF)-based distributed energy resource management system (DERMS) voltage regulation via recursive sensitivity estimation informed coordinated control of distributed PV inverters. The proposed approach leverages available grid and controllable DER measurements, eliminating reliance on system model information while being adaptive and robust to volatile operating conditions. A mean-corrected recursive ridge regression (MCRRR) algorithm is proposed for sensitivity estimation, continuously refining the sensitivity model through a closed-form solution. It effectively manages varying grid operating conditions, such as changes in power injections and topology reconfiguration, to facilitate a time-varying update of the Load Sensitivity Factors (LSF). The proposed approach is formulated as a linear programming (LP) problem and is thus scalable to larger-scale distribution systems. Its effectiveness and efficiency are demonstrated on a realistic distribution feeder with high PV penetrations in Southern California, USA.

14 SOLAR ENERGY↗

Multisublattice cluster expansion study of short-range ordering in iron-substituted strontium titanate

Owing to the challenges in obtaining realistic atomic configurations in large chemical phase spaces, it is not straightforward to describe structure–property relations in materials exhibiting configurational disorder. One example is iron-substituted strontium titanate (SrTi 1–x Fe x O 3–d , STF), a promising perovskite-derivative cathode material in solid oxide fuel cells that exhibits full solid solubility 0 ≤ x ≤ 1 and a tendency to exhibit short-range order. Here we demonstrate a multisublattice cluster expansion (CE) framework and apply it to STF across the full composition range. The CE approach is distinct from more traditional CE formulations in that clusters are defined explicitly by the chemical species distributed among multiple sublattices, rather than via cluster functions of occupation variables with decoration. The modified CE approach makes it easy to distinguish meaningful chemical interactions that are harder to extract from conventional CE, since for the latter chemical identity in a cluster is expressed as a product of site occupations. The least absolute shrinkage and selection operator (LASSO) is implemented as a regression analysis tool to select key clusters and avoid overfitting. We demonstrate this formulation on STF, and show that it can accurately predict configurational energies in comparison to conventional CE. From the key clusters, we identify that short-range ordering between substitutional Fe and oxygen vacancies (V O ) results in the formation of Fe–VO strings. In addition, we consider the stability of STF through CE-based Monte Carlo (MC) simulations and confirm the presence of superstructures that were previously observed in transmission electron microscopy. In this work, analysis of atomic configurations from MC samples reveals variations in the oxidation state of Fe atoms, which can be explained by the ordering tendency of Fe and V O . The cluster description and selection formalism described here may be applied to other disordered multisublattice systems for accurate and efficient material modeling.

36 MATERIALS SCIENCE↗