Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Transfer Learning-Based Independent Component Analysis

Understanding the underlying component structure is crucial for multivariate signal analysis. Among all the techniques that try to learn the latent structure, independent component analysis (ICA) is one of the most important and popular methods, which aims to extract independent components from multivariate signals and enables further analysis. For example, in electroencephalogram (EEG) analysis, artifacts filtering and disease detection are conducted based on the independent components of the signals. One critical challenge in existing ICA approaches is that the component extraction accuracy may degrade when the available data of a unit are limited. To address this issue, this paper proposes a transfer learning-based ICA method by innovatively transferring component distribution from a source domain, so that accurate component extraction results can be achieved even when only limited data are available in the target domain. To the best of our knowledge, this is the first work that leverages transfer learning to improve ICA accuracy with limited available data. In particular, we first extract all the independent components from the source domain by maximizing the log-likelihood function with a Newton-like method on a smooth manifold. Then for the target domain, the component with the largest negentropy is extracted in each round. To effectively leverage the knowledge from the source domain and to prevent the negative transfer, we try to find a component in the source domain that matches the component we are extracting. The probability density function of the matched component will then be used to improve the component extraction accuracy if such matched component can be found; otherwise, no knowledge will be transferred. Finally, numerical simulations and a case study with electrocardiogram (ECG) data are conducted, showing the effectiveness of the proposed method in transferring knowledge and reducing negative transfer.

42 ENGINEERING↗

Quantitative Power System Resilience Metrics and Evaluation Approach

Power system resilience is an emerging topic and plays an essential role in helping the power industry understand and respond to the increasing threats of extreme weather events. The first step of power system resilience analysis is to introduce metrics to quantify the resilience reasonably. Existing resilience metrics are typically restrained by the limited data for extreme event modeling and fall short in terms of physical interpretation and comparability. This paper develops novel quantitative metrics to evaluate power system resilience in pre- and post-event contexts. The developed metrics illustrate clear physical meanings and can be effectively used to compare resilience across different systems under different extreme events. Moreover, the developed metrics can be applied to both transmission and distribution systems. Simulation on a distribution system is employed to validate the effectiveness of the proposed resilience metrics and resilience evaluation approach.

power system resilience↗

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Enriching the physics program of the CMS experiment via data scouting and data parking

Specialized data-taking and data-processing techniques were introduced by the CMS experiment in Run 1 of the CERN LHC to enhance the sensitivity of searches for new physics and the precision of standard model measurements. These techniques, termed data scouting and data parking, extend the data-taking capabilities of CMS beyond the original design specifications. The novel data-scouting strategy trades complete event information for higher event rates, while keeping the data bandwidth within limits. Data parking involves storing a large amount of raw detector data collected by algorithms with low trigger thresholds to be processed when sufficient computational power is available to handle such data. The research program of the CMS Collaboration is greatly expanded with these techniques. The implementation, performance, and physics results obtained with data scouting and data parking in CMS over the last decade are discussed in this Report, along with new developments aimed at further improving low-mass physics sensitivity over the next years of data taking.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Perturbation of soil organic carbon induced by land-use change from primary forest

Abstract The impact of land-use change (LUC) on soil organic carbon (SOC) has been a wide concern of land management policymakers because CO 2 emissions induced by LUC have been the second largest carbon source worldwide. However, due to insufficient data quality and limited biome coverage, a global big picture of the impact of LUC on SOC is still not clear. This study conducted a meta-analysis on 288 independent observations sourced from 62 peer-reviewed papers to provide a global summary of the change in SOC after the conversion of primary forests into other land-use types. The conversion of primary forest to cropland resulted in the most severe SOC loss (−33.2%), followed by conversion into plantation forests (−22.3%) and secondary forests (−19.1%). Nonetheless, SOC increased by 9.1% after a conversion from primary forests into pasture. More SOC loss was found at sites with lower precipitation for primary forests converted to cropland and plantation forests. The SOC loss decreased consistently with increasing mean annual temperature (MAT) for all four types of LUC. Moreover, the loss of SOC tended to worsen over time when primary forests are converted to cropland or plantation forests. In contrast, SOC loss recovered over time following conversion to secondary forests. The gain of SOC gradually increased over time after conversion to pastures. To conclude, the changes in SOC are related not only to the land-use type but also to precipitation, temperature and turn years after LUC. Due to limited data, this study focuses on soil profiles within 30 cm depth, and future research should explore SOC dynamics induced by LUC at greater depths. Overall, cases of SOC loss of approximately 30% following deforestation were very common (except for conversion to pasture), and the results of this study show that the loss of SOC following LUC should be carefully considered and monitored in land management.

Zhang, Zhiyuan (ORCID:0000000223407001)↗

Summary of Pilot Project State Technical Assistance on Multi-Sector Analysis for Electric and Petroleum Fuels

The Oregon Energy Security Plan, (ODOE 2024) published in September 2024, builds a strong case for the state to give acute attention to the fuel supply chain. In December 2024, Pacific Northwest National Laboratory (PNNL) in partnership with Oregon Department of Energy (ODOE), announced a pilot project to conduct an analysis that synthesizes current and projected transportation fuel dynamics, supply chain risks, and risk comparators with relevant sectors, such as transportation electrification, sponsored by the Department of Energy’s (DOE) Office of Cybersecurity, Energy Security, and Emergency Response (CESER). The study, is intended to leverage existing modeling and frameworks from a recent 2024 sector coupling analysis supported by the DOEs Office of Electricity (OE) (B. Mitra, S. Pal, et al., Coupling of the Electricity and Transportation Sectors - Part I: Sector Overviews 2024) (B. Mitra, S. Pal and J. Reeve, et al. 2024). While the PNNL team set out to conduct a quantitative risk analysis driven by detailed data that synthesizes current and projected transportation fuel dynamics, supply chain risks, and risk comparators with relevant sectors. The intention was to provide an approach that could be extendable to other parts of the country. They encountered data limitations and adjusted their approach accordingly. This report summarizes PNNL's original plan for executing the study, including limitations for obtaining data requirements for fuel flows and interim products, as well as a risk matrix that can be used to identify supply chain risks.

02 PETROLEUM↗

Evaluating the limitations of Bayesian metabolic control analysis

Bayesian Metabolic Control Analysis (BMCA) is a promising framework for inferring metabolic control coefficients in data-limited scenarios, combining Bayesian inference with linear-logarithmic (lin-log) rate laws. These metabolic control coefficients quantify how changes in enzyme activities affect steady-state fluxes and metabolite concentrations across a metabolic network. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCC), and concentration control coefficients (CCC) under varying data availability conditions using three synthetic metabolic network models. We demonstrate that BMCA predictions are highly dependent on the inclusion of flux and enzyme concentration data, with the omission of these datasets leading to severe inaccuracies. In our synthetic, enzyme-perturbation datasets, external metabolite concentrations had minimal impact and, in some cases, their exclusion improved predictions; when external-nutrient perturbations were introduced and those concentrations were observed, gains were at most modest. Additionally, we find that posterior estimation with both ADVI and HMC can underestimate large-magnitude elasticities in our synthetic settings, with ADVI showing somewhat higher variance under strong up-regulation; thus, recovering |elasticity| ≳ 1.5 remains challenging regardless of the inference engine. ADVI also fails to accurately infer allosteric interactions, even when regulatory effects are strong. While BMCA maintains reasonable accuracy in partially recovering the rankings of the highest FCC values, its estimates of absolute values remain constrained by prior assumptions and data limitations. Our findings reveal the BMCA algorithm’s strengths and weaknesses, providing guidance on its application in metabolic engineering, and highlighting the need for methodological refinements to enhance its predictive capabilities.

59 BASIC BIOLOGICAL SCIENCES↗

Reference Correlations for the Density and Viscosity of Molten Alkali and Alkaline Earth Fluoride Salts

While there is a significant body of literature pertaining to thermophysical property measurements of molten salts, there is often a wide degree of variability among independent measurements of the same compounds. As such, the scientific community benefits greatly from an unbiased, independent assessment of duplicate datasets, so that reference correlations which describe these thermophysical properties as functions of temperature can be determined and then commonly used by researchers, scientists, and engineers. With regard to molten fluoride compounds, a significant time has elapsed since density and viscosity reference correlations have been determined; Janz conducted the most recent effort, in 1988, to provide reference correlations for the densities and viscosities of molten fluoride compounds via the National Standard Reference Data System coordinated by the National Bureau of Standards. Since then, new data have been published for molten fluoride compounds, and a new precedent has surfaced for putting forth reference correlations that involve fitting to multiple primary datasets. In this work, reference correlations are put forth for molten alkali and alkaline earth fluoride compounds in an effort to provide updated, improved correlations for general use. For molten alkali fluoride densities, estimated uncertainties with a 95% confidence interval are summarized as follows: LiF (0.63%), NaF (0.48%), KF (0.76%), RbF (0.93%), and CsF (0.75%). For molten alkaline earth fluoride densities, an estimated uncertainty was not able to be quantified for BeF 2 because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkaline earth fluorides: MgF 2 (1.5%), CaF 2 (0.92%), SrF 2 (1.6%), and BaF 2 (0.23%). For molten alkali fluoride viscosities, uncertainty was not able to be quantified for RbF and CsF because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkali fluorides: LiF (4.4%), NaF (3.0%), and KF (4.0%). For molten alkaline earth fluoride viscosities, limited consistent data resulted in the recommendation of single datasets (from literature) that are deemed to be the most trustworthy based on the quality of the underlying experimental studies.

Birri, A. [Oak Ridge National Laboratory (ORNL), O↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

Bayesian And Human Reliability Analysis (hra)-aided Method For The Reliability Analysis Of Software (bahamas)

The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.

Wang, Congjian (0000000207789927)↗

Pragmatic Uncertainty Quantification and Propagation in Inverse Estimation of Structural Dynamics Parameters given Material Property Uncertainties and Limited Sensor Data

In this report we demonstrate some relatively simple and inexpensive methods to effectively account for various sources of epistemic lack-of-knowledge type uncertainty in inverse problems. The demonstration problem involves inverse estimation of six parameters of a bolted joint that attaches a kettlebell shaped object to a thick plate. The parameters are efficiently inverted in a modal-based model calibration using gradient-based optimization. Two material properties of the kettlebell are treated as uncertain to within given epistemic uncertainty bounds. We apply and test interval and sparse-sample probabilistic approaches to account for uncertainty in the estimated parameters (and various scalar functionals of the parameters as generic quantities of interest, QOIs) due to uncertainties in the material properties. We also investigate the error effects of limited numbers of vibration sensors (accelerometers) on the kettlebell and plate, and therefore abbreviated excitation/response information in the parameter inversions. We propose and demonstrate a Leave-K-Sensors-Out “cross-prediction” UQ approach to estimate related uncertainties on the parameters and QOI functionals. We indicate how uncertainties from material properties and limited sensors are treated in a combined manner. The economical combined UQ approach involves just three to five samples (i.e. three to five inverse simulations), with no added complication or error/uncertainty from use of surrogate models for affordability. Finally, we describe a related economical UQ approach for handling potential parameter solution non-uniqueness and numerical optimization related precision uncertainties in the estimated parameter values. Indicated further research is identified.

36 MATERIALS SCIENCE↗

Phase Diagrams of Alloys and Their Hydrides via On-Lattice Graph Neural Networks and Limited Training Data

Efficient prediction of sampling-intensive thermodynamic properties is needed to evaluate material performance and permit high-throughput materials modeling for a diverse array of technology applications. To alleviate the prohibitive computational expense of high-throughput configurational sampling with density functional theory (DFT), surrogate modeling strategies like cluster expansion are many orders of magnitude more efficient but can be difficult to construct in systems with high compositional complexity. We therefore employ minimal-complexity graph neural network models that accurately predict and can even extrapolate to out-of-train distribution formation energies of DFT-relaxed structures from an ideal (unrelaxed) crystallographic representation. This enables the large-scale sampling necessary for various thermodynamic property predictions that may otherwise be intractable and can be achieved with small training data sets. Two exemplars, optimizing the thermodynamic stability of low-density high-entropy alloys and modulating the plateau pressure of hydrogen in metal alloys, demonstrate the power of this approach, which can be extended to a variety of materials discovery and modeling problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A data-driven multiscale model for reactive wetting simulations

Here, we describe a data-driven, multiscale technique to model reactive wetting of a silver–aluminum alloy on a Kovar™ (Fe-Ni-Co alloy) surface. We employ molecular dynamics simulations to elucidate the dependence of surface tension and wetting angle on the drop’s composition and temperature. A design of computational experiments is used to efficiently generate training data of surface tension and wetting angle from a limited number of molecular dynamics simulations. The simulation results are used to parameterize models of the material’s wetting properties and compute the uncertainty in the models due to limited data. The data-driven models are incorporated into an engineering-scale (continuum) model of a silver–aluminum sessile drop on a Kovar™ substrate. Model predictions of the wetting angle are compared with experiments of pure silver spreading on Kovar™ to quantify the model-form errors introduced by the limited training data versus the simplifications inherent in the molecular dynamics simulations. The paper presents innovations in the determination of “convergence” of noisy MD simulations before they are used to extract the wetting angle and surface tension, and the construction of their models which approximate physio-chemical processes that are left unresolved by the engineering-scale model. Together, these constitute a multiscale approach that integrates molecular-scale information into continuum scale models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluation of hydrograph separation techniques with uncertain end-member composition

Hydrograph separation is one of many approaches used to analyse shifts in source water contributions to stream flow resulting from climate change in remote watersheds. Understanding these shifts is vital, as shifts in source water contributions to a stream can shape water management decisions. Because remote watersheds are often inaccessible and have poorly characterized contributing water sources, or end-members, it is critical to understand the implications of using different hydrograph separation techniques in these data-limited environments. To explore the uncertainty associated with different techniques, results from two hydrograph separation techniques, mass balance and principle component analysis, were compared using 3 years of aqueous geochemical data from the East River watershed located in the Elk Mountains of Central Colorado. Solute concentrations of the end-members were characterized by both a limited set of direct chemical measurements of different sources and detailed seasonal instream chemistry to examine the influences of uncertain end-member compositions in a data-limited environment. Annual volumetric end-member contributions to stream flow had relatively good agreement across separation techniques. Large variations in time were observed in the hydrograph separations, depending on the end-member type, and estimated flow contributions varied between the selected solutes. End-member concentrations characterized by stream chemistry showed several limitations including a reduced number of distinguishable end-members and differences in timing of flow contributions. Here the results highlight the benefits of using multiple hydrograph separation techniques by providing a ‘weight-of-evidence’ approach to environments with limited end-member concentration data.

54 ENVIRONMENTAL SCIENCES↗