Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data-driven model validation for neutrino-nucleus cross section measurements

Neutrino-nucleus cross section measurements are needed to improve interaction modeling to meet the precision needs of neutrino experiments in efforts to measure oscillation parameters and search for physics beyond the Standard Model. We review the difficulties associated with modeling neutrino-nucleus interactions that lead to a dependence on event generators in oscillation analyses and cross section measurements alike. We then describe data-driven model validation techniques intended to address this model dependence. The method relies on utilizing various goodness-of-fit tests and the correlations between different observables and channels to probe the model for defects in the phase space relevant for the desired analysis. These techniques shed light on relevant mismodeling, allowing it to be detected before it begins to bias the cross section results. We compare more commonly used model validation methods which directly validate the model against alternative ones to these data-driven techniques and show their efficacy with fake data studies. These studies demonstrate that employing data-driven model validation in cross section measurements represents a reliable strategy to produce robust results that will stimulate the desired improvements to interaction modeling.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Combining physics-based and data-driven models for quantitatively accurate plasma profile prediction that extrapolates well; with application to DIII-D, AUG, and ITER tokamaks

For design, scenario planning, and control, ITER and all other envisioned tokamaks rely on a variety of statistical and physics-based models to extrapolate to unseen regimes; most notably from low plasma current to high. A 'meta-learning' methodology for combining the accuracy of data-driven models with the generalizability of physics-based models is described and tested, yielding a 5–10 percent improvement in performance beyond either alone for the task of extrapolating time-dependent plasma profile prediction from low- to high- plasma current DIII-D tokamak discharges. Meanwhile, it is shown that both machine learning models extrapolated far-distribution and state-of-the-art 'physics-based' profile predictors fare worse than merely assuming plasma profiles do not change from their initial values. Finally, a variety of other mechanisms for helping data-driven models generalize—transfer learning, adding contextual information from physics simulators, and adding data from the ASDEX Upgrade tokamak—are attempted for similar extrapolation tasks but, in the methodology used in this paper, yield no significant improvement beyond simple data-driven models. Results are summarized in figures 15 and 16.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Data-Driven Kinetic Reaction Networks for Separation Chemistry

Understanding complex, multistep chemical reactions at the molecular level is a major challenge whose solution would greatly benefit the design and optimization of numerous chemical processes. The separation of rare-earth (4f) and actinide (5f) elements is an example where improving our chemical understanding is important for designing and optimizing new chemistries, even with a limited number of observations. Here, in this work, we leverage data-driven artificial intelligence and machine-learning approaches to develop kinetic reaction networks that describe the liquid–liquid extraction mechanism of uranium using N,N-di-2-ethylhexyl-isobutyramide (DEHiBA). Specifically, we compare and contrast the properties of two classes of models: (1) purely data-driven models that are regularized using chemistry-agnostic, L1 regression and (2) chemistry-informed models that are regularized using relative reaction energies provided by quantum mechanical calculations. We observe that purely data-driven models are unbiased, simple, and accurate in their predictions of experimental measurements when provided with sufficient data but are difficult to fully constrain and interpret. In contrast, chemistry-informed models exhibit significantly improved chemical interpretability and consistency, providing a detailed description of the separation process while achieving high accuracy through ensemble averaging. Overall, the dominant species predicted to be extracted into the organic phase is UO 2 (NO 3 ) 2 (DEHiBA) 2 , agreeing with experimental slope analysis, thermodynamic modeling, EXAFS, and crystal structures. This work demonstrates that leveraging the fundamental structure of the problem can lead to efficient learning schemes that provide both accurate predictions and chemical insights at a low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data-Driven Voltage Regulation of Distribution Grid Using Nonlinear Autoregressive Model with Exogenous Inputs (NARX)

This article proposes data-driven control via a nonlinear autoregressive model with exogenous inputs (NARX) for real-time voltage regulation of a modified feeder using reactive power sources. Traditional voltage control strategies rely on rule-based heuristics or optimization techniques, which often require detailed system models and extensive computational resources. The NARX-based controller learns system dynamics from historical data and predicts optimal reactive power dispatch in real-time for voltage correction. The proposed approach is evaluated on a power system feeder model under varying load and network conditions. Simulation results demonstrate that the NARX-based controller achieves improved voltage regulation, offering higher adaptability to system fluctuations. This study highlights the potential of data-driven control for enhancing the reliability of power distribution networks.

Donge, Vrushabh [ORNL] (ORCID:0000000306062803)

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis

A Data-driven approach to Core Power distribution reconstruction in a Nuclear Reactor

This report presents the initial development of a data-driven approach for reconstructing the core power distribution in a nuclear reactor (power shape synthesis) using ex-core sensors. Traditional techniques rely on deploying a large number of detectors throughout the reactor core. However, this approach is not feasible for innovative reactor concepts like Advanced Reactors and Microreactors. First, the tight lattice pitch, designed to maximize power density, limits the space available for sensors. Secondly, the harsh operating conditions are not compatible with commercially available detectors. The method proposed in this work integrates high-fidelity modeling with data-driven techniques to accurately reconstruct power distribution across various reactor types, thereby reducing the reliance on in-core sensors. Purdue University Reactor One (PUR-1) was selected as the test case. The CAD model representing the latest configuration of the PUR-1 core was imported into the OpenMC simulation framework, and the model was built. Additionally, the previously developed MCNP6 model was updated. The two models were assessed against the data collected during an experimental campaign conducted in July 2024. Thirty gold foils were placed in three Irradiation Assemblies in PUR-1 core. Using the measured activity of the irradiated foils, the neutron flux at different core locations was estimated.

22 GENERAL STUDIES OF NUCLEAR REACTORS

MAD 3 (Material Data Driven Design) User Manual (v1.01)

MAD 3 (Material Data Driven Design) is a novel and unique software solution that provides initial plastic anisotropy of polycrystalline metals using crystallographic texture information, developed at Sandia National Laboratories. In this document, we describe the structure and functionality of the current MAD 3 software (v1.01).

36 MATERIALS SCIENCE

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Rare Lepton Decays and Differentiable Hadronization Models - From Signatures of New Physics to Data-driven Event Generation

This dissertation is partitioned into two parts: phenomenological studies focused on rare lepton decays as probes of heavy and light new physics, and the development of differentiable, data-driven hadronization models. Part I develops the phenomenology of new physics signatures stemming from rare charged lepton flavor violating decays probed by experiments at the intensity frontier. These include interactions mediated by both high-scale effective operators and light new physics, manifesting in multi-lepton final states ($\mu \to 5e$), elastic nuclear transitions ($\mu \to e$ conversion), baryon-number-violating muon capture, and time-dependent signals from ultralight dark matter ($\mu \to e \phi, \tau \to \ell \phi$). Part II develops two distinct strategies for advancing differentiable and data-driven hadronization models. One involves comprehensive reweighting frameworks for hadronization that enable efficient uncertainty estimation, facilitate parameter tuning, and interface naturally with differentiable programming paradigms. The other introduces machine-learning-based methods for extracting microscopic fragmentation dynamics directly from macroscopic observables through the deformation of existing models -- effectively providing solutions to the inverse problem of hadronization. Altogether, these studies advance the interpretability, flexibility, and precision of theoretical predictions for both high-intensity and high-energy experiments.

Menzo, Tony [Cincinnati U.] (ORCID:000000022013457

Data-Driven Digital Twin for Reliability Assessment of DC/DC Buck Converter

In commercial applications, the operation of DC/DC converters significantly impacts overall system performance and long-term reliability. This study introduces a data-driven digital twin (DT) approach for estimating critical degradation parameters of DC/DC BUCK converter under steady-state condition. Initially, a circuit-level MATLAB/Simulink digital model (DM C ) is refined against a hardware prototype’s switching model dataset using offline particle swarm optimization. The optimized digital model’s steady-state response is then verified with its average model response while varying the duty and load. Subsequently, degradation profiles are imposed on the inductor, capacitor, MOSFET in the DMC. A large dataset is generated from this model, allowing training, validation, and testing of machine learning (ML) models for component health regression tasks. The proposed method employs random forest ML models, achieving impressive regression results with a squared R value as high as 0.99978 and a root mean square error of 4.2× 10 –6 . The method is further validated on a medium power level DC/DC BUCK prototype with varying load conditions, and includes the analysis of MOSFET’s on-resistance under degradation conditions. This data-driven DT method shows promise for identifying parasitic degradation and ohmic loss parameters, enhancing converter reliability assessments in a non-invasive, generalized, and computationally efficient manner.

14 SOLAR ENERGY

Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?

Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties, including Monte Carlo dropout, deep ensemble, Bayesian neural networks, Gaussian Processes and conformal prediction. Lastly, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Efficient data-driven regression for reduced-order modeling of spatial pattern formation

We present an efficient data-driven regression approach for constructing reduced-order models (ROMs) of reaction-diffusion systems exhibiting pattern formation. The ROMs are learned non-intrusively from available training data of physically accurate numerical simulations. The method can be applied to general nonlinear systems through the use of polynomial model form, while not requiring knowledge of the underlying physical model, governing equations, or numerical solvers. The process of learning ROMs is posed as a low-cost least-squares problem in a reduced-order subspace identified via Proper Orthogonal Decomposition (POD). Numerical experiments on classical pattern-forming systems–including the Schnakenberg and Mimura–Tsujikawa models–demonstrate that higher-order surrogate models significantly improve prediction accuracy while maintaining low computational cost. The proposed method provides a flexible, non-intrusive model reduction framework, well suited for the analysis of complex spatio-temporal pattern formation phenomena.

Data-driven modeling

Crossing the Finish Line: Integration of Data-Driven Process Control for Maximization of Energy and Resource Efficiency in Advanced Water Resource Recovery Facilities

Improvements in process monitoring and control at water resource recovery facilities (WRRFs) could result in reductions in electricity consumption, chemical inputs, and greenhouse gas emissions, as well as improved energy recovery. Many current WRRF data collection, monitoring, and control approaches use 20th century process monitoring and control systems, which require large design safety factors to ensure reliability in the absence of more advanced, precise controls. Implementation of more modern data-driven control tools could lead to more efficient operations that provide intrinsic reliability with better overall process performance at full-scale. This project (1) developed and demonstrated data-driven process controls at full-scale facilities for five promising WRRF process technologies that provide whole-plant approaches and offer substantial energy and resource recovery benefits, and (2) created a Machine Learning (ML) Toolkit and an implementation guide of new process control approaches that walks users through each step of the ML workflow and illustrates the steps through case study examples.

54 ENVIRONMENTAL SCIENCES

Data-Driven Invertible Neural Surrogates of Atmospheric Transmission

We present Data-Driven Invertible Neural Surrogates of Atmospheric transmission, or DINSAT. DINSAT is a novel framework for inferring an atmospheric transmission profile from a spectral scene. This framework leverages a lightweight, physics-based simulator that is automatically tuned -- by virtue of autodifferentiation and differentiable programming -- to construct a surrogate atmospheric profile to model the observed data. The framework has utility in (i) performing atmospheric correction, (ii) recasting spectral data between various modalities (e.g. radiance and reflectance at the surface and at the sensor), and (iii) inferring atmospheric transmission profiles, such as absorbing bands and their relative magnitudes. We demonstrate the utility of these methods by performing a canonical atmospheric correction task for the purposes of further analysis - in this case, target detection within a scene.

Koch, James V.

A Practical Comparison of Data-Driven Prognostics Methods for Energy Systems

This study explores data-driven prognostics for nuclear power plant (NPP) condensers, focusing on tube fouling. We utilized the Asherah nuclear power plant simulator (ANS) to compare four methods: Random Forest (RF), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory Neural Network (LSTM). By simulating various fouling scenarios in the ANS, we generated data with different degradation rates under transient operations. The models were trained and tested on these data, with performance evaluated visually and numerically including uncertainty assessment. The LSTM model excelled, exhibiting minimal prediction noise and the most accurate remaining useful life estimates across all degradation levels. Its ability to capture long-term dependencies and produce cleaner outputs makes it a strong candidate, although accurate training data across the entire component lifespan are crucial. The RF model emerged as a robust alternative, providing reliable predictions with high confidence. The FCNN and SVR models, while less effective overall, showed potential under specific conditions. FCNN offers a less complex alternative to LSTM and might benefit from larger datasets. SVR excels in precision when the quality of the training data is high. Furthermore, this study highlights the operational benefits of advanced prognostics in the energy sector and emphasizes the need for further research in NPP condenser health management through real-life experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS