Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Comprehensive framework for data-driven model form discovery of the closure laws in thermal-hydraulics codes

The two-phase two-fluid model is a basis of many thermal-hydraulics codes used in design, licensing, and safety considerations of nuclear power plants. Thermal-hydraulics codes rely on the closure laws to close the system of conservation equations and describe the interactions between phases. These laws, derived from years of experimental investigations, are semi-empirical correlations that lack generality and have a limited range of applicability. Increase of computational power, availability of new experiments, and development of high-fidelity simulations has increased the number of validation data. The discrepancies between the code predictions and the validation data are a great source of knowledge. Missing physics that are not included in the model but are important for the considered phenomena can be discovered by propagating the information from the experimental results through the model. Furthermore, physics-discovered data-driven model form (P3DM) methodology integrates available integral effect tests and separate effects tests to determine the necessary corrections to the model form of the closure laws. In contrast to existing calibration techniques, the methodology modifies the functional form of the closure laws. Based on the functional form of the correction, the missing physics that were not included in the original model can be discovered. The methodology provides the alternative to the machine learning approach, in which the model is discovered in the form of the intractable black-box relation. In this work, the methodology was applied to the CTF subchannel code to improve the prediction of the two-phase flow phenomena.

42 ENGINEERING↗

MLMOD: Machine Learning Methods for Data-Driven Modeling in LAMMPS

MLMOD is a software package for incorporating machine learning approaches and models into simulations of microscale mechanics and molecular dynamics in LAMMPS. Recent machine learning approaches provide promising data-driven approaches for learning representations for system behaviors from experimental data and high fidelity simulations. The package facilitates learning and using data-driven models for (i) dynamics of the system at larger spatial-temporal scales (ii) interactions between system components, (iii) features yielding coarser degrees of freedom, and (iv) features for new quantities of interest characterizing system behaviors. MLMOD provides hooks in LAMMPS for (i) modeling dynamics and time-step integration, (ii) modeling interactions, and (iii) computing quantities of interest characterizing system states. The package allows for use of machine learning methods with general model classes including Neural Networks, Gaussian Process Regression, Kernel Models, and other approaches. Here we discuss our prototype C++/Python package, aims, and example usage. For related papers, examples, updates, and additional information see https://github.com/atzberg/mlmod and http://atzberger.org/.

97 MATHEMATICS AND COMPUTING↗

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING↗

Influence of initial conditions on data-driven model identification and information entropy for ideal mhd problems

Data-driven methods of model identification are able to discern governing dynamics of a system from data. Such methods are well suited to help us learn about systems with unpredictable evolution or systems with ambiguous governing dynamics given our current understanding. Many plasma problems of interest fall into these categories as there are a wide range of models that exist, however each model is only useful in a certain regime and often limited by computational complexity. To ensure data-driven methods align with theory, they must be consistent and predictable when acting on data whose governing dynamics are known. Weak Sparse Identification of Nonlinear Dynamics (WSINDy) is a recently developed data-driven method that has shown promise in learning governing dynamics from data with high noise levels [1]. This work examines how WSINDy acts on ideal MHD test problems as the initial conditions are varied and specifies limiting requirements for successful equation identification. Furthermore, it is hard to recover the governing dynamics from data that emphasize a single dominant behavior. In these low information cases, Shannon information entropy is able to pick up on the redundancies in the data that affect recoverability.

97 MATHEMATICS AND COMPUTING↗

Data-driven modeling to enhance municipal water demand estimates in response to dynamic climate conditions

Altered precipitation and temperature patterns from a changing climate will affect supply, demand, and overall municipal water system operations throughout the arid western U.S. While supply forecasts leverage hydrological models to connect climate influences with surface water availability, demand forecasts typically estimate water use independent of climate and other externalities. Stemming from an increased focus on seasonal water demand management, we use the Salt Lake City, Utah municipal water system as a test bed to assess model accuracy versus complexity trade-offs between simple climate-independent econometric-based models and complex climate-sensitive data-driven models to average to extreme wet and dry climate conditions—representative of a new climate normal. Here, the climate-independent model displayed low performance during extreme dry conditions with predictions exceeding 90% and 40% of the observed monthly and seasonal volumetric demands, respectively, which we attribute to insufficient model complexity. The climate-sensitive models displayed greater accuracy in all conditions, with an ordinary least squares model demonstrating a measurable reduction in prediction bias (3.4% vs. -27.3%) and RMSE (74.0 lpcd vs. 294 lpcd) compared to the climate-independent model. The climate-sensitive workflow increased model accuracy and characterized climate-demand interactions, demonstrating a novel tool to enhance water system management.

54 ENVIRONMENTAL SCIENCES↗

A data-driven model for thermodynamic properties of a steam generator under cycling operation

The varying electricity demand from coal power plants due to the intermittent nature of renewable sources leads to load-follow and on/off operations referred to as cycling. Cycling causes transients of properties such as pressure and temperature within various components of the steam generation system.These transients cause increased damage because of fatigue and creep-fatigue interactions shortening the life of components. An algorithm is developed to identify cycling operations from the gross power data. The data-driven model based on artificial neural networks (ANN) is developed using 10 years data from Coal Creek Station power plant located in North Dakota, USA to estimate properties of the steam generator components during cycling operations. Furthermore, the uniqueness of this model is the ability to predict component properties for the cycling as well as base-load operations and is reported for the first time. The ANN model estimates the component properties, for a given gross power profile and initial conditions, as they vary during cycling operations. As a representative example, the ANN estimates are presented for the superheater outlet pressure, reheater inlet temperature, and flue gas temperature at the air heater inlet. The changes in these variables as a function of the gross power over the time duration are compared with measurements to assess the predictive capability of the model. Mean square errors of 4.49E-04 for superheater outlet pressure, 1.62E-03 for reheater inlet temperature, and 4.14E-04 for flue gas temperature at the air heater inlet were observed.

01 COAL, LIGNITE, AND PEAT↗

Association Between Injection and Microseismicity in Geothermal Fields With Multiple Wells: Data-Driven Modeling of Rotokawa, New Zealand, and Húsmúli, Iceland

Understanding injection-induced microseismicity in geothermal systems can provide insight into reservoir connectedness. However, fault and reservoir complexity are difficult to represent in simple analytical models, which makes it difficult to discern clear relationships from incidental associations. Here, we have used data-driven models to study how fluid injection and microseismicity are related in the Rotokawa (New Zealand) and Hellisheiði (Iceland) geothermal fields. We tested two classes of model: (a) lagged linear regression of seismicity rate as a function of well injection rates; and (b) systematic extraction of injection time series features that are then evaluated for associations with the seismicity. These models allowed us to determine which wells had the greatest correlation with microseismicity and to explain this association in a reservoir context. Finally, exploring different data types and transformations, we were unable to establish a link between rapid changes in injection rate and seismicity spikes, as suggested by some theoretical models.

15 GEOTHERMAL ENERGY↗

Data-Driven Modeling and Control of Systems with Plasma-Surface Interactions (Final Technical Report)

This final technical report summarizes the activities and accomplishments in the period from February 2023 thru January 2026. The objective of the proposed research is to investigate the physical mechanisms and processes underlying the formation of structures and patterns in systems with plasma-surface interactions. In the past decades, there have been extensive studies on the interaction of glow discharges, dielectric barrier discharges, and arc discharges with confining or intervening surfaces. The advancement of the understanding of these phenomena is not only of fundamental scientific interest and relevance to the knowledge of the plasma state, but also with profound implications in various technological applications. The research will integrate theoretical, computational, and experimental work within an innovative framework of data assimilation, i.e., optimally combining model predictions with measurements. The scientific merit of this research has three aspects. Firstly, it extends the studies of plasma-surface interactions to systems with insulator surfaces and multi-layer systems, while existing studies are predominantly on electrode surfaces. Secondly, it expects to develop a novel data-driven modeling approach based on data assimilation to enhance the predictive and control capabilities, which could make transformative contributions to basic plasma research. Thirdly, it will shed new light on outstanding problems related to formation of patterns interfacing plasmas. This project also aims to launch an education and outreach initiative at Texas A&M University-Kingsville, a non-R1, minority-serving institution in South Texas. The initiative is structured as a four-tier pyramid. Tier one will be a webinar series for culture and capacity building to inform broader audience in the region about the research fields of plasma science and engineering. Tier two will be the creation and offering of an upper-level undergraduate course on introductory plasma physics, which will help with the recruitment for the upper tiers. On tier three, we will engage and mentor senior design students to conduct work toward the research goal of this project. There will also be a certificate program on general plasma science for undergrad and graduate students, part of which will be lab training at Princeton University. Tier four will be the supervision and mentoring of Ph.D. students. Therefore, this project will systematically expand the talent pipeline, broaden participation from communities historically and geographically underrepresented in DOE SC research portfolio, significantly improve the research and education capacity at the PI’s institution, and contribute to developing a diverse workforce in plasma science and engineering.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

Collaborative Research: Enabling multi-scale studies of magnetic reconnection with interpretable data-driven models

The development of accurate reduced descriptions and improved closures for magnetic reconnection is an important and a long‐standing challenge in plasma physics. The four‐fluid approach, and associated closures, that were investigated have the potential to improve the accuracy of plasma fluid models, capturing physical effects which would otherwise require a kinetic description. If successful, this approach could have an important impact for the modeling of laboratory and space plasmas. The major goals of this project were to develop new machine learning (ML) tools based on sparse and symbolic regression techniques, and to extract interpretable and generalizable reduced models (e.g., in the form of partial differential equations - PDEs) from data generated by first principles plasma simulations. Preserving interpretability of such data‐driven models is key to addressing the long‐standing theoretical and numerical challenges. Prior proof‐of‐principle studies have demonstrated the enormous potential of this approach, by recovering the well‐established hierarchy of plasma equations (from Vlasov to MHD) from data produced by particle‐in‐cell (PIC) simulations. Our goal in this project was to extend and apply these new tools to construct better kinetic closures for magnetic reconnection; to derive better models of particle injection and acceleration by this fundamental plasma process; and to use this understanding to accelerate the development of multi‐scale plasma algorithms. While our immediate focus was on the problem of magnetic reconnection, the tools that were will developed are general and applicable to other areas of plasma physics, and more broadly to many‐body phenomena. We anticipate that the development of these multi‐scale models will have a significant impact across different areas of plasma science, from fusion to space and astrophysical plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance analysis and comparison of data-driven models for predicting indoor temperature in multi-zone commercial buildings

Building thermal models, which characterize the properties of a building’s envelope and thermal mass, are essential for accurate indoor temperature and cooling/heating demand prediction. Because of their flexibility and ease of use, data-driven models are increasingly used. Here, this study compared and analyzed the performance of gray-box (resistance-capacitance) and black-box (recurrent neural network) models for predicting indoor air temperature in a real multi-zone commercial building. The developed resistance-capacitance model served as a benchmark model for which full sets of temporal data and building information were used as inputs. The recurrent neural network models were trained and tested assuming various available types and amounts of temporal data and known building physical information to investigate the effects of data and information availability. Feature importance analysis was conducted to select the key variables for different prediction targets under different scenarios. This research provides guidance in selecting an appropriate building thermal response modeling method based on the measured data availability, building physical information, and application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Rapid data-driven model reduction of nonlinear dynamical systems including chemical reaction networks using ℓ 1 -regularization

We develop a new data-driven paradigm for efficient model reduction of a broad class of nonlinear dynamical systems. Our model reduction method directly enables the interpretation of key components of the dynamical system, unlike traditional projection-based model reduction methods that focus on reducing computational complexity more than interpretability. Our method is not application specific and is simple to implement on nonlinear dynamical systems arising from a variety of different fields. It requires minimal parameterization using a single parameter to trade-off between model complexity and estimation error. We use a data-driven paradigm to formulate model reduction as an efficient convex optimization problem that scales polynomially in the original size of the complex system, enabling systems with as many as thousands of components to be reduced in a matter of minutes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Autonomous control for Heat-Pipe microreactor using Data-Driven model predictive control

To enable a self-regulating capability for heat pipe (HP) microreactors, an anticipatory control strategy achieved via model predictive control (MPC) could proactively respond to potential disturbances and deviations in operating setpoints. This paper demonstrates data-driven methods for predicting the distribution and transient of temperatures and heat fluxes at selected components and regions in a 37-HP system, based on which the optimal control actions in response to changes in user-defined setpoints can be found. We present the development and validation of linear state-space model, feedfoward, and recurrent neural networks. Here, we compare the performance of MPCs with different modeling approaches in terms of following setpoints for temperatures and averaged output heat fluxes. The accuracies of the three data-driven models are similar, but the control actions initiated by neural-network-based MPC can better adapt to drastic changes in setpoints yet generate the smallest errors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Large Destabilization of (TiVNb)-Based Hydrides via (Al, Mo) Addition: Insights from Experiments and Data-Driven Models

High-entropy alloys (HEAs) represent an interesting alloying strategy that can yield exceptional performance properties needed across a variety of technology applications, including hydrogen storage. Examples include ultrahigh volumetric capacity materials (BCC alloys → FCC dihydrides) with improved thermodynamics relative to conventional high-capacity metal hydrides (like MgH 2 ), but still further destabilization is needed to reduce operating temperature and increase system-level capacity. Here, in this work, we demonstrate efficient hydride destabilization strategies by synthesizing two new Al 0.05 (TiVNb) 0.95–x Mo x (x = 0.05, 0.10) compositions. We specifically evaluate the effect of molybdenum (Mo) addition on the phase structure, microstructure, hydrogen absorption, and desorption properties. Both alloys crystallize in a bcc structure with decreasing lattice parameters as the Mo content increases. The alloys can rapidly absorb hydrogen at 25 °C with capacities of 1.78 H/M (2.79 wt %) and 1.79 H/M (2.75 wt %) with increasing Mo content. Pressure-composition isotherms suggest a two-step reaction for hydrogen absorption to a final fcc dihydride phase. The experiments demonstrate that increasing Mo content results in a significant hydride destabilization, which is consistent with predictions from a gradient boosting tree data-driven model for metal hydride thermodynamics. Furthermore, improved desorption properties with increasing Mo content and reversibility were observed by in situ synchrotron X-ray diffraction, in situ neutron diffraction, and thermal desorption spectroscopy.

36 MATERIALS SCIENCE↗

Non-intrusive data-driven model reduction for differential–algebraic equations derived from lifting transformations

In this paper we present a non-intrusive data-driven approach for model reduction of nonlinear systems. The approach considers the particular case of nonlinear partial differential equations (PDEs) that form systems of partial differential–algebraic equations (PDAEs) when lifted to polynomial form. Such systems arise, for example, when the governing equations include Arrhenius reaction terms (e.g., in reacting flow models) and thermodynamic terms (e.g., the Helmholtz free energy terms in a phase-field solidification model). Using the known structured form of the lifted algebraic equations, the approach computes the reduced operators for the algebraic equations explicitly, using straightforward linear algebra operations on the basis matrices. The reduced operators for the differential equations are inferred from lifted snapshot data using operator inference, which solves a linear least squares regression problem. The approach is illustrated for the nonlinear model of solidification of a pure material. The lifting transformations reformulate the solidification PDEs as a system of PDAEs that have cubic structure. The operators of the lifted system for this solidification example have affine dependence on key process parameters, permitting us to learn a parametric reduced model with operator inference. Numerical experiments show the effectiveness of the resulting reduced models in capturing key aspects of the solidification dynamics.

42 ENGINEERING↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗

Data-Driven Model Predictive Control for Temperature Management of Heat Pipe Microreactor

To enable the self-regulating capability of heat pipe (HP) microreactors, an anticipatory control strategy through model predictive control (MPC) could proactively respond to potential disturbances and deviations in operating setpoints. However, a key factor prohibiting the widespread adoption of MPCs in nuclear applications is the effort and computational costs associated with learning and calibrating first-principles-based process models when the target system is complex and when there are gaps between modeled and target reactor systems. In this paper, we demonstrate data-driven MPC using three approaches for modeling the system dynamics, including a linear state-space model, feedforward neural network, and recurrent neural networks long short-term memory. We present the development and validation process of each model and compare the performance of data-driven MPCs in controlling the temperatures of selected HPs at the evaporator and condenser regions in a 37-HP-monolith system. Our results show that, qualitatively, all data-driven MPCs are producing similar control actions, while quantitatively, with artificial neural nets (especially feedforward neural nets), MPC can better follow drastic changes in setpoints with smallest errors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗