Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data-driven model construction for industrial asset decision boundary classification

In some embodiments, a system model construction platform may receive, from a system node data store, system node data associated with an industrial asset. The system model construction platform may automatically construct a data-driven, dynamic system model for the industrial asset based on the received system node data. A synthetic attack platform may then inject at least one synthetic attack into the data-driven, dynamic system model to create, for each of a plurality of monitoring nodes, a series of synthetic attack monitoring node values over time that represent simulated attacked operation of the industrial asset. The synthetic attack platform may store, in a synthetic attack space data source, the series of synthetic attack monitoring node values over time that represent simulated attacked operation of the industrial asset. This information may then be used, for example, along with normal operational data to construct a threat detection model for the industrial asset.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Modeling and Correction of Vehicle Dynamics

We develop a data-driven framework for learning and correcting nonautonomous vehicle dynamics. Physics-based vehicle models are often simplified for tractability and therefore exhibit inherent model-form uncertainty, motivating the need for data-driven correction. Moreover, nonautonomous dynamics are governed by time-dependent control inputs, which pose challenges in learning predictive models directly from temporal snapshot data. To address these, we reformulate the vehicle dynamics via a local parameterization of the time-dependent inputs, yielding a modified system composed ofa sequence of local parametric dynamical systems. Here, we approximate these parametric systems using two complementary approaches. First, we employ the dimension reduction and interpolation in parameter space (DRIPS) methodology to construct efficient linear surrogate models, equipped with lifted observable spaces and manifold-based operator interpolation. This enables data-efficient learning of vehicle models whose dynamics admit accurate linear representations in the lifted spaces. Second, for more strongly nonlinear systems, we employ flow map learning (FML), a deep neural network (DNN) approach that approximates the parametric evolution map without requiring special treatment of nonlinearities. We further extend FML with a transfer-learning-based model correction procedure, enabling the correction of misspecified prior models using only a sparse set of high-fidelity or experimental measurements, without assuming a prescribed form for the correction term. Through a suite of numerical experiments on unicycle, simplified bicycle, and slip-based bicycle models, we demonstrate that DRIPS offers robust and highly data-efficient learning of nonautonomous vehicle dynamics, while FML provides expressive nonlinear modeling and effective correction of model-form errors under severe data scarcity.

data-driven modeling↗

Comprehensive framework for data-driven model form discovery of the closure laws in thermal-hydraulics codes

The two-phase two-fluid model is a basis of many thermal-hydraulics codes used in design, licensing, and safety considerations of nuclear power plants. Thermal-hydraulics codes rely on the closure laws to close the system of conservation equations and describe the interactions between phases. These laws, derived from years of experimental investigations, are semi-empirical correlations that lack generality and have a limited range of applicability. Increase of computational power, availability of new experiments, and development of high-fidelity simulations has increased the number of validation data. The discrepancies between the code predictions and the validation data are a great source of knowledge. Missing physics that are not included in the model but are important for the considered phenomena can be discovered by propagating the information from the experimental results through the model. Furthermore, physics-discovered data-driven model form (P3DM) methodology integrates available integral effect tests and separate effects tests to determine the necessary corrections to the model form of the closure laws. In contrast to existing calibration techniques, the methodology modifies the functional form of the closure laws. Based on the functional form of the correction, the missing physics that were not included in the original model can be discovered. The methodology provides the alternative to the machine learning approach, in which the model is discovered in the form of the intractable black-box relation. In this work, the methodology was applied to the CTF subchannel code to improve the prediction of the two-phase flow phenomena.

42 ENGINEERING↗

MLMOD: Machine Learning Methods for Data-Driven Modeling in LAMMPS

MLMOD is a software package for incorporating machine learning approaches and models into simulations of microscale mechanics and molecular dynamics in LAMMPS. Recent machine learning approaches provide promising data-driven approaches for learning representations for system behaviors from experimental data and high fidelity simulations. The package facilitates learning and using data-driven models for (i) dynamics of the system at larger spatial-temporal scales (ii) interactions between system components, (iii) features yielding coarser degrees of freedom, and (iv) features for new quantities of interest characterizing system behaviors. MLMOD provides hooks in LAMMPS for (i) modeling dynamics and time-step integration, (ii) modeling interactions, and (iii) computing quantities of interest characterizing system states. The package allows for use of machine learning methods with general model classes including Neural Networks, Gaussian Process Regression, Kernel Models, and other approaches. Here we discuss our prototype C++/Python package, aims, and example usage. For related papers, examples, updates, and additional information see https://github.com/atzberg/mlmod and http://atzberger.org/.

97 MATHEMATICS AND COMPUTING↗

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING↗

Influence of initial conditions on data-driven model identification and information entropy for ideal mhd problems

Data-driven methods of model identification are able to discern governing dynamics of a system from data. Such methods are well suited to help us learn about systems with unpredictable evolution or systems with ambiguous governing dynamics given our current understanding. Many plasma problems of interest fall into these categories as there are a wide range of models that exist, however each model is only useful in a certain regime and often limited by computational complexity. To ensure data-driven methods align with theory, they must be consistent and predictable when acting on data whose governing dynamics are known. Weak Sparse Identification of Nonlinear Dynamics (WSINDy) is a recently developed data-driven method that has shown promise in learning governing dynamics from data with high noise levels [1]. This work examines how WSINDy acts on ideal MHD test problems as the initial conditions are varied and specifies limiting requirements for successful equation identification. Furthermore, it is hard to recover the governing dynamics from data that emphasize a single dominant behavior. In these low information cases, Shannon information entropy is able to pick up on the redundancies in the data that affect recoverability.

97 MATHEMATICS AND COMPUTING↗

Data-driven modeling to enhance municipal water demand estimates in response to dynamic climate conditions

Altered precipitation and temperature patterns from a changing climate will affect supply, demand, and overall municipal water system operations throughout the arid western U.S. While supply forecasts leverage hydrological models to connect climate influences with surface water availability, demand forecasts typically estimate water use independent of climate and other externalities. Stemming from an increased focus on seasonal water demand management, we use the Salt Lake City, Utah municipal water system as a test bed to assess model accuracy versus complexity trade-offs between simple climate-independent econometric-based models and complex climate-sensitive data-driven models to average to extreme wet and dry climate conditions—representative of a new climate normal. Here, the climate-independent model displayed low performance during extreme dry conditions with predictions exceeding 90% and 40% of the observed monthly and seasonal volumetric demands, respectively, which we attribute to insufficient model complexity. The climate-sensitive models displayed greater accuracy in all conditions, with an ordinary least squares model demonstrating a measurable reduction in prediction bias (3.4% vs. -27.3%) and RMSE (74.0 lpcd vs. 294 lpcd) compared to the climate-independent model. The climate-sensitive workflow increased model accuracy and characterized climate-demand interactions, demonstrating a novel tool to enhance water system management.

54 ENVIRONMENTAL SCIENCES↗

DEEP Solar: Data DrivEn Modeling and Analytics for Enhanced System Layer ImPlementation

Realizing the SETO 2030 mission of reducing solar energy costs to 3-5 c/kWh will require innovative enabling research on effective, cost-efficient integration of local PV within distribution systems. However, the intermittent and variable nature of PVs compels operators to impose conservative hosting capacity constraints. Given the extremely high variability of (intermittent and unpredictable) solar energy generation, relaxing the capacity constraints (which are currently around 15%) and achieving 100% or greater integration of renewables will require a fundamental transformation of the power grid via the utilization of exponentially larger amounts of AMI enabled fine-grained data. To address the challenges in increasing the penetration of renewable energy based DERs, this project envisions an Enhanced System Layer (ESL) at the distribution network level that is reliable, cost-effective and scalable to millions of Distributed Energy Resources (DERs)/devices. This includes developing: 1) Transformative and highly scalable machine learning based predictive analytics tools that plug into distribution system planning and provide real-time situational awareness at the distribution level for short and long-term operational planning. The tools will be built using novel data-driven energy models of millions of active nodes with AMI, 2) Adaptive stochastic analysis and optimization algorithms for real-time grid operations, 3) Dynamic Scenario Analysis using parallel Cloudenabled implementations with < 1 minute computational cycle times.

14 SOLAR ENERGY↗

A data-driven model for thermodynamic properties of a steam generator under cycling operation

The varying electricity demand from coal power plants due to the intermittent nature of renewable sources leads to load-follow and on/off operations referred to as cycling. Cycling causes transients of properties such as pressure and temperature within various components of the steam generation system.These transients cause increased damage because of fatigue and creep-fatigue interactions shortening the life of components. An algorithm is developed to identify cycling operations from the gross power data. The data-driven model based on artificial neural networks (ANN) is developed using 10 years data from Coal Creek Station power plant located in North Dakota, USA to estimate properties of the steam generator components during cycling operations. Furthermore, the uniqueness of this model is the ability to predict component properties for the cycling as well as base-load operations and is reported for the first time. The ANN model estimates the component properties, for a given gross power profile and initial conditions, as they vary during cycling operations. As a representative example, the ANN estimates are presented for the superheater outlet pressure, reheater inlet temperature, and flue gas temperature at the air heater inlet. The changes in these variables as a function of the gross power over the time duration are compared with measurements to assess the predictive capability of the model. Mean square errors of 4.49E-04 for superheater outlet pressure, 1.62E-03 for reheater inlet temperature, and 4.14E-04 for flue gas temperature at the air heater inlet were observed.

01 COAL, LIGNITE, AND PEAT↗

Association Between Injection and Microseismicity in Geothermal Fields With Multiple Wells: Data-Driven Modeling of Rotokawa, New Zealand, and Húsmúli, Iceland

Understanding injection-induced microseismicity in geothermal systems can provide insight into reservoir connectedness. However, fault and reservoir complexity are difficult to represent in simple analytical models, which makes it difficult to discern clear relationships from incidental associations. Here, we have used data-driven models to study how fluid injection and microseismicity are related in the Rotokawa (New Zealand) and Hellisheiði (Iceland) geothermal fields. We tested two classes of model: (a) lagged linear regression of seismicity rate as a function of well injection rates; and (b) systematic extraction of injection time series features that are then evaluated for associations with the seismicity. These models allowed us to determine which wells had the greatest correlation with microseismicity and to explain this association in a reservoir context. Finally, exploring different data types and transformations, we were unable to establish a link between rapid changes in injection rate and seismicity spikes, as suggested by some theoretical models.

15 GEOTHERMAL ENERGY↗

Data-Driven Modeling and Control of Systems with Plasma-Surface Interactions (Final Technical Report)

This final technical report summarizes the activities and accomplishments in the period from February 2023 thru January 2026. The objective of the proposed research is to investigate the physical mechanisms and processes underlying the formation of structures and patterns in systems with plasma-surface interactions. In the past decades, there have been extensive studies on the interaction of glow discharges, dielectric barrier discharges, and arc discharges with confining or intervening surfaces. The advancement of the understanding of these phenomena is not only of fundamental scientific interest and relevance to the knowledge of the plasma state, but also with profound implications in various technological applications. The research will integrate theoretical, computational, and experimental work within an innovative framework of data assimilation, i.e., optimally combining model predictions with measurements. The scientific merit of this research has three aspects. Firstly, it extends the studies of plasma-surface interactions to systems with insulator surfaces and multi-layer systems, while existing studies are predominantly on electrode surfaces. Secondly, it expects to develop a novel data-driven modeling approach based on data assimilation to enhance the predictive and control capabilities, which could make transformative contributions to basic plasma research. Thirdly, it will shed new light on outstanding problems related to formation of patterns interfacing plasmas. This project also aims to launch an education and outreach initiative at Texas A&M University-Kingsville, a non-R1, minority-serving institution in South Texas. The initiative is structured as a four-tier pyramid. Tier one will be a webinar series for culture and capacity building to inform broader audience in the region about the research fields of plasma science and engineering. Tier two will be the creation and offering of an upper-level undergraduate course on introductory plasma physics, which will help with the recruitment for the upper tiers. On tier three, we will engage and mentor senior design students to conduct work toward the research goal of this project. There will also be a certificate program on general plasma science for undergrad and graduate students, part of which will be lab training at Princeton University. Tier four will be the supervision and mentoring of Ph.D. students. Therefore, this project will systematically expand the talent pipeline, broaden participation from communities historically and geographically underrepresented in DOE SC research portfolio, significantly improve the research and education capacity at the PI’s institution, and contribute to developing a diverse workforce in plasma science and engineering.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

Collaborative Research: Enabling multi-scale studies of magnetic reconnection with interpretable data-driven models

The development of accurate reduced descriptions and improved closures for magnetic reconnection is an important and a long‐standing challenge in plasma physics. The four‐fluid approach, and associated closures, that were investigated have the potential to improve the accuracy of plasma fluid models, capturing physical effects which would otherwise require a kinetic description. If successful, this approach could have an important impact for the modeling of laboratory and space plasmas. The major goals of this project were to develop new machine learning (ML) tools based on sparse and symbolic regression techniques, and to extract interpretable and generalizable reduced models (e.g., in the form of partial differential equations - PDEs) from data generated by first principles plasma simulations. Preserving interpretability of such data‐driven models is key to addressing the long‐standing theoretical and numerical challenges. Prior proof‐of‐principle studies have demonstrated the enormous potential of this approach, by recovering the well‐established hierarchy of plasma equations (from Vlasov to MHD) from data produced by particle‐in‐cell (PIC) simulations. Our goal in this project was to extend and apply these new tools to construct better kinetic closures for magnetic reconnection; to derive better models of particle injection and acceleration by this fundamental plasma process; and to use this understanding to accelerate the development of multi‐scale plasma algorithms. While our immediate focus was on the problem of magnetic reconnection, the tools that were will developed are general and applicable to other areas of plasma physics, and more broadly to many‐body phenomena. We anticipate that the development of these multi‐scale models will have a significant impact across different areas of plasma science, from fusion to space and astrophysical plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance analysis and comparison of data-driven models for predicting indoor temperature in multi-zone commercial buildings

Building thermal models, which characterize the properties of a building’s envelope and thermal mass, are essential for accurate indoor temperature and cooling/heating demand prediction. Because of their flexibility and ease of use, data-driven models are increasingly used. Here, this study compared and analyzed the performance of gray-box (resistance-capacitance) and black-box (recurrent neural network) models for predicting indoor air temperature in a real multi-zone commercial building. The developed resistance-capacitance model served as a benchmark model for which full sets of temporal data and building information were used as inputs. The recurrent neural network models were trained and tested assuming various available types and amounts of temporal data and known building physical information to investigate the effects of data and information availability. Feature importance analysis was conducted to select the key variables for different prediction targets under different scenarios. This research provides guidance in selecting an appropriate building thermal response modeling method based on the measured data availability, building physical information, and application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data Driven Model Development for the Supersonic Semispan Transport (S(sup 4)T)

We investigate two common approaches to model development for robust control synthesis in the aerospace community; namely, reduced order aeroservoelastic modelling based on structural finite-element and computational fluid dynamics based aerodynamic models and a data-driven system identification procedure. It is shown via analysis of experimental Super- Sonic SemiSpan Transport (S4T) wind-tunnel data using a system identification approach it is possible to estimate a model at a fixed Mach, which is parsimonious and robust across varying dynamic pressures.

Kukreja, Sunil L.↗

Data Driven Model Development for the SuperSonic SemiSpan Transport (S(sup 4)T)

In this report, we will investigate two common approaches to model development for robust control synthesis in the aerospace community; namely, reduced order aeroservoelastic modelling based on structural finite-element and computational fluid dynamics based aerodynamic models, and a data-driven system identification procedure. It is shown via analysis of experimental SuperSonic SemiSpan Transport (S4T) wind-tunnel data that by using a system identification approach it is possible to estimate a model at a fixed Mach, which is parsimonious and robust across varying dynamic pressures.

Kukreja, Sunil L.↗

Rapid data-driven model reduction of nonlinear dynamical systems including chemical reaction networks using ℓ 1 -regularization

We develop a new data-driven paradigm for efficient model reduction of a broad class of nonlinear dynamical systems. Our model reduction method directly enables the interpretation of key components of the dynamical system, unlike traditional projection-based model reduction methods that focus on reducing computational complexity more than interpretability. Our method is not application specific and is simple to implement on nonlinear dynamical systems arising from a variety of different fields. It requires minimal parameterization using a single parameter to trade-off between model complexity and estimation error. We use a data-driven paradigm to formulate model reduction as an efficient convex optimization problem that scales polynomially in the original size of the complex system, enabling systems with as many as thousands of components to be reduced in a matter of minutes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗