Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Bayesian differential programming for robust systems identification under uncertainty

This paper presents a machine learning framework for Bayesian systems identification from noisy, sparse and irregular observations of nonlinear dynamical systems. The proposed method takes advantage of recent developments in differentiable programming to propagate gradient information through ordinary differential equation solvers and perform Bayesian inference with respect to unknown model parameters using Hamiltonian Monte Carlo sampling. This allows an efficient inference of the posterior distributions over plausible models with quantified uncertainty, while the use of sparsity-promoting priors enables the discovery of interpretable and parsimonious representations for the underlying latent dynamics. A series of numerical studies is presented to demonstrate the effectiveness of the proposed methods, including nonlinear oscillators, predator–prey systems and examples from systems biology. Taken together, our findings put forth a flexible and robust workflow for data-driven model discovery under uncertainty. All codes and data accompanying this article are available at https://bit.ly/34FOJMj .

Science & Technology - Other Topics↗

Thermodynamic Consistent Neural Networks for Learning Material Interfacial Mechanics

For multilayer materials in thin substrate systems, interfacial failure is one of the most challenges. The traction-separation relations (TSR) quantitatively describe the mechanical behavior of a material interface undergoing openings, which is critical to understand and predict interfacial failures under complex loadings. However, existing theoretical models have limitations on enough complexity and flexibility to well learn the real-world TSR from experimental observations. A neural network can fit well along with the loading paths but often fails to obey the laws of physics, due to a lack of experimental data and understanding of the hidden physical mechanism. In this paper, we propose a thermodynamic consistent neural network (TCNN) approach to build a data-driven model of the TSR with sparse experimental data. The TCNN leverages recent advances in physics-informed neural networks (PINN) that encode prior physical information into the loss function and efficiently train the neural networks using automatic differentiation. We investigate three thermodynamic consistent principles, i.e., positive energy dissipation, steepest energy dissipation gradient, and energy conservative loading path. All of them are mathematically formulated and embedded into a neural network model with a novel defined loss function. A real-world experiment demonstrates the superior performance of TCNN, and we find that TCNN provides an accurate prediction of the whole TSR surface and significantly reduces the violated prediction against the laws of physics.

Zhang, Jiaxin↗

Accelerating scientific discoveries through data-driven innovations

Developing artificial intelligence (AI) and machine learning (ML) methods that can accelerate scientific discoveries and advance science has become one of the important research directions for the AI/ML research community. It has been gaining increasing attention from researchers in diverse scientific areas, including biomedical science, materials science, climate science, physics, chemistry, and many others. Data-driven AI/ML innovations to enable reliable predictions and optimal decision making for scientific discoveries face several critical challenges, among which are high system complexity, large search space, incomplete knowledge, and small data, all of which demand novel strategies to effectively address them. Meeting these challenges and thereby accelerating scientific discoveries and industrial innovations, calls for research that can take full advantage of the latest advances in AI/ML to integrate data-driven techniques with scientific knowledge and is able to execute them in modern high-performance computing (HPC) environments at scale. This Patterns special collection "Accelerating scientific discoveries through data-driven innovations" features articles that showcase the promising roles of AI/ML and data-driven modeling in accelerating scientific discoveries and may inspire the next wave of data-driven innovations in various scientific domains.

97 MATHEMATICS AND COMPUTING↗

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees↗

A Hybrid Data-Driven and Model-Based Anomaly Detection Scheme for DER Operation

This paper proposes a hybrid data and model-based anomaly detection scheme to secure the operation of distributed energy resources (DERs) in distribution grids. Data-driven autoencoders are set up at the edge device level and they use local DER operational data as inputs. The abnormal statuses are detected by analyzing reconstruction errors. In parallel, modelbased state estimation (SE) is set up at the central level and it uses system-wide models and measurements as data inputs. The anomalies are identified by analyzing measurement residuals. The hybrid scheme preserves the benefits of both data-driven and model-based analyses and thus improves the robustness and the accuracy of anomaly detection. Numerical tests based on the model of a real distribution feeder in Southern California highlight the proposed scheme's effectiveness and benefits.

anomaly detection↗

Carbon-phosphorus cycle models overestimate CO 2 enrichment response in a mature Eucalyptus forest

The importance of phosphorus (P) in regulating ecosystem responses to climate change has fostered P-cycle implementation in land surface models, but their CO 2 effects predictions have not been evaluated against measurements. Here, we perform a data-driven model evaluation where simulations of eight widely used P-enabled models were confronted with observations from a long-term free-air CO 2 enrichment experiment in a mature, P-limited Eucalyptus forest. We show that most models predicted the correct sign and magnitude of the CO 2 effect on ecosystem carbon (C) sequestration, but they generally overestimated the effects on plant C uptake and growth. We identify leaf-to-canopy scaling of photosynthesis, plant tissue stoichiometry, plant belowground C allocation, and the subsequent consequences for plant-microbial interaction as key areas in which models of ecosystem C-P interaction can be improved. Together, this data-model intercomparison reveals data-driven insights into the performance and functionality of P-enabled models and adds to the existing evidence that the global CO 2 -driven carbon sink is overestimated by models.

54 ENVIRONMENTAL SCIENCES↗

Data-Enabled Fusion Technology (Final Scientific/Technical Report)

Advancing Scientific Understanding in Fusion Energy and Machine Learning This research represented a significant step forward in machine learning (ML) applications for fusion energy experiments. The project integrated advanced data-driven modeling, optimization techniques, and artificial intelligence to enhance the predictive capabilities and operational efficiency of plasma-based fusion systems. Specifically, tasks focused on ML-enhanced diagnostics, operator guidance tools, and predictive modeling helped improve the ability to interpret complex fusion experiments. Key areas of advancement included: 1) data-driven plasma control, i.e., using ML algorithms to optimize experimental conditions and classify plasma behaviors based on historical data; 2) spectroscopy and diagnostics, i.e., applying AI models to extract previously inaccessible insights from experimental spectroscopy data; and 3) configuration mapping and operator guidance, i.e., developing a predictive framework to assist scientists in identifying the most effective experimental parameters, reducing reliance on manual adjustments. By refining these ML-driven techniques, the project contributed to the broader scientific community’s understanding of plasma dynamics and fusion energy viability. Technical Effectiveness and Economic Feasibility The methods investigated demonstrated high technical effectiveness, as reflected in milestones assessing the predictive accuracy, performance, and optimization of fusion configurations. The development of an Operator Guidance Tool (OGT), for example, led to more precise control of plasma conditions by learning from experimental data and offering real-time adjustments. From an economic standpoint, DeFT provided: 1) the ability to reduce trial-and-error experimentation, which lowered operational costs; 2) improved data interpretation methods, which enabled more efficient resource allocation in large-scale fusion research projects; and 3) the automation of key diagnostic tasks, which reduced manual labor and human error, increasing overall efficiency. 13 The final assessments of predictive models and optimization strategies demonstrated that these approaches were scalable and could be implemented across multiple fusion energy research programs. Public Benefit and Societal Impact This project contributed directly to the broader goal of achieving sustainable and commercially viable fusion energy, which had profound implications for clean energy production and climate change mitigation. The integration of AI-driven solutions into fusion research: 1) sped up scientific discovery, accelerating progress towards achieving energy breakthroughs; 2) reduced the cost of experimentation, making fusion research more accessible; and 3) provided a framework for future AI applications in high-energy physics, benefiting adjacent fields like space exploration, material science, and renewable energy. Additionally, by fostering collaborations between AI researchers and plasma physicists, this project promoted interdisciplinary innovation that could lead to broader applications beyond fusion research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Long-term leaf C:N ratio change under elevated CO 2 and nitrogen deposition in China: Evidence from observations and process-based modeling

Climate change, elevating atmosphere CO 2 (eCO 2 ) and increased nitrogen deposition (iNDEP) are altering the biogeochemical interactions between plants, microbes and soils, which further modify plant leaf carbon-nitrogen (C:N) stoichiometry and their carbon assimilation capability. Many field experiments have observed large sensitivity of leaf C:N ratio to eCO 2 and iNDEP. However, the large-scale pattern of this sensitivity is still unclear, because eCO 2 and iNDEP drive leaf C:N ratio toward opposite directions, which are further compounded by the complex processes of nitrogen acquisition and plant-and-microbial nitrogen competition. Here, we attempt to map the leaf C:N ratio spatial variation in the past 5 decades in China with a combination of data-driven model and process-based modeling. These two approaches showed consistent results. Over different regions, we found that leaf C:N ratio had significant but uneven changes between 2 time periods (1960-1989 and 1990-2015): a 5% ± 8% increase for temperate grasslands in northern China, a 3% ± 6% increase for boreal grasslands in western China, and by contrast, a 7% ± 6% decrease for temperate forests in southern China, and a 3% ± 5% decrease for boreal forests in northeastern China. Additionally, the structural equation models indicated that the leaf C:N change was sensitive to ΔNDEP, ΔCO 2 and ΔMAT rather than ΔMAP and ecosystem types. In this work, process-based modeling suggested that iNDEP was the main source of soil mineral nitrogen change, dominating leaf C:N ratio change in most areas in China, while eCO 2 led to leaf C:N ratio increase in low iNDEP area. This study also indicates that the long-term leaf C:N ratio acclimation was dominated by climate constraint, especially temperature, but was constrained by soil N availability over decade scale.

54 ENVIRONMENTAL SCIENCES↗

Structured Neural Network Modeling for Developing Digital Twins Models of Hydropower Generation Units

Dynamic modeling is a key part in the development of digital twin (DT) for dynamic systems. This is true for hydropower systems, where whole system modeling including penstock, turbine and generators, etc is important in realizing actuate modeling for the real systems. On the other hand, in response to the large variations of the power demand due to increased penetration of renewables such as wind and solar, hydropower systems are now required to operate in a large power generation range. This situation triggers the nonlinear characteristics of the generation unit with respect to its models. As such, it is imperative to use data driven modeling such as neural networks to learn the nonlinear dynamics of the hydropower generation unit. To achieve this objective, this study constructs a modeling and learning algorithm integrated with multiple structured neural network models for the modeling of turbine shaft speed, penstock pressure, and generator power output based on the generator power control setpoint, field current, and field voltage. In addition, the study uses the hydropower data from Tacoma Public Utilities to train and validate the proposed neural network algorithm. The results have shown that this structured neural network modeling approach can learn the system dynamics effectively by using the real-time data collected from the hydropower system with the desired modeling results.

Wang, Hong↗

Learning Implicit Models of Complex Dynamical Systems From Partial Observations [Slides]

Conclusions: certain data-driven models do implicitly what physics simulations do explicitly; theoretical underpinnings in Koopman theory, delay-coordinate embeddings, and Mori-Zwanzig formalism; best approach for given system with finite data an open question; physics-informed machine learning (PIML) emerging framework for combining explicit and implicit modeling.

97 MATHEMATICS AND COMPUTING↗

Accelerating Multiphase Simulations With Denoising Diffusion Model Driven Initializations

This study introduces a hybrid fluid simulation approach that integrates generative diffusion models with physics‐based simulations, aiming at reducing the computational costs of flow simulations while still honoring all the physical properties of interest. Pore‐scale simulations enhance our understanding of applications such as assessing hydrogen and storage efficiency in underground reservoirs. Nevertheless, they are computationally expensive and the presence of non‐unique solutions can require multiple simulations within a single geometry. To overcome the computational cost hurdle, we propose a method that couples generative diffusion models and physics‐based simulations. While training the data‐driven model, we simultaneously generate initial conditions and perform physics‐based simulations using these. This integrated approach enables us to receive real‐time feedback on a single compute node equipped with both CPUs and GPUs. By efficiently managing these processes within a single compute node, we can continuously monitor performance and halt training once the model meets the specified criteria. To test our model, we generate realizations in a real Berea sandstone fracture which shows that our technique is up to 4.4 times faster than commonly used flow simulation initializations.

36 MATERIALS SCIENCE↗

Explainable Bayesian Neural Network for Probabilistic Transient Stability Analysis Considering Wind Energy

While several data-driven models have been developed for transient stability assessment, how to consider the uncertainties from load and renewable generations and provide interpretation of data-driven assessment results are still open. This paper proposes an explainable Bayesian Neural Network (BNN) for probabilistic transient stability assessment (TSA). By extracting the uncertainties from loads and wind farms, the BNN model can make a reliable prediction and quantify the prediction uncertainties. We also develop the Gradient Shap algorithm to make the global and local explanations for the probabilistic TSA model, a significant advantage over existing black-box data-driven methods. Numerical results on the modified IEEE 39-bus system show that the proposed method outperforms the existing methods in terms of prediction accuracy and uncertainty quantification capabilities. The explainability of the proposed method allows system operators to design preventive controls for enhancing system stability.

Bayesian Neural Network↗

Physics-informed State-space Neural Networks for transport phenomena

This work introduces Physics -informed State -space neural network Models (PSMs), a novel solution to achieving real-time optimization, flexibility, and fault tolerance in autonomous systems, particularly in transportdominated systems such as chemical, biomedical, and power plants. Traditional data -driven methods fall short due to a lack of physical constraints like mass conservation; PSMs address this issue by training deep neural networks with sensor data and physics -informing using components' Partial Differential Equations (PDEs), resulting in a physics -constrained, end -to -end differentiable forward dynamics model. Further, through two in silico experiments - a heated channel and a cooling system loop - we demonstrate that PSMs offer a more accurate approach than a purely data -driven model. In the former experiment, PSMs demonstrated significantly lower average root -mean -square errors across test datasets compared to a purely data -driven neural network, with reductions of 44 %, 48 %, and 94 % in predicting pressure, velocity, and temperature, respectively. Beyond accuracy, PSMs demonstrate a compelling multitask capability, making them highly versatile. In this work, we showcase two: supervisory control of a nonlinear system through a sequentially updated state -space representation and the proposal of a diagnostic algorithm using residuals from each of the PDEs. The former demonstrates PSMs' ability to handle constant and time -dependent constraints, while the latter illustrates their value in system diagnostics and fault detection.

42 ENGINEERING↗

Estimating Sediment Settling Velocities from a Theoretically Guided Data-Driven Approach

Sediment settling velocities are commonly estimated from analytical or process-based approaches. These approaches have theoretical constraints due to the incompletely resolved settling physics. A parametric data-driven approach was recently proposed without theoretical constraints, but it is limited by its mathematical assumptions. To overcome these limitations, here we apply a machine learning algorithm to an aggregated sediment settling experimental database and develops a nonparametric data-driven model to estimate the noncohesive sediment settling velocity in water. A cross-comparison against five process-based equations and a parametric data-driven equation demonstrates the higher accuracy and better consistency of the new model in estimating sediment settling velocities under various physical regimes. The new model also shows an easily implemented self-update capability by assimilating theoretical data derived from the process-based equations. The updated model, incorporating experimental and theoretical data of sediment settling processes, further improves the accuracy and reduces the uncertainty in estimating sediment settling velocities. This study demonstrates the capability of machine learning in sediment transport study and illustrates an alternative framework for other hydraulic engineering challenges.

42 ENGINEERING↗

The Double-edged Sword of Data-driven Super-Resolution: Adversarial Super-resolution Models

Data-driven super-resolution (SR) methods are often integrated into imaging pipelines as preprocessing steps to improve downstream tasks such as classification and detection. However, these SR models introduce a previously unexplored attack surface into imaging pipelines. In this paper, we present AdvSR, a framework demonstrating that adversarial behavior can be embedded directly into SR model weights during training, requiring no access to inputs at inference time. Unlike prior attacks that perturb inputs or rely on backdoor triggers, AdvSR operates entirely at the model level. By jointly optimizing for reconstruction quality and targeted adversarial outcomes, AdvSR produces models that appear benign under standard image quality metrics while inducing downstream misclassification. We evaluate AdvSR on three SR architectures (SRCNN, EDSR, SwinIR) paired with a YOLOv11 classifier and demonstrate that AdvSR models can achieve high attack success rates with minimal quality degradation. These findings highlight a new model-level threat for imaging pipelines, with implications for how practitioners source and validate models in safety-critical applications.

Sullivan, Haley [ORNL] (ORCID:0000000274069217)↗

A Hybrid Data-Driven and Model-Based Anomaly Detection Scheme for DER Operation: Preprint

This paper proposes a hybrid data and model-based anomaly detection for securing the operation of distributed energy resources (DERs) in distribution grids. Data-driven autoencoders (AE) are set up at the edge level by taking local DER data and detect anomalous operations by leveraging the reconstruction ability. In parallel, model-based state estimation (SE) is running at the system level by taking system models and measurements, the anomalies are identified by analyzing the measurements residual. The hybrid scheme preserves the benefits of both data-driven and model-based analysis and thus improves the robustness and accuracy of anomaly detection. It can be established by getting full use of the existing infrastructures in distribution grids. Numerical tests on a realistic distribution feeder in Southern California highlight the effectiveness as well as benefits of the proposed scheme.

anomaly detection↗

Scientific machine learning for modeling and simulating complex fluids

The formulation of rheological constitutive equations—models that relate internal stresses and deformations in complex fluids—is a critical step in the engineering of systems involving soft materials. While data-driven models provide accessible alternatives to expensive first-principles models and less accurate empirical models in many engineering disciplines, the development of similar models for complex fluids has lagged. The diversity of techniques for characterizing non-Newtonian fluid dynamics creates a challenge for classical machine learning approaches, which require uniformly structured training data. Consequently, early machine-learning based constitutive equations have not been portable between different deformation protocols or mechanical observables. Here, we present a data-driven framework that resolves such issues, allowing rheologists to construct learnable models that incorporate essential physical information, while remaining agnostic to details regarding particular experimental protocols or flow kinematics. These scientific machine learning models incorporate a universal approximator within a materially objective tensorial constitutive framework. By construction, these models respect physical constraints, such as frame-invariance and tensor symmetry, required by continuum mechanics. We demonstrate that this framework facilitates the rapid discovery of accurate constitutive equations from limited data and that the learned models may be used to describe more kinematically complex flows. This inherent flexibility admits the application of these “digital fluid twins” to a range of material systems and engineering problems. We illustrate this flexibility by deploying a trained model within a multidimensional computational fluid dynamics simulation—a task that is not achievable using any previously developed data-driven rheological equation of state.

Science & Technology - Other Topics↗

Automatic recognition system for document digitization in nuclear power plants

With the increasing number of data-driven models in nuclear applications, large volumes of numerical data are required to accurately model and predict the health status of a plant component. However, many historical operation logs that contain useful information are not fully utilized due to the lack of a systematic approach of digitization. To overcome this issue, this study proposes an automatic pipeline for extracting information from handwritten tabular documents collected from nuclear power plants. In our pipeline, we first denoise scanned documents with morphological operations, and then extract relevant parts from individual pages using both traditional computer vision and neural network methods. Handwriting recognition is applied to obtain text and numbers. As the most challenging step is how to crop only relevant information, the main focus of our paper is to detect tables and cells from scanned handwritten documents. Here we evaluate the efficiency and accuracy of our proposed method on handwritten operational reports obtained from a real-world case study. The results demonstrate the high accuracy and practicality of our proposed method.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗