Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Climate-invariant machine learning

Projecting climate change is a generalization problem: We extrapolate the recent past using physical models across past, present, and future climates. Current climate models require representations of processes that occur at scales smaller than model grid size, which have been the main source of model projection uncertainty. Recent machine learning (ML) algorithms hold promise to improve such process representations but tend to extrapolate poorly to climate regimes that they were not trained on. To get the best of the physical and statistical worlds, we propose a framework, termed “climate-invariant” ML, incorporating knowledge of climate processes into ML algorithms, and show that it can maintain high offline accuracy across a wide range of climate conditions and configurations in three distinct atmospheric models. Our results suggest that explicitly incorporating physical knowledge into data-driven models of Earth system processes can improve their consistency, data efficiency, and generalizability across climate regimes.

54 ENVIRONMENTAL SCIENCES↗

Survey on stochastic distribution systems: A full probability density function control theory with potential applications

Complex systems seen either in general engineering practice or economics are subjected to ever increased uncertainties that are mostly represented as random variables or parameters, and the characteristics of random variables are represented by their probability density functions (PDFs). Controlling their PDFs means to shape their stochastic distributions and in general it would provide a full treatment for system analysis and operational control and optimization. This leads to the development of stochastic distribution control (SDC) systems theory in the past decades, where the original aim of the controller design is to realize a shape control of the distributions of certain random variables in their PDFs sense for some engineering processes. Indeed, once the PDFs of these random variables or parameters are used to describe their distribution characters, the control task is to obtain control signals so that the output PDFs of stochastic systems are made to follow their target PDFs. The subject of SDC was initially originated for non-Gaussian stochastic control systems design but has found a wide spectrum of applications in general systems in terms of data-driven modeling, analysis, signal processing (filtering), data mining via multivariable statistics, decision-making (optimization) for systems subjected to uncertainties and even in economics. In this context, SDC constitutes an effective primer tool for complex system analysis, control and operational optimizations. In this review paper, a detailed survey of the developments on the research of SDC systems will be made together with their wide spectrum applications and future perspectives.

42 ENGINEERING↗

Deep-freeze graph training for latent learning

Scientific and engineering advances are primarily driven by multi-tier conceptual constructs and conditional theoretical frameworks. The theories allow predictions of hypothetical system responses, given a set of approximate conditions (ranges of applicability) imposed on latent parameters that cannot be measured directly. Learning to estimate the latent variables (Latent Learning) helps to pinpoint the anticipated range-edge anomalies and improves the confidence in interpretation, interpolation and extrapolation of limited experimental data. Due to high dimensionality and extreme non-linearity of the materials science problems, very large datasets are typically required for conventional data-driven model development. The vital experimental data collection, particularly on microstructural phases, is very challenging, which makes it difficult to compile a high-quality database. Incorporation of the domain knowledge into the computational graph structure, initialization and optimization processes presents a viable mechanism for developing accurate models, with limited datasets. Furthermore, this study successfully utilized the approach to build the Deep Freeze Graph (DeepFreG) by mapping known causality relationships and by digitizing empirical domain knowledge for Latent Learning (LL), with specific applications in materials science.

36 MATERIALS SCIENCE↗

Continental-Scale Controls on Hyporheic Respiration Revealed by Knowledge-Guided Machine Learning

Hyporheic zone sediments regulate organic matter turnover and in-stream respiration, yet controls on sediment respiration remain poorly constrained across heterogeneous river networks, limiting prediction of stream metabolism and carbon processing at continental scales. Here, we integrate observations from ~90 river corridors across the United States in the WHONDRS consortium with a knowledge-guided machine learning (KGML) framework that couples thermodynamic rate theory with machine learning to identify dominant controls on hyporheic respiration. Diagnostic analyses show that organic matter concentration and thermodynamic favorability define an upper bound on respiration potential, whereas biological catalytic capacity and physical accessibility jointly govern realized respiration rates through interaction effects. To represent unmeasurable accessibility constraints, we use the mechanistic model as a scaffold for KGML, allowing machine learning to target residual structure not explained by process theory. This hybrid framework improves predictive skill relative to both the mechanistic model alone and fully data-driven models while preserving interpretability. These results indicate that variability in hyporheic respiration is largely mechanistically structured and demonstrate how integrating process theory with explainable AI enhances predictive performance while enabling scalable synthesis of river corridor observations.

Zheng, Jianqiu↗

Automatic detection of low surface brightness galaxies from Sloan Digital Sky Survey images

ABSTRACT Low surface brightness (LSB) galaxies are galaxies with central surface brightness fainter than the night sky. Due to the faint nature of LSB galaxies and the comparable sky background, it is difficult to search LSB galaxies automatically and efficiently from large sky survey. In this study, we established the low surface brightness galaxies autodetect (LSBG-AD) model, which is a data-driven model for end-to-end detection of LSB galaxies from Sloan Digital Sky Survey (SDSS) images. Object-detection techniques based on deep learning are applied to the SDSS field images to identify LSB galaxies and estimate their coordinates at the same time. Applying LSBG-AD to 1120 SDSS images, we detected 1197 LSB galaxy candidates, of which 1081 samples are already known and 116 samples are newly found candidates. The B-band central surface brightness of the candidates searched by the model ranges from 22 to 24 mag arcsec−2, quite consistent with the surface brightness distribution of the standard sample. A total of 96.46 per cent of LSB galaxy candidates have an axial ratio (b/a) greater than 0.3, and 92.04 per cent of them have $fracDev\_r$ < 0.4, which is also consistent with the standard sample. The results show that the LSBG-AD model learns the features of LSB galaxies of the training samples well, and can be used to search LSB galaxies without using photometric parameters. Next, this method will be used to develop efficient algorithms to detect LSB galaxies from massive images of the next-generation observatories.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Sea Ice Predictability in E3SM with Machine Learning

Focal area(s): To improve predictions of sea ice in E3SM we propose to develop a hierarchy of data-driven models using observational and simulation data to investigate the most important Earth system drivers of sea ice variability and loss, develop surrogates that build on the reduced parameter space of important drivers, and, where appropriate, couple machine learning models with standard PDE models to capture important physical behavior at different scales. This work falls under Focal Area 2. Predictive modeling through the use of AI techniques.

54 ENVIRONMENTAL SCIENCES↗

Revolutionizing Materials Design: The Intersection of Quantum Mechanics and Data Modeling

The field of materials design is currently experiencing a notable evolution, driven by the convergence of sophisticated computational methodologies based on first principles and data-driven modeling approaches. I will review our recent endeavors employing AI/ML to expedite first-principles simulations and mitigate traditional methods' temporal and spatial limitations. Central to our efforts is developing and utilizing ML interatomic potentials (MLPs) across a diverse spectrum of materials. We show that MLPs serve as invaluable tools for navigating the complexities of the simulations, such as understanding the behavior of MgO at extreme environments of ~1 terapascal and temperatures >10,000 Kelvin. Moreover, we show that MLPs can provide precise details of the intricate dynamics governing the oxidation processes of binary alloy systems due to the competition between surface segregation and reconstruction tendencies. In summation, advancements in MLPs open the door to fresh possibilities in material modeling and, ultimately, discovery.

Saidi, Wissam↗

Evaluating performance of different generative adversarial networks for large-scale building power demand prediction

We report as an unsupervised-learning data-driven model, Generative Adversarial Networks (GANs) have recently attracted a lot of attention for various applications. There is potential to apply GANs for large-scale building power demand prediction, which is needed for power grid operation. However, there are many GAN variations and it is unclear which GAN is suitable for this application. To answer this question, this paper identifies five promising GANs (Original GAN, cGAN, SGAN, InfoGAN, and ACGAN) and evaluates their performance for predicting building power demand at a large scale. Physics-based building energy models are developed to generate training and reference data. A new evaluation indicator that combines accuracy and reproducibility is proposed to evaluate the performance of different GANs in predicting building power demand. The results show that SGAN and InfoGAN are not suitable because they cannot control the number of generated building samples for different building types. The prediction performance among the Original GAN, cGAN, and ACGAN can vary depending on training sample sizes and number of building types. If the training sample size is sufficiently large, Original GAN and cGAN can predict building power demand more accurately than ACGAN with the same number of samples. If training samples are limited, Original GAN provides better accuracy than cGAN and ACGAN. When the number of building types increase, the prediction accuracy increases for cGAN, decreases for ACGAN, and remains the same for Original GAN. As a result, cGAN and Original GAN are recommended for large-scale building power demand prediction.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Frame invariant neural network closures for Kraichnan turbulence

Numerical simulations of geophysical and atmospheric flows have to rely on parameterizations of subgrid scale processes due to their limited spatial resolution. Despite substantial progress in developing parameterization (or closure) models for subgrid scale (SGS) processes using physical insights and mathematical approximations, they remain imperfect and can lead to inaccurate predictions. In recent years, machine learning has been successful in extracting complex patterns from high-resolution spatio-temporal data, leading to improved parameterization models, and ultimately better coarse grid prediction. However, the inability to satisfy known physics and poor generalization hinders the application of these models for real-world problems. In this work, we put forth a frame invariant closure approach to improve the accuracy and generalizability of deep learning-based subgrid scale closure models by embedding physical symmetries directly into the structure of the neural network. Specifically, we utilized specialized layers within the convolutional neural network in such a way that desired constraints are theoretically guaranteed without the need for any regularization terms. We demonstrate our framework for a two-dimensional decaying turbulence test case mostly characterized by the forward enstrophy cascade. We show that our frame invariant SGS model (i) accurately predicts the subgrid scale source term, (ii) respects the physical symmetries such as translation, Galilean, and rotation invariance, and (iii) is numerically stable when implemented in coarse-grid simulation with generalization to different initial conditions and Reynolds number. This work opens up a possibility of connecting physics-based theories and data-driven modeling paradigms, and thus represents a promising step towards the development of physically consistent data-driven turbulence closure models.

42 ENGINEERING↗

Adaptively Learned Modeling for a Digital Twin of Hydropower Turbines with Application to a Pilot Testing System

In the development of a digital twin (DT) for hydropower turbines, dynamic modeling of the system (e.g., penstock, turbine, speed control) is crucial, along with all the necessary data interface, virtualization, and dashboard designs. Since the DT must mimic the actual dynamics of the hydropower turbine accurately, adaptive learning is required to train these dynamic models online so that the models in the DT can effectively follow the representation of the actual hydropower turbine dynamics accurately and reliably. This study presents an adaptive learning method for obtaining the hydropower turbine models for DT development of hydropower systems using the recursive least squares algorithm. To simplify the formulation, the hydropower turbine under consideration was assumed to operate near a fixed operating point, where the system dynamics can be well represented by a set of linear differential equations with constant parameters. In this context, the well-known six-coefficient model for the Francis turbine was formulated as the starting point to obtain input and output models for the turbine. Then, an adaptive learning mechanism was developed to learn model parameters using real-time data from a hydropower turbine testing system. This led to semi-physical modeling, in which first principles and data-driven modeling are integrated to produce dynamic models for DT development. Applications to a pilot system at the Norwegian University of Science and Technology (NTNU) were made, and the models learned adaptively using the data collected from the university’s pilot system. Desired modeling and validation results were obtained.

13 HYDRO ENERGY↗

Calibrating hypersonic turbulence flow models with the HIFiRE-1 experiment using data-driven machine-learned models.

In this paper we study the efficacy of combining machine-learning methods with projection-based model reduction techniques for creating data-driven surrogate models of computationally expensive, high-fidelity physics models. Such surrogate models are essential for many-query applications e.g., engineering design optimization and parameter estimation, where it is necessary to invoke the high-fidelity model sequentially, many times. Surrogate models are usually constructed for individual scalar quantities. However there are scenarios where a spatially varying field needs to be modeled as a function of the model’s input parameters. Here we develop a method to do so, using projections to represent spatial variability while a machine-learned model captures the dependence of the model’s response on the inputs. The method is demonstrated on modeling the heat flux and pressure on the surface of the HIFiRE-1 geometry in a Mach 7.16 turbulent flow. The surrogate model is then used to perform Bayesian estimation of freestream conditions and parameters of the SST (Shear Stress Transport) turbulence model embedded in the high-fidelity (Reynolds-Averaged Navier–Stokes) flow simulator, using shock-tunnel data. The paper provides the first-ever Bayesian calibration of a turbulence model for complex hypersonic turbulent flows. We find that the primary issues in estimating the SST model parameters are the limited information content of the heat flux and pressure measurements and the large model-form error encountered in a certain part of the flow.

42 ENGINEERING↗

Bayesian differential programming for robust systems identification under uncertainty

This paper presents a machine learning framework for Bayesian systems identification from noisy, sparse and irregular observations of nonlinear dynamical systems. The proposed method takes advantage of recent developments in differentiable programming to propagate gradient information through ordinary differential equation solvers and perform Bayesian inference with respect to unknown model parameters using Hamiltonian Monte Carlo sampling. This allows an efficient inference of the posterior distributions over plausible models with quantified uncertainty, while the use of sparsity-promoting priors enables the discovery of interpretable and parsimonious representations for the underlying latent dynamics. A series of numerical studies is presented to demonstrate the effectiveness of the proposed methods, including nonlinear oscillators, predator–prey systems and examples from systems biology. Taken together, our findings put forth a flexible and robust workflow for data-driven model discovery under uncertainty. All codes and data accompanying this article are available at https://bit.ly/34FOJMj .

Science & Technology - Other Topics↗

Thermodynamic Consistent Neural Networks for Learning Material Interfacial Mechanics

For multilayer materials in thin substrate systems, interfacial failure is one of the most challenges. The traction-separation relations (TSR) quantitatively describe the mechanical behavior of a material interface undergoing openings, which is critical to understand and predict interfacial failures under complex loadings. However, existing theoretical models have limitations on enough complexity and flexibility to well learn the real-world TSR from experimental observations. A neural network can fit well along with the loading paths but often fails to obey the laws of physics, due to a lack of experimental data and understanding of the hidden physical mechanism. In this paper, we propose a thermodynamic consistent neural network (TCNN) approach to build a data-driven model of the TSR with sparse experimental data. The TCNN leverages recent advances in physics-informed neural networks (PINN) that encode prior physical information into the loss function and efficiently train the neural networks using automatic differentiation. We investigate three thermodynamic consistent principles, i.e., positive energy dissipation, steepest energy dissipation gradient, and energy conservative loading path. All of them are mathematically formulated and embedded into a neural network model with a novel defined loss function. A real-world experiment demonstrates the superior performance of TCNN, and we find that TCNN provides an accurate prediction of the whole TSR surface and significantly reduces the violated prediction against the laws of physics.

Zhang, Jiaxin↗

Accelerating scientific discoveries through data-driven innovations

Developing artificial intelligence (AI) and machine learning (ML) methods that can accelerate scientific discoveries and advance science has become one of the important research directions for the AI/ML research community. It has been gaining increasing attention from researchers in diverse scientific areas, including biomedical science, materials science, climate science, physics, chemistry, and many others. Data-driven AI/ML innovations to enable reliable predictions and optimal decision making for scientific discoveries face several critical challenges, among which are high system complexity, large search space, incomplete knowledge, and small data, all of which demand novel strategies to effectively address them. Meeting these challenges and thereby accelerating scientific discoveries and industrial innovations, calls for research that can take full advantage of the latest advances in AI/ML to integrate data-driven techniques with scientific knowledge and is able to execute them in modern high-performance computing (HPC) environments at scale. This Patterns special collection "Accelerating scientific discoveries through data-driven innovations" features articles that showcase the promising roles of AI/ML and data-driven modeling in accelerating scientific discoveries and may inspire the next wave of data-driven innovations in various scientific domains.

97 MATHEMATICS AND COMPUTING↗

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees↗

A Hybrid Data-Driven and Model-Based Anomaly Detection Scheme for DER Operation

This paper proposes a hybrid data and model-based anomaly detection scheme to secure the operation of distributed energy resources (DERs) in distribution grids. Data-driven autoencoders are set up at the edge device level and they use local DER operational data as inputs. The abnormal statuses are detected by analyzing reconstruction errors. In parallel, modelbased state estimation (SE) is set up at the central level and it uses system-wide models and measurements as data inputs. The anomalies are identified by analyzing measurement residuals. The hybrid scheme preserves the benefits of both data-driven and model-based analyses and thus improves the robustness and the accuracy of anomaly detection. Numerical tests based on the model of a real distribution feeder in Southern California highlight the proposed scheme's effectiveness and benefits.

anomaly detection↗

Carbon-phosphorus cycle models overestimate CO 2 enrichment response in a mature Eucalyptus forest

The importance of phosphorus (P) in regulating ecosystem responses to climate change has fostered P-cycle implementation in land surface models, but their CO 2 effects predictions have not been evaluated against measurements. Here, we perform a data-driven model evaluation where simulations of eight widely used P-enabled models were confronted with observations from a long-term free-air CO 2 enrichment experiment in a mature, P-limited Eucalyptus forest. We show that most models predicted the correct sign and magnitude of the CO 2 effect on ecosystem carbon (C) sequestration, but they generally overestimated the effects on plant C uptake and growth. We identify leaf-to-canopy scaling of photosynthesis, plant tissue stoichiometry, plant belowground C allocation, and the subsequent consequences for plant-microbial interaction as key areas in which models of ecosystem C-P interaction can be improved. Together, this data-model intercomparison reveals data-driven insights into the performance and functionality of P-enabled models and adds to the existing evidence that the global CO 2 -driven carbon sink is overestimated by models.

54 ENVIRONMENTAL SCIENCES↗

Long-term leaf C:N ratio change under elevated CO 2 and nitrogen deposition in China: Evidence from observations and process-based modeling

Climate change, elevating atmosphere CO 2 (eCO 2 ) and increased nitrogen deposition (iNDEP) are altering the biogeochemical interactions between plants, microbes and soils, which further modify plant leaf carbon-nitrogen (C:N) stoichiometry and their carbon assimilation capability. Many field experiments have observed large sensitivity of leaf C:N ratio to eCO 2 and iNDEP. However, the large-scale pattern of this sensitivity is still unclear, because eCO 2 and iNDEP drive leaf C:N ratio toward opposite directions, which are further compounded by the complex processes of nitrogen acquisition and plant-and-microbial nitrogen competition. Here, we attempt to map the leaf C:N ratio spatial variation in the past 5 decades in China with a combination of data-driven model and process-based modeling. These two approaches showed consistent results. Over different regions, we found that leaf C:N ratio had significant but uneven changes between 2 time periods (1960-1989 and 1990-2015): a 5% ± 8% increase for temperate grasslands in northern China, a 3% ± 6% increase for boreal grasslands in western China, and by contrast, a 7% ± 6% decrease for temperate forests in southern China, and a 3% ± 5% decrease for boreal forests in northeastern China. Additionally, the structural equation models indicated that the leaf C:N change was sensitive to ΔNDEP, ΔCO 2 and ΔMAT rather than ΔMAP and ecosystem types. In this work, process-based modeling suggested that iNDEP was the main source of soil mineral nitrogen change, dominating leaf C:N ratio change in most areas in China, while eCO 2 led to leaf C:N ratio increase in low iNDEP area. This study also indicates that the long-term leaf C:N ratio acclimation was dominated by climate constraint, especially temperature, but was constrained by soil N availability over decade scale.

54 ENVIRONMENTAL SCIENCES↗