Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Data–Driven Velocity Model Evaluation Using K–Means Clustering

In this work, we develop a data-driven clustering method to evaluate a velocity model using surface wave velocity dispersion. This is done by first computing theoretical dispersion curves for 1-D velocity profiles of all the grid locations and then splitting the resulting dispersion curves into a certain number of groups via the K-means clustering. The observed dispersion curves are also clustered following the same procedure and the velocity model is assessed by comparing the spatial patterns obtained for the observed and synthetic data sets. The method is applied to evaluate two community velocity models in southern California, CVM-S4.26 and CVM-H15.1, using phase velocity maps derived for 3–16 s Rayleigh waves. We found a good correlation in the spatial distribution of clusters between the result of CVM-S4.26 and that of the observed data, suggesting that the CVM-S4.26 fits the observed dispersion maps better than the CVM-H15.1 in terms of features extracted from the clustering analysis.

58 GEOSCIENCES↗

Recurrent neural networks for short-term and long-term prediction of geothermal reservoirs

Accurate prediction of geothermal reservoir responses to alternative energy production scenarios is critical for optimizing the development of the underlying resources. While the conventional physics-based models offer a comprehensive prediction tool, data-driven models provide an efficient alternative to build fit-for-purpose predictive models by extracting and using the statistical patterns in the collected data to make predictions. The recurrent neural network (RNN) is a data-driven model that is commonly applied to predict time series sequences. This paper presents a variant of RNN that also utilizes the efficiency of convolutional neural networks (CNN) for the prediction of energy production from geothermal reservoirs. Specifically, a CNN–RNN architecture is developed that takes historical well controls as input (features) and their corresponding production response data as output (labels) to learn an input-output mapping that can predict the future well production responses/performance for any given future well control inputs. The model is paired with a labeling scheme to handle real field disturbances that create data gaps. In addition to the model structure, we introduce a thorough workflow for applying the model, which includes data pre-processing, feature selection, as well as different training strategies for short-term and long-term prediction. Finally, the performance and accuracy of the model are evaluated by applying it to multiple datasets, including a field reservoir model.

15 GEOTHERMAL ENERGY↗

Modular machine learning-based elastoplasticity: Generalization in the context of limited data

The development of highly accurate constitutive models for materials that undergo path-dependent processes continues to be a complex challenge in computational solid mechanics. Challenges arise both in considering the appropriate model assumptions and from the viewpoint of data availability, verification, and validation. Recently, data-driven modeling approaches have been proposed that aim to establish stress-evolution laws that avoid user-chosen functional forms by relying on machine learning representations and algorithms. However, these approaches not only require a significant amount of data but also need data that probes the full stress space with a variety of complex loading paths. Furthermore, they rarely enforce all necessary thermodynamic principles as hard constraints. Hence, they are in particular not suitable for low-data or limited-data regimes, where the first arises from the cost of obtaining the data and the latter from the experimental limitations of obtaining labeled data, which is commonly the case in engineering applications. In this work, we discuss a hybrid framework that can work on a variable amount of data by relying on the modularity of the elastoplasticity formulation where each component of the model can be chosen to be either a classical phenomenological or a data-driven model depending on the amount of available information and the complexity of the response. The method is tested on synthetic uniaxial data coming from simulations as well as cyclic experimental data for structural materials. The discovered material models are found to not only interpolate well but also allow for accurate extrapolation in a thermodynamically consistent manner far outside the domain of the training data. This ability to extrapolate from limited data was the main reason for the early and continued success of phenomenological models and the main shortcoming in machine learning-enabled constitutive modeling approaches. Training aspects and details of the implementation of these models into Finite Element simulations are discussed and analyzed.

42 ENGINEERING↗

Machine learning-based surrogate models and transfer learning for derivative free optimization of HT-PEM fuel cells

Widespread adoption of high-temperature polymer electrolyte membrane electrochemical systems, such as fuel cells (HT-PEMFCs), requires models and computational tools for accurate optimization and guiding new materials for enhancing performance and durability. In this contribution, knowledge-based modelling and data-driven modelling are combined using Few-Shot Learning and implementing an Automated Machine Learning framework for the generation of Machine Learning-based surrogate models. Applicability of the resulting model for derivative-free optimization is demonstrated. Additionally, a way of considering extrapolation in the optimization task is presented. Results show that although extrapolation is needed to achieve better solutions during optimization, it can be monitored and managed. As a result, tuning the electrode ionomer binder's properties, such as ionic conductivity, in the fuel cell represents a promising pathway for improving HT-PEMFC performance.

08 HYDROGEN↗

A data-driven framework for permeability prediction of natural porous rocks via microstructural characterization and pore-scale simulation

Understanding the microstructure–property relationships of porous media is of great practical significance, based on which macroscopic physical properties can be directly derived from measurable microstructural informatics. However, establishing reliable microstructure–property mappings in an explicit manner is difficult, due to the intricacy, stochasticity, and heterogeneity of porous microstructures. In this paper, a data-driven computational framework is presented to investigate the inherent microstructure–permeability linkage for natural porous rocks, where multiple techniques are integrated together, including microscopy imaging, stochastic reconstruction, microstructural characterization, pore-scale simulation, feature selection, and data-driven modeling. A large number of 3D digital rocks with a wide porosity range are acquired from microscopy imaging and stochastic reconstruction techniques. A broad variety of morphological descriptors are used to quantitatively characterize pore microstructures from different perspectives, and they compose the raw feature pool for feature selection. Here high-fidelity lattice Boltzmann simulations are conducted to resolve fluid flow passing through porous media, from which reliable permeability references are obtained. The optimal feature set that best represents permeability is identified through a performance-oriented feature selection process, upon which a cost-effective surrogate model is rapidly fitted to approximate the microstructure-permeability mapping via data-driven modeling. This surrogate model exhibits great advantages over empirical/analytical formulas in terms of prediction accuracy and generalization capacity, which can predict reliable permeability values spanning four orders of magnitude. Besides, feature selection also greatly enhances the interpretability of the data-driven prediction model, from which new insights into the mechanism of how microstructural characteristics determine intrinsic permeability are obtained.

58 GEOSCIENCES↗

An experimental database of cell performance for vanadium redox flow battery

The continual growth in energy demand has resulted in the deployment of renewable energy generators to reduce the impact of fossil fuel dependence. However, these generators often suffer from intermittency and require energy storage when there is over-generation and the subsequent release of this stored energy at high demand. One promising energy storage technology which can provide a solution to improve energy management and grid stability, is the redox flow battery. Among the numerous flow battery systems, vanadium redox flow battery is the most iconic solution to large scale energy storage, giving a more efficient link between energy production, especially from renewables, and energy demand. The aim of the current database is to characterize the performance of the cell design and to provide training/validation data for physical model or data-driven model. The database includes hundreds of experimental cell performance data of vanadium redox flow battery with various current densities for multiple charge-discharge cycles. All the cell parameters, chemical parameters, material parameters, operation parameters and thermodynamic parameters of the cell system are listed. Coulomb, voltaic and energy efficiencies are also provided. The database will be helpful for researchers in the field of redox flow batteries.

Gao, Peiyuan↗

A comparison of model validation approaches for echo state networks using climate model replicates

As global temperatures continue to rise, climate mitigation strategies such as stratospheric aerosol injections (SAI) are increasingly discussed, but the downstream effects of these strategies are not well understood. As such, there is interest in developing statistical methods to quantify the evolution of climate variable relationships during the time period surrounding an SAI. Feature importance applied to echo state network (ESN) models has been proposed as a way to understand the effects of SAI using a data-driven model. This approach depends on the ESN fitting the data well. If not, the feature importance may place importance on features that are not representative of the underlying relationships. Typically, time series prediction models such as ESNs are assessed using out-of-sample performance metrics that divide the times series into separate training and testing sets. However, this model assessment approach is geared towards forecasting applications and not scenarios such as the motivating SAI example where the objective is using a data driven model to capture variable relationships. Here, in this paper, we demonstrate a novel use of climate model replicates to investigate the applicability of the commonly used repeated hold-out model assessment approach for the SAI application. Simulations of an SAI are generated using a simplified climate model, and different initialization conditions are used to provide independent training and testing sets containing the same SAI event. The climate model replicates enable out-of-sample measures of model performance, which are compared to the single time series hold-out validation approach. For our case study, it is found that the repeated hold-out sample performance is comparable, but conservative, to the replicate out-of-sample performance when the training set contains enough time after the aerosol injection.

54 ENVIRONMENTAL SCIENCES↗

Active learning strategy for high fidelity short-term data-driven building energy forecasting

The quality of a data-driven model is heavily dependent on the quality of data. Data from building operation often have data bias problems, which means that the data sample is collected in a way that some members of the intended data population are less likely to be included than others. Data-driven energy forecasting models built on such data hence are biased and could lead to large forecasting errors. Active learning—an effective method to defying data bias—is rarely studied or applied in the area of data-driven building energy forecasting modeling. This paper attempts to fill this gap and explores the application of active learning in data-driven building energy forecasting. The developed strategy in this paper efficiently generate informative training data within a time budget and uses block design to passively consider weather disturbances. The developed active learning strategy is applied and evaluated in both virtual and real-building testbeds against traditional data-driven methods. Via these virtual and real-building evaluation cases, we have demonstrated that the data bias problem typically exists in building operation data is resolved by applying the developed active learning strategy. Furthermore, building energy forecasting models trained from data generated from the active learning strategy have shown improved performances in both model accuracy and model extendibility perspectives. The effectiveness of the block design module is also validated to effectively consider the impact of weather conditions on active learning design.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Neural lumped parameter differential equations with application in friction-stir processing

Lumped parameter methods aim to simplify the evolution of spatially-extended or continuous physical systems to that of a “lumped” element representative of the physical scales of the modeled system. For systems where the definition of a lumped element or its associated physics may be unknown, modeling tasks may be restricted to full-fidelity physics simulations. Here, in this work, we consider data-driven modeling tasks with limited point-wise measurements of otherwise continuous systems. We build upon the notion of the Universal Differential Equation (UDE) to construct data-driven models for reducing dynamics to that of a lumped parameter and inferring its properties. The flexibility of UDEs allow for composing various known physical priors suitable for application-specific modeling tasks, including lumped parameter methods. The motivating example for this work is the plunge and dwell stages for friction-stir welding; specifically, (i) mapping power input into the tool to a point-measurement of temperature and (ii) using this learned mapping for process control.

97 MATHEMATICS AND COMPUTING↗

Optimal Control of an Oscillating Surge Wave Energy Converter

During this project, we experimentally investigated the hydrodynamics and performance of a laboratory-scale oscillating surge wave energy converter (OSWEC).We looked at how flap buoyancy and driveline losses (primarily in the form of stiction) affected the dynamics and performance of the device. In addition, we assessed the influence of flap profile (rounded vs. square edges) on OSWEC hydrodynamics. Through this, we were able to develop a deeper understanding of OSWEC performance and provide guidance on strategies to counteract artifacts that may be present in laboratory models, but are absent in field-scale devices. To do this, we tested a laboratory-scale OSWEC in the Sea Wave Environmental Lab (SWEL) wave tank at the National Renewable Energy Laboratory (NREL). We ran several types of experiments to investigate the hydrodynamics and performance of the device. Overall, we achieved the overall goal of experimentally investigating the hydrodynamics and performance of this device. We discovered important and unexpected trends in performance, and collected time-resolved data to help us further investigate the underlying hydrodynamics responsible for these trends. In addition, we are currently using the time-resolved data from these experiments to build data-driven models of the dynamics, which can in turn be used to inform data-driven model predictive control of this device and address this objective in the future.

16 TIDAL AND WAVE POWER↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗

Data-driven surrogate modeling of hPIC ion energy-angle distributions for high-dimensional sensitivity analysis of plasma parameters' uncertainty

In this work, we present a data-driven strategy for effective construction of a surrogate model in high-dimensional parameter space for the ion energy-angle distribution (IEAD) output of hPIC simulations of plasma-surface interactions. The methodology is based on a bin-by-bin least-squares fitting of the IEAD in the parameter space. The fitting is performed in a transformed coordinate system to normalize the IEAD, and it employs sparse grids for sampling the parameter space to overcome sampling challenges in high dimensions. The surrogate model is significantly cheaper computationally than direct hPIC simulations yet maintains high fidelity to them, providing a fast emulator for hPIC simulations. Sensitivity analysis based on the surrogate model is utilized to characterize the dependence of the ion impact angle and energy moments on the physical parameters.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

PemNet: A Transfer Learning-Based Modeling Approach of High-Temperature Polymer Electrolyte Membrane Electrochemical Systems

Widespread adoption of high-temperature electrochemical systems such as polymer electrolyte membrane fuel cells (HT-PEMFCs) requires models and computational tools for accurate optimization and guiding new materials for enhancing fuel cell performance and durability. Furthermore, while robust and better suited for extrapolation, knowledge-based modeling has limitations as it is time-consuming and requires information about the system that is not always available (e.g., material properties and interfacial behavior between different materials). Data-driven modeling, on the other hand, is easier to implement but often necessitates large datasets that could be difficult to obtain. In this contribution, knowledge-based modeling and data-driven modeling are combined by implementing a few-shot learning (FSL) approach. A knowledge-based model originally developed for a HT-PEMFCs was used to generate simulated data (887,735 points) and used to pretrain a neural network source model tuned via a genetic algorithm-based AutoML. Then, experimental datasets from HT-PEMFCs with different materials and operating conditions (~50 points each) were used to train six target models via FSL. Models for the unseen data reached high accuracies in all cases (rRMSE < 10%).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Physics constrained learning for data-driven inverse modeling from sparse observations

Deep neural networks (DNN) have been used to model nonlinear relations between physical quantities. Those DNNs are embedded in physical systems described by partial differential equations (PDE) and trained by minimizing a loss function that measures the discrepancy between predictions and observations in some chosen norm. This loss function often includes the PDE constraints as a penalty term when only sparse observations are available. As a result, the PDE is only satisfied approximately by the solution. However, the penalty term typically slows down the convergence of the optimizer for stiff problems. We present a new approach that trains the embedded DNNs while numerically satisfying the PDE constraints. We develop an algorithm that enables differentiating both explicit and implicit numerical solvers in reverse-mode automatic differentiation. This allows the gradients of the DNNs and the PDE solvers to be computed in a unified framework. We demonstrate that our approach enjoys faster convergence and better stability in relatively stiff problems compared to the penalty method. Furthermore, our approach allows for the potential to solve and accelerate a wide range of data-driven inverse modeling, where the physical constraints are described by PDEs and need to be satisfied accurately.

97 MATHEMATICS AND COMPUTING↗

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora↗

Sharing is caring: An extensive analysis of parameter-based transfer learning for the prediction of building thermal dynamics

In recent years deep neural networks have been proposed as a lightweight data-driven model to capture high-dimensional, nonlinear physical processes to predict building thermal responses. However, the need of a large amount of data for the training process of deep neural networks clashes with the potential limited data availability in most existing or new buildings. Transfer learning aims to enhance the performance of a target learner exploiting knowledge from related and similar environments. This study conducted a suite of experiments that leveraged 250 data-driven models based on a synthetic dataset of a building archetype to study the influence of data availability, energy efficiency level, occupancy and climate for the transfer process of thermal dynamics. The performance of the transfer learning process was compared against a classical machine learning approach. Here, the results suggest that building thermal dynamics can be effectively transferred under the same climatic conditions, increasing performance when dealing with different occupancy schedules, efficiency levels and low data availability. Furthermore, the paper compares the performance of both transfer learning and machine learning approaches in an online fashion, to support the implementation in real-world deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ET-AL: Entropy-targeted active learning for bias mitigation in materials data

Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.

36 MATERIALS SCIENCE↗