Development of machine learning models for full field Reservoir Characterization
AGU 2021 Fall Meeting, Virtual, December 13-17, 2021.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
AGU 2021 Fall Meeting, Virtual, December 13-17, 2021.
We apply an interpretable Long Short-Term Memory (iLSTM) network for land-atmosphere carbon flux predictions based on time series observations of seven environmental variables. iLSTM enables interpretability of variable importance and variable-wise temporal importance to the prediction of targets by exploring internal network structures. The application results indicate that iLSTM not only improves prediction performance by capturing different dynamics of individual variables, but also reasonably interprets the different contribution of each variable to the target and its different temporal relevance to the target. This variable and temporal importance interpretation of iLSTM advances terrestrial ecosystem model development as well as our predictive understanding of the system.
Explore the source record for details and available documents.
In this study, we developed deep learning models that are capable of predicting spatio-temporal outputs of CO2 saturation, pressure, and brine production in a 3D saline storage reservoir over 30 years of continuous CO2 injection and a 50-year post-injection timeframe. To improve computational efficiency and maintain performance accuracy, the model framework involves ensembling multi-layer autoencoder networks that provide dimensionality reduction of geologic inputs with fully connected long short-term memory (LSTM) neural networks that generate time-series prediction. This study was presented as poster at the 2022 Carbon Management Project Review Meeting held in Pittsburgh, PA (August 15 – 19, 2022).
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The presentation was delivered at the 2022 American Institute of Chemical Engineers (AIChE) annual meeting.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
For presentation at the 2023 FECM/NETL Carbon Management Research Project Review Meeting, Pittsburgh, PA, August 28-September 1, 2023.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The data consists of a slide deck that was presented at the INFORMS 2023 conference. The presentation summarizes our approach to learning how to branch and compares our approach to the popular solver SCIP and a state-of-the-art ML-based branching rule.
Explore the source record for details and available documents.
Machine learning (ML) models have been applied to forecast solar energy; however, they often lack clarity of interpretability and underlying physics. This work addresses such challenges by developing a hierarchy of ML models that gradually introduce predictors to improve the forecast accuracy based on a physics-based framework. Three ML models (ARIMA, LSTM, and XGBoost) are examined and compared with four physics-informed persistence models reported in Part I and the simple persistence model to assess the improvement of different models. The 7-year measurements at the U.S. Department of Energy's Atmospheric Radiation Measurement's Southern Great Plains Central Facility site are used for forecasts and evaluations. The results reveal that the step-by-step introduction of predictors leads to different improvements for models at different hierarchical levels. Comparison of the ML models with persistence models shows that LSTM and XGBoost outperform all the persistence models, with LSTM having the overall best performance; however, ARIMA underperforms the four physics-informed persistence models. This study demonstrates the importance and utility of incorporating physics into ML models in improving forecast accuracy by introducing a hierarchy of physics-based predictors, distinguishing predictor contributions, and enhancing the ML interpretability. The combined use of Global Horizontal Irradiance (GHI) and Direct Normal Irradiance (DNI) significantly improves the forecast accuracy compared to using individual irradiances alone because the pair contains more information on cloud-radiation interactions.
Abstract We present an innovative approach called boosting Barlow Twins reduced order modeling (BBT‐ROM) to enhance the reliability of machine learning surrogate models for multiphase flow problems. BBT‐ROM builds upon Barlow Twins reduced order modeling that leverages self‐supervised learning to effectively handle linear and nonlinear manifolds by constructing well‐structured latent spaces of input parameters and output quantities. To address the challenge of high contrast data in multiphase flow problems due to injection wells and faults, we employ a boosting algorithm within BBT‐ROM. This algorithm sequentially trains a set of weak models (i.e., inaccurate models), improving prediction accuracy through ensemble learning. To evaluate the performance of BBT‐ROM, we conduct three three‐dimensional multiphase flow problems, including waterflooding and geologic carbon storage (GCS), with varying numbers of input parameter cases and model domain features. The results demonstrate that BBT‐ROM excels at predicting non‐wetting phase saturation (e.g., oil or saturation) and fluid pressure, with average relative errors ranging from 0.5% to 3%. Importantly, BBT‐ROM showcases robustness when faced with limited input parameter space during GCS testing.
Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.
This submission contains an ESRI map package (.mpk) with an embedded geodatabase for GIS resources used or derived in the Nevada Machine Learning project, meant to accompany the final report. The package includes layer descriptions, layer grouping, and symbology. Layer groups include: new/revised datasets (paleo-geothermal features, geochemistry, geophysics, heat flow, slip and dilation, potential structures, geothermal power plants, positive and negative test sites), machine learning model input grids, machine learning models (Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk) - supervised and unsupervised), original NV Play Fairway data and models, and NV cultural/reference data. See layer descriptions for additional metadata. Smaller GIS resource packages (by category) can be found in the related datasets section of this submission. A submission linking the full codebase for generating machine learning output models is available through the "Related Datasets" link on this page, and contains results beyond the top picks present in this compilation.