Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Lidar - ESRL WindCube 200s, Wasco Airport - Reviewed Data

The available "Readme" file introduces the basics of the Doppler lidar data and offers a detailed description of the variables present in the data files. For those with any further questions about the data and its interpretation, contact either Alan Brewer ( ) or Sunil Baidar ( ). It is highly recommended to discuss any planned use of the data with National Oceanic and Atmospheric Administration-Chemical Sciences Division (NOAA-CSD) scientists. For more information, refer to the Readme file: "noaa-esrl-wascolidar-readme-1.pdf."

17 WIND ENERGY↗

Lidar - ESRL WindCube 200s, Arlington Airport - Reviewed Data

The available "readme" file introduces the basics of the Doppler lidar data and offers a detailed description of the variables present in the data files. For those with any further questions about the data and its interpretation, contact either Alan Brewer ( ) or Sunil Baidar ( ). It is highly recommended to discuss any planned use of the data with National Oceanic and Atmospheric Administration-Chemical Sciences Division (NOAA-CSD) scientists. For more information, refer to the Readme file: "noaa-esrl-arlingtonlidar-readme-1.pdf."

17 WIND ENERGY↗

Lidar - ESRL WindCube 200s, Wasco Airport - Processed Data

The available "Readme" file introduces the basics of the Doppler lidar data and offers a detailed description of the variables present in the data files. For those with any further questions about the data and its interpretation, contact either Alan Brewer ( ) or Sunil Baidar ( ). It is highly recommended to discuss any planned use of the data with National Oceanic and Atmospheric Administration-Chemical Sciences Division (NOAA-CSD) scientists. For more information, refer to the Readme file: "noaa-esrl-wascolidar-readme.docx."

17 WIND ENERGY↗

Lidar - ESRL WindCube 200s, Arlington Airport - Processed Data

The available "readme" file introduces the basics of the Doppler lidar data and offers a detailed description of the variables present in the data files. If you have any further questions about the data and its interpretation, contact either Alan Brewer ( ) or Sunil Baidar ( ). It is highly recommended to discuss any planned use of the data with NOAA-CSD scientists. For more information, refer to the attached readme.

17 WIND ENERGY↗

The updated ITPA global H-mode confinement database: description and analysis

The multi-machine ITPA Global H-mode Confinement Database has been upgraded with new data from JET with the ITER-like wall and ASDEX Upgrade with the full tungsten wall. This paper describes the new database and presents results of regression analysis to estimate the global energy confinement scaling in H-mode plasmas using a standard power law. Various subsets of the database are considered, focusing on type of wall and divertor materials, confinement regime (all H-modes, ELMy H or ELM-free) and ITER-like constraints. Apart from ordinary least squares, two other, robust regression techniques are applied, which take into account uncertainty on all variables. Regression on data from individual devices shows that, generally, the confinement dependence on density and the power degradation are weakest in the fully metallic devices. Using the multi-machine scalings, predictions are made of the confinement time in a standard ELMy H-mode scenario in ITER. The uncertainty on the scaling parameters is discussed with a view to practically useful error bars on the parameters and predictions. One of the derived scalings for ELMy H-modes on an ITER-like subset is studied in particular and compared to the IPB98(y,2) confinement scaling in engineering and dimensionless form. Transformation of this new scaling from engineering variables to dimensionless quantities is shown to result in large error bars on the dimensionless scaling. Regression analysis in the space of dimensionless variables is therefore proposed as an alternative, yielding acceptable estimates for the dimensionless scaling. The new scaling, which is dimensionally correct within the uncertainties, suggests that some dependencies of confinement in the multi- machine database can be reconciled with parameter scans in individual devices. This includes vanishingly small dependence of confinement on line-averaged density and normalized plasma pressure (β), as well as a noticeable, positive dependence on effective atomic mass and plasma triangularity. Extrapolation of this scaling to ITER yields a somewhat lower confinement time compared to the IPB98(y, 2) prediction, possibly related to the considerably weaker dependence on major radius in the new scaling (slightly above linear). Further studies are needed to compare more flexible regression models with the power law used here. In addition, data from more devices concerning possible ‘hidden variables’ could help to determine their influence on confinement, while adding data in sparsely populated areas of the parameter space may contribute to further disentangling some of the global confinement dependencies in tokamak plasmas.

Database↗

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Data-driven Community-centered Resilient Assessment and Planning Toolkit for Nexus of Energy and Water (DCRAPT-NEW)

Urban areas, including Detroit and Pittsburgh, have suffered significant dual outages of the electrical and water infrastructure in the past decade due, in part, to the increasing number of extreme weather events. With increasing temperatures and rainfall intensity, these regions need to prepare for increasing extreme events through community-based energy and water resilience analysis, planning, and enhancement. This project developed a suite of open-source, open-access, community-centered, data-driven assessment and distributed energy resource (DER) and planning tools for energy and water resilience enhancement in urban areas. Through establishing a multi-level community awareness and engagement mechanism and a comprehensive collection of power outage and flooding data, an innovative group of community energy and water resilience assessment and planning tools have been developed for a wide range of users with differing and variable sets of data available to them. The developed tools include (1) DOE EAGLE-I data-driven, deep-learning assisted resilience assessment and DER planning tools at the county level with socioeconomic factors incorporated; (2) Utility annual power outage data-driven tools for long term resilience assessment and DER planning and 15-min power outage data-driven tools for short term resilience assessment and planning; (3) Detailed engineering tools for energy and water systems resilience assessment and planning when the system topology and component fragility curves are available; (4) Alternative Resiliency Metric Calculation that extracts and separates outage and restoration processes; and (5) Co-optimization tools that evaluate the resilience of the power and sewage system and allow users to conduct joint planning with energy and wastewater systems. The developed tools provide planners, decision-makers, and stakeholders with powerful capabilities to systematically evaluate system/community resilience and optimal and actionable guidance for enhancing resilience while prioritizing DER investments. The tools have been used and validated in Detroit and Pittsburgh and can be used in other areas of the nation. In addition, this project will (1) advance the knowledge and applications of machine-learning methods in analyzing and fusing different layers of information and generating meaningful data points such as generating rare weather events; (2) significantly improve the energy and water resilience of the identified communities in Detroit and Pittsburgh and prepare for more frequent and severe weather conditions; (3) help communities assess extreme weather event impacts and address short-term and long-term resilience-related issues The developed tools have been made public via GitHub and demonstrated to community stakeholders and utility companies via the two annual workshops and numerous community engagement meetings. The project outcomes are also disseminated through publications in various journals and conference proceedings, and presentations at top conferences.

13 HYDRO ENERGY↗

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder↗

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

Unsupervised probabilistic models for sequential Electronic Health Records

We develop an unsupervised probabilistic model for heterogeneous Electronic Health Record (EHR) data. Utilizing a mixture model formulation, our approach directly models sequences of arbitrary length, such as medications and laboratory results. This allows for subgrouping and incorporation of the dynamics underlying heterogeneous data types. The model consists of a layered set of latent variables that encode underlying structure in the data. These variables represent subject subgroups at the top layer, and unobserved states for sequences in the second layer. We train this model on episodic data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The resulting properties of the trained model generate novel insight from these complex and multifaceted data. In addition, we show how the model can be used to analyze sequences that contribute to assessment of mortality likelihood.

59 BASIC BIOLOGICAL SCIENCES↗

Using mixed methods to construct and analyze a participatory agent-based model of a complex Zimbabwean agro-pastoral system

Complex social-ecological systems can be difficult to study and manage. Simulation models can facilitate exploration of system behavior under novel conditions, and participatory modeling can involve stakeholders in developing appropriate management processes. Participatory modeling already typically involves qualitative structural validation of models with stakeholders, but with increased data and more sophisticated models, quantitative behavioral validation may be possible as well. In this study, we created a novel agent-basedmodel applied to a specific context: Zimbabwean non-governmental organization the Muonde Trust has been collecting data on their agro-pastoral system for the last 35 years and had concerns about land-use planning and the effectiveness of management interventions in the face of climate change. We collaboratively created an agent-based model of their system using their data archive, qualitatively calibrating it to the observed behavior of the real system without tuning any parameters to match specific quantitative outputs. We then behaviorally validated the model using quantitative community-based data and conducted a sensitivity analysis to determine the relative impact of underlying parameter assumptions, Indigenous management interventions, and different rainfall variation scenarios. We found that our process resulted in a model which was successfully structurally validated and sufficiently realistic to be useful for Muonde researchers as a discussion tool. The model was inconsistently behaviorally validated, however, with some model variables matching field data better than others. We observed increased model system instability due to increasing variability in underlying drivers (rainfall), and also due to management interventions that broke feedbacks between the components of the system. Interventions that smoothed year-to-year variation rather than exaggerating it tended to improve sustainability. The Muonde trust has used the model to successfully advocate to local leaders for changes in land-use planning policy that will increase the sustainability of their system.

54 ENVIRONMENTAL SCIENCES↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Investigation of Multiple Data Streams for Gearbox Bearing Fault Prediction Through Machine-Learning Models

Operations and maintenance (O&M) cost of wind plant accounts up to 30% of total energy cost, which can be reduced through continuous monitoring and successfully detecting incipient wind turbine failures. To accomplish this, condition monitoring and predictive maintenance systems are being implemented in wind industry to support O&M decision making. A wide range of approaches for condition monitoring and fault prediction have been developed. These approaches generally use historical data of wind turbines collected by Supervisory Control and Data Acquisition (SCADA) system to identify patterns that lead to failure. These SCADA data show the overall condition of a wind turbine and can be leveraged to detect when the turbine's performance is degrading and to identify if a fault is developing. However, it becomes challenging to predict the failure of a specific wind turbine gearbox bearing, because the SCADA data are often not directly linked to the component. To bridge the gap, we have investigated features calculated from SCADA data using physics-based models and the gearbox design over the years. The damaged metric we used in the physics domain is frictional energy. Combining these physics domain variables with SCADA data as inputs to various machine learning models for gearbox bearing fault prediction, we have demonstrated the benefits of leveraging both physics and data domain models. It was an attempt to improve frictional-energy-based damage metric by adding data domain inputs, as we had learned that the frictional-energy-based damage metric alone is not sufficient to single out failed bearings from healthy. As condition monitoring data (either vibration or oil debris data) has become available at more and more wind plants, we would like to evaluate whether by adding the condition monitoring data can help further improve the performance of frictional-energy-based damage metric for gearbox bearing fault prediction. Both cases by modeling through various machine learning algorithms are discussed in this study along with some observations.

fault prediction↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

GPS-Based Gamma Survey for Characterizing and Decommissioning NORM Sites - 20389

Gamma survey techniques are an especially powerful decommissioning tool at naturally occurring radioactive material (NORM) sites due to both the low cost to obtain data over a large spatial scale and the abundance of gamma emitters in the uranium and thorium decay series. Gamma surveys are executed by coupling a detector - most often a sodium iodide crystal - to a global positioning system (GPS), then reporting a location and gross gamma reading coincidentally to a data logger. Systems may be carried by workers or mounted to a car, all-terrain vehicle, or unmanned aerial system (UAS). The resulting data set provides a high-resolution but low precision map of the gamma radiation field over the area surveyed. Frequently this map is also correlated to soil concentrations of NORM radionuclides (most often, Ra-226) and/or exposure rate. Gamma survey parameters such as movement speed, transect spacing, and data logging frequency define the spatial resolution of the resulting surface, and can be optimized depending on the desired survey sensitivity. This paper examines gamma survey as a tool for decommissioning NORM sites and provides an overview of current gamma survey technology designed to improve the efficiency and effectiveness of the decommissioning process. Topics to be discussed in the paper include: - An overview of gamma survey systems, and the utility of different delivery vehicles depending on desired cost, desired spatial resolution, and site topography. - The influence of physical detector characteristics on detection sensitivity and survey planning. - The tradeoff between high-resolution and large spatial extent, but inherently uncertain data, and low-resolution, low spatial extent, but highly certain data, as well as the specific utility of each of these types of data during NORM facility decommissioning. - Confounding variables that may limit the utility of gamma survey at some sites (e.g., radon gas and spatial heterogeneity / hot spots), and methods to plan for and control these conditions. Results show that the confounding variables, such as radon and data output can greatly influence the overall data quality associated with the decommissioning process. In addition, the use of real-time and aerial survey platforms provides a method for ensuring proper spatial extent of the data. When applied thoughtfully, gamma survey is a powerful tool for detecting NORM radionuclides in the environment and a cost-effective technique for identifying areas requiring remediation. However, entities performing or using gamma survey as a decommissioning tool must be aware of both its advantages and its limitations before basing remediation or regulatory action on gamma survey results. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗