Data Curation for Basin-Scale Modeling in NRAP Phase III
National Risk Assessment Partnership Annual Technical Meeting, Virtual, May 3-4, 2022
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
National Risk Assessment Partnership Annual Technical Meeting, Virtual, May 3-4, 2022
For presentation at the Carbon Management Project Review Meeting, Pittsburgh, PA, August 28--September 1, 2023.
The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.
In this presentation, we will explore the Rucio experience with the Rubin Observatory experiment. Our discussion will cover several key areas: Scalability Tests: Insights into the performance and scalability evaluations of Rucio in the context of Rubin's data needs and what we have learned, especially with many small files. Role in Rubin's Data Curation: Rubin's Data Butler: An overview of how Rucio, along with with Rubin's Data Butler using Hermes-K, which involves message passing through Kafka, is integrated in the Rubin's data curation system. Monitoring and Support: Current status of Rucio and PostgreSQL monitoring and Rucio deployment and support within the Rubin environment. Tape RSE Implementation: Deal with the order of magnitude more files going to tape than HEP. Future Needs: An examination of Rubin's evolving requirements for Rucio services and how we plan to address them.
Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.
Title: Multidisciplinary Geotechnical Data Collection, Curation, and Analysis for Conformity with the Regulatory Framework for Geologic Carbon Storage in Wyoming, USA. Text: Construction and operation of wells for geologic sequestration of carbon dioxide necessitate that they are permitted under the Environmental Protection Agency’s Underground Injection Control Class VI requirements. Class VI wells conform to stringent requirements to ensure long-term safety and integrity of the storage site and the protection of Underground Sources of Drinking Water. Entities pursuing Class VI permitting must provide comprehensive geologic site characterization, including regional geologic structure and stratigraphy, aquifer information, reservoir and confining unit geomechanical properties, geochemical analyses, assessment of trapping capacity and mechanisms, and a variety of other of multidisciplinary geotechnical data. The Wyoming Class VI Site Characterization Database Project is focused on developing a geologic site characterization database of geotechnical information, which has been compiled and verified from established, public databases/entities and scientific literature to expedite Class VI permitting in Sweetwater County within the Greater Green River Basin of southern Wyoming. The preliminary suite of compiled data from 14,000 wells includes 8,000 wells with logs and 7,250 wells with formation tops, ~70 wells with core data (e.g., X-Ray diffraction, petrographic, and petrophysical data), ~2,500 water analyses, ~740 seismic events data, and ~520 bottom-hole temperature measurements. Future work on—and stemming from—this project will include new core analyses, calculation and interpolation of subsurface temperature gradients, mechanical earth models, geochemical simulations, storage capacity estimation, stratigraphic column generation and correlation, and construction of subsurface maps. Finally, this work will help to inspire and facilitate subsurface data compilation and curation beyond Sweetwater County, Wyoming.
This paper introduces a data-driven framework for combustor-focused, performance-based condition monitoring of gas turbines. Commercial condition monitoring systems typically generate huge amounts of data that make efficient onboard monitoring challenging. This paper focuses on quantifying combustor component degradation, using premixer centerbody degradation in a swirl stabilized combustor as a case study. The input for these analyses is acoustic pressure measurements acquired at various locations on the combustor. The diagnosis methodology is based on a classification framework and consists of 3 steps: 1) Data curation, 2) Feature Engineering, and 3) Diagnosis. Data curation ensures good quality of the data that is passed through the algorithm. Feature engineering deals with the extraction of the most informative features, from the most informative sensors, that can accurately capture the introduced fault. To perform diagnosis, the classification model is trained using experimentally acquired data and is then tested on a separate data set. The framework was able to achieve high classification accuracy (>99%) for training size as low as 30% of the total recorded observations. The low number of features required to achieve this accuracy suggests high potential for integration into existing onboard condition monitoring systems.
Designing alternative fuels for advanced compression ignition engines necessitates a predictive model for cetane number (CN). In this study, the physics-informed graph neural networks are introduced for a reliable CN prediction by considering molecular features pertinent to the physical properties of molecules that affect CN. The reliability of measured data is another key factor to consider for improving the predictive model. Various experimental instruments for measuring CN exist, including standard and non-standard methods. In this regard, a systematic data quality analysis was carried out for the total 630 CNs collected from literature and new measurements in this study using Advanced Fuel Ignition Delay Analyzer (AFIDA). The results from this data curation process were reflected in the model by imposing lower sample weights on the data coming from less reliable measurement techniques. This approach effectively maximized the prediction accuracy while incorporating data from all available sources. Using the sample weights decreased the mean absolute error (MAE) up to 0.8 CN units. The accuracy was also improved by introducing the CN-related physical properties (the number of hydrogen bond donors and acceptors); the test set MAE is 5.74 and 7.01 for the model with and without such properties, respectively. Investigating molecular structural effects on CN was also carried out to gain chemical insights into factors used to design new fuel candidates. The dimensionality reduction analysis of feature vectors showed a clear clustering in terms of functional groups and CN and the structural effect derived from the model was consistent with the physicochemical insights. Finally, this physics-informed model and data curation would be helpful for accurate CN prediction and inform rational fuel design.
For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.
For presentation at DOE Data Days, Livermore, CA, October 24-26, 2023.
Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.
Multimodal, curated data sets and nuisance event catalogs remain rare in the explosion monitoring community relative to curated seismic data sets. The source of this relative absence is the difficultly in deploying multimodal receivers that sense the seismic, acoustic, and other modalities from multiphysics sources. We provide such a data set in this study that delivers seismic, infrasound, and electromagnetic (magnetometer) sensor records collected over a two–week period, within 255 km of a 10 ton buried chemical explosion called DAG–4 that was located at 37.1146°, –116.0693° on 22 June 2019 21:06:19.88 UTC. This catalog includes 485 seismic, seismoacoustic, and infrasound–only events that an expert analyst manually built by reviewing waveforms from 29 seismic and infrasound sensors. Our data release includes waveforms from these 29 seismic, infrasound, and seismoacoustic stations and two magnetometer stations and their station metadata. We deliver these waveforms in NNSA KB Core CSS.w format (i4) with a corresponding wfdisc table that provides the header information. Here, we expect that this data set will provide a valuable, benchmark resource to develop signal processing algorithms and explosion monitoring methods against manual, human observations.
The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.
The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.
The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented or is currently implementing data standards and automated data pipelines for the following geothermal data types: 1) drilling data, 2) geospatial datasets, and 3) Distributed Acoustic Sensing (DAS) data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how the GDR team can improve this process.
The transformative impact of modern computational paradigms and technologies, such as high-performance computing, quantum computing, and cloud computing, has opened up profound new opportunities for scientific simulations. Scalable computational chemistry is one beneficiary of this technological progress. The main focus of this paper is on the performance of various quantum chemical formulations, ranging from low-order methods to high-accuracy approaches, implemented in different computational chemistry packages, such as NWChem, NWChemEx, SPEC, ExaChem, and FLOSIC codes on the Azure Quantum Element (AQE) Microsoft cloud services. We pay particular attention to the intricate workflows for performing composite chemistry simulations, associated data curation, and mechanisms for accuracy assessment, as defined by the enabling cloud Computational Chemistry as a Service (CCaaS). Our focus also extends to Arrows' automated workflow for high throughput simulations. Finally, we provide a perspective on the role of cloud computing in supporting the mission of leadership computational facilities (LCFs).
This presentation emphasizes the critical importance of data-centric AI. The limitations of model-centric AI when dealing with poor or insufficient data are highlighted, and it is illustrated how training models on inaccurate or noisy data leads to suboptimal results. This talk advocates for a hybrid approach that combines a focus on data quality and model parameters to achieve optimal results. The Open Energy Data Initiative (OEDI) is introduced as a valuable resource for obtaining high-quality energy-related datasets, hosting nearly 2,000 publicly accessible datasets, including 99 solar-related datasets, totaling over 2.7 petabytes of data. OEDI's data lakes enable users to query and work with data without extensive transfers. In conclusion, the significance of data-centric AI and adherence to data curation best practices is emphasized, positioning OEDI as a prime source of high-quality data for AI and machine learning in the renewable energy sector.
Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine