Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “geothermal machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Modeling Subsurface Performance of a Geothermal Reservoir Using Machine Learning

Geothermal power plants typically show decreasing heat and power production rates over time. Mitigation strategies include optimizing the management of existing wells—increasing or decreasing the fluid flow rates across the wells—and drilling new wells at appropriate locations. The latter is expensive, time-consuming, and subject to many engineering constraints, but the former is a viable mechanism for periodic adjustment of the available fluid allocations. In this study, we describe a new approach combining reservoir modeling and machine learning to produce models that enable such a strategy. Our computational approach allows us, first, to translate sets of potential flow rates for the active wells into reservoir-wide estimates of produced energy, and second, to find optimal flow allocations among the studied sets. In our computational experiments, we utilize collections of simulations for a specific reservoir (which capture subsurface characterization and realize history matching) along with machine learning models that predict temperature and pressure timeseries for production wells. We evaluate this approach using an “open-source” reservoir we have constructed that captures many of the characteristics of Brady Hot Springs, a commercially operational geothermal field in Nevada, USA. Selected results from a reservoir model of Brady Hot Springs itself are presented to show successful application to an existing system. In both cases, energy predictions prove to be highly accurate: all observed prediction errors do not exceed 3.68% for temperatures and 4.75% for pressures. In a cumulative energy estimation, we observe prediction errors that are less than 4.04%. A typical reservoir simulation for Brady Hot Springs completes in approximately 4 h, whereas our machine learning models yield accurate 20-year predictions for temperatures, pressures, and produced energy in 0.9 s. This paper aims to demonstrate how the models and techniques from our study can be applied to achieve rapid exploration of controlled parameters and optimization of other geothermal reservoirs.

15 GEOTHERMAL ENERGY↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

FlowDash Geothermal Energy Enhancer: Where is Next Geothermal Resource? Machine Learning + Multiple Datasets => Geothermal Exploration Indication?

This is the presentation delivered at the 2025 GEODE Datathon competition. GEODE is a consortium of experts that addresses technology and knowledge gaps in geothermal energy, leveraging technology and best practices from the oil and gas industry. NETL team was awarded the 1st place in the engineering track. 2025 GEODE Datathon had a total of 42 teams from top universities and several major industrial companies. This awarded work is founded on a robust idea and innovative approach that uses machine learning coupled to multiple datasets to visualize geothermal “sweet” spots/indications in Great Basin based on the data provided from the GEODE Datathon. The use case also leveraged other datasets and demonstrated insightful and valuable indications for geothermal exploration.

Geothermal energy, Machine learning, Multiple Data↗

Potential structures - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains shapefiles, geotiffs, and symbology for the revised-from-Play-Fairway potential structures/structural settings used in the Nevada Geothermal Machine Learning project. Layers include potential structural setting ellipses, centroids, and distance-to-centroid raster. A submission linking the full GitHub repository for our machine learning Jupyter Notebooks will appear in the related datasets section of this page once available.

15 GEOTHERMAL ENERGY↗

Machine Learning Model Geotiffs - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains geotiffs, supporting shapefiles and readmes for the inputs and output models of algorithms explored in the Nevada Geothermal Machine Learning project, meant to accompany the final report. Layers include: Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk), input rasters of feature sets, and positive/negative training sites. See readme .txt files and final report for additional metadata. A submission linking the full codebase for generating machine learning output models is available under "related resources" on this page.

15 GEOTHERMAL ENERGY↗

Geothermal Operational Optimization with Machine Learning

The Geothermal Operational Optimization with Machine Learning (GOOML) project has developed a generic and extensible component-based system modeling framework to study complex geothermal fields using a data-driven approach. Through building a digital twin of a geothermal steam field with the GOOML modeling framework, operators can analyze historical and forecasted power production, explore possible steam field configurations, and optimize real world operations, all in a cost-effective digital environment. The GOOML modeling software is based on a historical data-assimilation framework that uses first-principal thermodynamics to model steam field components using historical data, and a forecast framework that uses machine-learning-driven models of steam field components to predict future operations. This modeling framework creates countless new opportunities for digital exploration of steam field design and operations. To date, digital twins have been developed for several steam fields in New Zealand and the United States. These digital twins have been validated by comparing hindcast predictions against historical production data. Field design and operations have been explored using genetic optimization and reinforcement learning. Initial results show compelling and often surprising opportunities for improved design and operation of fields with 2 to 5 percent improvements in annual energy production. GOOML is driving a step-change in geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and a first-of-its-kind intelligent geothermal systems model.

40 EE - Geothermal Technologies Office (EE-4G)↗

GOOML - Real World Applications of Machine Learning in Geothermal Operations

GOOML (Geothermal Operational Optimization with Machine Learning) is a machine-learning based framework that enables geothermal power plant operators to explore optimization opportunities for their assets in an efficient and robust digital environment. Backed by real-world data sources, thermodynamic constraints and steamfield intelligence, the GOOML environment provides new tools to explore how to best operate steamfields as well as test new scenarios and configurations prior to implementation in the field. To prove the effectiveness of GOOML, we have undertaken optimization experiments using reinforcement learning (RL) to generate operational suggestions using a balance of mass-take targets, sustainability considerations and net generation. Our experiments use the GOOML construct to explore different field parameters and perform multiple reinforcement learning experiments. Like a comprehensive laboratory workbench, we can change out components of a steamfield to perform testing under a variety of conditions (restrict mass, increase pressure, reroute steam, etc.). This flexibility allows us to explore conditions that would require significant infrastructure changes in a real-world setting at a fraction of the cost and time in a digital environment. The results highlight the benefits of using digital twins and advanced data analytics for the geothermal industry.

forecasting↗

A New Modeling Framework for Geothermal Operational Optimization with Machine Learning (GOOML)

Geothermal power plants are excellent resources for providing low carbon electricity generation with high reliability. However, many geothermal power plants could realize significant improvements in operational efficiency from the application of improved modeling software. Increased integration of digital twins into geothermal operations will not only enable engineers to better understand the complex interplay of components in larger systems but will also enable enhanced exploration of the operational space with the recent advances in artificial intelligence (AI) and machine learning (ML) tools. Such innovations in geothermal operational analysis have been deterred by several challenges, most notably, the challenge in applying idealized thermodynamic models to imperfect as-built systems with constant degradation of nominal performance. This paper presents GOOML: a new framework for Geothermal Operational Optimization with Machine Learning. By taking a hybrid data-driven thermodynamics approach, GOOML is able to accurately model the real-world performance characteristics of as-built geothermal systems. Further, GOOML can be readily integrated into the larger AI and ML ecosystem for true state-of-the-art optimization. This modeling framework has already been applied to several geothermal power plants and has provided reasonably accurate results in all cases. Therefore, we expect that the GOOML framework can be applied to any geothermal power plant around the world.

15 GEOTHERMAL ENERGY↗

Machine Learning for Geothermal Resource Exploration in the Tularosa Basin, New Mexico

Geothermal energy is considered an essential renewable resource to generate flexible electricity. Geothermal resource assessments conducted by the U.S. Geological Survey showed that the southwestern basins in the U.S. have a significant geothermal potential for meeting domestic electricity demand. Within these southwestern basins, play fairway analysis (PFA), funded by the U.S. Department of Energy’s (DOE) Geothermal Technologies Office, identified that the Tularosa Basin in New Mexico has significant geothermal potential. This short communication paper presents a machine learning (ML) methodology for curating and analyzing the PFA data from the DOE’s geothermal data repository. The proposed approach to identify potential geothermal sites in the Tularosa Basin is based on an unsupervised ML method called non-negative matrix factorization with custom k-means clustering. This methodology is available in our open-source ML framework, GeoThermalCloud (GTC). Using this GTC framework, we discover prospective geothermal locations and find key parameters defining these prospects. Our ML analysis found that these prospects are consistent with the existing Tularosa Basin’s PFA studies. This instills confidence in our GTC framework to accelerate geothermal exploration and resource development, which is generally time-consuming.

15 GEOTHERMAL ENERGY↗

Python Codebase and Jupyter Notebooks - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

Git archive containing Python modules and resources used to generate machine-learning models used in the "Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada" project. This software is licensed as free to use, modify, and distribute with attribution. Full license details are included within the archive. See "documentation.zip" for setup instructions and file trees annotated with module descriptions.

Brown, Stephen↗

GeoThermalCloud: Machine Learning for Geothermal Resource Exploration

Geothermal is a renewable energy source that can provide reliable and flexible electricity generation for the world. In the past decade, the U.S. Geological Survey's resource assessments, Play Fairway Analyses (PFA), and GeoVision report by the U.S. Department of Energy's Geothermal Technologies Office provided insights on enormous untapped potential for geothermal energy to contribute to the U.S. domestic energy needs. The past studies identified that geothermal resources without surface expression (e.g., blind/hidden hydrothermal systems) comprise a huge potential. These blind systems can significantly increase power generation. But a primary challenge is locating and quantifying these hidden resources, which do not have any thermal manifestations on the surface. PFA has successfully identified some blind systems in the western USA (e.g., specific locations in the Great Basin region within Nevada). However, a comprehensive search for these blind systems can be time-consuming, expensive, and resource-intensive with a low probability of success. Accelerated discovery of these blind resources is needed with growing energy needs and higher chances of exploration success. Recent advances in machine learning (ML) have shown promise in shortening the timeline for this discovery. This paper presents a novel ML-based methodology for geothermal exploration towards PFA applications. Our methodology is provided through our open-source ML framework called GeoThermalCloud \url{https://github.com/SmartTensors/GeoThermalCloud.jl}. GeoThermalCloud uses a series of unsupervised, supervised, and physics-informed ML methods available in SmartTensors AI platform \url{https://github.com/SmartTensors}. Here, the presented analyses are performed using our unsupervised ML algorithm called NMF$k$, which is available in the SmartTensors AI platform. Our ML algorithm facilitates the discovery of new phenomena, hidden patterns, and mechanisms that helps us to make informed decisions. Moreover, the GeoThermalCloud enhances the collected PFA data and discovers signatures representative of geothermal resources. Through GeoThermalCloud, we were able to identify hidden patterns in the geothermal field data needed for the efficient discovery of blind systems. Crucial geothermal signatures often overlooked in traditional PFA are extracted using GeoThermalCloud and analyzed by the subject matter experts to provide ML-enhanced PFA, which is informative for efficient exploration. We applied our ML methodology on various open-source geothermal datasets within the U.S. (some of these are collected by past PFA work), and the results provide valuable insights on resource types within those explored regions. This ML-enhanced workflow makes GeoThermalCloud attractive for the geothermal community to improve existing datasets and extract valuable information often unnoticed during geothermal exploration.

machine learning (ML), geothermal energy↗

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗

GOOML (Geothermal Operational Optimization with Machine Learning) [SWR-23-01]

The Geothermal Operational Optimization with Machine Learning (GOOML) is a partnership between NREL and Upflow, NZ, awarded in response to the U.S. Department of Energy's Geothermal Technologies Office's Funding Opportunity Announcement (FOA) to expand the role of advanced analytics and automation in geothermal operations through machine learning. Partnering with industry (Contact Energy Limited ("Contact"), Ngati Tuwharetoa Geothermal Assets Limited ("NTGA"), Ormat Technologies Inc. ("Ormat") and Flow State Solutions Limited ("FSS"), GOOML was created to improve the operational efficiency of geothermal power plant steam fields through the analysis of historical operational data and the application of custom machine learning algorithms. NREL's contributions include machine learning, coding, and data management expertise as well as access to high-performance compute solutions. GOOML can increase geothermal operational efficiency through development of a digital system twin that can be utilized to provide optimal geothermal operating conditions for real-world geothermal fields. GOOML allows users to analyze field production histories in detail, develop models, and train machine learning algorithms to identify opportunities for increased geothermal efficiency, detect potential trouble, and allow predictive scenario modeling. Preliminary experiments have demonstrated a potential to increase total generation by as much as 12% through ML optimization of the utilization of existing steam field resources.

Buster, Grant↗