Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hidden (latent) signatures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗