Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Development of Gamma Background Radiation Digital Twin with Machine Learning Algorithms: Application of Unsupervised Machine Learning to Detection of Anomalies and Nuisances in Gamma Background Radiation Environmental Screening Data

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for development of a digital twin of gamma radiation background, and for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. In one segment of work, we developed a gamma background estimation model using a Longshort term memory (LSTM) network for one-step CPS time series prediction. The LSTM model was validated with two data sets of measurements from two independent NaI detectors positioned on a mobile platform. The data sets contained background radiation only and no orphan isotope sources. The LSTM model was constructed and tested using data from one of the detectors. Performance of the LSTM model was validate through one-step prediction of CPS time series of another NaI detector without re-training. This approach allows to create a digital twin for nuclear background estimation. Using LSTM, it could be possible to detect a source through subtraction of the estimated counts from the measured background. In another segment of work, we investigated detection of gamma emitting sources in the presence of complex background using unsupervised machine learning. Spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

54 ENVIRONMENTAL SCIENCES↗

Beyond the hubble sequence – exploring galaxy morphology with unsupervised machine learning

We explore unsupervised machine learning for galaxy morphology analyses using a combination of feature extraction with a vector-quantized variational autoencoder (VQ-VAE) and hierarchical clustering (HC). We propose a new methodology that includes: (1) consideration of the clustering performance simultaneously when learning features from images; (2) allowing for various distance thresholds within the HC algorithm; (3) using the galaxy orientation to determine the number of clusters. This set-up provides 27 clusters created with this unsupervised learning that we show are well separated based on galaxy shape and structure (e.g. Sérsic index, concentration, asymmetry, Gini coefficient). These resulting clusters also correlate well with physical properties such as the colour–magnitude diagram, and span the range of scaling relations such as mass versus size amongst the different machine-defined clusters. When we merge these multiple clusters into two large preliminary clusters to provide a binary classification, an accuracy of $\sim 87{{\ \rm per\ cent}}$ is reached using an imbalanced data set, matching real galaxy distributions, which includes 22.7 per cent early-type galaxies and 77.3 per cent late-type galaxies. Comparing the given clusters with classic Hubble types (ellipticals, lenticulars, early spirals, late spirals, and irregulars), we show that there is an intrinsic vagueness in visual classification systems, in particular galaxies with transitional features such as lenticulars and early spirals. Based on this, the main result in this work is not how well our unsupervised method matches visual classifications and physical properties, but that the method provides an independent classification that may be more physically meaningful than any visually based ones.

79 ASTRONOMY AND ASTROPHYSICS↗

Unveiling and Mapping Polymorphs in Fluorite Y2TiO5 Using 4D-STEM and Unsupervised Machine Learning

Y2TiO5 belongs to the Ln2TiO5 (Ln = lanthanide or Y) family of ceramic materials and exhibits a range of desirable material properties such as radiation tolerance, frustrated magnetism, and large dielectric constant. However, understanding the complex crystal structure of Y2TiO5 remains elusive, given that Y2TiO5 can adopt multiple polymorphs such as cubic, orthorhombic, and hexagonal phases within the lattice. In this work, we report a detailed structural analysis of Y2TiO5 using four-dimensional scanning transmission electron microscopy coupled with unsupervised machine learning. The pyrochlore nanodomains, characterized by the ordered arrangement of yttrium cations on the A site of their A2BO5 structure, are present within the matrix of a predominantly fluorite-structured Y2TiO5 along with a third polymorph, the hexagonal phase. The pyrochlore phase is found to form 2 nm boundary regions around hexagonal phase stacking faults, highlighting the potential influence of the hexagonal phase on the occurrence and distribution of the pyrochlore phase. Lastly, we identify a unique pyrochlore phase with asymmetric arrangement of cation ordering along a single planar direction. Our findings provide invaluable insights into the possible mechanisms stabilizing pyrochlore nanodomains within the fluorite lattice of Y2TiO5.

36 MATERIALS SCIENCE↗

Uncovering electronic and geometric descriptors of chemical activity for metal alloys and oxides using unsupervised machine learning

Here, we show that unsupervised machine learning (ML) using principal component analysis (PCA) provides a straightforward pathway for developing accurate and interpretable electronic-structure descriptors of the chemical and catalytic properties of materials. We demonstrate the approach by finding chemisorption descriptors for metal alloys and surface oxygens on metals and metal oxides. In both cases, the principal component (PC) descriptors yield ML models that predict the material’s chemical properties with competitive accuracy compared to ML models built using established descriptors. Importantly, interpreting the electronic-structure patterns captured by each PC descriptor via signal reconstruction suggests potential design motifs for future electronic-structure descriptor design and allows us to identify links between a material’s geometric and catalytic properties. Ultimately, we show that the unsupervised ML approach provides a route to find electronic-structure descriptors of the catalytic properties of materials that readily connect to geometric structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring variability in seasonal average and extreme precipitation using unsupervised machine learning.

Focal Area(s): We will use unsupervised machine learning methods to identify and quantify the influence of large scale natural modes of climate variability to gain insight into the observed and simulated seasonal average and extreme precipitation changes. Science Challenge: A recent paper, led by co-PI Mark Risser, finds that although much of the variability in seasonal average and extreme precipitation over CONUS is unforced, the effect of large-scale modes of circulation variability (such as ENSO, AMO, PNA, etc.) can be detected and attributed. However, it is unclear whether or not unsupervised learning methods can (a) replicate this finding or (b) yield insight into possible nonlinear behavior that was not captured in the initial statistical analysis. Further work would entail extending this framework to other global land areas.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised machine learning for unbiased chemical classification in X-ray absorption spectroscopy and X-ray emission spectroscopy

Here we report a comprehensive computational study of unsupervised machine learning for extraction of chemically relevant information in X-ray absorption near edge structure (XANES) and in valence-to-core X-ray emission spectra (VtC-XES) for classification of a broad ensemble of sulphorganic molecules. By progressively decreasing the constraining assumptions of the unsupervised machine learning algorithm, moving from principal component analysis (PCA) to a variational autoencoder (VAE) to t-distributed stochastic neighbour embedding (t-SNE), we find improved sensitivity to steadily more refined chemical information. Surprisingly, when embedding the ensemble of spectra in merely two dimensions, t-SNE distinguishes not just oxidation state and general sulphur bonding environment but also the aromaticity of the bonding radical group with 87% accuracy as well as identifying even finer details in electronic structure within aromatic or aliphatic sub-classes. We find that the chemical information in XANES and VtC-XES is very similar in character and content, although they unexpectedly have different sensitivity within a given molecular class. We also discuss likely benefits from further effort with unsupervised machine learning and from the interplay between supervised and unsupervised machine learning for X-ray spectroscopies. Our overall results, i.e., the ability to reliably classify without user bias and to discover unexpected chemical signatures for XANES and VtC-XES, likely generalize to other systems as well as to other one-dimensional chemical spectroscopies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discovery of peculiar radio morphologies with ASKAP using unsupervised machine learning

Abstract We present a set of peculiar radio sources detected using an unsupervised machine learning method. We use data from the Australian Square Kilometre Array Pathfinder (ASKAP) telescope to train a self-organizing map (SOM). The radio maps from three ASKAP surveys, Evolutionary Map of Universe pilot survey (EMU-PS), Deep Investigation of Neutral Gas Origins pilot survey (DINGO), and Survey With ASKAP of GAMA-09 + X-ray (SWAG-X), are used to search for the rarest or unknown radio morphologies. We use an extension of the SOM algorithm that implements rotation and flipping invariance on astronomical sources. The SOM is trained using the images of all ‘complex’ radio sources in the EMU-PS which we define as all sources catalogued as ‘multi-component’. The trained SOM is then used to estimate a similarity score for complex sources in all surveys. We select 0.5% of the sources that are most complex according to the similarity metric and visually examine them to find the rarest radio morphologies. Among these, we find two new odd radio circle (ORC) candidates and five other peculiar morphologies. We discuss multiwavelength properties and the optical/infrared counterparts of selected peculiar sources. In addition, we present examples of conventional radio morphologies including: diffuse emission from galaxy clusters, and resolved, bent-tailed, and FR-I and FR-II type radio galaxies. We discuss the overdense environment that may be the reason behind the circular shape of ORC candidates.

Astronomy & Astrophysics↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised machine learning discovery of structural units and transformation pathways from imaging data

We show that unsupervised machine learning can be used to learn chemical transformation pathways from observational Scanning Transmission Electron Microscopy (STEM) data. To enable this analysis, we assumed the existence of atoms, a discreteness of atomic classes, and the presence of an explicit relationship between the observed STEM contrast and the presence of atomic units. With only these postulates, we developed a machine learning method leveraging a rotationally invariant variational autoencoder (VAE) that can identify the existing molecular fragments observed within a material. The approach encodes the information contained in STEM image sequences using a small number of latent variables, allowing the exploration of chemical transformation pathways by tracing the evolution of atoms in the latent space of the system. The results suggest that atomically resolved STEM data can be used to derive fundamental physical and chemical mechanisms involved, by providing encodings of the observed structures that act as bottom-up equivalents of structural order parameters. The approach also demonstrates the potential of variational (i.e., Bayesian) methods in the physical sciences and will stimulate the development of more sophisticated ways to encode physical constraints in the encoder–decoder architectures and generative physical laws and causal relationships in the latent space of VAEs.

97 MATHEMATICS AND COMPUTING↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

AXEAP : a software package for X-ray emission data analysis using unsupervised machine learning

The Argonne X-ray Emission Analysis Package ( AXEAP ) has been developed to calibrate and process X-ray emission spectroscopy (XES) data collected with a two-dimensional (2D) position-sensitive detector. AXEAP is designed to convert a 2D XES image into an XES spectrum in real time using both calculations and unsupervised machine learning. AXEAP is capable of making this transformation at a rate similar to data collection, allowing real-time comparisons during data collection, reducing the amount of data stored from gigabyte-sized image files to kilobyte-sized text files. With a user-friendly interface, AXEAP includes data processing for non-resonant and resonant XES images from multiple edges and elements. AXEAP is written in MATLAB and can run on common operating systems, including Linux, Windows, and MacOS.

97 MATHEMATICS AND COMPUTING↗

Characterizing Drought Behavior in the Colorado River Basin Using Unsupervised Machine Learning

Drought is a pressing issue for the Colorado River Basin (CRB) due to the social and economic value of water resources in the region and the significant uncertainty of future drought under climate change. Here, we use climate simulations from various Earth System Models (ESMs) to force the Variable Infiltration Capacity hydrologic model and project multiple drought indicators for the sub-watersheds within the CRB. We apply an unsupervised machine learning (ML) based on Non-Negative Matrix Factorization using K-means clustering (NMFk) to synthesize the simulated historical, future, and change in drought indicators. The unsupervised ML approach can identify sub-watersheds where key changes to drought indicator behavior occur, including shifts in snowpack, snowmelt timing, precipitation, and evapotranspiration. While changes in future precipitation vary across ESMs, the results indicate that the Upper CRB will experience increasing evaporative demand and surface-water scarcity, with some locations experiencing a shift from a radiation-limited to a water-limited evaporation regime in the summer. Large shifts in peak runoff are observed in snowmelt-dominant sub-watersheds, with complete disappearance of the snowmelt signal for some sub-watersheds. The work demonstrates the utility of the NMFk algorithm to efficiently identify behavioral changes of drought indicators across space and time and to quickly analyze and interpret hydro climate model results.

54 ENVIRONMENTAL SCIENCES↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

Poster presented at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. This poster highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

This is the conference paper accompanying a poster presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. This work highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Uncertainty Quantification in CO2 Trapping Mechanisms: A Case Study of PUNQ-S3 Reservoir Model Using Representative Geological Realizations and Unsupervised Machine Learning

Evaluating uncertainty in CO2 injection projections often requires numerous high-resolution geological realizations (GRs) which, although effective, are computationally demanding. This study proposes the use of representative geological realizations (RGRs) as an efficient approach to capture the uncertainty range of the full set while reducing computational costs. A predetermined number of RGRs is selected using an integrated unsupervised machine learning (UML) framework, which includes Euclidean distance measurement, multidimensional scaling (MDS), and a deterministic K-means (DK-means) clustering algorithm. In the context of the intricate 3D aquifer CO2 storage model, PUNQ-S3, these algorithms are utilized. The UML methodology selects five RGRs from a pool of 25 possibilities (20% of the total), taking into account the reservoir quality index (RQI) as a static parameter of the reservoir. To determine the credibility of these RGRs, their simulation results are scrutinized through the application of the Kolmogorov–Smirnov (KS) test, which analyzes the distribution of the output. In this assessment, 40 CO2 injection wells cover the entire reservoir alongside the full set. The end-point simulation results indicate that the CO2 structural, residual, and solubility trapping within the RGRs and full set follow the same distribution. Simulating five RGRs alongside the full set of 25 GRs over 200 years, involving 10 years of CO2 injection, reveals consistently similar trapping distribution patterns, with an average value of Dmax of 0.21 remaining lower than Dcritical (0.66). Using this methodology, computational expenses related to scenario testing and development planning for CO2 storage reservoirs in the presence of geological uncertainties can be substantially reduced.

Mahjour, Seyed Kourosh↗

Harnessing interpretable and unsupervised machine learning to address big data from modern X-ray diffraction

The information content of crystalline materials becomes astronomical when collective electronic behavior and their fluctuations are taken into account. In the past decade, improvements in source brightness and detector technology at modern X-ray facilities have allowed a dramatically increased fraction of this information to be captured. Now, the primary challenge is to understand and discover scientific principles from big datasets when a comprehensive analysis is beyond human reach. We report the development of an unsupervised machine learning approach, X-ray diffraction (XRD) temperature clustering (X-TEC), that can automatically extract charge density wave order parameters and detect intraunit cell ordering and its fluctuations from a series of high-volume X-ray diffraction measurements taken at multiple temperatures. We benchmark X-TEC with diffraction data on a quasi-skutterudite family of materials, (Ca x Sr 1–x ) 3 Rh 4 Sn 13 , where a quantum critical point is observed as a function of Ca concentration. We apply X-TEC to XRD data on the pyrochlore metal, Cd 2 Re 2 O 7 , to investigate its two much-debated structural phase transitions and uncover the Goldstone mode accompanying them. We demonstrate how unprecedented atomic-scale knowledge can be gained when human researchers connect the X-TEC results to physical principles. Specifically, we extract from the X-TEC–revealed selection rules that the Cd and Re displacements are approximately equal in amplitude but out of phase. This discovery reveals a previously unknown involvement of 5d 2 Re, supporting the idea of an electronic origin to the structural order. Our approach can radically transform XRD experiments by allowing in operando data analysis and enabling researchers to refine experiments by discovering interesting regions of phase space on the fly.

36 MATERIALS SCIENCE↗