Engineering PapersSearch

DOE OSTI · 2564759

Interpretable machine learning models classify minerals via spectroscopy

Abstract

Developing methods to identify mineral species confidently and rapidly from Raman spectral analysis is critical to numerous fields. Traditionally, analysis relies on pattern matching the Raman spectrum of an unknown dataset with a supporting library of well-characterized spectral data, which may prove difficult for environmental samples that are poorly crystalline or phase mixtures. Here, we developed interpretable machine learning models that can classify uranium minerals by secondary oxyanion chemistry and other physicochemical properties based solely on Raman spectra. This new ML method produces a mineral profile of physical and chemical properties for an unknown sample and can rapidly classify or identify unknown minerals from Raman data, without the need for an exact pattern match in a spectral library. Training models are validated by 1. Strong correlation of high confidence model regions with published spectroscopic assignments and 2. Correct classification of a mineral not present in training data. Training data are from the Compendium of Uranium Raman and Infrared Experimental Spectra and available crystallographic information files within the open-source Smart Spectral Matching scientific framework. Physically meaningful classifier models can rapidly identify key structural and chemical information about unknown uranium minerals and the overall methodology is broadly applicable for mineral phases.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Smith, Robert [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000260581025), Spano, Tyler L. [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000165729722), McDonnell, Marshall [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000237132117), Drane, Lance [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000188081228), Gibbs, Ian [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)], Miskowiec, Andrew [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000203612614), Niedziela, Jennifer L. [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:000000022990923X), Shields, Ashley E. [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000210085242). 2025-05-06. Interpretable machine learning models classify minerals via spectroscopy. https://doi.org/10.1038/s41598-025-92686-2

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Predicting Adaptively Chosen Observables in Quantum Systems

Recent advances have demonstrated that 𝒪⁡(log 𝑀) measurements suffice to predict 𝑀 properties of arbitrarily large quantum many-body systems. However, these remarkable findings assume that the properties to be predicted are chosen independently of the data. This assumption can be violated in practice, where scientists adaptively select properties after looking at previous predictions. This work investigates the adaptive setting for three classes of observables: local, Pauli, and bounded-Frobenius-norm observables. We prove that Ω⁡(√𝑀) samples of an arbitrarily large unknown quantum state are necessary to predict expectation values of 𝑀 adaptively chosen local and Pauli observables, where the system size scales exponentially and polynomially in 𝑀, respectively. We also present computationally efficient algorithms that achieve this information-theoretic lower bound. In contrast, for bounded-Frobenius-norm observables, we devise an algorithm requiring only 𝒪⁡(log 𝑀) samples, independent of system size. These results highlight the potential pitfalls of adaptivity in analyzing data from quantum experiments and provide algorithmic tools to safeguard against erroneous predictions in quantum experiments.

Machine learning

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning