Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “non-negative matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Analysis of Interpretable Data Representations for 4D-STEM Using Unsupervised Learning

Abstract Understanding the structure of materials is crucial for engineering devices and materials with enhanced performance. Four-dimensional scanning transmission electron microscopy (4D-STEM) is capable of mapping nanometer-scale local crystallographic structure over micron-scale field of views. However, 4D-STEM datasets can contain tens of thousands of images from a wide variety of material structures, making it difficult to automate detection and classification of structures. Traditional automated analysis pipelines for 4D-STEM focus on supervised approaches, which require prior knowledge of the material structure and cannot describe anomalous or deviant structures. In this article, a pipeline for engineering 4D-STEM feature representations for unsupervised clustering using non-negative matrix factorization (NMF) is introduced. Each feature is evaluated using NMF and results are presented for both simulated and experimental data. It is shown that some data representations more reliably identify overlapping grains. Additionally, real space refinement is applied to identify spatially distinct sample regions, allowing for size and shape analysis to be performed. This work lays the foundation for improved analysis of nanoscale structural features in materials that deviate from expected crystallographic arrangement using 4D-STEM.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Node Distortion as a Tunable Mechanism for Negative Thermal Expansion in Metal–Organic Frameworks

Chemically functionalized series of metal–organic frameworks (MOFs), with subtle differences in local structure but divergent properties, provide a valuable opportunity to explore how local chemistry can be coupled to long-range structure and functionality. Using in situ synchrotron X-ray total scattering, with powder diffraction and pair distribution function (PDF) analysis, we investigate the temperature dependence of the local- and long-range structure of MOFs based on NU-1000, in which Zr 6 O 8 nodes are coordinated by different capping ligands (H 2 O/OH, Cl – ions, formate, acetylacetonate, and hexafluoroacetylacetonate). We show that the local distortion of the Zr 6 nodes depends on the lability of the ligand and contributes to a negative thermal expansion (NTE) of the extended framework. Using multivariate data analyses, involving non-negative matrix factorization (NMF), we demonstrate a new mechanism for NTE: progressive increase in the population of a smaller, distorted node state with increasing temperature leads to global contraction of the framework. The transformation between discrete node states is noncooperative and not ordered within the lattice, i.e., a solid solution of regular and distorted nodes. Density functional theory calculations show that removal of ligands from the node can lead to distortions consistent with the Zr···Zr distances observed in the experiment PDF data. Control of the node distortion imparted by the nonlinker ligand in turn controls the NTE behavior. Furthermore, these results reveal a mechanism to control the dynamic structure of MOFs based on local chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Robust quantification of the diamond nitrogen-vacancy center charge state via photoluminescence spectroscopy

Nitrogen vacancy (NV) centers in diamond are at the heart of many emerging quantum technologies, all of which require control over the NV charge state. Hence, methods for quantification of the relative photoluminescence intensities of the NV 0 and NV − charge states, i.e., a charge state ratio, are vital. Several approaches to quantify NV charge state ratios have been reported but are either limited to bulk-like NV diamond samples or yield qualitative results. We propose an NV charge state quantification protocol based on the determination of sample- and experimental setup-specific NV 0 and NV − reference spectra. The approach employs blue (400–470 nm) and green (480–570 nm) excitation to infer pure NV 0 and NV − spectra, which are then used to quantify NV charge state ratios in subsequent experiments via least squares fitting. We test our dual excitation protocol (DEP) for a bulk diamond NV sample and 20 and 100 nm nanodiamond particles and compare results with those obtained via other commonly used techniques such as zero-phonon line fitting and non-negative matrix factorization. We find that DEP can be employed across different samples and experimental setups and yields consistent and quantitative results for NV charge state ratios that are in agreement with our understanding of NV photophysics. By providing robust NV charge state quantification across sample types and measurement platforms, DEP will support the development of NV-based quantum technologies.

Color center laser spectroscopy↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

The AUREX cell: a versatile operando electrochemical cell for studying catalytic materials using X-ray diffraction, total scattering and X-ray absorption spectroscopy under working conditions

Understanding the structure–property relationship in electrocatalysts under working conditions is crucial for the rational design of novel and improved catalytic materials. This paper presents the Aarhus University reactor for electrochemical studies using X-rays (AUREX) operando electrocatalytic flow cell, designed as an easy-to-use versatile setup with a minimal background contribution and a uniform flow field to limit concentration polarization and handle gas formation. The cell has been employed to measure operando total scattering, diffraction and absorption spectroscopy as well as simultaneous combinations thereof on a commercial silver electrocatalyst for proof of concept. This combination of operando techniques allows for monitoring of the short-, medium- and long-range structure under working conditions, including an applied potential, liquid electrolyte and local reaction environment. The structural transformations of the Ag electrocatalyst are monitored with non-negative matrix factorization, linear combination analysis, the Pearson correlation coefficient matrix, and refinements in both real and reciprocal space. Upon application of an oxidative potential in an Ar-saturated aqueous 0.1 M KHCO 3 /K 2 CO 3 electrolyte, the face-centered cubic (f.c.c.) Ag gradually transforms first to a trigonal Ag 2 CO 3 phase, followed by the formation of a monoclinic Ag 2 CO 3 phase. A reducing potential immediately reverts the structure to the Ag (f.c.c.) phase. Following the electrochemical-reaction-induced phase transitions is of fundamental interest and necessary for understanding and improving the stability of electrocatalysts, and the operando cell proves a versatile setup for probing this. In addition, it is demonstrated that, when studying electrochemical reactions, a high energy or short exposure time is needed to circumvent beam-induced effects.

Frank, Sara (ORCID:0000000163218363)↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

Characterization of Precipitation-Induced Radon Progeny Deposition Events Using a City-Scale Sensor Network

Networks of radiation detectors provide a platform for real-time radioactive source detection and identification in urban environments. Detection algorithms in these systems must adapt to naturally-occurring changes in background, which requires well-characterized relationships between precipitation events and their corresponding radiological signature. Here, we present a quantitative and qualitative description of rain-induced radon progeny deposition events occurring in Chicago from September 2023 to February 2024. We measure ambient gamma radiation levels, precipitation rate, temperature, pressure, and relative humidity in a network of sensor nodes. For each identified precipitation period, we decompose spectra into static- and radon-associated components as defined by a non-negative matrix factorization (NMF) algorithm. We find a consistent power-law relationship between a precipitation-dependent peak of the radon progeny proxy (RPP) and the peak strength of the radon-associated NMF component for most precipitation events. We conduct a case study of a rainfall period with abnormally high levels of implied radon progeny concentration and describe its temporal and spatial evolution. We hypothesize that this phenomenon is due to the air mass path that intersects a uranium-rich region of Wyoming. Finally, we cluster precipitation events into three distinct categories. One category roughly corresponds to events with deep low-pressure systems and high relative radon concentration, while another is characteristic of light stratiform rain with slightly higher temperatures and intermediate relative radon concentration. The third category appears to contain weak-gradient or lake breeze convection showers with intermittent precipitation and low relative radon concentration. These findings suggest that radiological anomaly detection could be improved by training unique background models corresponding to each category of meteorological event.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data-Driven Optimization of Pixelated CdZnTe Spectrometers for Uranium Enrichment Assay

Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.

Gamma-ray detection↗

Data-Driven Performance Optimization of Gamma Spectrometers With Many Channels

In gamma spectrometers with variable spectroscopic performance across many channels (e.g., many pixels or voxels), a tradeoff exists between including data from successively worse-performing readout channels and increasing efficiency. Brute-force calculation of the optimal set of included channels is exponentially infeasible as the number of channels grows, and approximate methods are required. In this work, we present a data-driven framework for attempting to find near-optimal sets of included detector channels. The framework leverages non-negative matrix factorization (NMF) to learn the behavior of gamma spectra across the detector and clusters similarly-performing detector channels together. Performance comparisons are then made between spectra with channel clusters removed, which is more feasible than brute force. The framework is general and can be applied to arbitrary, user-defined performance metrics depending on the application. We apply this framework to optimizing gamma spectra measured by H3D M400 CdZnTe (CZT) spectrometers, which exhibit variable performance across their crystal volumes. In particular, we show several examples optimizing various performance metrics for uranium and plutonium gamma spectra in non-destructive assay (NDA) for nuclear safeguards, and explore trends in performance versus parameters such as clustering algorithm type. We also compare the NMF + clustering pipeline to several non-machine-learning (ML) algorithms, including several greedy algorithms. Although, we find that the NMF + clustering pipeline tends to find the best-performing set of detector voxels, significantly improving over the unoptimized spectra, but that a greedy accumulation of spectra segmented by detector depth can, in some cases, give similar performance improvements in much less computation time.

Energy resolution↗

LaFA

Latent Feature Attacks on Non-negative Matrix Factorization

Bhattarai, Manish↗

GeoThermalCloud: A Machine Learning Tool for Discovery, Exploration, and Development of Hidden Geothermal Resources

In this 25 minute presentation, we showcase our open source “GeoThermalCloud” tool for identifying hidden geothermal resources using a publicly available dataset for southwestern New Mexico. The presenters include Bulbul Ahmmed and Luke Frash. All of the visuals use source material from LA-UR approved publications and this work falls under the Earth Sciences DUSA. The code shown in this video is already released with LANL approval in open source format on GitHub and DockerHub. The audio in this video includes only material on the topics of geothermal energy and machine learning applied to geothermal energy. The primary machine learning method used is LANL’s Non-negative Matrix Factorization “NMFk” method. Modeling work also mentions LANL’s Geothermal Design Tool “GeoDT” which is another approved open source code that has been released by LANL. This work was performed for DOE Geothermal Technologies Office (DE-EE-3.1.8.1). The host for the released video is intended to be YouTube or a suitable perpetual data repository such as GDR.

15 GEOTHERMAL ENERGY↗

Machine Learning for Geothermal Resource Exploration in the Tularosa Basin, New Mexico

Geothermal energy is considered an essential renewable resource to generate flexible electricity. Geothermal resource assessments conducted by the U.S. Geological Survey showed that the southwestern basins in the U.S. have a significant geothermal potential for meeting domestic electricity demand. Within these southwestern basins, play fairway analysis (PFA), funded by the U.S. Department of Energy’s (DOE) Geothermal Technologies Office, identified that the Tularosa Basin in New Mexico has significant geothermal potential. This short communication paper presents a machine learning (ML) methodology for curating and analyzing the PFA data from the DOE’s geothermal data repository. The proposed approach to identify potential geothermal sites in the Tularosa Basin is based on an unsupervised ML method called non-negative matrix factorization with custom k-means clustering. This methodology is available in our open-source ML framework, GeoThermalCloud (GTC). Using this GTC framework, we discover prospective geothermal locations and find key parameters defining these prospects. Our ML analysis found that these prospects are consistent with the existing Tularosa Basin’s PFA studies. This instills confidence in our GTC framework to accelerate geothermal exploration and resource development, which is generally time-consuming.

15 GEOTHERMAL ENERGY↗

Algorithms for Spectral Decomposition with Applications to Optical Plume Anomaly Detection

The analysis of spectral signals for features that represent physical phenomenon is ubiquitous in the science and engineering communities. There are two main approaches that can be taken to extract relevant features from these high-dimensional data streams. The first set of approaches relies on extracting features using a physics-based paradigm where the underlying physical mechanism that generates the spectra is used to infer the most important features in the data stream. We focus on a complementary methodology that uses a data-driven technique that is informed by the underlying physics but also has the ability to adapt to unmodeled system attributes and dynamics. We discuss the following four algorithms: Spectral Decomposition Algorithm (SDA), Non-Negative Matrix Factorization (NMF), Independent Component Analysis (ICA) and Principal Components Analysis (PCA) and compare their performance on a spectral emulator which we use to generate artificial data with known statistical properties. This spectral emulator mimics the real-world phenomena arising from the plume of the space shuttle main engine and can be used to validate the results that arise from various spectral decomposition algorithms and is very useful for situations where real-world systems have very low probabilities of fault or failure. Our results indicate that methods like SDA and NMF provide a straightforward way of incorporating prior physical knowledge while NMF with a tuning mechanism can give superior performance on some tests. We demonstrate these algorithms to detect potential system-health issues on data from a spectral emulator with tunable health parameters.

Srivastava, Askok N.↗

Application of Atmospheric Gases and Particulate Matter to the Assessment of Urban Heat Island

Background: Urban heat island (UHI), where built areas are warmer compared to non-urban regions, increases human related diseases and mortality. A key challenge in UHI analysis is the designation of sites as urban or suburban/rural; however, the growing complexity of green spaces in urban areas and the predominance of the transportation sector in nonurban areas creates a dilemma for distinct delineation. Objectives: This study aims to utilize the variability of atmospheric components such as particulate matter (PM), inorganic gases, and volatile organic compounds (VOCs) as direct tracers of the degree of urbanization for ground-based measurements to fully comprehend UHI in convoluted regions with indistinct delineation of urban and nonurban environments. Methods: Atmospheric gases and aerosols were used as direct tracers of urbanization for UHI analysis. Inorganic gases and particulate matter were monitored in two sites in a southeastern US city with varying degrees of urbanization. VOCs were analyzed using a proton transfer reaction time-of-flight mass spectrometer. Results: The more-urbanized site exhibited warmer night conditions and elevated total oxidant levels, leading to the formation of nanometer-sized particles. Machine learning analysis revealed similar atmospheric pollutant profiles for both sites, suggesting comparable sources and variability. Biogenic VOCs were enhanced at the less-urbanized site; however, levels of anthropogenic aromatic VOCs were comparable for both sites. A comprehensive mass spectra analysis revealed distinct molecular backbones per site that further affirmed the applicability of VOCs as indicators of urbanization. Conclusion: This study concludes that VOCs provide more direct and accurate information than typical inorganic gases and PM parameters for characterizing the degree of urbanization. Further exploration of VOCs can enhance our understanding of UHI dynamics and its interaction with vegetation in urban green spaces.

Air quality sensor↗

Mapping the proteogenomic landscape enables prediction of drug response in acute myeloid leukemia

Acute myeloid leukemia is a poor prognosis cancer commonly stratified by genetic aberrations, but these mutations are often heterogeneous and don’t always predict therapeutic response. Here we combine transcriptomic, proteomic, and phosphoproteomic datasets with ex vivo drug sensitivity data to help understand the underlying pathophysiology of AML beyond mutations. We measured the proteome and phosphoproteome of 210 patients and combined them with genomics and transcriptomic measurements to identify four proteogenomic subtypes that complemented existing genetic subtypes. We then built a predictor to classify samples into subtypes based on 147 molecular features and mapped them to a ‘landscape’. Each region of this landscape corresponded to specific drug response patterns. We then built a drug response prediction model to identify drugs that target distinct subtypes. We can ultimately use these models to predict drug treatment response and prioritize treatments. Finally, we extended our models and mapped a series of cell lines representing various stages of quizartinib resistance into our subtype landscape, predicting and experimentally validating a switch in sensitivity to venetoclax to panobinostat, two drugs with very different mechanisms than quizartinib. Our results show how multi-omics data together with drug sensitivity data can inform therapy stratification and drug combinations in AML.

59 BASIC BIOLOGICAL SCIENCES↗

A Latent-Variable Formulation of the Poisson Canonical Polyadic Tensor Model: Maximum Likelihood Estimation and Fisher Information

We establish parameter inference for the Poisson canonical polyadic (PCP) tensor model through a latent-variable formulation. Our approach exploits the observation that any random PCP tensor can be derived by marginalizing an unobservable random tensor of one dimension larger. The loglikelihood of this larger dimensional tensor, referred to as the “complete” loglikelihood, is comprised of multiple rank one PCP loglikelihoods. Using this methodology, we first derive maximum likelihood estimators for the PCP model and demonstrate that several existing algorithms for fitting non-negative matrix and tensor factorizations are Expectation-Maximization algorithms. Next, we derive the observed and expected Fisher information matrices for the PCP model. The Fisher information provides us crucial insights into the well-posedness of the tensor model, such as the role that tensor rank plays in identifiability and indeterminacy. For the special case of rank one PCP models, we demonstrate that these results are greatly simplified.

97 MATHEMATICS AND COMPUTING↗