Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “non-negative matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

71 records · Page 4

Data-Driven Performance Optimization of Gamma Spectrometers With Many Channels

In gamma spectrometers with variable spectroscopic performance across many channels (e.g., many pixels or voxels), a tradeoff exists between including data from successively worse-performing readout channels and increasing efficiency. Brute-force calculation of the optimal set of included channels is exponentially infeasible as the number of channels grows, and approximate methods are required. In this work, we present a data-driven framework for attempting to find near-optimal sets of included detector channels. The framework leverages non-negative matrix factorization (NMF) to learn the behavior of gamma spectra across the detector and clusters similarly-performing detector channels together. Performance comparisons are then made between spectra with channel clusters removed, which is more feasible than brute force. The framework is general and can be applied to arbitrary, user-defined performance metrics depending on the application. We apply this framework to optimizing gamma spectra measured by H3D M400 CdZnTe (CZT) spectrometers, which exhibit variable performance across their crystal volumes. In particular, we show several examples optimizing various performance metrics for uranium and plutonium gamma spectra in non-destructive assay (NDA) for nuclear safeguards, and explore trends in performance versus parameters such as clustering algorithm type. We also compare the NMF + clustering pipeline to several non-machine-learning (ML) algorithms, including several greedy algorithms. Although, we find that the NMF + clustering pipeline tends to find the best-performing set of detector voxels, significantly improving over the unoptimized spectra, but that a greedy accumulation of spectra segmented by detector depth can, in some cases, give similar performance improvements in much less computation time.

Energy resolution↗

LaFA

Latent Feature Attacks on Non-negative Matrix Factorization

Bhattarai, Manish↗

Machine Learning Model Geotiffs - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains geotiffs, supporting shapefiles and readmes for the inputs and output models of algorithms explored in the Nevada Geothermal Machine Learning project, meant to accompany the final report. Layers include: Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk), input rasters of feature sets, and positive/negative training sites. See readme .txt files and final report for additional metadata. A submission linking the full codebase for generating machine learning output models is available under "related resources" on this page.

15 GEOTHERMAL ENERGY↗

GIS Resource Compilation Map Package - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains an ESRI map package (.mpk) with an embedded geodatabase for GIS resources used or derived in the Nevada Machine Learning project, meant to accompany the final report. The package includes layer descriptions, layer grouping, and symbology. Layer groups include: new/revised datasets (paleo-geothermal features, geochemistry, geophysics, heat flow, slip and dilation, potential structures, geothermal power plants, positive and negative test sites), machine learning model input grids, machine learning models (Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk) - supervised and unsupervised), original NV Play Fairway data and models, and NV cultural/reference data. See layer descriptions for additional metadata. Smaller GIS resource packages (by category) can be found in the related datasets section of this submission. A submission linking the full codebase for generating machine learning output models is available through the "Related Datasets" link on this page, and contains results beyond the top picks present in this compilation.

15 GEOTHERMAL ENERGY↗

GeoThermalCloud: A Machine Learning Tool for Discovery, Exploration, and Development of Hidden Geothermal Resources

In this 25 minute presentation, we showcase our open source “GeoThermalCloud” tool for identifying hidden geothermal resources using a publicly available dataset for southwestern New Mexico. The presenters include Bulbul Ahmmed and Luke Frash. All of the visuals use source material from LA-UR approved publications and this work falls under the Earth Sciences DUSA. The code shown in this video is already released with LANL approval in open source format on GitHub and DockerHub. The audio in this video includes only material on the topics of geothermal energy and machine learning applied to geothermal energy. The primary machine learning method used is LANL’s Non-negative Matrix Factorization “NMFk” method. Modeling work also mentions LANL’s Geothermal Design Tool “GeoDT” which is another approved open source code that has been released by LANL. This work was performed for DOE Geothermal Technologies Office (DE-EE-3.1.8.1). The host for the released video is intended to be YouTube or a suitable perpetual data repository such as GDR.

15 GEOTHERMAL ENERGY↗

Machine Learning for Geothermal Resource Exploration in the Tularosa Basin, New Mexico

Geothermal energy is considered an essential renewable resource to generate flexible electricity. Geothermal resource assessments conducted by the U.S. Geological Survey showed that the southwestern basins in the U.S. have a significant geothermal potential for meeting domestic electricity demand. Within these southwestern basins, play fairway analysis (PFA), funded by the U.S. Department of Energy’s (DOE) Geothermal Technologies Office, identified that the Tularosa Basin in New Mexico has significant geothermal potential. This short communication paper presents a machine learning (ML) methodology for curating and analyzing the PFA data from the DOE’s geothermal data repository. The proposed approach to identify potential geothermal sites in the Tularosa Basin is based on an unsupervised ML method called non-negative matrix factorization with custom k-means clustering. This methodology is available in our open-source ML framework, GeoThermalCloud (GTC). Using this GTC framework, we discover prospective geothermal locations and find key parameters defining these prospects. Our ML analysis found that these prospects are consistent with the existing Tularosa Basin’s PFA studies. This instills confidence in our GTC framework to accelerate geothermal exploration and resource development, which is generally time-consuming.

15 GEOTHERMAL ENERGY↗

Neural Network Approaches for Mobile Spectroscopic Gamma-Ray Source Detection

Artificial neural networks (ANNs) for performing spectroscopic gamma-ray source identification have been previously introduced, primarily for applications in controlled laboratory settings. To understand the utility of these methods in scenarios and environments more relevant to nuclear safety and security, this work examines the use of ANNs for mobile detection, which involves highly variable gamma-ray background, low signal-to-noise ratio measurements, and low false alarm rates. Simulated data from a 2” × 4” × 16” NaI(Tl) detector are used in this work for demonstrating these concepts, and the minimum detectable activity (MDA) is used as a performance metric in assessing model performance.In addition to examining simultaneous detection and identification, binary spectral anomaly detection using autoencoders is introduced in this work, and benchmarked using detection methods based on Non-negative Matrix Factorization (NMF) and Principal Component Analysis (PCA). On average, the autoencoder provides a 12% and 23% improvement over NMF- and PCA-based detection methods, respectively. Additionally, source identification using ANNs is extended to leverage temporal dynamics by means of recurrent neural networks, and these time-dependent models outperform their time-independent counterparts by 17% for the analysis examined here. The paper concludes with a discussion on tradeoffs between the ANN-based approaches and the benchmark methods examined here.

Bilton, Kyle J. (ORCID:0000000184553689)↗

Algorithms for Spectral Decomposition with Applications to Optical Plume Anomaly Detection

The analysis of spectral signals for features that represent physical phenomenon is ubiquitous in the science and engineering communities. There are two main approaches that can be taken to extract relevant features from these high-dimensional data streams. The first set of approaches relies on extracting features using a physics-based paradigm where the underlying physical mechanism that generates the spectra is used to infer the most important features in the data stream. We focus on a complementary methodology that uses a data-driven technique that is informed by the underlying physics but also has the ability to adapt to unmodeled system attributes and dynamics. We discuss the following four algorithms: Spectral Decomposition Algorithm (SDA), Non-Negative Matrix Factorization (NMF), Independent Component Analysis (ICA) and Principal Components Analysis (PCA) and compare their performance on a spectral emulator which we use to generate artificial data with known statistical properties. This spectral emulator mimics the real-world phenomena arising from the plume of the space shuttle main engine and can be used to validate the results that arise from various spectral decomposition algorithms and is very useful for situations where real-world systems have very low probabilities of fault or failure. Our results indicate that methods like SDA and NMF provide a straightforward way of incorporating prior physical knowledge while NMF with a tuning mechanism can give superior performance on some tests. We demonstrate these algorithms to detect potential system-health issues on data from a spectral emulator with tunable health parameters.

Srivastava, Askok N.↗

Correlations Between Panoramic Imagery and Gamma-Ray Background in an Urban Area

When searching for radiological sources in an urban area, a vehicle-borne detector system will often measure complex, varying backgrounds primarily from natural gamma-ray sources. Much work has been focused on developing spectral algorithms that retain sensitivity and minimize the false-positive rate even in the presence of such spectral and temporal variability. However, information about the environment surrounding the detector system might also provide useful clues about the expected background, which if incorporated into an algorithm, could improve performance. Recent work has focused on extensive measuring and modeling of urban areas with the goal of understanding how these complex backgrounds arise. This work presents an analysis of panoramic video images and gamma-ray background data collected in Oakland, California, by the radiological multisensor analysis platform (RadMAP) vehicle. Features were extracted from the panoramic images by semantically labeling the images and then convolving the labeled regions with the detector response. A linear model was used to relate the image-derived features to gamma-ray spectral features obtained using nonnegative matrix factorization (NMF) under different regularizations. Here we find some gamma-ray background features correlate strongly with image-derived features that measure the response-adjusted solid angle subtended by sky and buildings, and we discuss the implications for the development of future, contextually aware detection algorithms.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Application of Atmospheric Gases and Particulate Matter to the Assessment of Urban Heat Island

Background: Urban heat island (UHI), where built areas are warmer compared to non-urban regions, increases human related diseases and mortality. A key challenge in UHI analysis is the designation of sites as urban or suburban/rural; however, the growing complexity of green spaces in urban areas and the predominance of the transportation sector in nonurban areas creates a dilemma for distinct delineation. Objectives: This study aims to utilize the variability of atmospheric components such as particulate matter (PM), inorganic gases, and volatile organic compounds (VOCs) as direct tracers of the degree of urbanization for ground-based measurements to fully comprehend UHI in convoluted regions with indistinct delineation of urban and nonurban environments. Methods: Atmospheric gases and aerosols were used as direct tracers of urbanization for UHI analysis. Inorganic gases and particulate matter were monitored in two sites in a southeastern US city with varying degrees of urbanization. VOCs were analyzed using a proton transfer reaction time-of-flight mass spectrometer. Results: The more-urbanized site exhibited warmer night conditions and elevated total oxidant levels, leading to the formation of nanometer-sized particles. Machine learning analysis revealed similar atmospheric pollutant profiles for both sites, suggesting comparable sources and variability. Biogenic VOCs were enhanced at the less-urbanized site; however, levels of anthropogenic aromatic VOCs were comparable for both sites. A comprehensive mass spectra analysis revealed distinct molecular backbones per site that further affirmed the applicability of VOCs as indicators of urbanization. Conclusion: This study concludes that VOCs provide more direct and accurate information than typical inorganic gases and PM parameters for characterizing the degree of urbanization. Further exploration of VOCs can enhance our understanding of UHI dynamics and its interaction with vegetation in urban green spaces.

Air quality sensor↗

Mapping the proteogenomic landscape enables prediction of drug response in acute myeloid leukemia

Acute myeloid leukemia is a poor prognosis cancer commonly stratified by genetic aberrations, but these mutations are often heterogeneous and don’t always predict therapeutic response. Here we combine transcriptomic, proteomic, and phosphoproteomic datasets with ex vivo drug sensitivity data to help understand the underlying pathophysiology of AML beyond mutations. We measured the proteome and phosphoproteome of 210 patients and combined them with genomics and transcriptomic measurements to identify four proteogenomic subtypes that complemented existing genetic subtypes. We then built a predictor to classify samples into subtypes based on 147 molecular features and mapped them to a ‘landscape’. Each region of this landscape corresponded to specific drug response patterns. We then built a drug response prediction model to identify drugs that target distinct subtypes. We can ultimately use these models to predict drug treatment response and prioritize treatments. Finally, we extended our models and mapped a series of cell lines representing various stages of quizartinib resistance into our subtype landscape, predicting and experimentally validating a switch in sensitivity to venetoclax to panobinostat, two drugs with very different mechanisms than quizartinib. Our results show how multi-omics data together with drug sensitivity data can inform therapy stratification and drug combinations in AML.

59 BASIC BIOLOGICAL SCIENCES↗

Python Codebase and Jupyter Notebooks - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

Git archive containing Python modules and resources used to generate machine-learning models used in the "Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada" project. This software is licensed as free to use, modify, and distribute with attribution. Full license details are included within the archive. See "documentation.zip" for setup instructions and file trees annotated with module descriptions.

Brown, Stephen↗

A Latent-Variable Formulation of the Poisson Canonical Polyadic Tensor Model: Maximum Likelihood Estimation and Fisher Information

We establish parameter inference for the Poisson canonical polyadic (PCP) tensor model through a latent-variable formulation. Our approach exploits the observation that any random PCP tensor can be derived by marginalizing an unobservable random tensor of one dimension larger. The loglikelihood of this larger dimensional tensor, referred to as the “complete” loglikelihood, is comprised of multiple rank one PCP loglikelihoods. Using this methodology, we first derive maximum likelihood estimators for the PCP model and demonstrate that several existing algorithms for fitting non-negative matrix and tensor factorizations are Expectation-Maximization algorithms. Next, we derive the observed and expected Fisher information matrices for the PCP model. The Fisher information provides us crucial insights into the well-posedness of the tensor model, such as the role that tensor rank plays in identifiability and indeterminacy. For the special case of rank one PCP models, we demonstrate that these results are greatly simplified.

97 MATHEMATICS AND COMPUTING↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Correlated mechanochemical maps of Arabidopsis thaliana primary cell walls using atomic force microscope infrared spectroscopy

Spatial heterogeneity in composition and organisation of the primary cell wall affects the mechanics of cellular morphogenesis. However, directly correlating cell wall composition, organisation and mechanics has been challenging. To overcome this barrier, we applied atomic force microscopy coupled with infrared (AFM-IR) spectroscopy to generate spatially correlated maps of chemical and mechanical properties for paraformaldehyde-fixed, intact Arabidopsis thaliana epidermal cell walls. AFM-IR spectra were deconvoluted by non-negative matrix factorisation (NMF) into a linear combination of IR spectral factors representing sets of chemical groups comprising different cell wall components. This approach enables quantification of chemical composition from IR spectral signatures and visualisation of chemical heterogeneity at nanometer resolution. Cross-correlation analysis of the spatial distribution of NMFs and mechanical properties suggests that the carbohydrate composition of cell wall junctions correlates with increased local stiffness. Together, our work establishes new methodology to use AFM-IR for the mechanochemical analysis of intact plant primary cell walls.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced relaxed physical factorization preconditioner for coupled poromechanics

The relaxed physical factorization (RPF) preconditioner is a recent algorithm allowing for the efficient and robust solution to the block linear systems arising from the three-field displacement-velocity-pressure formulation of coupled poromechanics. For its application, however, it is necessary to invert blocks with the algebraic form C^ = (C + βFF T ), where C is a symmetric positive definite matrix, FF T a rank-deficient term, and β a real non-negative coefficient. The inversion of C^, performed in an inexact way, can become unstable for large values of β, as it usually occurs at some stages of a full poromechanical simulation. In this work, we propose a family of algebraic techniques to stabilize the inexact solve with C^. This strategy can prove useful in other problems as well where such an issue might arise, such as augmented Lagrangian preconditioning techniques for Navier-Stokes or incompressible elasticity. First, we introduce an iterative scheme obtained by a natural splitting of matrix C^. Second, we develop a technique based on the use of a proper projection operator annihilating the near-kernel modes of C^. Both approaches give rise to a novel class of preconditioners denoted as Enhanced RPF (ERPF). Furthermore, effectiveness and robustness of the proposed algorithms are demonstrated in both theoretical benchmarks and real-world large-size applications, outperforming the native RPF preconditioner.

97 MATHEMATICS AND COMPUTING↗