Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

Chemical and elemental mapping of spent nuclear fuel sections by soft X-ray spectromicroscopy

Soft X-ray spectromicroscopy at the O K -edge, U N 4,5 -edges and Ce M 4,5 -edges has been performed on focused ion beam sections of spent nuclear fuel for the first time, yielding chemical information on the sub-micrometer scale. To analyze these data, a modification to non-negative matrix factorization (NMF) was developed, in which the data are no longer required to be non-negative, but the non-negativity of the spectral components and fit coefficients is largely preserved. The modified NMF method was utilized at the O K -edge to distinguish between two components, one present in the bulk of the sample similar to UO 2 and one present at the interface of the sample which is a hyperstoichiometric UO 2+ x species. The species maps are consistent with a model of a thin layer of UO 2+ x over the entire sample, which is likely explained by oxidation after focused ion beam (FIB) sectioning. In addition to the uranium oxide bulk of the sample, Ce measurements were also performed to investigate the oxidation state of that fission product, which is the subject of considerable interest. Analysis of the Ce spectra shows that Ce is in a predominantly trivalent state, with a possible contribution from tetravalent Ce. Atom probe analysis was performed to provide confirmation of the presence and localization of Ce in the spent fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Fast Active-Set Thresholding Method for Nonnegative Least Squares

Nonnegative Least Squares (NNLS) is a fundamental constrained optimization problem encountered in many applications such as image deblurring, signal processing, nonnegative matrix factorization, magnetic microscopy, and hyperspectral imaging. Active-set based methods are a common class of algorithms for solving NNLS which identify the optimal variable set of the NNLS solution. They do so by iteratively solving a series of unconstrained least squares problems, identifying which variables violate the nonnegativity constraints, and then swapping variables in/out of consideration until the optimal set of variables is found. Several variations improving upon this method exist in the literature. In this work, we propose an active-set swap heuristic which further improves upon existing active-set based methods for NNLS. Our optimizations are based upon adding multiple variables to the passive set within a threshold of the smallest gradient value and removing variables within a similar threshold of the closest boundary constraint. We leverage these optimizations to yield a Fast Active-Set Thresholding NNLS (FAST-NNLS) algorithm which significantly outperforms the existing state-of-the-art NNLS algorithms for a wide range of problems. Rigorous convergence guarantees are proven for the proposed method. We demonstrate the effectiveness of our proposed method on multiple synthetic datasets and two realworld text analysis applications. In doing so, we present the most comprehensive NNLS solver comparison in the literature to date.

Cobb, Benjamin [Georgia Institute of Technology]↗

User Role Identification in Software Vulnerability Discussions over Social Networks

Understanding and early awareness of software vulnerabilities is vital for preventing and mitigating potential impacts from cybersecurity events. One step toward early characterization of software vulnerabilities may involve analyzing discussion and spread of information in online social networks. Prior work has used information from such discussions over multiple online forums to develop dynamic networks among users followed by analysis of structure, spread, and information evolution. In this work, we advance the state-of-the-art by focusing on data-driven learning of types, roles, and transition of roles exhibited by users over time. In social networks, users take on particular roles based on their actions and structure of the network. Identifying “meaningful” roles can help separate potential users of interest from the larger community, and identify patterns in a network. We will identify and compare roles found in online forums (e.g., Twitter) using techniques such as feature-based Non-negative Matrix Factorization coupled with topological and influence-based measures of centrality. Since users’ activities change over time, we also analyze role evolution in dynamic networks.

Jones, Rebecca D.↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗

Characterization of Precipitation-Induced Radon Progeny Deposition Events Using a City-Scale Sensor Network

Networks of radiation detectors provide a platform for real-time radioactive source detection and identification in urban environments. Detection algorithms in these systems must adapt to naturally-occurring changes in background, which requires well-characterized relationships between precipitation events and their corresponding radiological signature. Here, we present a quantitative and qualitative description of rain-induced radon progeny deposition events occurring in Chicago from September 2023 to February 2024. We measure ambient gamma radiation levels, precipitation rate, temperature, pressure, and relative humidity in a network of sensor nodes. For each identified precipitation period, we decompose spectra into static- and radon-associated components as defined by a non-negative matrix factorization (NMF) algorithm. We find a consistent power-law relationship between a precipitation-dependent peak of the radon progeny proxy (RPP) and the peak strength of the radon-associated NMF component for most precipitation events. We conduct a case study of a rainfall period with abnormally high levels of implied radon progeny concentration and describe its temporal and spatial evolution. We hypothesize that this phenomenon is due to the air mass path that intersects a uranium-rich region of Wyoming. Finally, we cluster precipitation events into three distinct categories. One category roughly corresponds to events with deep low-pressure systems and high relative radon concentration, while another is characteristic of light stratiform rain with slightly higher temperatures and intermediate relative radon concentration. The third category appears to contain weak-gradient or lake breeze convection showers with intermittent precipitation and low relative radon concentration. These findings suggest that radiological anomaly detection could be improved by training unique background models corresponding to each category of meteorological event.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data-Driven Optimization of Pixelated CdZnTe Spectrometers for Uranium Enrichment Assay

Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.

Gamma-ray detection↗

Data-Driven Performance Optimization of Gamma Spectrometers With Many Channels

In gamma spectrometers with variable spectroscopic performance across many channels (e.g., many pixels or voxels), a tradeoff exists between including data from successively worse-performing readout channels and increasing efficiency. Brute-force calculation of the optimal set of included channels is exponentially infeasible as the number of channels grows, and approximate methods are required. In this work, we present a data-driven framework for attempting to find near-optimal sets of included detector channels. The framework leverages non-negative matrix factorization (NMF) to learn the behavior of gamma spectra across the detector and clusters similarly-performing detector channels together. Performance comparisons are then made between spectra with channel clusters removed, which is more feasible than brute force. The framework is general and can be applied to arbitrary, user-defined performance metrics depending on the application. We apply this framework to optimizing gamma spectra measured by H3D M400 CdZnTe (CZT) spectrometers, which exhibit variable performance across their crystal volumes. In particular, we show several examples optimizing various performance metrics for uranium and plutonium gamma spectra in non-destructive assay (NDA) for nuclear safeguards, and explore trends in performance versus parameters such as clustering algorithm type. We also compare the NMF + clustering pipeline to several non-machine-learning (ML) algorithms, including several greedy algorithms. Although, we find that the NMF + clustering pipeline tends to find the best-performing set of detector voxels, significantly improving over the unoptimized spectra, but that a greedy accumulation of spectra segmented by detector depth can, in some cases, give similar performance improvements in much less computation time.

Energy resolution↗

Distributed-Memory Parallel JointNMF

Joint Nonnegative Matrix Factorization (JointNMF) is a hybrid method for mining information from datasets that contain both feature and connection information. We propose distributed-memory parallelizations of three algorithms for solving the JointNMF problem based on Alternating Nonnegative Least Squares, Projected Gradient Descent, and Projected Gauss-Newton. We extend well-known communication-avoiding algorithms using a single processor grid case to our coupled case on two processor grids. We demonstrate the scalability of the algorithms on up to 960 cores (40 nodes) with 60% parallel efficiency. The more sophisticated Alternating Nonnegative Least Squares (ANLS) and Gauss-Newton variants outperform the first-order gradient descent method in reducing the objective on large-scale problems. We perform a topic modelling task on a large corpus of academic papers that consists of over 37 million paper abstracts and nearly a billion citation relationships, demonstrating the utility and scalability of the methods.

Eswar, Srinivas↗

LaFA

Latent Feature Attacks on Non-negative Matrix Factorization

Bhattarai, Manish↗

Impaired cardiac glycolysis and glycogen depletion are linked to poor myocardial outcomes in juvenile male swine with metabolic syndrome and ischemia

Abstract Obesity continues to rise in the juveniles and obese children are more likely to develop metabolic syndrome (MetS) and related cardiovascular disease. Unfortunately, effective prevention and long‐term treatment options remain limited. We determined the juvenile cardiac response to MetS in a swine model. Juvenile male swine were fed either an obesogenic diet, to induce MetS, or a lean diet, as a control (LD). Myocardial ischemia was induced with surgically placed ameroid constrictor on the left circumflex artery. Physiological data were recorded and at 22 weeks of age the animals underwent a terminal harvest procedure and myocardial tissue was extracted for total metabolic and proteomic LC/MS–MS, RNA‐seq analysis, and data underwent nonnegative matrix factorization for metabolic signatures. Significantly altered in MetS versus. LD were the glycolysis‐related metabolites and enzymes. In MetS compared with LD Glycogen synthase 1 (GYS1)‐glycogen phosphorylases (PYGM/PYGL) expression disbalance resulted in a loss of myocardial glycogen. Our findings are consistent with the concept that transcriptionally driven myocardial changes in glycogen and glucose metabolism‐related enzymes lead to a deficiency of their metabolite products in MetS. This abnormal energy metabolism provides insight into the pathogenesis of the juvenile heart in MetS. This study reveals that MetS and ischemia diminishes ATP availability in the myocardium via altering the glucose‐G6P‐pyruvate axis at the level of metabolites and gene expression of related enzymes. The observed severe glycogen depletion in MetS coincides with disbalance in expression of GYS1 and both PYGM and PYGL. This altered energy substrate metabolism is a potential target of pharmacological agents for improving juvenile myocardial function in MetS and ischemia.

Broadwin, Mark↗

Machine Learning Model Geotiffs - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains geotiffs, supporting shapefiles and readmes for the inputs and output models of algorithms explored in the Nevada Geothermal Machine Learning project, meant to accompany the final report. Layers include: Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk), input rasters of feature sets, and positive/negative training sites. See readme .txt files and final report for additional metadata. A submission linking the full codebase for generating machine learning output models is available under "related resources" on this page.

15 GEOTHERMAL ENERGY↗

GIS Resource Compilation Map Package - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains an ESRI map package (.mpk) with an embedded geodatabase for GIS resources used or derived in the Nevada Machine Learning project, meant to accompany the final report. The package includes layer descriptions, layer grouping, and symbology. Layer groups include: new/revised datasets (paleo-geothermal features, geochemistry, geophysics, heat flow, slip and dilation, potential structures, geothermal power plants, positive and negative test sites), machine learning model input grids, machine learning models (Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk) - supervised and unsupervised), original NV Play Fairway data and models, and NV cultural/reference data. See layer descriptions for additional metadata. Smaller GIS resource packages (by category) can be found in the related datasets section of this submission. A submission linking the full codebase for generating machine learning output models is available through the "Related Datasets" link on this page, and contains results beyond the top picks present in this compilation.

15 GEOTHERMAL ENERGY↗

Randomized Algorithms for Scientific Computing (RASC)

Randomized algorithms have propelled advances in artificial intelligence (AI) and represent a foundational research area in advancing AI for Science. Future advancements in DOE Office of Science priority areas such as climate science, astrophysics, fusion, advanced materials, combustion, and quantum computing all require randomized algorithms for surmounting challenges of complexity, robustness, and scalability. Advances in data collection and numerical simulation have changed the dynamics of scientific research and motivate the need for randomized algorithms. For instance, advances in imaging technologies such as X-ray ptychography, electron microscopy, electron energy loss spectroscopy, or adaptive optics lattice light-sheet microscopy collect hyperspectral imaging and scattering data in terabytes, at breakneck speed enabled by state-of-the-art detectors. The data collection is exceptionally fast compared with its analysis. Likewise, advances in high-performance architectures have made exascale computing a reality and changed the economies of scientific computing in the process. Floating-point operations that create data are essentially free in comparison with data movement. Thus far, most approaches have focused on creating faster hardware. Ironically, this faster hardware has exacerbated the problem by making data still easier to create. Under such an onslaught, scientists often resort to heuristic deterministic sampling schemes (e.g., low-precision arithmetic, sampling every nth element) and sacrifice potentially valuable accuracy. Dramatically better results can be achieved via randomized algorithms, reducing the data size as much as or more than naive deterministic subsampling can achieve, while retaining the high accuracy of computing on the full data set. By randomized algorithms we mean those algorithms that employ some form of randomness in internal algorithmic decisions to accelerate time to solution, increase scalability, or improve reliability. Examples include matrix sketching for solving large-scale least-squares problems (see Figure 1) and stochastic gradient descent for training machine learning models. We are not recommending heuristic methods but rather randomized algorithms that have certificates of correctness and probabilistic guarantees of optimality and near-optimality. Such approaches can be useful beyond acceleration, for example, in understanding how to avoid measure zero worst-case scenarios that plague methods such as QR matrix factorization.

97 MATHEMATICS AND COMPUTING↗

Applying novel analytical tools for analyzing multidimensional secondary organic aerosol measurements

In the atmosphere, secondary organic aerosols (SOA) are often the major components of fine particulate matter and interact with clouds and radiation. SOA comprises a mixture of thousands of organic compounds. There is tremendous complexity and uncertainty in understanding SOA formation, since it is formed by oxidation and gas to particle conversion of a variety of sources: natural biogenic, anthropogenic (vehicles, cooking coal combustion) and biomass burning. The Aerosol Mass Spectrometer (AMS) produces multidimensional chemical information about SOA but analyzing this data to understand SOA sources relies on time consuming analyses (~months to years) such as the positive matrix factorization (PMF). PMF also becomes difficult for aircraft data where signal to noise ratio is weaker. There is a critical need to develop fast machine learning techniques that can analytically provide information about SOA sources using AMS data on the same timescales as the data is being collected (~minutes). We apply a machine learning supervised classification approach: the multinomial logistic regression to rapidly classify AMS data obtained from aircraft measurements.

47 OTHER INSTRUMENTATION↗

GeoThermalCloud: A Machine Learning Tool for Discovery, Exploration, and Development of Hidden Geothermal Resources

In this 25 minute presentation, we showcase our open source “GeoThermalCloud” tool for identifying hidden geothermal resources using a publicly available dataset for southwestern New Mexico. The presenters include Bulbul Ahmmed and Luke Frash. All of the visuals use source material from LA-UR approved publications and this work falls under the Earth Sciences DUSA. The code shown in this video is already released with LANL approval in open source format on GitHub and DockerHub. The audio in this video includes only material on the topics of geothermal energy and machine learning applied to geothermal energy. The primary machine learning method used is LANL’s Non-negative Matrix Factorization “NMFk” method. Modeling work also mentions LANL’s Geothermal Design Tool “GeoDT” which is another approved open source code that has been released by LANL. This work was performed for DOE Geothermal Technologies Office (DE-EE-3.1.8.1). The host for the released video is intended to be YouTube or a suitable perpetual data repository such as GDR.

15 GEOTHERMAL ENERGY↗