Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Fast Active-Set Thresholding Method for Nonnegative Least Squares

Nonnegative Least Squares (NNLS) is a fundamental constrained optimization problem encountered in many applications such as image deblurring, signal processing, nonnegative matrix factorization, magnetic microscopy, and hyperspectral imaging. Active-set based methods are a common class of algorithms for solving NNLS which identify the optimal variable set of the NNLS solution. They do so by iteratively solving a series of unconstrained least squares problems, identifying which variables violate the nonnegativity constraints, and then swapping variables in/out of consideration until the optimal set of variables is found. Several variations improving upon this method exist in the literature. In this work, we propose an active-set swap heuristic which further improves upon existing active-set based methods for NNLS. Our optimizations are based upon adding multiple variables to the passive set within a threshold of the smallest gradient value and removing variables within a similar threshold of the closest boundary constraint. We leverage these optimizations to yield a Fast Active-Set Thresholding NNLS (FAST-NNLS) algorithm which significantly outperforms the existing state-of-the-art NNLS algorithms for a wide range of problems. Rigorous convergence guarantees are proven for the proposed method. We demonstrate the effectiveness of our proposed method on multiple synthetic datasets and two realworld text analysis applications. In doing so, we present the most comprehensive NNLS solver comparison in the literature to date.

Cobb, Benjamin [Georgia Institute of Technology]↗

User Role Identification in Software Vulnerability Discussions over Social Networks

Understanding and early awareness of software vulnerabilities is vital for preventing and mitigating potential impacts from cybersecurity events. One step toward early characterization of software vulnerabilities may involve analyzing discussion and spread of information in online social networks. Prior work has used information from such discussions over multiple online forums to develop dynamic networks among users followed by analysis of structure, spread, and information evolution. In this work, we advance the state-of-the-art by focusing on data-driven learning of types, roles, and transition of roles exhibited by users over time. In social networks, users take on particular roles based on their actions and structure of the network. Identifying “meaningful” roles can help separate potential users of interest from the larger community, and identify patterns in a network. We will identify and compare roles found in online forums (e.g., Twitter) using techniques such as feature-based Non-negative Matrix Factorization coupled with topological and influence-based measures of centrality. Since users’ activities change over time, we also analyze role evolution in dynamic networks.

Jones, Rebecca D.↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗

Characterization of Precipitation-Induced Radon Progeny Deposition Events Using a City-Scale Sensor Network

Networks of radiation detectors provide a platform for real-time radioactive source detection and identification in urban environments. Detection algorithms in these systems must adapt to naturally-occurring changes in background, which requires well-characterized relationships between precipitation events and their corresponding radiological signature. Here, we present a quantitative and qualitative description of rain-induced radon progeny deposition events occurring in Chicago from September 2023 to February 2024. We measure ambient gamma radiation levels, precipitation rate, temperature, pressure, and relative humidity in a network of sensor nodes. For each identified precipitation period, we decompose spectra into static- and radon-associated components as defined by a non-negative matrix factorization (NMF) algorithm. We find a consistent power-law relationship between a precipitation-dependent peak of the radon progeny proxy (RPP) and the peak strength of the radon-associated NMF component for most precipitation events. We conduct a case study of a rainfall period with abnormally high levels of implied radon progeny concentration and describe its temporal and spatial evolution. We hypothesize that this phenomenon is due to the air mass path that intersects a uranium-rich region of Wyoming. Finally, we cluster precipitation events into three distinct categories. One category roughly corresponds to events with deep low-pressure systems and high relative radon concentration, while another is characteristic of light stratiform rain with slightly higher temperatures and intermediate relative radon concentration. The third category appears to contain weak-gradient or lake breeze convection showers with intermittent precipitation and low relative radon concentration. These findings suggest that radiological anomaly detection could be improved by training unique background models corresponding to each category of meteorological event.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data-Driven Optimization of Pixelated CdZnTe Spectrometers for Uranium Enrichment Assay

Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.

Gamma-ray detection↗

Data-Driven Performance Optimization of Gamma Spectrometers With Many Channels

In gamma spectrometers with variable spectroscopic performance across many channels (e.g., many pixels or voxels), a tradeoff exists between including data from successively worse-performing readout channels and increasing efficiency. Brute-force calculation of the optimal set of included channels is exponentially infeasible as the number of channels grows, and approximate methods are required. In this work, we present a data-driven framework for attempting to find near-optimal sets of included detector channels. The framework leverages non-negative matrix factorization (NMF) to learn the behavior of gamma spectra across the detector and clusters similarly-performing detector channels together. Performance comparisons are then made between spectra with channel clusters removed, which is more feasible than brute force. The framework is general and can be applied to arbitrary, user-defined performance metrics depending on the application. We apply this framework to optimizing gamma spectra measured by H3D M400 CdZnTe (CZT) spectrometers, which exhibit variable performance across their crystal volumes. In particular, we show several examples optimizing various performance metrics for uranium and plutonium gamma spectra in non-destructive assay (NDA) for nuclear safeguards, and explore trends in performance versus parameters such as clustering algorithm type. We also compare the NMF + clustering pipeline to several non-machine-learning (ML) algorithms, including several greedy algorithms. Although, we find that the NMF + clustering pipeline tends to find the best-performing set of detector voxels, significantly improving over the unoptimized spectra, but that a greedy accumulation of spectra segmented by detector depth can, in some cases, give similar performance improvements in much less computation time.

Energy resolution↗

Distributed-Memory Parallel JointNMF

Joint Nonnegative Matrix Factorization (JointNMF) is a hybrid method for mining information from datasets that contain both feature and connection information. We propose distributed-memory parallelizations of three algorithms for solving the JointNMF problem based on Alternating Nonnegative Least Squares, Projected Gradient Descent, and Projected Gauss-Newton. We extend well-known communication-avoiding algorithms using a single processor grid case to our coupled case on two processor grids. We demonstrate the scalability of the algorithms on up to 960 cores (40 nodes) with 60% parallel efficiency. The more sophisticated Alternating Nonnegative Least Squares (ANLS) and Gauss-Newton variants outperform the first-order gradient descent method in reducing the objective on large-scale problems. We perform a topic modelling task on a large corpus of academic papers that consists of over 37 million paper abstracts and nearly a billion citation relationships, demonstrating the utility and scalability of the methods.

Eswar, Srinivas↗

LaFA

Latent Feature Attacks on Non-negative Matrix Factorization

Bhattarai, Manish↗

Impaired cardiac glycolysis and glycogen depletion are linked to poor myocardial outcomes in juvenile male swine with metabolic syndrome and ischemia

Abstract Obesity continues to rise in the juveniles and obese children are more likely to develop metabolic syndrome (MetS) and related cardiovascular disease. Unfortunately, effective prevention and long‐term treatment options remain limited. We determined the juvenile cardiac response to MetS in a swine model. Juvenile male swine were fed either an obesogenic diet, to induce MetS, or a lean diet, as a control (LD). Myocardial ischemia was induced with surgically placed ameroid constrictor on the left circumflex artery. Physiological data were recorded and at 22 weeks of age the animals underwent a terminal harvest procedure and myocardial tissue was extracted for total metabolic and proteomic LC/MS–MS, RNA‐seq analysis, and data underwent nonnegative matrix factorization for metabolic signatures. Significantly altered in MetS versus. LD were the glycolysis‐related metabolites and enzymes. In MetS compared with LD Glycogen synthase 1 (GYS1)‐glycogen phosphorylases (PYGM/PYGL) expression disbalance resulted in a loss of myocardial glycogen. Our findings are consistent with the concept that transcriptionally driven myocardial changes in glycogen and glucose metabolism‐related enzymes lead to a deficiency of their metabolite products in MetS. This abnormal energy metabolism provides insight into the pathogenesis of the juvenile heart in MetS. This study reveals that MetS and ischemia diminishes ATP availability in the myocardium via altering the glucose‐G6P‐pyruvate axis at the level of metabolites and gene expression of related enzymes. The observed severe glycogen depletion in MetS coincides with disbalance in expression of GYS1 and both PYGM and PYGL. This altered energy substrate metabolism is a potential target of pharmacological agents for improving juvenile myocardial function in MetS and ischemia.

Broadwin, Mark↗

Machine Learning Model Geotiffs - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains geotiffs, supporting shapefiles and readmes for the inputs and output models of algorithms explored in the Nevada Geothermal Machine Learning project, meant to accompany the final report. Layers include: Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk), input rasters of feature sets, and positive/negative training sites. See readme .txt files and final report for additional metadata. A submission linking the full codebase for generating machine learning output models is available under "related resources" on this page.

15 GEOTHERMAL ENERGY↗

GIS Resource Compilation Map Package - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains an ESRI map package (.mpk) with an embedded geodatabase for GIS resources used or derived in the Nevada Machine Learning project, meant to accompany the final report. The package includes layer descriptions, layer grouping, and symbology. Layer groups include: new/revised datasets (paleo-geothermal features, geochemistry, geophysics, heat flow, slip and dilation, potential structures, geothermal power plants, positive and negative test sites), machine learning model input grids, machine learning models (Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk) - supervised and unsupervised), original NV Play Fairway data and models, and NV cultural/reference data. See layer descriptions for additional metadata. Smaller GIS resource packages (by category) can be found in the related datasets section of this submission. A submission linking the full codebase for generating machine learning output models is available through the "Related Datasets" link on this page, and contains results beyond the top picks present in this compilation.

15 GEOTHERMAL ENERGY↗

Randomized Algorithms for Scientific Computing (RASC)

Randomized algorithms have propelled advances in artificial intelligence (AI) and represent a foundational research area in advancing AI for Science. Future advancements in DOE Office of Science priority areas such as climate science, astrophysics, fusion, advanced materials, combustion, and quantum computing all require randomized algorithms for surmounting challenges of complexity, robustness, and scalability. Advances in data collection and numerical simulation have changed the dynamics of scientific research and motivate the need for randomized algorithms. For instance, advances in imaging technologies such as X-ray ptychography, electron microscopy, electron energy loss spectroscopy, or adaptive optics lattice light-sheet microscopy collect hyperspectral imaging and scattering data in terabytes, at breakneck speed enabled by state-of-the-art detectors. The data collection is exceptionally fast compared with its analysis. Likewise, advances in high-performance architectures have made exascale computing a reality and changed the economies of scientific computing in the process. Floating-point operations that create data are essentially free in comparison with data movement. Thus far, most approaches have focused on creating faster hardware. Ironically, this faster hardware has exacerbated the problem by making data still easier to create. Under such an onslaught, scientists often resort to heuristic deterministic sampling schemes (e.g., low-precision arithmetic, sampling every nth element) and sacrifice potentially valuable accuracy. Dramatically better results can be achieved via randomized algorithms, reducing the data size as much as or more than naive deterministic subsampling can achieve, while retaining the high accuracy of computing on the full data set. By randomized algorithms we mean those algorithms that employ some form of randomness in internal algorithmic decisions to accelerate time to solution, increase scalability, or improve reliability. Examples include matrix sketching for solving large-scale least-squares problems (see Figure 1) and stochastic gradient descent for training machine learning models. We are not recommending heuristic methods but rather randomized algorithms that have certificates of correctness and probabilistic guarantees of optimality and near-optimality. Such approaches can be useful beyond acceleration, for example, in understanding how to avoid measure zero worst-case scenarios that plague methods such as QR matrix factorization.

97 MATHEMATICS AND COMPUTING↗

Applying novel analytical tools for analyzing multidimensional secondary organic aerosol measurements

In the atmosphere, secondary organic aerosols (SOA) are often the major components of fine particulate matter and interact with clouds and radiation. SOA comprises a mixture of thousands of organic compounds. There is tremendous complexity and uncertainty in understanding SOA formation, since it is formed by oxidation and gas to particle conversion of a variety of sources: natural biogenic, anthropogenic (vehicles, cooking coal combustion) and biomass burning. The Aerosol Mass Spectrometer (AMS) produces multidimensional chemical information about SOA but analyzing this data to understand SOA sources relies on time consuming analyses (~months to years) such as the positive matrix factorization (PMF). PMF also becomes difficult for aircraft data where signal to noise ratio is weaker. There is a critical need to develop fast machine learning techniques that can analytically provide information about SOA sources using AMS data on the same timescales as the data is being collected (~minutes). We apply a machine learning supervised classification approach: the multinomial logistic regression to rapidly classify AMS data obtained from aircraft measurements.

47 OTHER INSTRUMENTATION↗

GeoThermalCloud: A Machine Learning Tool for Discovery, Exploration, and Development of Hidden Geothermal Resources

In this 25 minute presentation, we showcase our open source “GeoThermalCloud” tool for identifying hidden geothermal resources using a publicly available dataset for southwestern New Mexico. The presenters include Bulbul Ahmmed and Luke Frash. All of the visuals use source material from LA-UR approved publications and this work falls under the Earth Sciences DUSA. The code shown in this video is already released with LANL approval in open source format on GitHub and DockerHub. The audio in this video includes only material on the topics of geothermal energy and machine learning applied to geothermal energy. The primary machine learning method used is LANL’s Non-negative Matrix Factorization “NMFk” method. Modeling work also mentions LANL’s Geothermal Design Tool “GeoDT” which is another approved open source code that has been released by LANL. This work was performed for DOE Geothermal Technologies Office (DE-EE-3.1.8.1). The host for the released video is intended to be YouTube or a suitable perpetual data repository such as GDR.

15 GEOTHERMAL ENERGY↗

Domain Aware Deep-learning Algorithms Integrated with Scientific-computing Technologies (DADAIST)

This technical report summarized the contribution of the DADAIST project funded by the Data Model Convergence Initiative via the Laboratory Directed Research and Development (LDRD) investments at Pacific Northwest National Laboratory (PNNL). Specifically, we report the development of the NeuroMANCER (Neural Modules with Adaptive Nonlinear Constraints and Efficient Regularizations), a new open-source Scientific Machine Learning library for formulating and solving parametric constrained optimization problems, physics-informed system identification, and parametric optimal control problems. NeuroMANCER is using differentiable programming to combine modern data-driven models and optimization modeling language into a coherent algorithmic and software framework. NeuroMANCER is a Pytorch-based framework and adopts much of its philosophy focused on research and development, rapid prototyping, and streamlined deployment. Strong emphasis is given to extensibility, interoperability with the PyTorch ecosystem, and quick adaptability to custom domain problems. Neuromancer repository contains a comprehensive library of differentiable modules, including custom activation functions, matrix factorizations, deep learning architectures, neural differential equations, differential equation solvers, implicit layers such as iterative solvers, high-level API for symbolic expressions, API for modeling and control of dynamical systems, and extensive set of tutorial code examples in the form of python scripts and jupyter notebooks.

97 MATHEMATICS AND COMPUTING↗

Observations of ozone, acyl peroxy nitrates, and their precursors during summer 2019 at Carlsbad Caverns National Park, New Mexico

Carlsbad Caverns National Park (CAVE) is located in southeastern New Mexico and is adjacent to the Permian Basin, one of the most productive oil and natural gas (O&G) production regions in the United States. Since 2018, ozone (O 3 ) at CAVE has frequently exceeded the 70 ppbv 8-hour National Ambient Air Quality Standard. We examine the influence of regional emissions on O 3 formation using observations of O 3 , nitrogen oxides (NO x = NO + NO 2 ), a suite of volatile organic compounds (VOCs), peroxyacetyl nitrate (PAN), and peroxypropionyl nitrate (PPN). Elevated O 3 and its precursors are observed when the wind is from the southeast, the direction of the Permian Basin. We identify 13 days during the July 25 to September 5, 2019 study period when the maximum daily 8-hour average (MDA8) O 3 exceeded 65 ppbv; MDA8 O 3 exceeded 70 ppbv on 5 of these days. The results of a positive matrix factorization (PMF) analysis are used to identify and attribute source contributions of VOCs and NO x . On days when the winds are from the southeast, there are larger contributions from factors associated with primary O&G emissions; and, on high O 3 days, there is more contribution from factors associated with secondary photochemical processing of O&G emissions. The observed ratio of VOCs to NO x is consistently high throughout the study period, consistent with NO x -limited O 3 production. Finally, all high O 3 days coincide with elevated acyl peroxy nitrate abundances with PPN to PAN ratios > 0.15 ppbv ppbv -1 indicating that anthropogenic VOC precursors, and often alkanes specifically, dominate the photochemistry. Implications: The results above strongly indicate NO x -sensitive photochemistry at Carlsbad Caverns National Park indicating that reductions in NO x emissions should drive reductions in O 3 . However, the NO x -sensitivity is largely driven by emissions of NO x into a VOC-rich environment, and a high PPN:PAN ratio and its relationship to O 3 indicate substantial influence from alkanes in the regional photochemistry. Thus, simultaneous reductions in emissions of NO x and non-methane VOCs from the oil and gas sector should be considered for reducing O 3 at Carlsbad Caverns National Park. Reductions in non-methane VOCs will have the added benefit of reducing formation of other secondary pollutants and air toxics.

54 ENVIRONMENTAL SCIENCES↗

Source characterization of volatile organic compounds at Carlsbad Caverns National Park

Carlsbad Caverns National Park (CAVE), located in southeastern New Mexico, experiences elevated ground-level ozone (O 3 ) exceeding the National Ambient Air Quality Standard (NAAQS) of 70 ppbv. It is situated adjacent to the Permian Basin, one of the largest oil and gas (O&G) producing regions in the US. In 2019, the Carlsbad Caverns Air Quality Study (CarCavAQS) was conducted to examine impacts of different sources on ozone precursors, including nitrogen oxides (NO x ) and volatile organic compounds (VOCs). Here, we use positive matrix factorization (PMF) analysis of speciated VOCs to characterize VOC sources at CAVE during the study. Seven factors were identified. Three factors composed largely of alkanes and aromatics with different lifetimes were attributed to O&G development and production activities. VOCs in these factors were typical of those emitted by O&G operations. Associated residence time analyses (RTA) indicated their contributions increased in the park during periods of transport from the Permian Basin. These O&G factors were the largest contributor to VOC reactivity with hydroxyl radicals (62%). Two PMF factors were rich in photochemically generated secondary VOCs; one factor contained species with shorter atmospheric lifetimes and one with species with longer lifetimes. RTA of the secondary factors suggested impacts of O&G emissions from regions farther upwind, such as Eagle Ford Shale and Barnett Shale formations. The last two factors were attributed to alkenes likely emitted from vehicles or other combustion sources in the Permian Basin and regional background VOCs, respectively. Implications: Carlsbad Caverns National Park experiences ground-level ozone exceeding the National Ambient Air Quality Standard. Volatile organic compounds are critical precursors to ozone formation. Measurements in the Park identify oil and gas production and development activities as the major contributors to volatile organic compounds. Emissions from the adjacent Permian Basin contributed to increases in primary species that enhanced local ozone formation. Observations of photochemically generated compounds indicate that ozone was also transported from shale formations and basins farther upwind. Therefore, emission reductions of volatile organic compounds from oil and gas activities are important for mitigating elevated O 3 in the region.

54 ENVIRONMENTAL SCIENCES↗

OR22-Neuromorphic Rad Detector-PD3Ra (Final Report)

In unattended monitoring scenarios, automated radiation detection algorithms must be able to detect low signal-to-noise ratio (SNR) anomalies in a potentially dynamic and noisy background and report these anomalies in a timely fashion. Dynamic and noisy backgrounds complicate the use of simple gross-counting algorithms because they can lead to either high false positive rates or low sensitivity. Algorithms that use the entire spectrum have been the most successful in this area; notable examples are the NSCRAD algorithm developed at Pacific Northwest National Laboratory and recently the nonnegative matrix factorization approach developed at Lawrence Berkeley National Laboratory (LBNL). These approaches use either spectral regions of interest or spectral decomposition to detect threat isotopes in the background.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗