Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Quantum annealing algorithms for Boolean tensor networks

Abstract Quantum annealers manufactured by D-Wave Systems, Inc., are computational devices capable of finding high-quality heuristic solutions of NP-hard problems. In this contribution, we explore the potential and effectiveness of such quantum annealers for computing Boolean tensor networks. Tensors offer a natural way to model high-dimensional data commonplace in many scientific fields, and representing a binary tensor as a Boolean tensor network is the task of expressing a tensor containing categorical (i.e., $$\{0, 1\}$$ { 0 , 1 } ) values as a product of low dimensional binary tensors. A Boolean tensor network is computed by Boolean tensor decomposition, and it is usually not exact. The aim of such decomposition is to minimize the given distance measure between the high-dimensional input tensor and the product of lower-dimensional (usually three-dimensional) tensors and matrices representing the tensor network. In this paper, we introduce and analyze three general algorithms for Boolean tensor networks: Tucker, Tensor Train, and Hierarchical Tucker networks. The computation of a Boolean tensor network is reduced to a sequence of Boolean matrix factorizations, which we show can be expressed as a quadratic unconstrained binary optimization problem suitable for solving on a quantum annealer. By using a novel method we introduce called parallel quantum annealing, we demonstrate that Boolean tensor’s with up to millions of elements can be decomposed efficiently using a DWave 2000Q quantum annealer.

97 MATHEMATICS AND COMPUTING↗

Local structure elucidation of tungsten-substituted vanadium dioxide (V$$_{1-x}$$W$$_x$$O$$_2$$)

Abstract Initially, vanadium dioxide seems to be an ideal first-order phase transition case study due to its deceptively simple structure and composition, but upon closer inspection there are nuances to the driving mechanism of the metal-insulator transition (MIT) that are still unexplained. In this study, a local structure analysis across a bulk powder tungsten-substitution series is utilized to tease out the nuances of this first-order phase transition. A comparison of the average structure to the local structure using synchrotron x-ray diffraction and total scattering pair-distribution function methods, respectively, is discussed as well as comparison to bright field transmission electron microscopy imaging through a similar temperature-series as the local structure characterization. Extended x-ray absorption fine structure fitting of thin film data across the substitution-series is also presented and compared to bulk. Machine learning technique, non-negative matrix factorization, is applied to analyze the total scattering data. The bulk MIT is probed through magnetic susceptibility as well as differential scanning calorimetry. The findings indicate the local transition temperature ( $$T_c$$ T c ) is less than the average $$T_c$$ T c supporting the Peierls-Mott MIT mechanism, and demonstrate that in bulk powder and thin-films, increasing tungsten-substitution instigates local V-oxidation through the phase pathway VO $$_2\, \rightarrow$$ 2 → V $$_6$$ 6 O $$_{13} \, \rightarrow$$ 13 → V $$_2$$ 2 O $$_5$$ 5 .

Wilson, Catrina E. (ORCID:0000000173397318)↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Robust quantification of the diamond nitrogen-vacancy center charge state via photoluminescence spectroscopy

Nitrogen vacancy (NV) centers in diamond are at the heart of many emerging quantum technologies, all of which require control over the NV charge state. Hence, methods for quantification of the relative photoluminescence intensities of the NV 0 and NV − charge states, i.e., a charge state ratio, are vital. Several approaches to quantify NV charge state ratios have been reported but are either limited to bulk-like NV diamond samples or yield qualitative results. We propose an NV charge state quantification protocol based on the determination of sample- and experimental setup-specific NV 0 and NV − reference spectra. The approach employs blue (400–470 nm) and green (480–570 nm) excitation to infer pure NV 0 and NV − spectra, which are then used to quantify NV charge state ratios in subsequent experiments via least squares fitting. We test our dual excitation protocol (DEP) for a bulk diamond NV sample and 20 and 100 nm nanodiamond particles and compare results with those obtained via other commonly used techniques such as zero-phonon line fitting and non-negative matrix factorization. We find that DEP can be employed across different samples and experimental setups and yields consistent and quantitative results for NV charge state ratios that are in agreement with our understanding of NV photophysics. By providing robust NV charge state quantification across sample types and measurement platforms, DEP will support the development of NV-based quantum technologies.

Color center laser spectroscopy↗

Source Characterization of Volatile Organic Compounds at Carlsbad Caverns National Park

Carlsbad Caverns National Park (CAVE), located in southeastern New Mexico, experiences elevated ground-level ozone (O 3 ) exceeding the National Ambient Air Quality Standard (NAAQS) of 70 ppbv. It is situated adjacent to the Permian Basin, one of the largest oil and gas (O&G) producing regions in the US. In 2019, the Carlsbad Caverns Air Quality Study (CarCavAQS) was conducted to examine impacts of different sources on ozone precursors, including nitrogen oxides (NO x ) and volatile organic compounds (VOCs). Here, we use positive matrix factorization (PMF) analysis of speciated VOCs to characterize VOC sources at CAVE during the study. Seven factors were identified. Three factors composed largely of alkanes and aromatics with different lifetimes were attributed to O&G development and production activities. VOCs in these factors were typical of those emitted by O&G operations. Associated residence time analyses (RTA) indicated their contributions increased in the park during periods of transport from the Permian Basin. These O&G factors were the largest contributor to VOC reactivity with hydroxyl radicals (62%). Two PMF factors were rich in photochemically generated secondary VOCs; one factor contained species with shorter atmospheric lifetimes and one with species with longer lifetimes. RTA of the secondary factors suggested impacts of O&G emissions from regions farther upwind, such as Eagle Ford Shale and Barnett Shale formations. The last two factors were attributed to alkenes likely emitted from vehicles or other combustion sources in the Permian Basin and regional background VOCs, respectively.

54 ENVIRONMENTAL SCIENCES↗

Observations of ozone, acyl peroxy nitrates, and their precursors during summer 2019 at Carlsbad Caverns National Park, New Mexico

Carlsbad Caverns National Park (CAVE) is located in southeastern New Mexico and is adjacent to the Permian Basin, one of the most productive oil and natural gas (O&G) production regions in the United States. Since 2018, ozone (O 3 ) at CAVE has frequently exceeded the 70 ppbv 8-hour National Ambient Air Quality Standard. We examine the influence of regional emissions on O 3 formation using observations of O 3 , nitrogen oxides (NO x = NO + NO 2 ), a suite of volatile organic compounds (VOCs), peroxyacetyl nitrate (PAN), and peroxypropionyl nitrate (PPN). Elevated O 3 and its precursors are observed when the wind is from the southeast, the direction of the Permian Basin. We identify 13 days during the July 25 to September 5, 2019 study period when the maximum daily 8-hour average (MDA8) O 3 exceeded 65 ppbv; MDA8 O 3 exceeded 70 ppbv on 5 of these days. The results of a positive matrix factorization (PMF) analysis are used to identify and attribute source contributions of VOCs and NO x . On days when the winds are from the southeast, there are larger contributions from factors associated with primary O&G emissions; and, on high O 3 days, there is more contribution from factors associated with secondary photochemical processing of O&G emissions. The observed ratio of VOCs to NOx is consistently high throughout the study period, consistent with NO x -limited O 3 production. Finally, all high O 3 days coincide with elevated acyl peroxy nitrate abundances with PPN to PAN ratios > 0.15 ppbv ppbv -1 indicating that anthropogenic VOC precursors, and often alkanes specifically, dominate the photochemistry.

54 ENVIRONMENTAL SCIENCES↗

Source apportionment of airborne volatile organic compounds near unconventional oil and gas development

Oil and natural gas (ONG) extraction emits volatile organic compounds (VOCs). Certain VOCs are identified as hazardous air pollutants (HAPS) while others contribute to ozone formation. This study examines the impact of ONG operations on VOC levels during the development of multi-well ONG pads in suburban Broomfield, Colorado. From October 2018 to December 2020, weekly VOC measurements were taken at 18 sites across the area. These included spots near well pads, in adjacent neighborhoods, and at a background site, covering various stages of well pad development including drilling, hydraulic fracturing, flowback, and production. Analysis using Positive Matrix Factorization (PMF) identified six factors, including combustion, background/biogenic sources, light and complex alkanes, drilling activities, and ONG acetylene. Factors linked to local ONG activities exhibited clear temporal and spatial correlations with Broomfield well development. Benzene source analysis revealed distinct contribution gradients, with ONG-related sources notably influencing areas near the well pads, particularly in pre-production. ONG-related weekly benzene contributions varied from 9% to 63% at a community background site and 18% to 89% in a neighborhood close to a well pad.

54 ENVIRONMENTAL SCIENCES↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

NASMDR: a framework for miRNA-drug resistance prediction using efficient neural architecture search and graph isomorphism networks

Abstract As a frontier field of individualized therapy, microRNA (miRNA) pharmacogenomics facilitates the understanding of different individual responses to certain drugs and provides a reasonable reference for clinical treatment. However, the known drug resistance-associated miRNAs are not yet sufficient to support precision medicine. Although existing methods are effective, they all focus on modelling miRNA-drug resistance interaction graphs, making their performance bounded by the interaction density. In this study, we propose a framework for miRNA-drug resistance prediction through efficient neural architecture search and graph isomorphism networks (NASMDR). NASMDR uses attribute information instead of the commonly used interactive graph information. In the cross-validation experiment, the proposed framework can achieve an AUC of 0.9468 on the ncDR dataset, which is 2.29% higher than the state-of-the-art method. In addition, we propose a novel sequence characterization approach, k-mer Sparse Nonnegative Matrix Factorization (KSNMF). The results show that NASMDR provides novel insights for integrating efficient neural architecture search and graph isomorphic networks into a unified framework to predict drug resistance-related miRNAs. The codes for NASMDR are available at https://github.com/kaizheng-academic/NASMDR.

Zheng, Kai↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

Chemical and elemental mapping of spent nuclear fuel sections by soft X-ray spectromicroscopy

Soft X-ray spectromicroscopy at the O K -edge, U N 4,5 -edges and Ce M 4,5 -edges has been performed on focused ion beam sections of spent nuclear fuel for the first time, yielding chemical information on the sub-micrometer scale. To analyze these data, a modification to non-negative matrix factorization (NMF) was developed, in which the data are no longer required to be non-negative, but the non-negativity of the spectral components and fit coefficients is largely preserved. The modified NMF method was utilized at the O K -edge to distinguish between two components, one present in the bulk of the sample similar to UO 2 and one present at the interface of the sample which is a hyperstoichiometric UO 2+ x species. The species maps are consistent with a model of a thin layer of UO 2+ x over the entire sample, which is likely explained by oxidation after focused ion beam (FIB) sectioning. In addition to the uranium oxide bulk of the sample, Ce measurements were also performed to investigate the oxidation state of that fission product, which is the subject of considerable interest. Analysis of the Ce spectra shows that Ce is in a predominantly trivalent state, with a possible contribution from tetravalent Ce. Atom probe analysis was performed to provide confirmation of the presence and localization of Ce in the spent fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Fast Active-Set Thresholding Method for Nonnegative Least Squares

Nonnegative Least Squares (NNLS) is a fundamental constrained optimization problem encountered in many applications such as image deblurring, signal processing, nonnegative matrix factorization, magnetic microscopy, and hyperspectral imaging. Active-set based methods are a common class of algorithms for solving NNLS which identify the optimal variable set of the NNLS solution. They do so by iteratively solving a series of unconstrained least squares problems, identifying which variables violate the nonnegativity constraints, and then swapping variables in/out of consideration until the optimal set of variables is found. Several variations improving upon this method exist in the literature. In this work, we propose an active-set swap heuristic which further improves upon existing active-set based methods for NNLS. Our optimizations are based upon adding multiple variables to the passive set within a threshold of the smallest gradient value and removing variables within a similar threshold of the closest boundary constraint. We leverage these optimizations to yield a Fast Active-Set Thresholding NNLS (FAST-NNLS) algorithm which significantly outperforms the existing state-of-the-art NNLS algorithms for a wide range of problems. Rigorous convergence guarantees are proven for the proposed method. We demonstrate the effectiveness of our proposed method on multiple synthetic datasets and two realworld text analysis applications. In doing so, we present the most comprehensive NNLS solver comparison in the literature to date.

Cobb, Benjamin [Georgia Institute of Technology]↗

User Role Identification in Software Vulnerability Discussions over Social Networks

Understanding and early awareness of software vulnerabilities is vital for preventing and mitigating potential impacts from cybersecurity events. One step toward early characterization of software vulnerabilities may involve analyzing discussion and spread of information in online social networks. Prior work has used information from such discussions over multiple online forums to develop dynamic networks among users followed by analysis of structure, spread, and information evolution. In this work, we advance the state-of-the-art by focusing on data-driven learning of types, roles, and transition of roles exhibited by users over time. In social networks, users take on particular roles based on their actions and structure of the network. Identifying “meaningful” roles can help separate potential users of interest from the larger community, and identify patterns in a network. We will identify and compare roles found in online forums (e.g., Twitter) using techniques such as feature-based Non-negative Matrix Factorization coupled with topological and influence-based measures of centrality. Since users’ activities change over time, we also analyze role evolution in dynamic networks.

Jones, Rebecca D.↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗