Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PARAFAC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

PARAFAC_T1

This is a method of performing trilinear analysis on large data sets using a modification of the PARAFAC-ALS algorithm. It iteratively decomposes the data matrix into a core matrix and three loading matrices based on the Tucker1 model. The algorithm is particularly useful for data sets that are too large to upload into a computer?s main memory. While the performance advantage in utilizing our algorithm is dependent on the number of data elements and dimensions of the data array, we have seen a significant performance improvement over operating PARAFAC-ALS on the full data set. SAND2020-12649 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Van Benthem, Mark↗

Investigating the impacts of solid phase extraction on dissolved organic matter optical signatures and the pairing with high‐resolution mass spectrometry data across a freshwater stream network

Abstract Advancing our understanding of dissolved organic matter (DOM) chemistry in aquatic systems necessitates the integration of data streams from multiple analytical platforms. Some measurements require pretreatment with solid phase extraction (SPE), while others are performed directly on whole water samples. Evidence has suggested that SPE will be biased against select DOM fractions, leading to concerns over the ability to establish data linkages across platforms with variable needs for SPE pretreatment, such as those from optical measurements and those that provide high‐resolution molecular information. Here, we directly addressed this concern by assessing the impact of SPE on DOM optical properties through excitation–emission matrices with parallel factor analysis (PARAFAC) for 47 samples across a stream network within a single watershed reflective of variable DOM sources. PARAFAC data was further paired with molecular information obtained by Fourier transform ion cyclotron resonance mass spectrometry (FTICR‐MS). A comparison of PARAFAC models first revealed no systematic qualitative differences in major components between whole water DOM and DOM isolated by SPE (SPE‐DOM); however, quantitative biases against select components were observed. Further linkages with FTICR‐MS data revealed that the molecular fingerprint associated with each PARAFAC component was consistent between the whole water DOM and SPE‐DOM. Our results suggest that bulk scale linkages across these analytical platforms could be inferred irrespective of the observed quantitative biases resulting from SPE for samples within this example watershed. This work represents a key step toward the systematic evaluation of linkages between optical and high‐resolution mass spectrometry datasets in freshwater lotic environments.

59 BASIC BIOLOGICAL SCIENCES↗

Data and scripts associated with a manuscript investigating impacts of solid phase extraction on freshwater organic matter optical signatures and mass spectrometry pairing

This data package is associated with the publication “Investigating the impacts of solid phase extraction on dissolved organic matter optical signatures and the pairing with high-resolution mass spectrometry data in a freshwater system” submitted to “Limnology and Oceanography: Methods.” This data is an extension of the River Corridor and Watershed Biogeochemistry SFA’s Spatial Study 2021 (https://doi.org/10.15485/1898914). Other associated data and field metadata can be found at the link provided. The goal of this manuscript is to assess the impact of solid phase extraction (SPE) on the ability to pair ultra-high resolution mass spectrometry data collected from SPE extracts with optical properties collected on ambient stream samples. Forty-seven samples collected from within the Yakima River Basin, Washington were analyzed dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), absorbance, and fluorescence. Samples were subsequently concentrated with SPE and reanalyzed for each measurement. The extraction efficiency for the DOC and common optical indices were calculated. In addition, SPE samples were subject to ultra-high resolution mass spectrometry and compared with the ambient and SPE generated optical data. Finally, in addition to this cross-platform inter-comparison, we further performed and intra-comparison among the high-resolution mass spectrometry data to determine the impact of sample preparation on the interpretability of results. Here, the SPE samples were prepared at 40 milligrams per liter (mg/L) based on the known DOC extraction efficiency of the samples (ranging from ~30 to ~75%) compared to the common practice of assuming the DOC extraction efficiency of freshwater samples at 60%. This data package folder consists of one main data folder with one subfolder (Data_Input). The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) final data summary output from processing script; and (5) the processing script. The R-markdown processing script (SPE_Manuscript_Rmarkdown_Data_Package.rmd) contains all code needed to reproduce manuscript statistics and figures (with the exception of that stated below). The Data_Input folder has two subfolders: (1) FTICR and (2) Optics. Additionally, the Data_Input folder contains dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data (SPS_NPOC_Summary.csv) and relevant supporting Solid Phase Extraction Volume information (SPS_SPE_Volumes.csv). Methods information for the optical and FTICR data is embedded in the header rows of SPS_EEMs_Methods.csv and SPS_FTICR_Methods.csv, respectively. In addition, the data dictionary (SPS_SPE_dd.csv), file level metadata (SPS_SPE_flmd.csv), and methods codes (SPS_SPE_Methods_codes.csv) are provided. The FTICR subfolder contains all raw FTICR data as well as instructions for processing. In addition, post processed FTICR molecular information (Processed_FTICRMS_Mol.csv) and sample data (Processed_FTICRMS_Data.csv) is provided that can be directly read into R with the associated R-markdown file. The Optics subfolder contains all Absorbance and Fluorescence Spectra. Fluorescence spectra have been blank corrected, inner filter corrected, and undergone scatter removal. In addition, this folder contains Matlab code used to make a portion of Figure 1 within the manuscript, derive various spectral parameters used within the manuscript, and used for parallel factor analysis (PARAFAC) modeling. Spectral indices (SPS_SpectralIndices.csv) and PARAFAC outputs (SPS_PARAFAC_Model_Loadings.csv and SPS_PARAFAC_Sample_Scores.csv) are directly read into the associated R-markdown file. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Active Thermography Based on Tensor Rank Decomposition

Principal Component Thermography applies Singular Value Decomposition (SVD) to post-process data that are derived from active thermographic inspections. SVD provides useful compression of the data and allows for better understanding of substructure and indications of potential damage. In the standard approach, SVD is applied to a certain reshaping of a three-dimensional data stack into a two-dimensional array. This work applies the CANDECOMP-PARAFAC (CP) tensor rank decomposition directly to the three-dimensional data to avoid the initial reshaping step in order to begin to develop an inspection method that can more accurately detect defects in non-homogeneous and anisotropic materials. Tests against simulated data that compare the CP decomposition method with traditional Principal Component Thermography based on SVD are described. Finally, the method of Proper Generalized Decomposition (PGD) is used to derive the CP decomposition, and its performance against other algorithms is also discussed.

Thermography↗

Organic matter distribution in the icy environments of Taylor Valley, Antarctica

Glaciers can accumulate and release organic matter affecting the structure and function of associated terrestrial and aquatic ecosystems. Here we analyzed 18 ice cores collected from six locations in Taylor Valley (McMurdo Dry Valleys), Antarctica to determine the spatial abundance and quality of organic matter, and the spatial distribution of bacterial density and community structure from the terminus of the Taylor Glacier to the coast (McMurdo Sound). Our results showed that dissolved and particulate organic carbon (DOC and POC) concentrations in the ice core samples increased from the Taylor Glacier to McMurdo Sound, a pattern also shown by bacterial cell density. Fluorescence Excitation Emission Matrices Spectroscopy (EEMs) and multivariate parallel factor (PARAFAC) modeling identified one humic-like (C1) and one protein-like (C2) component in ice cores whose fluorescent intensities all increased from the Polar Plateau to the coast. The fluorescence index showed that the bioavailability of dissolved organic matter (DOM) also decreased from the Polar Plateau to the coast. Partial least squares path modeling analysis revealed that bacterial abundance was the main positive biotic factor influencing both the quantity and quality of organic matter. Marine aerosol influenced the spatial distribution of DOC more than katabatic winds in the ice cores. Certain bacterial taxa showed significant correlations with DOC and POC concentrations. Collectively, our results show the tight connectivity among organic matter spatial distribution, bacterial abundance and meteorology in the McMurdo Dry Valley ecosystem.

54 ENVIRONMENTAL SCIENCES↗

Manuscript Workflows from and Processed Organic Matter Composition of Experimentally Burned Open Air and Muffle Furnace Vegetation Chars across Differing Burn Severity and Feedstock Types from Pacific Northwest, USA (v3)

This dataset includes processed organic matter chemistry data from an experimental study designed to compare how the chemical composition of organic matter changes across different burn conditions and vegetation materials representative of major land cover types of the Pacific Northwest, USA. Chars were created in a closed muffle furnace or on an open burn table from four different feedstock species representing vegetation commonly impacted by fire regimes across the Pacific Northwest, USA. Source data and associated metadata (including methods and geospatial information) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1894135 (Grieger et al. 2022). This dataset provides processing scripts and processed data for both solid and dissolved phase organic matter characterization data from experimentally generated chars. These processed data can be used to compare how different burn conditions may influence resultant organic matter chemistry and help further our understanding of potential biogeochemical impacts on river corridors post-fire. The processed data were subsequently analyzed; and the results and ecological implications of the findings were published in peer-reviewed manuscripts. The scripts and workflows used to develop the manuscripts are also included in this data package.This data package was originally published June 2024. It was updated September 2024 (new and modified files) and in January 2025 (modified files). See the change history section in the readme for more details.This dataset is comprised of one data package readme, one data dictionary (dd), one file level metadata (flmd), and folders containing (A) processed data; (B) general processing scripts; and (C) additional folders with specific manuscript analysis scripts and processed data. Step-by-step instructions to assist the user in recreating the workflow used to generate the results in the manuscripts is also provided. The processed data folder includes (1) a folder of processed Parallel Factor Analysis (PARAFAC) and spectra indices outputs from excitation emissions matrix (EEM) fluorescence and absorbance data; (2) a folder of processed solid state carbon-13 (13-C NMR) integrals; (3) folder of high resolution characterization of organic matter via 21 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory) processed data outputs from Formultitude (https://github.com/PNNL-Comp-Mass-Spec/Formultitude), blank corrections and data aggregation, and calculated molecular indices. All files are .pdf, .csv, .html, .Rmd, .R, or .RData.

54 ENVIRONMENTAL SCIENCES↗

Data for Kim et al., "Variations in the optical and molecular composition of dissolved organic matter exported from coastal wetlands"

Knowledge about sources and composition of marsh-derived dissolved organic matter (DOM) is critical for understanding the role of marshes in coastal biogeochemical cycling and the fate of marsh-derived DOM in the ocean. To investigate tidal variability in composition of marsh-derived DOM, Kim et al. examined the optical and molecular characteristics of hourly surface water samples at three tidal creeks in the Chesapeake Bay. Groundwater samples along the terrestrial landscape gradient as well as estuarine water from the adjacent estuary at each site were also collected to help resolve sources of surface water DOM. Samples were collected in summer 2024 at three sites – SWH: Sweet Hall Marsh, GCW: Kirkpatrick Marsh, and GWI: Goodwin Islands – which are part of synoptic sites in the Chesapeake Bay region of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales - Field, Measurements, and Experiments) project. Surface water samples were collected hourly over a 48-hour period at each site. Groundwater and estuarine water samples were collected once. This dataset includes- Surface water depth and salinity- Dissolved organic carbon (DOC) and total dissolved nitrogen (TDN) concentrations- Optical indices and relative composition of parallel factor analysis (PARAFAC) components- High resolution mass spectrometry data.

54 ENVIRONMENTAL SCIENCES↗

Getting to the Core of PARAFAC2, A Nonnegative Approach

In this paper, the authors present a novel method of performing PARAFAC2 factorization of three-way data using a compact representation of that data. In the standard PARAFAC2 algorithm, two modes of the data are recovered directly during the decomposition while the third mode is returned as a transformation matrix, which is then used to rotate sets of orthogonal third-mode basis factors into interpretable factors. In our new method, the data are first decomposed into a core matrix and orthogonal factor loading matrices in the first two modes as well as sets of orthogonal factors in the third mode (as in standard PARAFAC2). The core matrix is then decomposed using a the standard PARAFAC2 strategy to produce transformation matrices in all three modes. The algorithm is particularly useful for very large data sets and essentially permits imposition of nonnegativity in all three modes.

97 MATHEMATICS AND COMPUTING↗

Linking Hydrology and Dissolved Organic Matter Characteristics in a Subtropical Wetland: A Long-Term Study of the Florida Everglades

Dissolved organic matter (DOM) acts as an important biogeochemical component of aquatic ecosystems that controls nutrient cycling, influences water quality, and links terrestrial and oceanic carbon pools, yet long-term studies of how changing environmental drivers alter its abundance and composition are rare. Using a ten-year, spatially explicit dataset from Everglades National Park (ENP), a globally significant wetland, we investigated the relationships between DOM quality/quantity and hydrologic/climatic drivers along two contrasting marsh-estuarine transects based on generalized linear modeling and a cumulative sums analysis. Analyses revealed distinct spatial, seasonal, and interannual patterns in variability of DOC and optical properties. Landscape-scale seasonal patterns showed an enrichment in microbial-like and protein-like DOM during the dry season relative to the wet season. We found that, while some compositional constituents varied with the solar calendar, responsive to temperature and photoperiod, others varied with the hydrologic calendar. Independent water level and discharge effects indicated strong hydrologic control on DOM quality that differed between the two transects, evidencing differences in their connectivity to areas of high agricultural activity. Across all sites, a significant long-term increasing trend in the fluorescence index was observed, associated with a positive correlation with precipitation and also potential changes in agricultural inputs, with other features associated with drought and hurricanes. Lastly, the cumulative sums analysis revealed differences between the two transects in the sensitivity of DOM composition to decreased water levels associated with 30-year climate scenarios, with the less hydrologically dynamic transect exhibiting greater potential sensitivity.

54 ENVIRONMENTAL SCIENCES↗

Large Scale Tensor Factorization via Parallel Sketches

Tensor factorization methods have recently gained increased popularity. A key feature that renders tensors attractive is the ability to directly model multi-relational data. In this work, we propose ParaSketch, a parallel tensor factorization algorithm that enables massive parallelism, to deal with large tensors. The idea is to compress the large tensor into multiple small tensors, decompose each small tensor in parallel, and combine the results to reconstruct the desired latent factors. Prior art in this direction entails potentially very high complexity in the (Gaussian) compression and final combining stages. Adopting sketching matrices for compression, the proposed method enjoys a dramatic reduction in compression complexity, and features a much lighter combining step. Moreover, theoretical analysis shows that the compressed tensors inherit latent identifiability under mild conditions, hence establishing correctness of the overall approach. Numerical experiments corroborate the theory and demonstrate the effectiveness of the proposed algorithm.

block term decomposition↗

Generalized Canonical Polyadic Tensor Decomposition

Tensor decomposition is a fundamental unsupervised machine learning method in data science, with applications including network analysis and sensor data processing. This work develops a generalized canonical polyadic (GCP) low-rank tensor decomposition that allows other loss functions besides squared error. For instance, we can use logistic loss or Kullback--Leibler divergence, enabling tensor decomposition for binary or count data. We present a variety of statistically motivated loss functions for various scenarios. We provide a generalized framework for computing gradients and handling missing data that enables the use of standard optimization methods for fitting the model. Furthermore, we demonstrate the flexibility of the GCP decomposition on several real-world examples including interactions in a social network, neural activity in a mouse, and monthly rainfall measurements in India.

97 MATHEMATICS AND COMPUTING↗