Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “FTICR”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data and scripts associated with “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” (v2)

This data package is associated with the publication “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” submitted to Biogeochemistry by Ryan et al., 2024 (DOI: https://doi.org/10.1007/s10533-024-01169-5). This study aims to investigate fundamental and transferable drivers of dissolved organic matter (DOM) diversity across five nested watersheds within the contiguous United States. DOM diversity was explored using ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). The samples and the unprocessed FTICR-MS data used in this study are publicly available on the Environmental System Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository (see DOIs below). The data for the Willamette, Gunnison, Connecticut, and Deschutes basins were collected as part of a collaboration between the Watershed Rules of Life (WROL) project and Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS). The data for the Yakima River basin (YRB) was collected by the PNNL River Corridor SFA. The raw, unprocessed FTICR-MS data with additional (meta)data can be found at doi:10.15485/1895159 for WROL samples and doi:10.15485/1898912 for YRB samples. This data package contains the processed data used in the associated manuscript. This package also contains ancillary geospatial, hydrological, and geochemical information that supports the interpretation of the FTICR-MS data within Ryan et al., 2024. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/rcsfa-RC4-WROL-YRB_DOM_Diversity. This data package was originally published August 2024. It was updated January 2025 (modified files). See the change history in the readme more details. At the directory level, the data package is comprised of three folders: (1) data, (2) output, and (3) src; and five additional files including the data dictionary (file ending in "_dd.csv”) and file-level metadata (file ending in “_flmd.csv”). The “src” folder contains the scripts used to process the FTICR data, conduct the analyses, and produce the manuscript figures. The inputs for these scripts are in the “data” folder and the returned outputs in the “output” folder. Inputs include temporal and spatial metadata associated with the sampling efforts, processed FTICR data, and total and normalized putative biochemical transformations per sample. Outputs include cleaned and combined data presented as tables, descriptive statistics, and plots. The file-level metadata file lists all files contained in this data package and descriptions for each. The data dictionary describes the units and definitions for each tabular data column or row header.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript investigating impacts of solid phase extraction on freshwater organic matter optical signatures and mass spectrometry pairing

This data package is associated with the publication “Investigating the impacts of solid phase extraction on dissolved organic matter optical signatures and the pairing with high-resolution mass spectrometry data in a freshwater system” submitted to “Limnology and Oceanography: Methods.” This data is an extension of the River Corridor and Watershed Biogeochemistry SFA’s Spatial Study 2021 (https://doi.org/10.15485/1898914). Other associated data and field metadata can be found at the link provided. The goal of this manuscript is to assess the impact of solid phase extraction (SPE) on the ability to pair ultra-high resolution mass spectrometry data collected from SPE extracts with optical properties collected on ambient stream samples. Forty-seven samples collected from within the Yakima River Basin, Washington were analyzed dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), absorbance, and fluorescence. Samples were subsequently concentrated with SPE and reanalyzed for each measurement. The extraction efficiency for the DOC and common optical indices were calculated. In addition, SPE samples were subject to ultra-high resolution mass spectrometry and compared with the ambient and SPE generated optical data. Finally, in addition to this cross-platform inter-comparison, we further performed and intra-comparison among the high-resolution mass spectrometry data to determine the impact of sample preparation on the interpretability of results. Here, the SPE samples were prepared at 40 milligrams per liter (mg/L) based on the known DOC extraction efficiency of the samples (ranging from ~30 to ~75%) compared to the common practice of assuming the DOC extraction efficiency of freshwater samples at 60%. This data package folder consists of one main data folder with one subfolder (Data_Input). The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) final data summary output from processing script; and (5) the processing script. The R-markdown processing script (SPE_Manuscript_Rmarkdown_Data_Package.rmd) contains all code needed to reproduce manuscript statistics and figures (with the exception of that stated below). The Data_Input folder has two subfolders: (1) FTICR and (2) Optics. Additionally, the Data_Input folder contains dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data (SPS_NPOC_Summary.csv) and relevant supporting Solid Phase Extraction Volume information (SPS_SPE_Volumes.csv). Methods information for the optical and FTICR data is embedded in the header rows of SPS_EEMs_Methods.csv and SPS_FTICR_Methods.csv, respectively. In addition, the data dictionary (SPS_SPE_dd.csv), file level metadata (SPS_SPE_flmd.csv), and methods codes (SPS_SPE_Methods_codes.csv) are provided. The FTICR subfolder contains all raw FTICR data as well as instructions for processing. In addition, post processed FTICR molecular information (Processed_FTICRMS_Mol.csv) and sample data (Processed_FTICRMS_Data.csv) is provided that can be directly read into R with the associated R-markdown file. The Optics subfolder contains all Absorbance and Fluorescence Spectra. Fluorescence spectra have been blank corrected, inner filter corrected, and undergone scatter removal. In addition, this folder contains Matlab code used to make a portion of Figure 1 within the manuscript, derive various spectral parameters used within the manuscript, and used for parallel factor analysis (PARAFAC) modeling. Spectral indices (SPS_SpectralIndices.csv) and PARAFAC outputs (SPS_PARAFAC_Model_Loadings.csv and SPS_PARAFAC_Sample_Scores.csv) are directly read into the associated R-markdown file. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript investigating dissolved organic matter and microbial community linkages across seven globally distributed rivers

This data package is associated with the publication “Meta-metabolome ecology reveals that geochemistry and microbial functional potential are linked to organic matter development across seven rivers” submitted to Science of the Total Environment. This data package includes the data necessary to replicate the analyses presented within the manuscript to investigate dissolved organic matter (DOM) development across broad spatial distances and within divergent biomes. Specifically, we included the Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data, geochemistry data, annotated metagenomic data, and results from ecological null modeling analyses in this data package. Additionally, we included the scripts necessary to generate the figures from the manuscript. Complete metagenomic data associated with this data package can be found at the National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. This dataset consists of (1) four folders; (2) a file-level metadata (flmd) file; (3) a data dictionary (dd) file; (4) a factor sheet describing samples; and (5) a readme. The FTICR Data folder contains (1) the processed Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data; (2) a transformation-weighted characteristics dendrogram generated from the FTICR-MS data; and (3) the script used to generate all FTICR-MS related figures. The Geochemical Data folder contains (1) the single geochemistry data file and (2) the R script responsible for generating associated figures. The Metagenomic Data folder contains (1) annotation information across different levels; (2) carbohydrate active enzyme (CAZyme) information from the dbCAN database (Yin et al., 2012); (3) phylogenetic tree data (FASTAs, alignments, and tree file); and (4) the scripts necessary to analyze all of these data and generate figures. The Null Modeling Data folder contains (1) data generated during null modeling for each river and all rivers combined and (2) the R scripts necessary to process the data. All files are .csv, .pdf, .tsv, .tre, .faa, .afa, .tree, or .R.

54 ENVIRONMENTAL SCIENCES↗

Surface Water Disinfection Byproducts and Organic Matter Characterization Data Associated with: “Disinfection byproducts formed during drinking water treatment reveal an export control point for dissolved organic matter in a subalpine headwater stream”

This dataset is associated with the publication “Disinfection byproducts formed during drinking water treatment reveal an export control point for dissolved organic matter in a subalpine headwater stream” published in Water Research X (Leonard et al. 2022; https://doi.org/10.1016/j.wroa.2022.100144). The associated study analyzed temporal trends from the Town of Crested Butte water treatment facility and synoptic sampling at Coal Creek in Crested Butte, Colorado, US. This work demonstrates how drinking water quality archives combined with synoptic sampling and targeted analyses can be used to identify and understand export control points for dissolved organic matter. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) metadata and international geo-sample number (IGSN) mapping file; (4) dissolved organic carbon (DOC), ultraviolet absorbance at 254 nanometers (UV254), total nitrogen (TN), and specific ultraviolet absorbance (SUVA) data; (5) disinfection bioproduct formation potential (DBP-FP) data; (6) readme; (7) methods codes; (8) water collection protocol; (9) folder of high resolution characterization of organic matter via 12 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory); and (10) folder of excitation emissions matrix (EEM) spectra. The FTICR folder contains a file of DOC (measured as non-purgeable organic carbon; NPOC) used for FTICR sample preparation. The FTICR folder also contains three subfolders: one subfolder containing the raw .xml data files, one containing the processed data, and the other containing instructions for using Formularity (https://omics.pnl.gov/software/formularity) and an R script to process the data based on the user's specific needs. All files are .csv, .pdf, .dat, .R, .ref, or .xml

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts associated with “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling.”

This data package is associated with the publication “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling” submitted to Geoscientific Model Development (Muller et al., 2024). In this manuscript, organic matter chemistry and thermodynamics are directly connected to reactive transport simulators through the newly developed Lambda-PFLOTRAN (Parallel Reactive Flow and Transport model) workflow tool that succinctly incorporates organic matter chemistry data generated from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) into reaction networks to simulate aerobic respiration of the organic matter and the resulting biogeochemistry. Lambda-PFLOTRAN is a python-based workflow, executed through a Jupyter Notebook interface, that digests raw FTICR-MS data, develops a representative reaction network based on substrate-explicit thermodynamic modeling (also termed lambda modeling due to its key thermodynamic parameter λ used therein), and completes a biogeochemical simulation with the open source, reactive flow, and transport code PFLOTRAN. This data package contains Jupyter Notebook based workflows for two test cases for running biogeochemical simulations of organic matter oxidation identified by FTICR-MS. It contains four primary folders (workflow, data, src, and analysis), a file-level metadata file (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_flmd.csv) that lists all the files contained in this data package with a short description of each, and a data dictionary (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_dd.csv) file that describes the tabular column headers. The ‘workflow’ folder contains the Jupyter Notebook based workflows for running the lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘data’ folder contains the FTICR-MS data, initial conditions, and incubation data for test cases 1 and 2 in folders titled ‘WHONDRS’ and ‘Colloids’, respectively. The data folder also has a ‘Database’ folder containing a reaction network for bulk organic matter (assumed to be CH2O) and a general database for PFLOTRAN (hanford_rxn_network). The CH2O reaction network defines bulk organic matter oxidation. Biogeochemical simulations are completed for both the lambda binned organic matter and bulk organic matter reaction networks. The ‘hanford_rxn_network’ database includes information required for PFLTORAN simulations including ion size, molar mass, and charge of the aqueous species, gases, and minerals phases. The ‘src’ folder contains python source codes for performing lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘analysis’ folder contains outputs from the test cases 1 and 2 including lambda analysis, PFLOTRAN runs and the calibration results.

54 ENVIRONMENTAL SCIENCES↗

Water chemistry in flume channel and hyporheic zone (i.e., porewater) associated with: “Rethinking Aerobic Respiration in the Hyporheic Zone Under Variation in Carbon and Nitrogen Stoichiometry”

Dissolved oxygen (DO), total organic carbon (TOC), total nitrogen (TN), molecular data for organic matter, and biochemical reactions for surface water and porewater (i.e., hyporheic zone) collected from a water recirculating flume located at the University of Texas, Austin. The flume contained real river water from Lower Colorado River(Austin, TX) and clean sand. Hyporheic exchange in the flume was induced through The study aims to understand relationships between aerobic metabolism of organic matter and molecular characteristics of organic matter, such as thermodynamic signature and nitrogen content, through the extent of the hyporheic zone at 10 cm- resolution, and through time. During the experiment, organic matter (dry leaves) was added to the flume and removed after 24 hours. The water samples were collected before the addition of leaves, at the time of removal of leaves, and at hour 72. The water samples were analyzed using ultrahigh resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) and total organic carbon (TOC) and total nitrogen (TN) analysis. Dissolved oxygen content throughout the surface water and the hyporheic zone of the flume was measured with a large planar optode. This data package is associated with the publication ’ Rethinking Aerobic Respiration in the Hyporheic Zone Under Variation in Carbon and Nitrogen Stoichiometry’ published in Environmental Science and Technology (Turețcaia et al., 2023 https://doi.org/10.1021/acs.est.3c04765). The dataset is comprised of five folders (1) Diss_O2_pic, (2) input_files (3) output_files; (4) python_code; and (5) R_code . Diss_O2_pic contains siximages of dissolved oxygen distribution in a bedform at hours 0, 24, and 72 of the experiment conducted in a large recirculation flume. Images are in separate R and G channels (i.e., RGB). The input_files contains (1) a csv file with FTICR peaks identified within each sample, (2) a csv file with molecular information pertinent to FTICR data with Gibbs free energy calculations adjusted for environmental temperature, (3) a csv file containing concentrations of non-purgeable organic carbon measured throughout the experiment , (4) a csv file containing concentrations of total nitrogen measured throughout the experiment, (5) a csv file containing total biochemical reactions (i.e., transformations) identified in the dataset, (6) a csv containing transformation profiles, and (7) a csv file containing transformations with formulas, and (8) a jpg file with schematic representation of locations for sample collection. The output_files contains (1) and xlsx file containing percent biochemical reactions containing nitrogen identified across all 39 sample, (2) a csv file of merged FTICR data and molecular information files, (3) a csv files containing average Gibbs free energy within sampling domains and at each sampling location, (4) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 0, (5) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 24, (6) a csv file with average concentrations of dissolved oxygen across sampling locations at hour 72, (7) a csv file with percent chemical classes identified across sampling locations at hour 0, (8) a csv file with percent chemical classes identified across sampling locations at hour 24, (9) a csv file with percent chemical classes identified across sampling locations at hour 72, and (10) a csv file containing percent nitrogen containing biochemical reactions identified across sampling locations at hours 0, 24, and 72. The python_code contains seven ipynb files which are Jupyter Notebooks used for data analysis and figures generation. The R_code contains 3 R files with R code used for data analysis and figures generation. This data package contains the processed data used in the associated manuscript. This data has not been previously published.

54 ENVIRONMENTAL SCIENCES↗

Organic Matter Concentration and Composition in November 2021 and April 2022 from 12 Streams Impacted by the 2020 Holiday Farm Fire (v2)

This dataset represents results from a field study aiming to understand storm induced transport of pyrogenic materials to streams impacted by varying degrees of burn severity. Time series samples were collected at 5 sites within the McKenzie River Watershed (Oregon, USA) whose catchment were each completely engulfed by the 2020 Holiday Farm Fire. An additional 7 sites were sampled once during the storm. The samples were collected during storm events in November 2020, January 2021, November 2021, and April 2022. Samples were characterized for benezenepolycarboxylic acids (BPCA), ultra-high resolution mass spectrometry, dissolved organic carbon and optics (absorbance and fluorescence). Fourier-transform ion cyclotron resonance mass spectrometry (FTICR) and dissolved organic carbon data from the November 2020 (referred to as “EWEB_2020”) sampling can be found in a separate data package (doi: 10.15485/1869708). NOTE: The 2020 samples were run on FTICR-MS in two unique instances. The first run can be found in the previous data package (EWEB_2020). The second run is included in this data package. These samples were run for a second time so that the data were more directly interoperable with the other samples in this data package. We have not done any investigation into the differences/similarities between these datasets and the previously ran/published data in the other data package. This data package was originally published in November 2024. It was updated in April 2025 (v2; new and modified files). See the change history section below for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset contains (1) file-level metadata; (2) data dictionary; (3) data package readme; (4) metadata; (5) methods information; (6) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data; (7) excitation emission matrix (EEM) methods; and (8) a sub-folder with processed EEM data (9) benzene polycarboxylic acid (BPCA) concentration data; (10) Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) methods; and (11) folder of high-resolution characterization of organic matter via 12 Tesla FTICR-MS generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The EEMs sub-folder contains two additional folders; the Absorbance and Fluorescence folders which contain the processed EEMs absorbance and fluorescence data respectively. This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from 7 Perennial and 7 Intermittent Streams across San Antonio, Texas (v3)

This dataset supports a broader study examining the effects of intermittency on sediment respiration. The dataset provides sediment and surface water geochemistry and in situ sensor data from 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). Related data were collected and will be published separately in collaboration with A. Veach. The data package was originally published in April 2025. It was updated in June 2025 (v2; modified and new files) and September 2025 (v3; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) sediment grain size data; (4) sediment iron (II) data and averages; (5) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment percent carbon and nitrogen; (11) sediment X-ray diffraction (XRD) data; (12) gravimetric moisture and averages; (13) a subfolder with sediment incubation respiration data, scripts, and plots; (14) surface water and sediment FTICR methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: The data processing methods for FTICR described in “v3_WHONDRS_AV1_Methods_Codes.csv” mistakenly indicate that users should process the data in Formultitude. The corrected description should read: “Both unprocessed and processed data are provided to allow users flexibility in data processing. Instructions and scripts for processing the data using CoreMS are included.” CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package.

54 ENVIRONMENTAL SCIENCES↗

Post-fire time series of sensor and geochemistry sample data from surface water, groundwater, precipitation, soil, and vegetation across Oak Creek watershed, Washington

This dataset supports a broader study examining wildfire impacts on hydrologic connectivity across 5 sites within the Oak Creek watershed and the resulting biogeochemical impacts. Stream sites were selected using the Advanced Terrestrial Simulator (ATS) hydrologic model to identify locations with varying groundwater contributions and hydrologic responses across different burn severity scenarios. The Retreat Fire burned from July 23 to August 2 in 2024, affecting the five study sites at varying burn severities. Each site is equipped with YSI EXO2 sondes logging sub-hourly throughout the year, and grab samples are collected approximately every six weeks. YSI sondes are used to measure temporally resolved proxies for groundwater inputs (specific conductivity) and organic matter (fluorescent dissolved organic matter; fDOM) along with basic water quality and depth. Grab samples of surface water, groundwater, and precipitation are analyzed for water stable isotopes and conductivity to understand endmembers for hydrologic mixing Grab samples of surface water, groundwater, soil water, and litter/vegetation/soil leachates are analyzed for organic matter composition measured by Fourier-Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) to understand organic matter dynamics. Game camera photos are provided in a separate data package available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3018598. Future versions of this dataset will include time series data from YSI EXO2 sondes (fDOM, dissolved oxygen, temperature, depth, specific conductance, turbidity, pH), BaroTROLL sensors (air temperature and barometric pressure), rain gauges (precipitation), and data from the soil and vegetation samples. Because this study is ongoing, this data package will be updated regularly to include newly collected data and the additional data types. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data; (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) a data checks report; (5) file-level metadata; (6) data dictionary; (7) field metadata; (8) readme; (9) international generic sample number (IGSN) mapping file; and (10) field protocols. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) stable water isotopes and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Biogeochemistry↗

WHONDRS River Corridor Surface Water Metabolites and Geochemistry from Global Sites

This dataset supports a broader study examining the character of organic matter that may be delivered to subsurface sediments via hydrologic exchange. To implement the global survey, free stream sampling kits were provided to interested volunteers throughout the world. Samples were collected with minimal constraints in terms of location, but following strict protocols, and shipped for metabolomic analysis via Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). In addition, basic geochemistry analyses (e.g., dissolved organic matter concentration) were conducted, standardized photos of each field system were taken, and extensive metadata were captured. Sampling began in 2018 and is ongoing as of 2025. This dataset is comprised of one folders of field photos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; and (7) a subfolder with sample data. The sample data subfolder contains (1) surface water dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) methods codes; (3) surface water FTICR methods; and (4) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains three subfolders, one containing the.xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, or .png. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

Biogeochemistry↗

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions↗

Organic matter concentration and composition of experimentally burned open air and muffle furnace vegetation chars across differing burn severity and feedstock types from Pacific Northwest, USA (v4).

This dataset represents results from an experimental study designed to compare how the chemical composition of organic matter changes across different burn conditions and feedstock materials. The dataset provides both solid and dissolved phase bulk concentration and organic matter characterization data from experimentally generated chars. Chars were created in a closed muffle furnace or on an open burn table from four different feedstock species representing vegetation commonly impacted by fire regimes across the Pacific Northwest, USA. This data can be used to compare how different burn conditions may influence resultant organic matter chemistry and help further our understanding of potential biogeochemical impacts on river corridors post-fire. This dataset is comprised of one data package readme, one data dictionary (dd), one file level metadata (flmd), fourteen burn table videos, burn table video metadata and three folders containing (A) data; (B) metadata and protocols; and (C) photos. The folder names and the file name of the data package readme include a version number which will be updated with future iterations of this data package. The data folder includes (1) solid carbon and solid nitrogen; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) and total dissolved nitrogen (TN); (3) pH; (4) thermocouple time series temperature; (5) methods codes; (6) installation methods; (7) excitation emissions matrix (EEM) methods information; (8) a folder of excitation emissions matrix (EEM) fluorescence and absorbance spectra in dissolved organic matter and EEMs processing instructions; (9) solid state carbon-13 and solution state phosphorus nuclear magnetic resonance (13-C NMR and 31-P NMR) data and methods; (10) benzene polycarboxylic acid (BPCA) concentration and stable isotope data; (11) FTICR-MS methods; (12) Inductively coupled plasma (ICP) data for total calcium, magnesium, iron, aluminum, potassium, phosphorus, sodium, and sulfur along with sodium hydroxide-ethylenediaminetetraacetic acid (sodium hydroxide-EDTA) extractable calcium, magnesium, iron, aluminum, potassium, phosphorus, and sulfur; (13) a folder of phosphorus, carbon, and nitrogen X-ray absorption near edge structure (P-XANES, N-XANES, C-XANES) data for samples and standards; (14) P-XANES, N-XANES, C-XANES methods; (15) molybdate reactive phosphorus; and (16) folder of high resolution characterization of organic matter via 21 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The FTICR folder contains .txt data files and a subfolder containing instruments for using Formularity (https://omics.pnl.gov/software/formularity) and an R script to process the data based on the user's specific needs. The metadata and protocols folder includes (1) international geo-sample number (IGSN) mapping file (2) burn and laboratory metadata; (3) burn protocol; (4) laboratory protocol; (5) vegetation collection metadata; and (6) vegetation collection protocol. The folder contains photos of the solid chars. All files are .csv, .txt, .pdf, .jpg, .jpeg, .R, .ref, or .mp4. The data package was originally published October 2022 (v1). It was updated April 2023 (v2; new data files), September 2023 (v3; new and corrected data files), and September 2024 (v4; new and added/updated files). Metadata files were also updated to reflect these changes. See the change history section in the readme for more details.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS Surface Water and Sediment Geochemistry and Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon (v2)

This dataset supports a broader study developing conceptual models for river corridor critical zone processes across spatial scales and was generated in collaboration with the HJ Andrews River Corridor Critical Zone Workshop in 2025. The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen) from 48 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Some of the sites have been impacted by the Holiday Farm Fire and the Lookout Fire in 2020 and 2023, respectively. Related data were collected as part of the workshop and will be published separately in collaboration with other workshop attendees and available at http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. Related genomic data can be found on the National Center for Biotechnology Information (NCBI) under BioProject PRJNA1503030 (see critical details section below for more information). Additional related data collected in 2016 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377027 and http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1-2019 (Ward et al., 2019). This data package was originally published in March 2026. It was updated in August 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos, (2) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, (3) a data checks report, (4) a folder of sample data, (5) file-level metadata, (6) data dictionary, (7) field metadata, (8) readme, (9) international generic sample number (IGSN) mapping file; and (10) field protocol. The sample data subfolder contains surface water and sediment (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages, (2) total dissolved nitrogen data and averages, (3) methods codes, (4) FTICR-MS methods; and (5) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the CoreMS processed data and seven subfolders, thee containing .xml files for each sample type (sediment, surface water and blank samples), three containing the sediment CoreMS output files for each sample type (sediment, surface water and blank samples), and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, .json, .jpg, or .jpeg.

Biogeochemistry↗