Engineering PapersSearch

SEARCH · Engineering Papers

Results for “ESS-DIVE Sample ID and Metadata Reporting Format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Temporal Study 2022-2024: Sample-Based Surface Water Dissolved Inorganic Carbon, Dissolved Organic Carbon, Total Nitrogen, Stable Isotopes, and Total Suspended Solids from across Multiple Watersheds in the Yakima River Basin, Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides geochemistry data generated from samples collected at bi-weekly or monthly intervals at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sensor data from 2022-2024 will be published separately. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) dissolved inorganic carbon (DIC) and averages; (6) dissolved organic carbon (DOC; reported as non-purgeable organic carbon; NPOC) and averages; (7) total dissolved nitrogen (TN) and averages; (8) total suspended solids (TSS); (9) stable isotopes; (10) surface water sampling protocol; (11) sensor protocol; (12) methods codes; and (13) international generic sample number (IGSN) mapping file. All files are .csv or .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. For data and scripts associated with "Shifts in rain-snow partitioning drive faster water transit times in the US Pacific Northwest" (Butler et al., 2026), go to https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3025481

18-O

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES

Surface water and groundwater FTICR-MS, NPOC, and TN from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama

This dataset supports a broader study examining wetland hydrobiogeochemical responses to flood disturbance and the subsequent impacts on watershed nutrient export. The study was designed following ICON (integrated, coordinated, open, and networked) principles. Samples were collected from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama in August 2024 and February 2025, during the dry and wet season, respectively. The contents include geochemistry (dissolved organic carbon measured as non-purgeable organic carbon; total dissolved nitrogen) and organic matter characterization (FTICR-MS). Related water level data from the same locations can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253. Additional geochemistry will be published in a separate data package. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; (7) the field protocol; and (8) a subfolder with sample data. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total nitrogen data and averages; (3) methods codes; and (4) a subfolder of 12 Tesla (12T) Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES

Pyrogenic Organic Matter Laboratory Experiment: Aerobic Respiration and Geochemistry from Variably Inundated Stream Sediments (v3)

This dataset supports a broader study examining the effects of variable inundation and pyrogenic organic matter on ecosystem respiration. The dataset provides data generated from a laboratory batch experiment investigating the interaction between variable inundation conditions (wet and dry sediment) and pyrogenic organic matter (burned and unburned treatments). The contents include time series dissolved oxygen, sediment geochemistry data, and field metadata (including qualitative information on instream and river corridor characteristics). This data package was originally published in November 2025. It was updated in April 2026 (v2; new and modified files) and May 2026 (v3; modified files). See the change history section in the readme for more details For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) international generic sample number (IGSN) mapping file; (5) readme; (6) field protocol; (7) sample name metadata; (8) an environmental context picture for the dry and inundated sampling locations; and (9) a subfolder with sample data from the sediment incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) gravimetric moisture; (4) partial pressure and production rates of carbon dioxide, methane, and nitrous oxide; (5) field wet sediment mass, dry sediment mass, water mass, and field wet sediment volume in incubation and sediment NPOC/TN vials; (6) methods codes; (7) respiration rates, pH, and temperature from after the incubation, raw time series dissolved oxygen and temperature, and a subfolder containing associated plots and scripts; (8) ions; (9) FTICR-MS methods; and (10) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains the CoreMS processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, or .jpg.

54 ENVIRONMENTAL SCIENCES

Post-fire time series of sensor and geochemistry sample data from surface water, groundwater, precipitation, soil, and vegetation across Oak Creek watershed, Washington

This dataset supports a broader study examining wildfire impacts on hydrologic connectivity across 5 sites within the Oak Creek watershed and the resulting biogeochemical impacts. Stream sites were selected using the Advanced Terrestrial Simulator (ATS) hydrologic model to identify locations with varying groundwater contributions and hydrologic responses across different burn severity scenarios. The Retreat Fire burned from July 23 to August 2 in 2024, affecting the five study sites at varying burn severities. Each site is equipped with YSI EXO2 sondes logging sub-hourly throughout the year, and grab samples are collected approximately every six weeks. YSI sondes are used to measure temporally resolved proxies for groundwater inputs (specific conductivity) and organic matter (fluorescent dissolved organic matter; fDOM) along with basic water quality and depth. Grab samples of surface water, groundwater, and precipitation are analyzed for water stable isotopes and conductivity to understand endmembers for hydrologic mixing Grab samples of surface water, groundwater, soil water, and litter/vegetation/soil leachates are analyzed for organic matter composition measured by Fourier-Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) to understand organic matter dynamics. Game camera photos are provided in a separate data package available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3018598. Future versions of this dataset will include time series data from YSI EXO2 sondes (fDOM, dissolved oxygen, temperature, depth, specific conductance, turbidity, pH), BaroTROLL sensors (air temperature and barometric pressure), rain gauges (precipitation), and data from the soil and vegetation samples. Because this study is ongoing, this data package will be updated regularly to include newly collected data and the additional data types. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data; (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) a data checks report; (5) file-level metadata; (6) data dictionary; (7) field metadata; (8) readme; (9) international generic sample number (IGSN) mapping file; and (10) field protocols. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) stable water isotopes and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Biogeochemistry

WHONDRS Surface Water and Sediment Geochemistry and Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon (v2)

This dataset supports a broader study developing conceptual models for river corridor critical zone processes across spatial scales and was generated in collaboration with the HJ Andrews River Corridor Critical Zone Workshop in 2025. The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen) from 48 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Some of the sites have been impacted by the Holiday Farm Fire and the Lookout Fire in 2020 and 2023, respectively. Related data were collected as part of the workshop and will be published separately in collaboration with other workshop attendees and available at http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. Related genomic data can be found on the National Center for Biotechnology Information (NCBI) under BioProject PRJNA1503030 (see critical details section below for more information). Additional related data collected in 2016 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377027 and http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1-2019 (Ward et al., 2019). This data package was originally published in March 2026. It was updated in August 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos, (2) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, (3) a data checks report, (4) a folder of sample data, (5) file-level metadata, (6) data dictionary, (7) field metadata, (8) readme, (9) international generic sample number (IGSN) mapping file; and (10) field protocol. The sample data subfolder contains surface water and sediment (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages, (2) total dissolved nitrogen data and averages, (3) methods codes, (4) FTICR-MS methods; and (5) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the CoreMS processed data and seven subfolders, thee containing .xml files for each sample type (sediment, surface water and blank samples), three containing the sediment CoreMS output files for each sample type (sediment, surface water and blank samples), and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, .json, .jpg, or .jpeg.

Biogeochemistry

WHONDRS River Corridor Surface Water Metabolites and Geochemistry from Global Sites

This dataset supports a broader study examining the character of organic matter that may be delivered to subsurface sediments via hydrologic exchange. To implement the global survey, free stream sampling kits were provided to interested volunteers throughout the world. Samples were collected with minimal constraints in terms of location, but following strict protocols, and shipped for metabolomic analysis via Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). In addition, basic geochemistry analyses (e.g., dissolved organic matter concentration) were conducted, standardized photos of each field system were taken, and extensive metadata were captured. Sampling began in 2018 and is ongoing as of 2025. This dataset is comprised of one folders of field photos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; and (7) a subfolder with sample data. The sample data subfolder contains (1) surface water dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) methods codes; (3) surface water FTICR methods; and (4) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains three subfolders, one containing the.xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, or .png. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

Biogeochemistry

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions

Soil nitrogen mineralization rates, nutrient stocks, stable isotopes, and water volumetric measurements across terrestrial-aquatic interfaces from three wetlands at the Tanglewood Biological Station, Alabama

This dataset supports a broader study investigating wetland hydrologic and biogeochemical responses to inundation events. Soil samples were collected across four sampling events along terrestrial-aquatic gradients at three wetland sites located within the Tanglewood Biological Station in Alabama from April 2024 to June 2025. The contents in this data package include soil in-situ nitrogen mineralization rates (measurements of net nitrification, net ammonification, and net mineralization), nutrient stocks (total carbon, total nitrogen, and organic matter), stable isotopes (carbon and nitrogen), and water volumetric measurements (water-filled pore space). Water level data related to each wetland location can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253 (Kirker et al., 2024), related water geochemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3001967 (Forbes et al., 2025), and related surface water sediment chemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377325 (Molina Serpas et al., 2026). In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata and international generic sample numbers (IGSNs); (4) readme; (5) the field protocol; and (6) a subfolder with sample data. The sample data subfolder contains (1) net nitrification rate, (2) net ammonification rate, (3) areal net mineralization rate, (4) percent organic matter, (5) water-filled pore space, (6) total carbon content, (7) total nitrogen content, (8) stable carbon isotope (delta carbon-13), and (9) stable nitrogen isotope (delta nitrogen-15), and (10) methods codes. All files are .csv or .pdf.

13-C

Data for "Depth of nutrient uptake by deep-rooted plants is regulated by water availability"

The data set consists of strontium (Sr) isotope ratios (87Sr/86Sr), water isotopes, soil cation concentrations, soil water potential sensor data, and results of 87Sr/86Sr mixing model. The plant canopy size files include the dataset of canopy dimension of sagebrush, lupine, and sunflower. The soil and plant ICPMS (Inductively Coupled Plasma Mass Spectrometry) data file includes both of 87Sr/86Sr, and cation concentration dataset from soil exchangeable pool, apatite pool, silicate extract, atmospheric rain deposition, and plant leaf and stem tissues. The plant dendrochronology file includes the dendrochronogical ring width of several sagebrush, and dendrochemical sample data includes the 87Sr/86Sr for each separated growth ring. The modeling result gives the proportion of nutrient sources of each plants (based on their 87Sr/86Sr in leaf tissues and growth rings) from atmospheric deposition and mineral weathering. Soil water potential data includes continuous collection of soil water potential dataset at 2 depths (30 cm and 60 cm, from Nov 24 - Jun 25) of the sampling site. All the samples were collected from 2 sampling campaign June and July 2023, and rain water is a separate sampling from Aug - Sept 2023, at north-facing hillslope near pumphouse site. The data showed that the depth of cation nutrient acquisition is thus tightly coupled with, and likely determined by, water availability in soil, saprolite and bedrock. The enhanced uptake of cations and water from regions of mineral weathering could confer plant and ecosystem resilience during low water years and may impact the rate of bedrock weathering and watershed chemistry during drought. This dataset includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type; a location metadata file (locations.csv); and a samples metadata file (samples.csv). All files are provided as comma-separated values (CSV) files (.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

SPRUCE: Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2024

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2024. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored at Oak Ridge National Laboratory and may be available for further analysis. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B00VQ. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR. Note: Only dried and ground material from C Cores are available for new analysis.

Birkebak, Joshua [ORNL] (ORCID:0009000955611494)

SPRUCE Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2025

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2025. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored in the SPRUCE archive and may be available for further analysis by request. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B05LW. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR.

EARTH SCIENCE > BIOSPHERE > ECOSYSTEMS > TERRESTRI

Soil physical and chemical measurements for topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes soil physical and chemical measurements from topsoils collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. The soil measurements include soil bulk density, soil volumetric water content, soil microbial biomass C (Carbon), N (Nitrogen) and C:N (C to N ratio), soil DNA yield, soil total extractable organic C, soil total extractable N, soil extractable nitrate, soil extractable ammonium, soil dissolved inorganic N, soil dissolved organic N, soil pH, soil TOC400 (total organic carbon at 400°C), soil ROC (residual oxidizable carbon), soil TIC (total inorganic carbon), soil TOC (total organic carbon), soil TC (total carbon), soil N, soil OM (organic matter) loss on ignition. Additional associated site metadata can be found in the ESS-DIVE package 10.15485/1618130. The dataset includes (1) 2018_NEON_soil_physical_chemical_measurements.csv: soil physical and chemical measurements indexed by soil sample IGSNs; (2) samples.csv: sample metadata file used to register International Generic Sample Numbers (IGSNs); (3) flmd.csv: file level metadata file; and (4) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS (Catchment Hydrology and

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA

Surface water nitrogen and sediment potential nitrate reduction rates, nutrient stocks, and stable isotopes from nine wetlands at the Tanglewood Biological Station, Alabama

This dataset supports a broader study investigating wetland hydrologic and biogeochemical responses to inundation disturbances. Bimonthly surface water and sediment sampling events were conducted at nine wetland sites situated within the Tanglewood Biological Station in Alabama from April 2023 to February 2024. The contents included in the data package include surface water nitrogen (nitrogen oxides and ammonium) and sediment potential nitrate reduction rates (measured as potential denitrification and dissimilatory nitrate reduction to ammonium processing), nutrient stocks (total carbon, total nitrogen, and organic matter), and stable isotopes (carbon and nitrogen). Water level data related to each wetland location can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253 (Kirker et al., 2024) and related water geochemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3001967 (Forbes et al., 2025). In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata and international generic sample numbers (IGSNs); (4) readme; (5) the field protocol; and (6) a subfolder with sample data. The sample data subfolder contains (1) sediment potential denitrification rate, (2) sediment potential dissimilatory nitrate reduction to ammonium (DNRA) rate, (3) sediment total carbon and nitrogen content, (4) sediment stable isotopes (delta nitrogen-15 and delta carbon-13), (5) sediment percent organic matter, (6) surface water nitrous oxides, (7) surface water ammonium, and (8) methods codes. All files are .csv or .pdf.

Ammonium

Porewater chemistry in Typha-dominated brackish tidal marsh, PIE LTER, Plum Island Sound, MA, July 2022–September 2024

This dataset contains profile measurements of porewater constituents taken on 3-4 days across the growing seasons in 2022, 2023, and 2024 in a tidal brackish marsh within the Plum Island Ecosystems Long Term Ecological Research site (PIE LTER), located in the Plum Island Sound, Massachusetts (MA). Measurements were taken to monitor changes in porewater chemistry induced by seasonal saltwater intrusion at the site. Samples were taken in two locations: one was close to the creek bank and the other in the marsh interior. Water was sampled from 2-5 depths between the surface to 50cm using a sipper consisting of a hollow stainless steel rod with an opening at the end similar to that described in (Berg & McGlathery, 2001). The rod was pushed into the sediment to the desired depth, typically every 10cm, and water samples were taken by syringe. Water was not obtained at all depths. Samples were preserved and analyzed in the lab. Metadata files Typha_porewater_sipper_dd.csv and Typha_porewater_sipper_flmd.csv contain detailed information on data variables, sampling and QA/QC methods, and site location.

54 ENVIRONMENTAL SCIENCES

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES