Engineering PapersSearch

SEARCH · Engineering Papers

Results for “ESS-DIVE Sample ID and Metadata Reporting Format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2

SPRUCE Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2025

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2025. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored in the SPRUCE archive and may be available for further analysis by request. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B05LW. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR.

EARTH SCIENCE > BIOSPHERE > ECOSYSTEMS > TERRESTRI

Soil physical and chemical measurements for topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes soil physical and chemical measurements from topsoils collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. The soil measurements include soil bulk density, soil volumetric water content, soil microbial biomass C (Carbon), N (Nitrogen) and C:N (C to N ratio), soil DNA yield, soil total extractable organic C, soil total extractable N, soil extractable nitrate, soil extractable ammonium, soil dissolved inorganic N, soil dissolved organic N, soil pH, soil TOC400 (total organic carbon at 400°C), soil ROC (residual oxidizable carbon), soil TIC (total inorganic carbon), soil TOC (total organic carbon), soil TC (total carbon), soil N, soil OM (organic matter) loss on ignition. Additional associated site metadata can be found in the ESS-DIVE package 10.15485/1618130. The dataset includes (1) 2018_NEON_soil_physical_chemical_measurements.csv: soil physical and chemical measurements indexed by soil sample IGSNs; (2) samples.csv: sample metadata file used to register International Generic Sample Numbers (IGSNs); (3) flmd.csv: file level metadata file; and (4) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS (Catchment Hydrology and

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA

Surface water nitrogen and sediment potential nitrate reduction rates, nutrient stocks, and stable isotopes from nine wetlands at the Tanglewood Biological Station, Alabama

This dataset supports a broader study investigating wetland hydrologic and biogeochemical responses to inundation disturbances. Bimonthly surface water and sediment sampling events were conducted at nine wetland sites situated within the Tanglewood Biological Station in Alabama from April 2023 to February 2024. The contents included in the data package include surface water nitrogen (nitrogen oxides and ammonium) and sediment potential nitrate reduction rates (measured as potential denitrification and dissimilatory nitrate reduction to ammonium processing), nutrient stocks (total carbon, total nitrogen, and organic matter), and stable isotopes (carbon and nitrogen). Water level data related to each wetland location can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253 (Kirker et al., 2024) and related water geochemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3001967 (Forbes et al., 2025). In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata and international generic sample numbers (IGSNs); (4) readme; (5) the field protocol; and (6) a subfolder with sample data. The sample data subfolder contains (1) sediment potential denitrification rate, (2) sediment potential dissimilatory nitrate reduction to ammonium (DNRA) rate, (3) sediment total carbon and nitrogen content, (4) sediment stable isotopes (delta nitrogen-15 and delta carbon-13), (5) sediment percent organic matter, (6) surface water nitrous oxides, (7) surface water ammonium, and (8) methods codes. All files are .csv or .pdf.

Ammonium

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from 7 Perennial and 7 Intermittent Streams across San Antonio, Texas (v3)

This dataset supports a broader study examining the effects of intermittency on sediment respiration. The dataset provides sediment and surface water geochemistry and in situ sensor data from 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). Related data were collected and will be published separately in collaboration with A. Veach. The data package was originally published in April 2025. It was updated in June 2025 (v2; modified and new files) and September 2025 (v3; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) sediment grain size data; (4) sediment iron (II) data and averages; (5) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment percent carbon and nitrogen; (11) sediment X-ray diffraction (XRD) data; (12) gravimetric moisture and averages; (13) a subfolder with sediment incubation respiration data, scripts, and plots; (14) surface water and sediment FTICR methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: The data processing methods for FTICR described in “v3_WHONDRS_AV1_Methods_Codes.csv” mistakenly indicate that users should process the data in Formultitude. The corrected description should read: “Both unprocessed and processed data are provided to allow users flexibility in data processing. Instructions and scripts for processing the data using CoreMS are included.” CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package.

54 ENVIRONMENTAL SCIENCES

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES

Porewater chemistry in Typha-dominated brackish tidal marsh, PIE LTER, Plum Island Sound, MA, July 2022–September 2024

This dataset contains profile measurements of porewater constituents taken on 3-4 days across the growing seasons in 2022, 2023, and 2024 in a tidal brackish marsh within the Plum Island Ecosystems Long Term Ecological Research site (PIE LTER), located in the Plum Island Sound, Massachusetts (MA). Measurements were taken to monitor changes in porewater chemistry induced by seasonal saltwater intrusion at the site. Samples were taken in two locations: one was close to the creek bank and the other in the marsh interior. Water was sampled from 2-5 depths between the surface to 50cm using a sipper consisting of a hollow stainless steel rod with an opening at the end similar to that described in (Berg & McGlathery, 2001). The rod was pushed into the sediment to the desired depth, typically every 10cm, and water samples were taken by syringe. Water was not obtained at all depths. Samples were preserved and analyzed in the lab. Metadata files Typha_porewater_sipper_dd.csv and Typha_porewater_sipper_flmd.csv contain detailed information on data variables, sampling and QA/QC methods, and site location.

54 ENVIRONMENTAL SCIENCES

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES

Meteorological and Soil Data from Ecohydrology Sensor Towers at Pump House and Snodgrass Mountain in East River Watershed, Colorado, 2019-2025

This data package includes hourly meteorological and soil sensor data at eight ecohydrology monitoring sites in East River Watershed, Colorado as part of the Watershed Function Scientific Focus Area (WFSFA) research led by Lawrence Berkeley National Lab (LBNL). Four field sites were located on the hillslope of East River (ER) near Pump House (PH) at Mount Crested Butte (ER-PHS1 to 4), and the other four are in the Snodgrass Mountain (SG) area (SG-EHS5 to 8). In terms of vegetation cover, three sites are in montane grasslands (ER-PHS1, ER-PHS2, and SG-EHS5), three are below evergreen conifer canopy (ER-PHS3, SG-EHS6, and SG-EHS7), and two are below deciduous aspen canopy (ER-PHS4 and SG-EHS8). The monitoring period began in October 2019 at the East River sites, in October 2020 at SG-EHS5 and SG-EHS6, and in October 2021 at SG-EHS7 and SG-EHS8. In September 2024, all four East River sites were fully retired. The four Snodgrass Mountain sites remain active. Each site is equipped with a comprehensive suite of meteorological sensors on a tripod and soil sensors that measure weather, energy fluxes, and soil variables. This data package includes measurements from ten different types of sensors and up to thirteen individual sensors per site, including (1) a weather station (measurement height ranges from 2.8~3.8 meters (m) above ground), (2) a quantum sensor for photosynthetic active radiation (PAR) (2.4~3.3m), (3) a net radiometer (1.7~2.1m), (4) an infrared radiometer (1.6~2.2m), (5) a sonic distance sensor (1.5~1.9m), (6) a soil carbon dioxide (CO2) flux chamber (0m), (7) a soil heat flux plate (-0.05m below ground), (8) a soil oxygen sensor (-0.3m), (9) a soil water potential sensor (-0.3m), and (10) soil water content sensors at 3~4 depths (-1.15 ~ -0.1m). A total of twenty-three variables is reported in this data package, including (1) atmospheric variables: air temperature (TA), atmospheric pressure (PA), vapor pressure (VP), and vapor pressure deficit (VPD), (2) precipitation variables: rain precipitation (P) and snow depth (D_SNOW), (3) energy fluxes variables: four-component net radiation (NETRAD) (shortwave/longwave incoming/outgoing radiation, SW_IN, SW_OUT, LW_IN, LW_OUT), photosynthetic photon flux density (PPFD), and soil heat flux (G), (4) soil variables: soil water content (SWC), soil water potential (SWP), soil temperature (TS), soil bulk electrical conductivity (COND_SOIL), and soil gaseous oxygen concentration (O2_SOIL), (5) wind variables: two-dimensional wind speed (WS), gust speed (WS_MAX), and wind direction (WD), and (6) surface variables: surface infrared temperature (T_CANOPY) and soil CO2 flux (CO2_SOIL). Please see the Methods section for data processing and QA/QC steps taken to generate the hourly datasets. The following files are included in this data package (notes on version: v{x}-{y}, where x is the metadata version, and y is the data version, when applicable): (1) “metadata_site_v{x}-{y}.csv” - a site metadata file that summarizes location information of all sites, including site ID, description, coordinates, timeframe, elevation, and vegetation cover, (2) “metadata_instrument_v{x}-{y}.csv” - an instrument metadata file that summarizes sensor information of all sites, including sensor manufacturer and model, measurement height, and sampling and averaging interval of all variables, (3) "data_{SITE_ID}_v{x}-{y}.csv" - eight data files that contain hourly data of each site indicated by {SITE_ID} in the filename, (4) “/figure/data_{SITE_ID}_v{x}-{y}.png" - eight figures that help visualize data of each site indicated by {SITE_ID} in the filename, (5) “/photo/*” - photos of each site indicated by {SITE_ID} in the filename, and (6) four file level metadata (flmd.csv) and data dictionary (*_dd.csv) files that summarize file, header, column, and variable information of all files. Notes: (1) Measurement height: Each variable name is followed by conventional positional qualifiers “H_V_R”, where H indicates the relative horizontal positions of that specific variable, V the vertical positions, and R the replicates. In this data package, only the vertical qualifier V varies, and V increases from the highest vertical position (V=1) to the lowest. Variables with the same qualifier are not necessarily measured by the same sensor, and the same variable with the same qualifier across different sites are not necessarily measured at the same height. Please refer to “metadata_instrument.csv” for the sensor information and measurement heights, and whether a variable is measured below the canopy. (2) Variable availability: Snow depth is not available at ER-PHS3 and SG-EHS7. SWC, soil temperature, and soil bulk EC at the deepest depth (<-1m) are not available at SG-EHS6 and SG-EHS7. The missing value code for numeric variables is -9999, except for SWP. For SWP, the missing value code is +9999, because SWP values are negative. (3) Sampling frequency: Please refer to “metadata_instrument.csv” for the increase of sampling frequency of some variables from 30-min to 1-min at ER-PHS1 to 4 in July 2020. (4) Sensors: While the methods of each sensor are not detailed, all sensors are commercially available, and their methods can be found in their manuals. Please refer to “metadata_instrument.csv” for the sensor manufacturer and model information. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES