Engineering PapersSearch

SEARCH · Engineering Papers

Results for “metagenomes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Metagenomes and metagenome-assembled genomes from microbial communities in a biological nutrient removal plant operated at Hamptons Road Sanitation District (HRSD) with high and low dissolved oxygen conditions

Aeration is a major cost at biological nutrient removal (BNR) plants. We report on microbial communities in a pilot-scale BNR system before and after a dissolved oxygen transition from 2.5 to 0.2 mg/L implemented over 18 months. Four PacBio metagenomes and 316 metagenome-assembled genomes are announced.

dissolved oxygen

Metagenomes and metagenome-assembled genomes from a nutrient removal plant at Los Angeles County Sanitation Districts (LACSD) that transitioned from high to low dissolved oxygen

Operating biological nutrient removal (BNR) wastewater treatment plants with low dissolved oxygen (DO) conditions can reduce energy costs. We report on five metagenomes and 492 metagenome-assembled genomes (MAGs) obtained from samples collected at the Pomona water reclamation plant before and after a DO reduction from 3.5 to 0.7 mg/L.

dissolved oxygen

Metagenomes and Metagenome-Assembled Genomes from Microbial Communities in a Biological Nutrient Removal Plant Operated at Hamptons Road Sanitation District (HRSD) with High and Low Dissolved Oxygen Conditions

In this study, we aimed to evaluate Biological Nutrient Removal (BNR) and investigate microbial community changes as the dissolved oxygen is reduced in the aerated portions of wastewater treatment trains. We present a dataset of Metagenome-Assembled Genomes (MAGs) obtained from activated sludge collected from the Hamptons Road Sanitation District (HRSD) BNR plant at the beginning of operation, when the DO was high, and at the end of operation, when the DO was low.

Genomics

Metagenomes and Metagenome-Assembled Genomes from Microbial Communities in a Biological Nutrient Removal Plant Operated at Los Angeles County Sanitation District (LACSD) with High and Low Dissolved Oxygen Conditions

In this study, we aimed to evaluate Biological Nutrient Removal (BNR) and investigate microbial community changes as the dissolved oxygen is reduced in the aerated portions of wastewater treatment trains. We present a dataset of Metagenome-Assembled Genomes (MAGs) obtained from activated sludge collected from the Los Angeles County Sanitation District (LACSD) BNR plant at the beginning of operation, when the DO was high, and at the end of operation, when the DO was low.

Genomics

Metagenome-assembled genomes measured at 3 depths during snowmelt period in East River, CO (March, May, and June, September 2017)

Snowmelt is a critical biogeochemical period that accounts for large nitrogen (N) export events from high-elevation watersheds. Soil microbial populations bloom and immobilize N during snowmelt, yet the population size crashes in spring, which releases a pulse of soil N. We sought to discover the N sources fueling this microbial bloom and determine the fate of N following microbial die-off. Here, focusing on the snowmelt period within a headwater catchment of the Upper Colorado River Basin (East River, CO), we deployed strain-resolved metagenomics to identify the metabolic pathways and processes that mobilize soil N during and after snowmelt. Soil metagenome samples were taken from 6 snowpits from 3 depths (0-5cm, 5-15cm, >15cm) at 4 time points during snowmelt period (March 2017, May 2017, and June 2017, September 2017) generating 48 metagenomes. We reconstructed 474 metagenome-assembled genomes (MAGs) across all metagenomes.All 48 metagenomes were sequenced at JGI and raw data can be found under JGI (Joint Genome Institute) GOLD Study Gs0135149. Metagenome assemblies from IMG under the same study were used for genome binning. This dataset (1) a zip file of 474 MAGs (as fasta files, Gs0135149_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0135149.kml), (4) metagenome metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (metagenomes.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of LBNL (Lawrence Berkeley National Laboratory) TES (Terrestrial Ecosystem Science) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from soil depth profiles collected from 2014 to 2021 from three paired control and warming plots. We collected soil samples across a range of depth profiles (spanning surface to 90 cm deep) from three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. 101 soil metagenomes were sequenced at JGI (Joint Genome Institute) and UCSF (University of California San Francisco) Center for Advanced Technology and can be found under the JGI (Joint Genome Institute) GOLD (Genomes Online Database) Sequencing project Gs0151586 and NCBI (National Center for Biotechnology Information) Projects PRJNA1225762 and PRJEB39497. Metagenomes were assembled using JGI (Joint Genome Institute) Metagenome Workflow (10.1128/mSystems.00804-20). For each metagenome, the assembled contigs were binned into genomes using 3 binning algorithms (cocacola, metabat, and maxbin) and the resulting bins were consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>50%) and contamination (<25%), and dereplicated at 99% ANI (average nucleotide identity) using dRep (https://github.com/MrOlm/drep).The dataset includes a zip file of 2321 MAG (Metagenome Assembled Genome) fasta files, the accession numbers for the underlying metagenomes, and a csv file with MAG (Metagenome Assembled Genome) quality metrics and taxonomic classification (GTDB -Genome Taxonomy Database-RS220). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (June to October 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken at three time points from June 12, 2019 to October 23,2019 at a location (PTT1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples were collected from 60 to 180 cm below surface every 30cm for microbial analyses through metagenomic sequencing. 15 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 780 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (August 2015)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken August 29, 2015 at a location (KB1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples from a deep soil pit were collected from 0 to 234 cm depth below surface at discrete depths every ~10-20 cm for microbial analyses. 13 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores).This dataset includes a zip file of 2216 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June to October 2020)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June to October 2020 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 30 cm depth below surface to just above the cobble layer (~190-250 cm depth) at discrete depths every 40 cm for microbial analyses. A total of 35 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2848 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken September 2019 at one locations (OBJ1) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 6 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2562 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES