Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “JGI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Genomes to Structure and Function Workshop Report 2022

The goal of the U.S. Department of Energy (DOE) Biological and Environmental Research (BER) Program is to achieve a predictive understanding of complex biological, earth, and environmental systems with the aim of advancing the nation’s energy and infrastructure security. (https://www.energy.gov/science/ ber/biological-and-environmental-research). To pursue this goal, collaborations among experts in diverse research areas that lead to multidisciplinary projects are indispensable. The roles of DOE’s User Facilities, which offer unique and powerful resources for such research projects, are evolving, and expectations for the facilities are increasing. To respond to Users’ needs, the Joint Genome Institute (JGI) and Environmental Molecular Sciences Laboratory (EMSL) initiated the Facilities Integrating Collaborations for User Science (FICUS) program in 2014. This collaboration has grown into a popular and successful program, advancing more than 100 multidisciplinary projects to date. Similarly, the new interFacility collaborations among the JGI, EMSL, and User resources for BER structural biology and imaging at the Basic Energy Science (BES) Program’s synchrotron and neutron facilities are becoming essential for cutting-edge transdisciplinary science. To further explore the need for the BER research community to combine genomic, functional, and structural approaches to advance their research, an organizing committee was formed to develop and jointly host a 3-part workshop. The committee’s members represented seven DOE National Laboratory User Facilities (Appendix 1 lists the members). The “Genomes to Structure and Function” virtual workshop (see Appendices 2–5) was composed of three sessions. The first session, titled “Molecular Structures” (October 27– 28, 2021), highlighted diverse integrative experimental and computational approaches correlating structural data with sequencing and functional information, as well as predicting protein structures to model complex biological systems. The second session, “Intracellular Organization, and Material Synthesis and Decomposition” (December 15–16, 2021), covered imaging methods for observing, quantifying, and manipulating biosystems. The third session, “Imaging the Rhizosphere and Cellular Organization” (January 26–27, 2022) emphasized advanced and non-invasive imaging techniques applied to plant root-microbe-soil interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Trimming and Decontamination of Metagenomic Data can Significantly Impact Assembly and Binning Metrics, Phylogenomic and Functional Analysis

Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.

Whitham, Jason M.↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

A k-mer based approach for classifying viruses without taxonomy identifies viral associations in human autism and plant microbiomes

Viruses are an underrepresented taxa in the study and identification of microbiome constituents; however, they play an essential role in health, microbiome regulation, and transfer of genetic material. Only a few thousand viruses have been isolated, sequenced, and assigned a taxonomy, which limits the ability to identify and quantify viruses in the microbiome. Additionally, the vast diversity of viruses represents a challenge for classification, not only in constructing a viral taxonomy, but also in identifying similarities between a virus’ genotype and its phenotype. However, the diversity of viral sequences can be leveraged to classify their sequences in metagenomic and metatranscriptomic samples, even if they do not have a taxonomy. To identify and quantify viruses in transcriptomic and genomic samples, we developed a dynamic programming algorithm for creating a classification tree out of 715,672 metagenome viruses. To create the classification tree, we clustered proportional similarity scores generated from the k-mer profiles of each of the metagenome viruses to create a database of metagenomic viruses. The resulting Kraken2 database of the metagenomic viruses can be found here: https://www.osti.gov/biblio/1615774 and is compatible with Kraken2. We then integrated the viral classification database with databases created with genomes from NCBI for use with ParaKraken (a parallelized version of Kraken provided in Supplemental Zip 1), a metagenomic/transcriptomic classifier. To illustrate the breadth of our utility for classifying metagenome viruses, we analyzed data from a plant metagenome study identifying genotypic and compartment specific differences between two Populus genotypes in three different compartments. We also identified a significant increase in abundance of eight viral sequences in post mortem brains in a human metatranscriptome study comparing Autism Spectrum Disorder patients and controls. We also show the potential accuracy for classifying viruses by utilizing both the JGI and NCBI viral databases to identify the uniqueness of viral sequences. Finally, we validate the accuracy of viral classification with NCBI databases containing viruses with taxonomy to identify pathogenic viruses in known COVID-19 and cassava brown streak virus infection samples. Our method represents the compulsory first step in better understanding the role of viruses in the microbiome by allowing for a more complete identification of sequences without taxonomy. Better classification of viruses will improve identifying associations between viruses and their hosts as well as viruses and other microbiome members. Despite the lack of taxonomy, this database of metagenomic viruses can be used with any tool that utilizes a taxonomy, such as Kraken, for accurate classification of viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Twenty-five years of Genomes OnLine Database (GOLD): data updates and new features in v.9

We report the Genomes OnLine Database (GOLD) (https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute (DOE-JGI) continues to maintain its role as one of the flagship genomic metadata repositories of the world. The ever-increasing number of projects and metadata are freely available to the user community world-wide. GOLD’s metadata is consumed by scientists and remains an important source for large-scale comparative genomics analysis initiatives. Encouraged by this active user engagement and growth, GOLD has continued to add new components and capabilities. The new features such as a public Application Programming Interface (API) and Ecosystem landing page as well as the growth of different entities in this current GOLD v.9 edition are described in detail in this manuscript.

59 BASIC BIOLOGICAL SCIENCES↗

Metatranscriptomics sheds light on the links between the functional traits of fungal guilds and ecological processes in forest soil ecosystems

Soil fungi belonging to different functional guilds, such as saprotrophs, pathogens, and mycorrhizal symbionts, play key roles in forest ecosystems. To date, no study has compared the actual gene expression of these guilds in different forest soils. We used metatranscriptomics to study the competition for organic resources by these fungal groups in boreal, temperate, and Mediterranean forest soils. Using a dedicated mRNA annotation pipeline combined with the JGI MycoCosm database, we compared the transcripts of these three fungal guilds, targeting enzymes involved in C- and N mobilization from plant and microbial cell walls. Genes encoding enzymes involved in the degradation of plant cell walls were expressed at a higher level in saprotrophic fungi than in ectomycorrhizal and pathogenic fungi. However, ectomycorrhizal and saprotrophic fungi showed similarly high expression levels of genes encoding enzymes involved in fungal cell wall degradation. Transcripts for N-related transporters were more highly expressed in ectomycorrhizal fungi than in other groups. Here, we showed that ectomycorrhizal and saprotrophic fungi compete for N in soil organic matter, suggesting that their interactions could decelerate C cycling. Metatranscriptomics provides a unique tool to test controversial ecological hypotheses and to better understand the underlying ecological processes involved in soil functioning and carbon stabilization.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Web of Registries Search (WoRS) v1.0.0

The Web of Registries Search (WoRs) is a web based software application that enables users to search for publicly available biological parts using keywords or sequence fragments. For the initial version (1.0.0) of the application, WoRS targets 10 sources of biological part data: the GenBank NIH genetic sequence database (https://www.ncbi.nlm.nih.gov/genbank/), the iGem parts registry (parts.igem.org), the Addgene plasmid repository (https://www.addgene.org/), and 7 Inventory of Composable Elements (ACS Synbio, JGI, JBEI, JBEI Public, ABF, SynBerc, ABF Public) registry instances. WoRs has built-in automated web scrapers which extract data from sources that do not have a public or well-defined application programming interface (API). They extract as much public data as they can find and create a searchable index to speed up searches. Included in the indexed information is the source of the information.

Plahar, Hector↗

IMG Annotation Pipeline (IMGAP) v5.1.13

The IMG Annotation Pipeline is a collection of Bash and Python scripts to control a workflow for structural and functional annotation of prokaryotic genomes, metagenomes, and metatranscriptomes. The bash scripts in general control the overall workflow and are wrappers around 3rd party executables (not included in repo) that predict features or functions. Whereas the Python scripts do post-processing of raw output in terms of filtering or format transformation and in some cases contain some logic for picking the correct predictions or resolving overlaps. The pipeline is tailored to produce results required by IMG (https://img.jgi.doe.gov/) and is executed on every dataset submitted to IMG via https://img.jgi.doe.gov/submit. These consist of internal genomes, metagenomes and metatranscriptomes sequenced and assembled at the JGI, as well as datasets submitted by external users (non-lab/JGI affiliates).

Huntemann, Marcel↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (August 2015)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken August 29, 2015 at a location (KB1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples from a deep soil pit were collected from 0 to 234 cm depth below surface at discrete depths every ~10-20 cm for microbial analyses. 13 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores).This dataset includes a zip file of 2216 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (June to October 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken at three time points from June 12, 2019 to October 23,2019 at a location (PTT1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples were collected from 60 to 180 cm below surface every 30cm for microbial analyses through metagenomic sequencing. 15 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 780 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June to October 2020)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June to October 2020 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 30 cm depth below surface to just above the cobble layer (~190-250 cm depth) at discrete depths every 40 cm for microbial analyses. A total of 35 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2848 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken September 2019 at one locations (OBJ1) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 6 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2562 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June 2018 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 12 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 1233 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗