Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metagenomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DOE JGI Metagenome Workflow

The DOE Joint Genome Institute (JGI) Metagenome Workflow performs metagenome data processing, including assembly; structural, functional, and taxonomic annotation; and binning of metagenomic data sets that are subsequently included into the Integrated Microbial Genomes and Microbiomes (IMG/M) (I.-M. A. Chen, K. Chu, K. Palaniappan, A. Ratner, et al., Nucleic Acids Res, 49:D751–D763, 2021, https://doi.org/10.1093/nar/gkaa939) comparative analysis system and provided for download via the JGI data portal (https://genome.jgi.doe.gov/portal/). This workflow scales to run on thousands of metagenome samples per year, which can vary by the complexity of microbial communities and sequencing depth. Here, we describe the different tools, databases, and parameters used at different steps of the workflow to help with the interpretation of metagenome data available in IMG and to enable researchers to apply this workflow to their own data. We use 20 publicly available sediment metagenomes to illustrate the computing requirements for the different steps and highlight the typical results of data processing. The workflow modules for read filtering and metagenome assembly are available as a workflow description language (WDL) file (https://code.jgi.doe.gov/BFoster/jgi_meta_wdl). The workflow modules for annotation and binning are provided as a service to the user community at https://img.jgi.doe.gov/submit and require filling out the project and associated metadata descriptions in the Genomes OnLine Database (GOLD) (S. Mukherjee, D. Stamatis, J. Bertsch, G. Ovchinnikova, et al., Nucleic Acids Res, 49:D723–D733, 2021, https://doi.org/10.1093/nar/gkaa983).

59 BASIC BIOLOGICAL SCIENCES↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

Supervised extraction of near-complete genomes from metagenomic samples: A new service in PATRIC

Large amounts of metagenomically-derived data are submitted to PATRIC for analysis. In the future, we expect even more jobs submitted to PATRIC will use metagenomic data. One in-demand use case is the extraction of near-complete draft genomes from assembled contigs of metagenomic origin. The PATRIC metagenome binning service utilizes the PATRIC database to furnish a large, diverse set of reference genomes. We provide a new service for supervised extraction and annotation of high-quality, near-complete genomes from metagenomically-derived contigs. Reference genomes are assigned to putative draft genome bins based on the presence of single-copy universal marker roles in the sample, and contigs are sorted into these bins by their similarity to reference genomes in PATRIC. Each set of binned contigs represents a draft genome that will be annotated by RASTtk in PATRIC. A structured-language binning report is provided containing quality measurements and taxonomic information about the contig bins. The PATRIC metagenome binning service emphasizes extraction of high-quality genomes for downstream analysis using other PATRIC tools and services. Due to its supervised nature, the binning service is not appropriate for mining novel or extremely low-coverage genomes from metagenomic samples.

59 BASIC BIOLOGICAL SCIENCES↗

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Subsurface hydrocarbon degradation strategies in low- and high-sulfate coal seam communities identified with activity-based metagenomics

Environmentally relevant metagenomes and BONCAT-FACS derived translationally active metagenomes from Powder River Basin coal seams were investigated to elucidate potential genes and functional groups involved in hydrocarbon degradation to methane in coal seams with high- and low-sulfate levels. An advanced subsurface environmental sampler allowed the establishment of coal-associated microbial communities under in situ conditions for metagenomic analyses from environmental and translationally active populations. Metagenomic sequencing demonstrated that biosurfactants, aerobic dioxygenases, and anaerobic phenol degradation pathways were present in active populations across the sampled coal seams. In particular, results suggested the importance of anaerobic degradation pathways under high-sulfate conditions with an emphasis on fumarate addition. Under low-sulfate conditions, a mixture of both aerobic and anaerobic pathways was observed but with a predominance of aerobic dioxygenases. The putative low-molecular-weight biosurfactant, lichysein, appeared to play a more important role compared to rhamnolipids. The methods used in this study—subsurface environmental samplers in combination with metagenomic sequencing of both total and translationally active metagenomes—offer a deeper and environmentally relevant perspective on community genetic potential from coal seams poised at different redox conditions broadening the understanding of degradation strategies for subsurface carbon.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic Methods for Addressing NASA's Planetary Protection Policy Requirements on Future Missions: A Workshop Report

Molecular biology methods and technologies have advanced substantially over the past decade. These new molecular methods should be incorporated among the standard tools of planetary protection (PP) and could be validated for incorporation by 2026. To address the feasibility of applying modern molecular techniques to such an application, NASA conducted a technology workshop with private industry partners, academics, and government agency stakeholders, along with NASA staff and contractors. The technical discussions and presentations of the Multi-Mission Metagenomics Technology Development Workshop focused on modernizing and supplementing the current PP assays. The goals of the workshop were to assess the state of metagenomics and other advanced molecular techniques in the context of providing a validated framework to supplement the bacterial endospore-based NASA Standard Assay and to identify knowledge and technology gaps. In particular, workshop participants were tasked with discussing metagenomics as a stand-alone technology to provide rapid and comprehensive analysis of total nucleic acids and viable microorganisms on spacecraft surfaces, thereby allowing for the development of tailored and cost-effective microbial reduction plans for each hardware item on a spacecraft. Workshop participants recommended metagenomics approaches as the only data source that can adequately feed into quantitative microbial risk assessment models for evaluating the risk of forward (exploring extraterrestrial planet) and back (Earth harmful biological) contamination. Participants were unanimous that a metagenomics workflow, in tandem with rapid targeted quantitative (digital) PCR, represents a revolutionary advance over existing methods for the assessment of microbial bioburden on spacecraft surfaces. The workshop highlighted low biomass sampling, reagent contamination, and inconsistent bioinformatics data analysis as key areas for technology development. Finally, it was concluded that implementing metagenomics as an additional workflow for addressing concerns of NASA's robotic mission will represent a dramatic improvement in technology advancement for PP and will benefit future missions where mission success is affected by backward and forward contamination.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic strategies identify diverse integron–integrase and antibiotic resistance genes in the Antarctic environment

The objective of this study is to identify and analyze integrons and antibiotic resistance genes (ARGs) in samples collected from diverse sites in terrestrial Antarctica. Integrons were studied using two independent methods. One involved the construction and analysis of intI gene amplicon libraries. In addition, we sequenced 17 metagenomes of microbial mats and soil by high-throughput sequencing and analyzed these data using the IntegronFinder program. As expected, the metagenomic analysis allowed for the identification of novel predicted intI integrases and gene cassettes (GCs), which mostly encode unknown functions. However, some intI genes are similar to sequences previously identified by amplicon library analysis in soil samples collected from non-Antarctic sites. ARGs were analyzed in the metagenomes using ABRIcate with CARD database and verified if these genes could be classified as GCs by IntegronFinder. We identified 53 ARGs in 15 metagenomes, but only four were classified as GCs, one in MTG12 metagenome (Continental Antarctica), encoding an aminoglycoside-modifying enzyme (AAC(6´)acetyltransferase) and the other three in CS1 metagenome (Maritime Antarctica). One of these genes encodes a class D β-lactamase (blaOXA-205) and the other two are located in the same contig. One is part of a gene encoding the first 76 amino acids of aminoglycoside adenyltransferase (aadA6), and the other is a qacG2 gene.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (June to October 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken at three time points from June 12, 2019 to October 23,2019 at a location (PTT1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples were collected from 60 to 180 cm below surface every 30cm for microbial analyses through metagenomic sequencing. 15 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 780 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

A consensus protocol for the recovery of mercury methylation genes from metagenomes

Abstract Mercury (Hg) methylation genes ( hgcAB ) mediate the formation of the toxic methylmercury and have been identified from diverse environments, including freshwater and marine ecosystems, Arctic permafrost, forest and paddy soils, coal‐ash amended sediments, chlor‐alkali plants discharges and geothermal springs. Here we present the first attempt at a standardized protocol for the detection, identification and quantification of hgc genes from metagenomes. Our Hg‐cycling microorganisms in aquatic and terrestrial ecosystems (Hg‐MATE) database, a catalogue of hgc genes, provides the most accurate information to date on the taxonomic identity and functional/metabolic attributes of microorganisms responsible for Hg methylation in the environment. Furthermore, we introduce “marky‐coco”, a ready‐to‐use bioinformatic pipeline based on de novo single‐metagenome assembly, for easy and accurate characterization of hgc genes from environmental samples. We compared the recovery of hgc genes from environmental metagenomes using the marky‐coco pipeline with an approach based on coassembly of multiple metagenomes. Our data show similar efficiency in both approaches for most environments except those with high diversity (i.e., paddy soils) for which a coassembly approach was preferred. Finally, we discuss the definition of true hgc genes and methods to normalize hgc gene counts from metagenomes.

59 BASIC BIOLOGICAL SCIENCES↗

BinaRena: a dedicated interactive platform for human-guided exploration and binning of metagenomes

Background: Exploring metagenomic contigs and “binning” them into metagenome-assembled genomes (MAGs) are essential for the delineation of functional and evolutionary guilds within microbial communities. Despite the advances in automated binning algorithms, their capabilities in recovering MAGs with accuracy and biological relevance are so far limited. Researchers often find that human involvement is necessary to achieve representative binning results. This manual process however is expertise demanding and labor intensive, and it deserves to be supported by software infrastructure. Results: We present BinaRena, a comprehensive and versatile graphic interface dedicated to aiding human operators to explore metagenome assemblies via customizable visualization and to associate contigs with bins. Contigs are rendered as an interactive scatter plot based on various data types, including sequence metrics, coverage profiles, taxonomic assignments, and functional annotations. Various contig-level operations are permitted, such as selection, masking, highlighting, focusing, and searching. Binning plans can be conveniently edited, inspected, and compared visually or using metrics including silhouette coefficient and adjusted Rand index. Completeness and contamination of user-selected contigs can be calculated in real time. In demonstration of BinaRena’s usability, we show that it facilitated biological pattern discovery, hypothesis generation, and bin refinement in a complex tropical peatland metagenome. It enabled isolation of pathogenic genomes within closely related populations from the gut microbiota of diarrheal human subjects. It significantly improved overall binning quality after curating results of automated binners using a simulated marine dataset. Conclusions: BinaRena is an installation-free, dependency-free, client-end web application that operates directly in any modern web browser, facilitating ease of deployment and accessibility for researchers of all skill levels. The program is hosted at https://github.com/qiyunlab/binarena, together with documentation, tutorials, example data, and a live demo. It effectively supports human researchers in intuitive interpretation and fine tuning of metagenomic data.

59 BASIC BIOLOGICAL SCIENCES↗

Time-series metagenomics reveals changing protistan ecology of a temperate dimictic lake

Abstract Background Protists, single-celled eukaryotic organisms, are critical to food web ecology, contributing to primary productivity and connecting small bacteria and archaea to higher trophic levels. Lake Mendota is a large, eutrophic natural lake that is a Long-Term Ecological Research site and among the world’s best-studied freshwater systems. Metagenomic samples have been collected and shotgun sequenced from Lake Mendota for the last 20 years. Here, we analyze this comprehensive time series to infer changes to the structure and function of the protistan community and to hypothesize about their interactions with bacteria. Results Based on small subunit rRNA genes extracted from the metagenomes and metagenome-assembled genomes of microeukaryotes, we identify shifts in the eukaryotic phytoplankton community over time, which we predict to be a consequence of reduced zooplankton grazing pressures after the invasion of a invasive predator (the spiny water flea) to the lake. The metagenomic data also reveal the presence of the spiny water flea and the zebra mussel, a second invasive species to Lake Mendota, prior to their visual identification during routine monitoring. Furthermore, we use species co-occurrence and co-abundance analysis to connect the protistan community with bacterial taxa. Correlation analysis suggests that protists and bacteria may interact or respond similarly to environmental conditions. Cryptophytes declined in the second decade of the timeseries, while many alveolate groups (e.g., ciliates and dinoflagellates) and diatoms increased in abundance, changes that have implications for food web efficiency in Lake Mendota. Conclusions We demonstrate that metagenomic sequence-based community analysis can complement existing efforts to monitor protists in Lake Mendota based on microscopy-based count surveys. We observed patterns of seasonal abundance in microeukaryotes in Lake Mendota that corroborated expectations from other systems, including high abundance of cryptophytes in winter and diatoms in fall and spring, but with much higher resolution than previous surveys. Our study identified long-term changes in the abundance of eukaryotic microbes and provided context for the known establishment of an invasive species that catalyzes a trophic cascade involving protists. Our findings are important for decoding potential long-term consequences of human interventions, including invasive species introduction.

59 BASIC BIOLOGICAL SCIENCES↗

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Trimming and Decontamination of Metagenomic Data can Significantly Impact Assembly and Binning Metrics, Phylogenomic and Functional Analysis

Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.

Whitham, Jason M.↗

Improved Microbial Community Characterization of 16S rRNA via Metagenome Hybridization Capture Enrichment

Environmental microbial diversity is often investigated from a molecular perspective using 16S ribosomal RNA (rRNA) gene amplicons and shotgun metagenomics. While amplicon methods are fast, low-cost, and have curated reference databases, they can suffer from amplification bias and are limited in genomic scope. In contrast, shotgun metagenomic methods sample more genomic regions with fewer sequence acquisition biases, but are much more expensive (even with moderate sequencing depth) and computationally challenging. Here, we develop a set of 16S rRNA sequence capture baits that offer a potential middle ground with the advantages from both approaches for investigating microbial communities. These baits cover the diversity of all 16S rRNA sequences available in the Greengenes (v. 13.5) database, with no sequence having <78% sequence identity to at least one bait for all segments of 16S. The use of our baits provide comparable results to 16S amplicon libraries and shotgun metagenomic libraries when assigning taxonomic units from 16S sequences within the metagenomic reads. We demonstrate that 16S rRNA capture baits can be used on a range of microbial samples (i.e., mock communities and rodent fecal samples) to increase the proportion of 16S rRNA sequences (average > 400-fold) and decrease analysis time to obtain consistent community assessments. Furthermore, our study reveals that bioinformatic methods used to analyze sequencing data may have a greater influence on estimates of community composition than library preparation method used, likely due in part to the extent and curation of the reference databases considered. Thus, enriching existing aliquots of shotgun metagenomic libraries and obtaining modest numbers of reads from them offers an efficient orthogonal method for assessment of bacterial community composition.

59 BASIC BIOLOGICAL SCIENCES↗