Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metagenomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Integrating chromatin conformation information in a self-supervised learning model improves metagenome binning

Metagenome binning is a key step, downstream of metagenome assembly, to group scaffolds by their genome of origin. Although accurate binning has been achieved on datasets containing multiple samples from the same community, the completeness of binning is often low in datasets with a small number of samples due to a lack of robust species co-abundance information. In this study, we exploited the chromatin conformation information obtained from Hi-C sequencing and developed a new reference-independent algorithm, Metagenome Binning with Abundance and Tetra-nucleotide frequencies—Long Range (metaBAT-LR), to improve the binning completeness of these datasets. This self-supervised algorithm builds a model from a set of high-quality genome bins to predict scaffold pairs that are likely to be derived from the same genome. Then, it applies these predictions to merge incomplete genome bins, as well as recruit unbinned scaffolds. We validated metaBAT-LR’s ability to bin-merge and recruit scaffolds on both synthetic and real-world metagenome datasets of varying complexity. Benchmarking against similar software tools suggests that metaBAT-LR uncovers unique bins that were missed by all other methods.

59 BASIC BIOLOGICAL SCIENCES↗

Viromes outperform total metagenomes in revealing the spatiotemporal patterns of agricultural soil viral communities

Abstract Viruses are abundant yet understudied members of soil environments that influence terrestrial biogeochemical cycles. Here, we characterized the dsDNA viral diversity in biochar-amended agricultural soils at the preplanting and harvesting stages of a tomato growing season via paired total metagenomes and viral size fraction metagenomes (viromes). Size fractionation prior to DNA extraction reduced sources of nonviral DNA in viromes, enabling the recovery of a vaster richness of viral populations (vOTUs), greater viral taxonomic diversity, broader range of predicted hosts, and better access to the rare virosphere, relative to total metagenomes, which tended to recover only the most persistent and abundant vOTUs. Of 2961 detected vOTUs, 2684 were recovered exclusively from viromes, while only three were recovered from total metagenomes alone. Both viral and microbial communities differed significantly over time, suggesting a coupled response to rhizosphere recruitment processes and/or nitrogen amendments. Viral communities alone were also structured along an 18 m spatial gradient. Overall, our results highlight the utility of soil viromics and reveal similarities between viral and microbial community dynamics throughout the tomato growing season yet suggest a partial decoupling of the processes driving their spatial distributions, potentially due to differences in dispersal, decay rates, and/or sensitivities to soil heterogeneity.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome

Bacteriophages have important roles in the ecology of the human gut microbiome but are under-represented in reference databases. To address this problem, we assembled the Metagenomic Gut Virus catalogue that comprises 189,680 viral genomes from 11,810 publicly available human stool metagenomes. Over 75% of genomes represent double-stranded DNA phages that infect members of the Bacteroidia and Clostridia classes. Based on sequence clustering we identified 54,118 candidate viral species, 92% of which were not found in existing databases. The Metagenomic Gut Virus catalogue improves detection of viruses in stool metagenomes and accounts for nearly 40% of CRISPR spacers found in human gut Bacteria and Archaea. We also produced a catalogue of 459,375 viral protein clusters to explore the functional potential of the gut virome. This revealed tens of thousands of diversity-generating retroelements, which use error-prone reverse transcription to mutate target genes and may be involved in the molecular arms race between phages and their bacterial hosts.

59 BASIC BIOLOGICAL SCIENCES↗

CheckV assesses the quality and completeness of metagenome-assembled viral genomes

Abstract Millions of new viral sequences have been identified from metagenomes, but the quality and completeness of these sequences vary considerably. Here we present CheckV, an automated pipeline for identifying closed viral genomes, estimating the completeness of genome fragments and removing flanking host regions from integrated proviruses. CheckV estimates completeness by comparing sequences with a large database of complete viral genomes, including 76,262 identified from a systematic search of publicly available metagenomes, metatranscriptomes and metaviromes. After validation on mock datasets and comparison to existing methods, we applied CheckV to large and diverse collections of metagenome-assembled viral sequences, including IMG/VR and the Global Ocean Virome. This revealed 44,652 high-quality viral genomes (that is, >90% complete), although the vast majority of sequences were small fragments, which highlights the challenge of assembling viral genomes from short-read metagenomes. Additionally, we found that removal of host contamination substantially improved the accurate identification of auxiliary metabolic genes and interpretation of viral-encoded functions.

59 BASIC BIOLOGICAL SCIENCES↗

Improvement of eukaryotic protein predictions from soil metagenomes

During the last decades, metagenomics has highlighted the diversity of microorganisms from environmental or host-associated samples. Most metagenomics public repositories use annotation pipelines tailored for prokaryotes regardless of the taxonomic origin of contigs. Consequently, eukaryotic contigs with intrinsically different gene features, are not optimally annotated. Using a bioinformatics pipeline, we have filtered 7.9 billion contigs from 6,872 soil metagenomes in the JGI’s IMG/M database to identify eukaryotic contigs. We have re-annotated genes using eukaryote-tailored methods, yielding 8 million eukaryotic proteins and over 300,000 orphan proteins lacking homology in public databases. Comparing the gene predictions we made with initial JGI ones on the same contigs, we confirmed our pipeline improves eukaryotic proteins completeness and contiguity in soil metagenomes. The improved quality of eukaryotic proteins combined with a more comprehensive assignment method yielded more reliable taxonomic annotation. This dataset of eukaryotic soil proteins with improved completeness, quality and taxonomic annotation reliability is of interest for any scientist aiming at studying the composition, biological functions and gene flux in soil communities involving eukaryotes.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic features of bioburden serve as outcome indicators in combat extremity wounds

Abstract Battlefield injury management requires specialized care, and wound infection is a frequent complication. Challenges related to characterizing relevant pathogens further complicates treatment. Applying metagenomics to wounds offers a comprehensive path toward assessing microbial genomic fingerprints and could indicate prognostic variables for future decision support tools. Wound specimens from combat-injured U.S. service members, obtained during surgical debridements before delayed wound closure, were subjected to whole metagenome analysis and targeted enrichment of antimicrobial resistance genes. Results did not indicate a singular, common microbial metagenomic profile for wound failure, instead reflecting a complex microenvironment with varying bioburden diversity across outcomes. Genus-level Pseudomonas detection was associated with wound failure at all surgeries. A logistic regression model was fit to the presence and absence of antimicrobial resistance classes to assess associations with nosocomial pathogens. A. baumannii detection was associated with detection of genomic signatures for resistance to trimethoprim, aminoglycosides, bacitracin, and polymyxin. Machine learning classifiers were applied to identify wound and microbial variables associated with outcome. Feature importance rankings averaged across models indicated the variables with the largest effects on predicting wound outcome, including an increase in P. putida sequence reads. These results describe the microbial genomic determinants in combat wound bioburden and demonstrate metagenomic investigation as a comprehensive tool for providing information toward aiding treatment of combat-related injuries.

59 BASIC BIOLOGICAL SCIENCES↗

Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses

ABSTRACT Although long-read sequencing has enabled obtaining high-quality and complete genomes from metagenomes, many challenges still remain to completely decompose a metagenome into its constituent prokaryotic and viral genomes. This study focuses on decomposing an estuarine metagenome to obtain a more accurate estimate of microbial diversity. To achieve this, we developed a new bead-based DNA extraction method, a novel bin refinement method, and obtained 150 Gbp of Nanopore sequencing. We estimate that there are ~500 bacterial and archaeal species in our sample and obtained 68 high-quality bins (>90% complete, <5% contamination, ≤5 contigs, contig length of >100 kbp, and all ribosomal and tRNA genes). We also obtained many contigs of picoeukaryotes, environmental DNA of larger eukaryotes such as mammals, and complete mitochondrial and chloroplast genomes and detected ~40,000 viral populations. Our analysis indicates that there are only a few strains that comprise most of the species abundances. IMPORTANCE Ocean and estuarine microbiomes play critical roles in global element cycling and ecosystem function. Despite the importance of these microbial communities, many species still have not been cultured in the lab. Environmental sequencing is the primary way the function and population dynamics of these communities can be studied. Long-read sequencing provides an avenue to overcome limitations of short-read technologies to obtain complete microbial genomes but comes with its own technical challenges, such as needed sequencing depth and obtaining high-quality DNA. We present here new sampling and bioinformatics methods to attempt decomposing an estuarine microbiome into its constituent genomes. Our results suggest there are only a few strains that comprise most of the species abundances from viruses to picoeukaryotes, and to fully decompose a metagenome of this diversity requires 1 Tbp of long-read sequencing. We anticipate that as long-read sequencing technologies continue to improve, less sequencing will be needed.

Lui, Lauren M.↗

Accelerating large scale de novo metagenome assembly using GPUs

Metagenomic workflows involve studying uncultured microorganisms directly from the environment. These environmental samples when processed by modern sequencing machines yield large and complex datasets that exceed the capabilities of metagenomic software. The increasing sizes and complexities of datasets make a strong case for exascale-capable metagenome assemblers. However, the underlying algorithmic motifs are not well suited for GPUs. This poses a challenge since the majority of next-generation supercomputers will rely primarily on GPUs for computation. In this paper we present the first of its kind GPU-Accelerated implementation of the local assembly approach that is an integral part of a widely used large-scale metagenome assembler, MetaHipMer. Local assembly uses algorithms that induce random memory accesses and non-deterministic workloads, which make GPU offloading a challenging task. Our GPU implementation outperforms the CPU version by about 7x and boosts the performance of MetaHipMer by 42% when running on 64 Summit nodes.

Awan, Muaaz Gul↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

Genome-Resolved Metagenomics of Nitrogen Transformations in the Switchgrass Rhizosphere Microbiome on Marginal Lands

Switchgrass (Panicum virgatum L.) remains the preeminent American perennial (C4) bioenergy crop for cellulosic ethanol, that could help displace over a quarter of the US current petroleum consumption. Intriguingly, there is often little response to nitrogen fertilizer once stands are established. The rhizosphere microbiome plays a critical role in nitrogen cycling and overall plant nutrient uptake. We used high-throughput metagenomic sequencing to characterize the switchgrass rhizosphere microbial community before and after a nitrogen fertilization event for established stands on marginal land. We examined community structure and bulk metabolic potential, and resolved 29 individual bacteria genomes via metagenomic de novo assembly. Community structure and diversity were not significantly different before and after fertilization; however, the bulk metabolic potential of carbohydrate-active enzymes was depleted after fertilization. We resolved 29 metagenomic assembled genomes, including some from the ‘most wanted’ soil taxa such as Verrucomicrobia, Candidate phyla UBA10199, Acidobacteria (rare subgroup 23), Dormibacterota, and the very rare Candidatus Eisenbacteria. The Dormibacterota (formally candidate division AD3) we identified have the potential for autotrophic CO utilization, which may impact carbon partitioning and storage. Our study also suggests that the rhizosphere microbiome may be involved in providing associative nitrogen fixation (ANF) via the novel diazotroph Janthinobacterium to switchgrass.

60 APPLIED LIFE SCIENCES↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

Activity-based, genome-resolved metagenomics uncovers key populations and pathways involved in subsurface conversions of coal to methane

Microbial metabolisms and interactions that facilitate subsurface conversions of recalcitrant carbon to methane are poorly understood. We deployed an in situ enrichment device in a subsurface coal seam in the Powder River Basin (PRB), USA, and used BONCAT-FACS-Metagenomics to identify translationally active populations involved in methane generation from a variety of coal-derived aromatic hydrocarbons. From the active fraction, high-quality metagenome-assembled genomes (MAGs) were recovered for the acetoclastic methanogen, Methanothrix paradoxum, and a novel member of the Chlorobi with the potential to generate acetate via the Pta-Ack pathway. Members of the Bacteroides and Geobacter also encoded Pta-Ack and together, all four populations had the putative ability to degrade ethylbenzene, phenylphosphate, phenylethanol, toluene, xylene, and phenol. Metabolic reconstructions, gene analyses, and environmental parameters also indicated that redox fluctuations likely promote facultative energy metabolisms in the coal seam. The active "Chlorobi PRB" MAG encoded enzymes for fermentation, nitrate reduction, and multiple oxygenases with varying binding affinities for oxygen. "M. paradoxum PRB" encoded an extradiol dioxygenase for aerobic phenylacetate degradation, which was also present in previously published Methanothrix genomes. Finaly, these observations outline underlying processes for bio-methane from subbituminous coal by translationally active populations and demonstrate activity-based metagenomics as a powerful strategy in next generation physiology to understand ecologically relevant microbial populations.

59 BASIC BIOLOGICAL SCIENCES↗

Vitamin interdependencies predicted by metagenomics-informed network analyses and validated in microbial community microcosms

Abstract Metagenomic or metabarcoding data are often used to predict microbial interactions in complex communities, but these predictions are rarely explored experimentally. Here, we use an organism abundance correlation network to investigate factors that control community organization in mine tailings-derived laboratory microbial consortia grown under dozens of conditions. The network is overlaid with metagenomic information about functional capacities to generate testable hypotheses. We develop a metric to predict the importance of each node within its local network environments relative to correlated vitamin auxotrophs, and predict that a Variovorax species is a hub as an important source of thiamine. Quantification of thiamine during the growth of Variovorax in minimal media show high levels of thiamine production, up to 100 mg/L. A few of the correlated thiamine auxotrophs are predicted to produce pantothenate, which we show is required for growth of Variovorax , supporting that a subset of vitamin-dependent interactions are mutualistic. A Cryptococcus yeast produces the B-vitamin pantothenate, and co-culturing with Variovorax leads to a 90-130-fold fitness increase for both organisms. Our study demonstrates the predictive power of metagenome-informed, microbial consortia-based network analyses for identifying microbial interactions that underpin the structure and functioning of microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking second and third-generation sequencing platforms for microbial metagenomics

Shotgun metagenomic sequencing is a common approach for studying the taxonomic diversity and metabolic potential of complex microbial communities. Current methods primarily use second generation short read sequencing, yet advances in third generation long read technologies provide opportunities to overcome some of the limitations of short read sequencing. Here, we compared seven platforms, encompassing second generation sequencers (Illumina HiSeq 300, MGI DNBSEQ-G400 and DNBSEQ-T7, ThermoFisher Ion GeneStudio S5 and Ion Proton P1) and third generation sequencers (Oxford Nanopore Technologies MinION R9 and Pacific Biosciences Sequel II). We constructed three uneven synthetic microbial communities composed of up to 87 genomic microbial strains DNAs per mock, spanning 29 bacterial and archaeal phyla, and representing the most complex and diverse synthetic communities used for sequencing technology comparisons. Our results demonstrate that third generation sequencing have advantages over second generation platforms in analyzing complex microbial communities, but require careful sequencing library preparation for optimal quantitative metagenomic analysis. Our sequencing data also provides a valuable resource for testing and benchmarking bioinformatics software for metagenomics.

59 BASIC BIOLOGICAL SCIENCES↗

ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data

Background: Metagenomics and metatranscriptomics studies provide valuable insight into the composition and function of microbial populations from diverse environments, however the data processing pipelines that rely on mapping reads to gene catalogs or genome databases for cultured strains yield results that underrepresent the genes and functional potential of uncultured microbes. Recent improvements in sequence assembly methods have eased the reliance on genome databases, thereby allowing the recovery of genomes from uncultured microbes. However, configuring these tools, linking them with advanced binning and annotation tools, and maintaining provenance of the processing continues to be challenging for researchers. Results: Here we present ATLAS, a software package for customizable data processing from raw sequence reads to functional and taxonomic annotations using state-of-the-art tools to assemble, annotate, quantify, and bin metagenome and metatranscriptome data. Genome-centric resolution and abundance estimates are provided for each sample in a dataset. ATLAS is written in Python and the workflow implemented in Snakemake; it operates in a Linux environment, and is compatible with Python 3.5+ and Anaconda 3+ versions. The source code for ATLAS is freely available, distributed under a BSD-3 license. Conclusions: ATLAS provides a user-friendly, modular and customizable Snakemake workflow for metagenome and metatranscriptome data processing; it is easily installable with conda and maintained as open-source on GitHub at https://github.com/metagenome-atlas/atlas.

59 BASIC BIOLOGICAL SCIENCES↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

A method for achieving complete microbial genomes and improving bins from metagenomics data

Metagenomics facilitates the study of the genetic information from uncultured microbes and complex microbial communities. Assembling complete genomes from metagenomics data is difficult because most samples have high organismal complexity and strain diversity. Some studies have attempted to extract complete bacterial, archaeal, and viral genomes and often focus on species with circular genomes so they can help confirm completeness with circularity. However, less than 100 circularized bacterial and archaeal genomes have been assembled and published from metagenomics data despite the thousands of datasets that are available. Circularized genomes are important for (1) building a reference collection as scaffolds for future assemblies, (2) providing complete gene content of a genome, (3) confirming little or no contamination of a genome, (4) studying the genomic context and synteny of genes, and (5) linking protein coding genes to ribosomal RNA genes to aid metabolic inference in 16S rRNA gene sequencing studies. We developed a semi-automated method called Jorg to help circularize small bacterial, archaeal, and viral genomes using iterative assembly, binning, and read mapping. In addition, this method exposes potential misassemblies from k-mer based assemblies. We chose species of the Candidate Phyla Radiation (CPR) to focus our initial efforts because they have small genomes and are only known to have one ribosomal RNA operon. In addition to 34 circular CPR genomes, we present one circular Margulisbacteria genome, one circular Chloroflexi genome, and two circular megaphage genomes from 19 public and published datasets. We demonstrate findings that would likely be difficult without circularizing genomes, including that ribosomal genes are likely not operonic in the majority of CPR, and that some CPR harbor diverged forms of RNase P RNA. Code and a tutorial for this method is available at https://github.com/lmlui/Jorg and is available on the DOE Systems Biology KnowledgeBase as a beta app.

59 BASIC BIOLOGICAL SCIENCES↗