Engineering PapersSearch

Engineering topics

Roux, Simon

Publications and source records attributed to Roux, Simon.

At least 19 records

Mobile genetic elements shape microbial diversity and functions in thawing permafrost soils

Ecosystems are shaped by communities of microorganisms whose niches and impacts depend on functional profiles influenced by gene gains and losses. Culture-based experiments demonstrate that mobile genetic elements (MGEs) can mediate gene flux, but quantitative understanding of these dynamics in natural systems remains limited. Here we develop and apply a systematic, meta-omic framework to investigate MGEs in a complex natural system using an 8-year soil time series collected at Stordalen Mire, in Sweden’s thawing permafrost margin. In this climate-critical peatland, we identify ~2.1 million MGE recombinases across 89 microbial phyla and assess ecological distributions, affected functions, past mobility and current activity. This revealed an active mobilome that shapes natural genetic diversity via differential impacts on major phyla and affects a wide range of functions, including metabolic genes involved in carbon flux and nutrient cycling. These findings and this analytic framework suggest avenues towards a better understanding of MGE diversity, activity, mobility and impacts across ecosystems.

Biological and medical sciences

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B

Viromics approaches for the study of viral diversity and ecology in microbiomes

Viruses are found across all ecosystems and infect every type of organism on Earth. Traditional culture-based methods have proven insufficient to explore this viral diversity at scale, driving the development of viromics, the sequence-based analysis of uncultivated viruses. Viromics approaches have been particularly useful for studying viruses of microorganisms, which can act as crucial regulators of microbiomes across ecosystems. They have already revealed the broad geographic distribution of viral communities and are progressively uncovering the expansive genetic and functional diversity of the global virome. Moving forward, large-scale viral ecogenomics studies combined with new experimental and computational approaches to identify virus activity and host interactions will enable a more complete characterization of global viral diversity and its effects.

Ecology

A call for caution in the biological interpretation of viral auxiliary metabolic genes

Virus-encoded auxiliary metabolic genes (AMGs) are non-essential genes that increase viral fitness by maintaining or manipulating host metabolism during infection. AMGs are intriguing from an evolutionary perspective, as most viral genomes are highly compact and have limited coding capacity for accessory genes. Advances in viral (meta)genomics have expanded the detection of putative AMGs from viruses in diverse environments. However, this has also led to many instances of misannotation due to the limitations of annotation tools, resulting in misinterpretations about the roles of some viral genes. Here, we highlight studies that support claims about AMGs with more than just function predictions for guidance on best practices. We then propose the adoption of an expanded, inclusive view of all genes auxiliary to core viral functions with the term ‘auxiliary viral genes’ (AVGs), alongside an associated eco-evolutionary framework for considering the types of analyses that can better support claims made about AVGs.

Environmental microbiology

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant

Virus species names have been standardized; virus names remain unchanged

Virus taxonomy, comprising classification and nomenclature, is regulated by the International Committee on Taxonomy of Viruses (ICTV). Taxon names are standardized to facilitate recognition and communication, with defined suffixes for each rank (e.g., the names of orders, families, and genera end in -virales, -viridae, and -virus, respectively). However, until recently, a standard format for species names was lacking. In 2021, following extensive discussion and community consultation, the ICTV decided to adopt a standardized binomial (Linnaean) format for virus species names, consisting of the genus name followed by a "freeform" species epithet. Previously assigned virus species names that were non-compliant with the binomial format have been fully updated. In contrast to taxon names regulated by the ICTV, the names of viruses, or "common" names, such as yellow fever virus or human immunodeficiency virus, are not under the remit of the ICTV and have not been changed.

Zerbini, F Murilo

Population ecology and biogeochemical implications of ssDNA and dsDNA viruses along a permafrost thaw gradient

Anthropogenic-driven climate change is accelerating permafrost thaw, threatening to release vast carbon stores through increased microbial activity. While microbial roles are increasingly studied, the contributions of viruses remain largely unexplored, in part due to soil-associated technical challenges that have hindered their detection and characterization. Here, we applied an optimized virion enrichment workflow along a permafrost thaw gradient, identifying 9,963 viral populations (vOTUs), including single- and double-stranded DNA viruses, with 99.9% novelty compared to other soils. Hosts were predicted for 38% of vOTUs, spanning nine archaeal, and 36 bacterial phyla, 22% of which were linked to metagenome-assembled genomes, including key carbon-cycling taxa. Genomic analyses revealed 811 putative auxiliary metabolic genes (AMGs) from 658 vOTUs, nearly half involved in carbon processing. These included 59 glycoside hydrolases (GH) across nine GH families, 45 for monosaccharide degradation, and seven involved in short-chain fatty acid and C1 metabolism, linking viruses to both early and late stages of carbon turnover. Additionally, six vOTUs carried racD, which may stabilize microbial necromass and promote long-term carbon storage. Viral and AMG functional diversity increased with thaw stage, indicating that viruses might participate in a broadening range of microbial metabolic processes as permafrost thaws. These findings expand our understanding of virus contributions in microbial carbon processing and suggest their important role in deciphering soil carbon fate under changing climate conditions.

Biological and medical sciences

Breaking the reproducibility barrier with standardized protocols for plant–microbiome research

Inter-laboratory replicability is crucial yet challenging in microbiome research. Leveraging microbiomes to promote soil health and plant growth requires understanding underlying molecular mechanisms using reproducible experimental systems. In a global collaborative effort involving five laboratories, we aimed to help advance reproducibility in microbiome studies by testing our ability to replicate synthetic community assembly experiments. Our study compared fabricated ecosystems constructed using two different synthetic bacterial communities, the model grass Brachypodium distachyon, and sterile EcoFAB 2.0 devices. All participating laboratories observed consistent inoculum-dependent changes in plant phenotype, root exudate composition, and final bacterial community structure, where Paraburkholderia sp. OAS925 could dramatically shift microbiome composition. Comparative genomics and exudate utilization linked the pH-dependent colonization ability of Paraburkholderia, which was further confirmed with motility assays. The study provides detailed protocols, benchmarking datasets, and best practices to help advance replicable science and inform future multi-laboratory reproducibility studies.

Novak, Vlastimil

Microbial Metagenomes Across a Complete Phytoplankton Bloom Cycle: High-Resolution Sampling Every 4 Hours Over 22 Days

In May and June of 2021, marine microbial samples were collected for DNA sequencing in East Sound, WA, USA every 4 hours for 22 days. This high temporal resolution sampling effort captured the last 3 days of a Rhizosolenia sp. bloom, the initiation and complete bloom cycle of Chaetoceros socialis (8 days), and the following bacterial bloom (2 days). Metagenomes were completed on the time series, and the dataset includes 128 size-fractionated microbial samples (0.22–1.2 µm), providing gene abundances for the dominant members of bacteria, archaea, and viruses. This dataset also has time-matched nutrient analyses, flow cytometry data, and physical parameters of the environment at a single point of sampling within a coastal ecosystem that experiences regular bloom events, facilitating a range of modeling efforts that can be leveraged to understand microbial community structure and their influences on the growth, maintenance, and senescence of phytoplankton blooms.

59 BASIC BIOLOGICAL SCIENCES

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES

Tapping the treasure trove of atypical phages

With advancements in genomics technologies, a vast diversity of ‘atypical’ phages, that is, with single-stranded DNA or RNA genomes, are being uncovered from different ecosystems. Though these efforts have revealed the existence and prevalence of these nonmodel phages, computational approaches often fail to associate these phages with their specific bacterial host(s), while the lack of methods to isolate these phages has limited our ability to characterize infectivity pathways and new gene function. In this review, we call for the development of generalizable experimental methods to better capture this understudied viral diversity via isolation and study them through gene-level characterization and engineering. Establishing a diverse set of new ‘atypical’ phage model systems has the potential to provide many new biotechnologies, including potential uses of these atypical phages in halting the spread of antibiotic resistance and engineering of microbial communities for beneficial outcomes.

59 BASIC BIOLOGICAL SCIENCES

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri

SpacerExtractor v0.9

The SpacerExtractor tool is meant to robustly identify and extract CRISPR spacers from metagenome short reads. Working from a database of known CRISPR repeats, SpacerExtractor quickly scans short reads for the corresponding repeat sequences, extract the potential spacer between two repeats, apply several quality control, denoising, and clustering steps, and provides a full non-redundant complement of spacers for each detected repeat. Because of the high variability observed at CRISPR loci, this read mining approach typically recovers a much larger diversity of spacers than can be found in assembled contigs. SpacerExtractor also includes commands to run CRISPR-Cas Typer on a new set of genomes or MAGs, and add newly predicted repeats to the repeat database.

Bushnell, Brian