Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “viral genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Re-Evaluation of African Swine Fever Genotypes Based on p72 Sequences Reveals the Existence of Only Six Distinct p72 Groups

The African swine fever virus (ASFV) is currently causing a world-wide pandemic of a highly lethal disease in domestic swine and wild boar. Currently, recombinant ASF live-attenuated vaccines based on a genotype II virus strain are commercially available in Vietnam. With 25 reported ASFV genotypes in the literature, it is important to understand the molecular basis and usefulness of ASFV genotyping, as well as the true significance of genotypes in the epidemiology, transmission, evolution, control, and prevention of ASFV. Historically, genotyping of ASFV was used for the epidemiological tracking of the disease and was based on the analysis of small fragments that represent less than 1% of the viral genome. The predominant method for genotyping ASFV relies on the sequencing of a fragment within the gene encoding the structural p72 protein. Genotype assignment has been accomplished through automated phylogenetic trees or by comparing the target sequence to the most closely related genotyped p72 gene. To evaluate its appropriateness for the classification of genotypes by p72, we reanalyzed all available genomic data for ASFV. We conclude that the majority of p72-based genotypes, when initially created, were neither identified under any specific methodological criteria nor correctly compared with the already existing ASFV genotypes. Based on our analysis of the p72 protein sequences, we propose that the current twenty-five genotypes, created exclusively based on the p72 sequence, should be reduced to only six genotypes. To help differentiate between the new and old genotype classification systems, we propose that Arabic numerals (1, 2, 8, 9, 15, and 23) be used instead of the previously used Roman numerals. Furthermore, we discuss the usefulness of genotyping ASFV isolates based only on the p72 gene sequence.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of Selenium and Vitamin E Deficiency on Zika Virus Pathogenesis and Immune Response in Mice

Micronutrient status is recognized to influence host susceptibility to viral infections, yet its impact on Zika virus (ZIKV) pathogenesis remains incompletely understood. We investigated the effects of dietary selenium and combined selenium plus vitamin E deficiency on ZIKV infection outcomes in a type I interferon α/β receptor knockout (Ifnar1 −/− ) murine model. Mice maintained on deficient diets exhibited significantly lower neutralizing antibody titers and reduced levels of key antiviral cytokines (IFN-γ, TNF-α, IFN-α, IFN-β, IL-12p70, CCL5) compared to controls. Correspondingly, higher viral RNA loads were detected in the brains of double-deficient mice, which also experienced greater weight loss and increased mortality. Deep sequencing revealed no major differences in overall viral genome diversity across diet groups; however, specific mutations, including V330L and D67E in the E gene, and V360I in the NS3 gene, were enriched or detected in nutritionally deficient animals. These findings suggest that antioxidant micronutrient deficiency impairs both humoral and cellular immune responses to ZIKV, potentially facilitating enhanced neuroinvasion. While the functional consequences of the identified mutations warrant further investigation, our results underscore the importance of adequate micronutrient intake for optimal antiviral defense. Further studies are needed to clarify the epidemiological significance of these observations.

Biological and medical sciences↗

Resistance of virus to extinction on bottleneck passages: study of a decaying and fluctuating pattern of fitness loss

RNA viruses display high mutation rates and their populations replicate as dynamic and complex mutant distributions, termed viral quasispecies. Repeated genetic bottlenecks, which experimentally are carried out through serial plaque-to-plaque transfers of the virus, lead to fitness decrease (measured here as diminished capacity to produce infectious progeny). Here we report an analysis of fitness evolution of several low fitness foot-and-mouth disease virus clones subjected to 50 plaque-to-plaque transfers. Unexpectedly, fitness decrease, rather than being continuous and monotonic, displayed a fluctuating pattern, which was influenced by both the virus and the state of the host cell as shown by effects of recent cell passage history. The amplitude of the fluctuations increased as fitness decreased, resulting in a remarkable resistance of virus to extinction. Whereas the frequency distribution of fitness in control (independent) experiments follows a log-normal distribution, the probability of fitness values in the evolving bottlenecked populations fitted a Weibull distribution. We suggest that multiple functions of viral genomic RNA and its encoded proteins, subjected to high mutational pressure, interact with cellular components to produce this nontrivial, fluctuating pattern.

Serial Passage↗

Cell and molecular biology of simian virus 40: implications for human infections and disease

Simian virus 40 (SV40), a polyomavirus of rhesus macaque origin, was discovered in 1960 as a contaminant of polio vaccines that were distributed to millions of people from 1955 through early 1963. SV40 is a potent DNA tumor virus that induces tumors in rodents and transforms many types of cells in culture, including those of human origin. This virus has been a favored laboratory model for mechanistic studies of molecular processes in eukaryotic cells and of cellular transformation. The viral replication protein, named large T antigen (T-ag), is also the viral oncoprotein. There is a single serotype of SV40, but multiple strains of virus exist that are distinguishable by nucleotide differences in the regulatory region of the viral genome and in the part of the T-ag gene that encodes the protein's carboxyl terminus. Natural infections in monkeys by SV40 are usually benign but may become pathogenic in immunocompromised animals, and multiple tissues can be infected. SV40 can replicate in certain types of simian and human cells. SV40-neutralizing antibodies have been detected in individuals not exposed to contaminated polio vaccines. SV40 DNA has been identified in some normal human tissues, and there are accumulating reports of detection of SV40 DNA and/or T-ag in a variety of human tumors. This review presents aspects of replication and cell transformation by SV40 and considers their implications for human infections and disease pathogenesis by the virus. Critical assessment of virologic and epidemiologic data suggests a probable causative role for SV40 in certain human cancers, but additional studies are necessary to prove etiology.

Review↗

Methods for determining the genetic affinity of microorganisms and viruses

Selecting which sub-sequences in a database of nucleic acid such as 16S rRNA are highly characteristic of particular groupings of bacteria, microorganisms, fungi, etc. on a substantially phylogenetic tree. Also applicable to viruses comprising viral genomic RNA or DNA. A catalogue of highly characteristic sequences identified by this method is assembled to establish the genetic identity of an unknown organism. The characteristic sequences are used to design nucleic acid hybridization probes that include the characteristic sequence or its complement, or are derived from one or more characteristic sequences. A plurality of these characteristic sequences is used in hybridization to determine the phylogenetic tree position of the organism(s) in a sample. Those target organisms represented in the original sequence database and sufficient characteristic sequences can identify to the species or subspecies level. Oligonucleotide arrays of many probes are especially preferred. A hybridization signal can comprise fluorescence, chemiluminescence, or isotopic labeling, etc.; or sequences in a sample can be detected by direct means, e.g. mass spectrometry. The method's characteristic sequences can also be used to design specific PCR primers. The method uniquely identifies the phylogenetic affinity of an unknown organism without requiring prior knowledge of what is present in the sample. Even if the organism has not been previously encountered, the method still provides useful information about which phylogenetic tree bifurcation nodes encompass the organism.

Fox, George E.↗

Visualization of conformational changes and membrane remodeling leading to genome delivery by viral class-II fusion machinery

Chikungunya virus (CHIKV) is a human pathogen that delivers its genome to the host cell cytoplasm through endocytic low pH-activated membrane fusion mediated by class-II fusion proteins. Though structures of prefusion, icosahedral CHIKV are available, structural characterization of virion interaction with membranes has been limited. Here, we have used cryo-electron tomography to visualize CHIKV’s complete membrane fusion pathway, identifying key intermediary glycoprotein conformations coupled to membrane remodeling events. Using sub-tomogram averaging, we elucidate features of the low pH-exposed virion, nucleocapsid and full-length E1-glycoprotein’s post-fusion structure. Contrary to class-I fusion systems, CHIKV achieves membrane apposition by protrusion of extended E1-glycoprotein homotrimers into the target membrane. The fusion process also features a large hemifusion diaphragm that transitions to a wide pore for intact nucleocapsid delivery. Our analyses provide comprehensive ultrastructural insights into the class-II virus fusion system function and direct mechanistic characterization of the fundamental process of protein-mediated membrane fusion.

59 BASIC BIOLOGICAL SCIENCES↗

Illuminating the pathways to carbon liberation: a systems approach to characterizing the consequential unknowns of carbon transformation and loss from thawing permafrost peatlands (Final Report)

The IsoGenie3 Project delivered new systems-level insights into carbon cycling in thawing permafrost landscapes, with an emphasis on methane and carbon dioxide emissions. From >200 samples from the site collected over a decade, co-analyzed for geochemistry and microbiology, the team recovered ~1,500 assembled microbial genomes and ~1,900 viral population genomes, revealing appreciable genetic novelty - from a new highly abundant bacterial phylum, to novel methane consumers and their activities, to rampant viral novelty. IsoGenie3 linked these organisms to carbon compound transformations (which define the cycling of organic matter in soils, and the loss of the greenhouse gases carbon dioxide and methane), and saw that the microbes at each stage of permafrost thaw had different genetic potential to degrade categories of compounds, expressed that genetic potential differently, and actually transformed carbon compounds into greenhouse gases in different ways. IsoGenie 3 identified that some of the thaw-stage differences were due to plant-microbiome relationships; the plant species across the thaw gradient contributed different carbon compounds into the soil, and hosted distinct microbiota (differing among parts of plants as well as species). Lastly, microbes in the saturated post-thaw conditions appeared likely to contribute to the mobilization and toxification of mercury released during thaw. In parallel with ongoing field sampling and analysis, hypotheses arising from field observations were tested via lab incubation experiments. When communities are taken out of their native habitats, they behave differently, and the team first rigorously quantified the magnitude of this effect on microbiome composition and functional capacity, organic matter composition, and gas production; overall the main system processes were maintained in the lab incubations under the conditions tested. Further, the microbial data could inform geochemical reaction network models of those processes. Then, the team ran experiments with additions of compounds, varying temperature, and “live” vs. “dead” peat (the latter having been gamma irradiated, with a few additional variants to control for methodological artifacts). From these, we (a) determined the importance of plant-derived soluble phenolic compounds in bogs’ extraordinary recalcitrance of organic matter, and carbon gas emissions skewed to carbon dioxide; (b) proposed an abiotic ‘tanning’ mechanism, which could contribute to Sphagnum’s inhibitory effect on anaerobic decomposition through alteration of N availability. IsoGenie3 illuminated longer-term and landscape-scale interactions of permafrost thaw and carbon cycling, advancing knowledge of the drivers of methane dynamics not only across in the permafrost-associated peatland (where hydrology and plant communities dictate microbiomes) but also their interconnected lakes (where sediment carbon quality and resident microbiota are determined by position within lake, and lake features). By leveraging observations of site methane dynamics extending well before this project, the team was able to construct a 44-year portrait of the interplay of permafrost thaw, hydrology, vegetation dynamics, and carbon gas emissions, and the doubling of the fully-thawed fens over this time. From the detailed study of this focal site, IsoGenie3 also aimed to improve model representation of these kinds of sites and processes. To improve predictions of methane transformations, we incorporated acetate and isotope dynamics into the ‘DNDC’ biogeochemistry model. In addition, recovered genomes were grouped into ‘functional groups’, i.e. the genomes that perform a specific function of interest, then used to parameterize maximum growth rate and optimum growth temperature (via signatures in their sequence composition) for the BioCrunch model. The BioCrunch model was then in turn used to test the impact of increasing functional resolution of the microbes, on the carbon gas emissions. Lastly for modeling, the ecosys model was parameterized from the microbial and other data, and used to evaluate drivers of e.g. change in methane emissions. Finally, this project also led to the development of a range of new methods and tools, a new metric of organic matter decomposability, as well as a graph-database solution to multidisciplinary data storage and querying. This project’s ongoing analyses at our focal site also contributed to broader advancements in understanding elements of genetic plasticity and methane metabolism, climate change microbiology and community assembly, global peatland geochemistry and Arctic lakes’ roles in climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Latent Viruses: A Space Travel Hazard??

A major issue associated with long-duration space flight is the possibility of infectious disease causing an unacceptable medical risk to crew members. Our proposal is designed to gain information that addresses several issues outlined in the Immunology/Infectious disease critical path. The major hypothesis addressed is that space flight causes alterations in the immune system that may allow latent viruses which are endogenous in the human population to reactivate and shed to higher levels than normal which can affect the health of crew members during a long term space-flight mission. We will initially focus our studies on the human herpesviruses and human polyomaviruses which are important pathogens known to establish latent infections in the human population. Both primary infection and reactivation from latent infection with this group of viruses can cause a variety of illnesses that result in morbidity and occasionally mortality of infected individuals. Effective vaccines exist for only one of the eight known human herpesviruses and the vaccine itself can still reactivate from latent infection. Available antivirals are of limited use and are effective against only a few of the human herpesviruses. Although most individuals display little if any clinical consequences from latent infection, events which alter immune function such as immunosuppressive therapy following solid organ transplantation are known to increase the risk of developing complications as a result of latent virus reactivation. This proposal will measure both the frequency and magnitude of viral shedding and genome loads in the blood from humans participating in activities that serve as ground based models of space flight conditions. Our initial goal is to develop sensitive quantitative competitive PCR- based assays (QC-PCR) to detect the herpesvirus Epstein-Barr virus (EBV), and the polyomaviruses SV40, BKV, and JCV. Using these assays we will establish baseline patterns of viral genome load in the blood and viral shedding from normal volunteers in a longitudinal study over I year in length. As a comparison, we will measure patterns of viral genome loads and shedding from individuals who are severely immunosuppressed, in whom herpesvirus reactivation or primary infection with a herpesvirus is known to cause complications. In addition, we will proceed to testing ground based analogs in collaboration with Dr. Duane Pierson (Lyndon B. Johnson Space Center). This will include measuring samples obtained from individuals living and working in the extreme environment of Antarctica. We expect to detect viral shedding or reactivation from most of the test groups, although the magnitude of shedding or reactivation cannot be predicted. The data accumulated from studies in this proposal should allow us to evaluate whether events that simulate certain aspects of space flight reactivate viral infections severe enough in nature that they may compromise the success of long-term space flight missions. These studies will also provide a foundation to monitor viral reactivation and shedding from crew members participating in actual space flight missions. We will present data showing the establishment of our QC-PCR assay for detection of EBV.

Ling, P. D.↗

VIBES: a workflow for annotating and visualizing viral sequences integrated into bacterial genomes

Abstract Bacteriophages are viruses that infect bacteria. Many bacteriophages integrate their genomes into the bacterial chromosome and become prophages. Prophages may substantially burden or benefit host bacteria fitness, acting in some cases as parasites and in others as mutualists. Some prophages have been demonstrated to increase host virulence. The increasing ease of bacterial genome sequencing provides an opportunity to deeply explore prophage prevalence and insertion sites. Here we present VIBES (Viral Integrations in Bacterial genomES), a workflow intended to automate prophage annotation in complete bacterial genome sequences. VIBES provides additional context to prophage annotations by annotating bacterial genes and viral proteins in user-provided bacterial and viral genomes. The VIBES pipeline is implemented as a Nextflow-driven workflow, providing a simple, unified interface for execution on local, cluster and cloud computing environments. For each step of the pipeline, a container including all necessary software dependencies is provided. VIBES produces results in simple tab-separated format and generates intuitive and interactive visualizations for data exploration. Despite VIBES’s primary emphasis on prophage annotation, its generic alignment-based design allows it to be deployed as a general-purpose sequence similarity search manager. We demonstrate the utility of the VIBES prophage annotation workflow by searching for 178 Pf phage genomes across 1072 Pseudomonas spp. genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-Resolved Metaproteomics Decodes the Microbial and Viral Contributions to Coupled Carbon and Nitrogen Cycling in River Sediments

Rivers have a significant role in global carbon and nitrogen cycles, serving as a nexus for nutrient transport between terrestrial and marine ecosystems. Although rivers have a small global surface area, they contribute substantially to worldwide greenhouse gas emissions through microbially mediated processes within the river hyporheic zone. Despite this importance, research linking microbial and viral communities to specific biogeochemical reactions is still nascent in these sediment environments. To survey the metabolic potential and gene expression underpinning carbon and nitrogen biogeochemical cycling in river sediments, we collected an integrated data set of 33 metagenomes, metaproteomes, and paired metabolomes. We reconstructed over 500 microbial metagenome-assembled genomes (MAGs), which we dereplicated into 55 unique, nearly complete medium- and high-quality MAGs spanning 12 bacterial and archaeal phyla. We also reconstructed 2,482 viral genomic contigs, which were dereplicated into 111 viral MAGs (vMAGs) of >10 kb in size. As a result of integrating gene expression data with geochemical and metabolite data, we created a conceptual model that uncovered new roles for microorganisms in organic matter decomposition, carbon sequestration, nitrogen mineralization, nitrification, and denitrification. We show how these metabolic pathways, integrated through shared resource pools of ammonium, carbon dioxide, and inorganic nitrogen, could ultimately contribute to carbon dioxide and nitrous oxide fluxes from hyporheic sediments. Further, by linking viral MAGs to these active microbial hosts, we provide some of the first insights into viral modulation of river sediment carbon and nitrogen cycling.

54 ENVIRONMENTAL SCIENCES↗

Dispersal, habitat filtering, and eco-evolutionary dynamics as drivers of local and global wetland viral biogeography

Abstract Wetlands store 20–30% of the world’s soil carbon, and identifying the microbial controls on these carbon reserves is essential to predicting feedbacks to climate change. Although viral infections likely play important roles in wetland ecosystem dynamics, we lack a basic understanding of wetland viral ecology. Here 63 viral size-fraction metagenomes (viromes) and paired total metagenomes were generated from three time points in 2021 at seven fresh- and saltwater wetlands in the California Bodega Marine Reserve. We recovered 12,826 viral population genomic sequences (vOTUs), only 4.4% of which were detected at the same field site two years prior, indicating a small degree of population stability or recurrence. Viral communities differed most significantly among the seven wetland sites and were also structured by habitat (plant community composition and salinity). Read mapping to a new version of our reference database, PIGEONv2.0 (515,763 vOTUs), revealed 196 vOTUs present over large geographic distances, often reflecting shared habitat characteristics. Wetland vOTU microdiversity was significantly lower locally than globally and lower within than between time points, indicating greater divergence with increasing spatiotemporal distance. Viruses tended to have broad predicted host ranges via CRISPR spacer linkages to metagenome-assembled genomes, and increased SNP frequencies in CRISPR-targeted major tail protein genes suggest potential viral eco-evolutionary dynamics in response to both immune targeting and changes in host cell receptors involved in viral attachment. Together, these results highlight the importance of dispersal, environmental selection, and eco-evolutionary dynamics as drivers of local and global wetland viral biogeography.

Environmental Sciences & Ecology↗

RNA nanotechnology to build a dodecahedral genome of single-stranded RNA virus

The quest for artificial RNA viral complexes with authentic structure while being non-replicative is on its way for the development of viral vaccines. RNA viruses contain capsid proteins that interact with the genome during morphogenesis. The sequence and properties of the protein and genome determine the structure of the virus. For example, the Pariacoto virus ssRNA genome assembles into a dodecahedron. Virus-inspired nanotechnology has progressed remarkably due to the unique structural and functional properties of viruses, which can inspire the design of novel nanomaterials. RNA is a programmable biopolymer able to self-assemble sophisticated 3D structures with rich functionalities. RNA dodecahedrons mimicking the Pariacoto virus quasi-icosahedral genome structures were constructed from both native and 2'-F modified RNA oligos. The RNA dodecahedron easily self-assembled using the stable pRNA three-way junction of bacteriophage phi29 as building blocks. The RNA dodecahedron cage was further characterized by cryo-electron microscopy and atomic force microscopy, confirming the spontaneous and homogenous formation of the RNA cage. The reported RNA dodecahedron cage will likely provide further studies on the mechanisms of interaction of the capsid protein with the viral genome while providing a template for further construction of the viral RNA scaffold to add capsid proteins for the assembly of the viral nucleocapsid as a model. Understanding the self-assembly and RNA folding of this RNA cage may offer new insights into the 3D organization of viral RNA genomes. Finally, the reported RNA cage also has the potential to be explored as a novel virus-inspired nanocarrier.

59 BASIC BIOLOGICAL SCIENCES↗

Lytic archaeal viruses infect abundant primary producers in Earth’s crust

The continental subsurface houses a major portion of life’s abundance and diversity, yet little is known about viruses infecting microbes that reside there. Here, we use a combination of metagenomics and virus-targeted direct-geneFISH (virusFISH) to show that highly abundant carbon-fixing organisms of the uncultivated genus Candidatus Altiarchaeum are frequent targets of previously unrecognized viruses in the deep subsurface. Analysis of CRISPR spacer matches display resistances of Ca. Altiarchaea against eight predicted viral clades, which show genomic relatedness across continents but little similarity to previously identified viruses. Based on metagenomic information, we tag and image a putatively viral genome rich in protospacers using fluorescence microscopy. VirusFISH reveals a lytic lifestyle of the respective virus and challenges previous predictions that lysogeny prevails as the dominant viral lifestyle in the subsurface. CRISPR development over time and imaging of 18 samples from one subsurface ecosystem suggest a sophisticated interplay of viral diversification and adapting CRISPR-mediated resistances of Ca. Altiarchaeum. We conclude that infections of primary producers with lytic viruses followed by cell lysis potentially jump-start heterotrophic carbon cycling in these subsurface ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of hidden N4-like viruses and their interactions with hosts

The N4-like viruses, which were recently assigned to the novel viral family Schitoviridae in 2021, belong to a podoviral-like viral lineage and possess conserved genomic characteristics and a unique replication mechanism. Despite their significance, our understanding of N4-like viruses is primarily based on viral isolates. To address this knowledge gap, this study has established a comprehensive N4-like viral data sets comprising 342 high-quality N4-like viruses/proviruses (144 viral isolates, 158 uncultured viruses, and 40 integrated N4-like proviruses). These viruses were classified into 97 subfamilies (89 of which are newly identified), 148 genera (100 of which are newly identified), and 253 species (177 of which are newly identified). The study reveals that N4-like viruses inhibit the polar region, oligotrophic open oceans, and the human gut, where they infect various bacterial lineages, such as Alpha/Beta/Gamma/Epsilon-proteobacteria in the Proteobacteria phylum. Although N4-like viral endogenization appears to be prevalent in Proteobacteria, it has also been observed in Firmicutes. Additionally, the phylogenetic analysis has identified evolutionary divergence within the hallmark genes of N4-like viruses, indicating a complex origin of the different conserved parts of viral genomes. Moreover, 1,101 putative auxiliary metabolic genes (AMGs) were identified in the N4-like viral pan-proteome, which mainly participate in nucleotide and cofactor/vitamin metabolisms. Of these AMGs, 27 were found to be associated with virulence, suggesting their potential involvement in the spread of bacterial pathogenicity. The findings of this study are significant, as N4-like viruses represent a unique viral lineage with a distinct replication mechanism and a conserved core genome. This work has resulted in a comprehensive global map of the entire N4-like viral lineage, including information on their distribution in different biomes, evolutionary divergence, genomic diversity, and the potential for viral-mediated host metabolic reprogramming. As such, this work significantly contributes to our understanding of the ecological function and viral-host interactions of bacteriophages.

60 APPLIED LIFE SCIENCES↗

NCBI’s Virus Discovery Codeathon: Building “FIVE” —The Federated Index of Viral Experiments API Index

Viruses represent important test cases for data federation due to their genome size and the rapid increase in sequence data in publicly available databases. However, some consequences of previously decentralized (unfederated) data are lack of consensus or comparisons between feature annotations. Unifying or displaying alternative annotations should be a priority both for communities with robust entry representation and for nascent communities with burgeoning data sources. To this end, during this three-day continuation of the Virus Hunting Toolkit codeathon series (VHT-2), a new integrated and federated viral index was elaborated. This Federated Index of Viral Experiments (FIVE) integrates pre-existing and novel functional and taxonomy annotations and virus–host pairings. Variability in the context of viral genomic diversity is often overlooked in virus databases. As a proof-of-concept, FIVE was the first attempt to include viral genome variation for HIV, the most well-studied human pathogen, through viral genome diversity graphs. As per the publication of this manuscript, FIVE is the first implementation of a virus-specific federated index of such scope. FIVE is coded in BigQuery for optimal access of large quantities of data and is publicly accessible. Many projects of database or index federation fail to provide easier alternatives to access or query information. To this end, a Python API query system was developed to enhance the accessibility of FIVE.

59 BASIC BIOLOGICAL SCIENCES↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

iPHoP: An integrated machine learning framework to maximize host prediction for metagenome-derived viruses of archaea and bacteria

The extraordinary diversity of viruses infecting bacteria and archaea is now primarily studied through metagenomics. While metagenomes enable high-throughput exploration of the viral sequence space, metagenome-derived sequences lack key information compared to isolated viruses, in particular host association. Different computational approaches are available to predict the host(s) of uncultivated viruses based on their genome sequences, but thus far individual approaches are limited either in precision or in recall, i.e., for a number of viruses they yield erroneous predictions or no prediction at all. Here, we describe iPHoP, a two-step framework that integrates multiple methods to reliably predict host taxonomy at the genus rank for a broad range of viruses infecting bacteria and archaea, while retaining a low false discovery rate. Based on a large dataset of metagenome-derived virus genomes from the IMG/VR database, we illustrate how iPHoP can provide extensive host prediction and guide further characterization of uncultivated viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Three-dimensional structure of a flavivirus dumbbell RNA reveals molecular details of an RNA regulator of replication

Abstract Mosquito-borne flaviviruses (MBFVs) including dengue, West Nile, yellow fever, and Zika viruses have an RNA genome encoding one open reading frame flanked by 5′ and 3′ untranslated regions (UTRs). The 3′ UTRs of MBFVs contain regions of high sequence conservation in structured RNA elements known as dumbbells (DBs). DBs regulate translation and replication of the viral RNA genome, functions proposed to depend on the formation of an RNA pseudoknot. To understand how DB structure provides this function, we solved the x-ray crystal structure of the Donggang virus DB to 2.1Å resolution and used structural modeling to reveal the details of its three-dimensional fold. The structure confirmed the predicted pseudoknot and molecular modeling revealed how conserved sequences form a four-way junction that appears to stabilize the pseudoknot. Single-molecule FRET suggests that the DB pseudoknot is a stable element that can regulate the switch between translation and replication during the viral lifecycle by modulating long-range RNA conformational changes.

Akiyama, Benjamin M.↗