Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Impact of short-read sequencing on the misassembly of a plant genome

Abstract Background Availability of plant genome sequences has led to significant advances. However, with few exceptions, the great majority of existing genome assemblies are derived from short read sequencing technologies with highly uneven read coverages indicative of sequencing and assembly issues that could significantly impact any downstream analysis of plant genomes. In tomato for example, 0.6% (5.1 Mb) and 9.7% (79.6 Mb) of short-read based assembly had significantly higher and lower coverage compared to background, respectively. Results To understand what the causes may be for such uneven coverage, we first established machine learning models capable of predicting genomic regions with variable coverages and found that high coverage regions tend to have higher simple sequence repeat and tandem gene densities compared to background regions. To determine if the high coverage regions were misassembled, we examined a recently available tomato long-read based assembly and found that 27.8% (1.41 Mb) of high coverage regions were potentially misassembled of duplicate sequences, compared to 1.4% in background regions. In addition, using a predictive model that can distinguish correctly and incorrectly assembled high coverage regions, we found that misassembled, high coverage regions tend to be flanked by simple sequence repeats, pseudogenes, and transposon elements. Conclusions Our study provides insights on the causes of variable coverage regions and a quantitative assessment of factors contributing to plant genome misassembly when using short reads and the generality of these causes and factors should be tested further in other species.

59 BASIC BIOLOGICAL SCIENCES↗

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley↗

Genetic and behavioral adaptation of Candida parapsilosis to the microbiome of hospitalized infants revealed by in situ genomics, transcriptomics, and proteomics

Background Candida parapsilosis is a common cause of invasive candidiasis, especially in newborn infants, and infections have been increasing over the past two decades. C. parapsilosis has been primarily studied in pure culture, leaving gaps in understanding of its function in a microbiome context. Results. Here, we compare five unique C. parapsilosis genomes assembled from premature infant fecal samples, three of which are newly reconstructed, and analyze their genome structure, population diversity, and in situ activity relative to reference strains in pure culture. All five genomes contain hotspots of single nucleotide variants, some of which are shared by strains from multiple hospitals. A subset of environmental and hospital-derived genomes share variants within these hotspots suggesting derivation of that region from a common ancestor. Four of the newly reconstructed C. parapsilosis genomes have 4 to 16 copies of the gene RTA3, which encodes a lipid translocase and is implicated in antifungal resistance, potentially indicating adaptation to hospital antifungal use. Time course metatranscriptomics and metaproteomics on fecal samples from a premature infant with a C. parapsilosis blood infection revealed highly variable in situ expression patterns that are distinct from those of similar strains in pure cultures. For example, biofilm formation genes were relatively less expressed in situ, whereas genes linked to oxygen utilization were more highly expressed, indicative of growth in a relatively aerobic environment. In gut microbiome samples, C. parapsilosis co-existed with Enterococcus faecalis that shifted in relative abundance over time, accompanied by changes in bacterial and fungal gene expression and proteome composition. Conclusions The results reveal potentially medically relevant differences in Candida function in gut vs. laboratory environments, and constrain evolutionary processes that could contribute to hospital strain persistence and transfer into premature infant microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗

Insights into the physiological and genomic characterization of three bacterial isolates from a highly alkaline, terrestrial serpentinizing system

The terrestrial serpentinite-hosted ecosystem known as “The Cedars” is home to a diverse microbial community persisting under highly alkaline (pH ~ 12) and reducing (Eh < -550 mV) conditions. This extreme environment presents particular difficulties for microbial life, and efforts to isolate microorganisms from The Cedars over the past decade have remained challenging. Herein, we report the initial physiological assessment and/or full genomic characterization of three isolates: Paenibacillus sp. Cedars (‘Paeni-Cedars’), Alishewanella sp. BS5-314 (‘Ali-BS5-314’), and Anaerobacillus sp. CMMVII (‘Anaero-CMMVII’). Paeni-Cedars is a Gram-positive, rod-shaped, mesophilic facultative anaerobe that grows between pH 7–10 (minimum pH tested was 7), temperatures 20–40°C, and 0–3% NaCl concentration. The addition of 10–20 mM CaCl 2 enhanced growth, and iron reduction was observed in the following order, 2-line ferrihydrite > magnetite > serpentinite ~ chromite ~ hematite. Genome analysis identified genes for flavin-mediated iron reduction and synthesis of a bacillibactin-like, catechol-type siderophore. Ali-BS5-314 is a Gram-negative, rod-shaped, mesophilic, facultative anaerobic alkaliphile that grows between pH 10–12 and temperatures 10–40°C, with limited growth observed 1–5% NaCl. Nitrate is used as a terminal electron acceptor under anaerobic conditions, which was corroborated by genome analysis. The Ali-BS5-314 genome also includes genes for benzoate-like compound metabolism. Anaero-CMMVII remained difficult to cultivate for physiological studies; however, growth was observed between pH 9–12, with the addition of 0.01–1% yeast extract. Anaero-CMMVII is a probable oxygen-tolerant anaerobic alkaliphile with hydrogenotrophic respiration coupled with nitrate reduction, as determined by genome analysis. Based on single-copy genes, ANI, AAI and dDDH analyses, Paeni-Cedars and Ali-BS5-314 are related to other species (P. glucanolyticus and A. aestuarii, respectively), and Anaero-CMMVII represents a new species. The characterization of these three isolates demonstrate the range of ecophysiological adaptations and metabolisms present in serpentinite-hosted ecosystems, including mineral reduction, alkaliphily, and siderophore production.

59 BASIC BIOLOGICAL SCIENCES↗

Hybridization History and Repetitive Element Content in the Genome of a Homoploid Hybrid, Yucca gloriosa (Asparagaceae)

Hybridization in plants results in phenotypic and genotypic perturbations that can have dramatic effects on hybrid physiology, ecology, and overall fitness. Hybridization can also perturb epigenetic control of transposable elements, resulting in their proliferation. Understanding the mechanisms that maintain genomic integrity after hybridization is often confounded by changes in ploidy that occur in hybrid plant species. Homoploid hybrid species, which have no change in chromosome number relative to their parents, offer an opportunity to study the genomic consequences of hybridization in the absence of change in ploidy. Yucca gloriosa (Asparagaceae) is a young homoploid hybrid species, resulting from a cross between Yucca aloifolia and Yucca filamentosa. Previous analyses of ~11 kb of the chloroplast genome and nuclear-encoded microsatellites implicated a single Y. aloifolia genotype as the maternal parent of Y. gloriosa. Using whole genome resequencing, we assembled chloroplast genomes from 41 accessions of all three species to re-assess the hybrid origins of Y. gloriosa. We further used re-sequencing data to annotate transposon abundance in the three species and mRNA-seq to analyze transcription of transposons. The chloroplast phylogeny and haplotype analysis suggest multiple hybridization events contributing to the origin of Y. gloriosa, with both parental species acting as the maternal donor. Transposon abundance at the superfamily level was significantly different between the three species; the hybrid was frequently intermediate to the parental species in TE superfamily abundance or appeared more similar to one or the other parent. In only one case—Copia LTR transposons—did Y. gloriosa have a significantly higher abundance relative to either parent. Expression patterns across the three species showed little increased transcriptional activity of transposons, suggesting that either no transposon release occurred in Y. gloriosa upon hybridization, or that any transposons that were activated via hybridization were rapidly silenced. The identification and quantification of transposon families paired with expression evidence paves the way for additional work seeking to link epigenetics with the important trait variation seen in this homoploid hybrid system.

59 BASIC BIOLOGICAL SCIENCES↗

MaizeMine: A Data Mining Warehouse for the Maize Genetics and Genomics Database

MaizeMine is the data mining resource of the Maize Genetics and Genome Database (MaizeGDB; http://maizemine.maizegdb.org). It enables researchers to create and export customized annotation datasets that can be merged with their own research data for use in downstream analyses. MaizeMine uses the InterMine data warehousing system to integrate genomic sequences and gene annotations from the Zea mays B73 RefGen_v3 and B73 RefGen_v4 genome assemblies, Gene Ontology annotations, single nucleotide polymorphisms, protein annotations, homologs, pathways, and precomputed gene expression levels based on RNA-seq data from the Z. mays B73 Gene Expression Atlas. MaizeMine also provides database cross references between genes of alternative gene sets from Gramene and NCBI RefSeq. MaizeMine includes several search tools, including a keyword search, built-in template queries with intuitive search menus, and a QueryBuilder tool for creating custom queries. The Genomic Regions search tool executes queries based on lists of genome coordinates, and supports both the B73 RefGen_v3 and B73 RefGen_v4 assemblies. The List tool allows you to upload identifiers to create custom lists, perform set operations such as unions and intersections, and execute template queries with lists. When used with gene identifiers, the List tool automatically provides gene set enrichment for Gene Ontology (GO) and pathways, with a choice of statistical parameters and background gene sets. With the ability to save query outputs as lists that can be input to new queries, MaizeMine provides limitless possibilities for data integration and meta-analysis.

59 BASIC BIOLOGICAL SCIENCES↗

De Novo Assembly and Annotation of 11 Diverse Shrub Willow (Salix) Genomes Reveals Novel Gene Organization in Sex-Linked Regions

Poplar and willow species in the Salicaceae are dioecious, yet have been shown to use different sex determination systems located on different chromosomes. Willows in the subgenus Vetrix are interesting for comparative studies of sex determination systems, yet genomic resources for these species are still quite limited. Only a few annotated reference genome assemblies are available, despite many species in use in breeding programs. Here we present de novo assemblies and annotations of 11 shrub willow genomes from six species. Copy number variation of candidate sex determination genes within each genome was characterized and revealed remarkable differences in putative master regulator gene duplication and deletion. We also analyzed copy number and expression of candidate genes involved in floral secondary metabolism, and identified substantial variation across genotypes, which can be used for parental selection in breeding programs. Lastly, we report on a genotype that produces only female descendants and identified gene presence/absence variation in the mitochondrial genome that may be responsible for this unusual inheritance.

59 BASIC BIOLOGICAL SCIENCES↗

The Near-Gapless Penicillium fuscoglaucum Genome Enables the Discovery of Lifestyle Features as an Emerging Post-Harvest Phytopathogen

Penicillium spp. occupy many diverse biological niches that include plant pathogens, opportunistic human pathogens, saprophytes, indoor air contaminants, and those selected specifically for industrial applications to produce secondary metabolites and lifesaving antibiotics. Recent phylogenetic studies have established Penicillium fuscoglaucum as a synonym for Penicillium commune, which is an indoor air contaminant and toxin producer and can infect apple fruit during storage. During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of five scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum and contained a lineage-specific, intact cyclopiazonic acid (CPA) gene cluster. For comparison, a highly virulent postharvest apple pathogen, P. expansum strain TDL 12.1, was included and showed a similar growth pattern in culture to our Pf_T2 isolate but was far more aggressive in apple fruit than P. fuscoglaucum. The genome of Pf_T2 serves as a major improvement over existing resources, has superior annotation, and can inform forthcoming omics-based work and functional genetic studies to probe secondary metabolite production and disparities in aggressiveness during apple fruit decay.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Genomics and Transcriptomics Analyses Reveal Divergent Plant Biomass-Degrading Strategies in Fungi

Plant biomass is one of the most abundant renewable carbon sources, which holds great potential for replacing current fossil-based production of fuels and chemicals. In nature, fungi can efficiently degrade plant polysaccharides by secreting a broad range of carbohydrate-active enzymes (CAZymes), such as cellulases, hemicellulases, and pectinases. Due to the crucial role of plant biomass-degrading (PBD) CAZymes in fungal growth and related biotechnology applications, investigation of their genomic diversity and transcriptional dynamics has attracted increasing attention. In this project, we systematically compared the genome content of PBD CAZymes in six taxonomically distant species, Aspergillus niger, Aspergillus nidulans, Penicillium subrubescens, Trichoderma reesei, Phanerochaete chrysosporium, and Dichomitus squalens, as well as their transcriptome profiles during growth on nine monosaccharides. Considerable genomic variation and remarkable transcriptomic diversity of CAZymes were identified, implying the preferred carbon source of these fungi and their different methods of transcription regulation. In addition, the specific carbon utilization ability inferred from genomics and transcriptomics was compared with fungal growth profiles on corresponding sugars, to improve our understanding of the conversion process. This study enhances our understanding of genomic and transcriptomic diversity of fungal plant polysaccharide-degrading enzymes and provides new insights into designing enzyme mixtures and metabolic engineering of fungi for related industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

Comparison of Auxenochlorella protothecoides and Chlorella spp. Chloroplast Genomes: Evidence for Endosymbiosis and Horizontal Virus-like Gene Transfer

Resequencing of the chloroplast genome (cpDNA) of Auxenochlorella protothecoides UTEX 25 was completed (GenBank Accession no. KC631634.1), revealing a genome size of 84,576 base pairs and 30.8% GC content, consistent with features reported for the previously sequenced A. protothecoides 0710, (GenBank Accession no. KC843975). The A. protothecoides UTEX 25 cpDNA encoded 78 predicted open reading frames, 32 tRNAs, and 4 rRNAs, making it smaller and more compact than the cpDNA genome of C. variabilis (124,579 bp) and C. vulgaris (150,613 bp). By comparison, the compact genome size of A. protothecoides was attributable primarily to a lower intergenic sequence content. The cpDNA coding regions of all known Chlorella species were found to be organized in conserved colinear blocks, with some rearrangements. The Auxenochlorella and Chlorella species genome structure and composition were similar, and of particular interest were genes influencing photosynthetic efficiency, i.e., chlorophyll synthesis and photosystem subunit I and II genes, consistent with other biofuel species of interest. Phylogenetic analysis revealed that Prototheca cutis is the closest known A. protothecoides relative, followed by members of the genus Chlorella. The cpDNA of A. protothecoides encodes 37 genes that are highly homologous to representative cyanobacteria species, including rrn16, rrn23, and psbA, corroborating a well-recognized symbiosis. Several putative coding regions were identified that shared high nucleotide sequence identity with virus-like sequences, suggestive of horizontal gene transfer. Despite these predictions, no corresponding transcripts were obtained by RT-PCR amplification, indicating they are unlikely to be expressed in the extant lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Conserved gene clusters in bacterial genomes provide further support for the primacy of RNA

Five complete bacterial genome sequences have been released to the scientific community. These include four (eu)Bacteria, Haemophilus influenzae, Mycoplasma genitalium, M. pneumoniae, and Synechocystis PCC 6803, as well as one Archaeon, Methanococcus jannaschii. Features of organization shared by these genomes are likely to have arisen very early in the history of the bacteria and thus can be expected to provide further insight into the nature of early ancestors. Results of a genome comparison of these five organisms confirm earlier observations that gene order is remarkably unpreserved. There are, nevertheless, at least 16 clusters of two or more genes whose order remains the same among the four (eu)Bacteria and these are presumed to reflect conserved elements of coordinated gene expression that require gene proximity. Eight of these gene orders are essentially conserved in the Archaea as well. Many of these clusters are known to be regulated by RNA-level mechanisms in Escherichia coli, which supports the earlier suggestion that this type of regulation of gene expression may have arisen very early. We conclude that although the last common ancestor may have had a DNA genome, it likely was preceded by progenotes with an RNA genome.

Non-NASA Center↗

Genome-wide functional screens enable the prediction of high activity CRISPR-Cas9 and -Cas12a guides in Yarrowia lipolytica

Abstract Genome-wide functional genetic screens have been successful in discovering genotype-phenotype relationships and in engineering new phenotypes. While broadly applied in mammalian cell lines and in E. coli , use in non-conventional microorganisms has been limited, in part, due to the inability to accurately design high activity CRISPR guides in such species. Here, we develop an experimental-computational approach to sgRNA design that is specific to an organism of choice, in this case the oleaginous yeast Yarrowia lipolytica . A negative selection screen in the absence of non-homologous end-joining, the dominant DNA repair mechanism, was used to generate single guide RNA (sgRNA) activity profiles for both SpCas9 and LbCas12a. This genome-wide data served as input to a deep learning algorithm, DeepGuide, that is able to accurately predict guide activity. DeepGuide uses unsupervised learning to obtain a compressed representation of the genome, followed by supervised learning to map sgRNA sequence, genomic context, and epigenetic features with guide activity. Experimental validation, both genome-wide and with a subset of selected genes, confirms DeepGuide’s ability to accurately predict high activity sgRNAs. DeepGuide provides an organism specific predictor of CRISPR guide activity that with retraining could be applied to other fungal species, prokaryotes, and other non-conventional organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Major proliferation of transposable elements shaped the genome of the soybean rust pathogen Phakopsora pachyrhizi

With >7000 species the order of rust fungi has a disproportionately large impact on agriculture, horticulture, forestry and foreign ecosystems. The infectious spores are typically dikaryotic, a feature unique to fungi in which two haploid nuclei reside in the same cell. A key example is Phakopsora pachyrhizi, the causal agent of Asian soybean rust disease, one of the world’s most economically damaging agricultural diseases. Despite P. pachyrhizi’s impact, the exceptional size and complexity of its genome prevented generation of an accurate genome assembly. Here, we sequence three independent P. pachyrhizi genomes and uncover a genome up to 1.25 Gb comprising two haplotypes with a transposable element (TE) content of ~93%. We study the incursion and dominant impact of these TEs on the genome and show how they have a key impact on various processes such as host range adaptation, stress responses and genetic plasticity.

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive gene loss in the common bean pan-genome during range expansion and domestication

The common bean ( Phaseolus vulgaris L.) is a crucial legume crop and an ideal evolutionary model to study adaptive diversity in wild and domesticated populations. Here, we present a common bean pan-genome based on five high-quality genomes and whole-genome reads representing 339 genotypes. It reveals ~234 Mb of additional sequences containing 6,905 protein-coding genes missing from the reference, constituting 49% of all presence/absence variants (PAVs). More non-synonymous mutations are found in PAVs than core genes, probably reflecting the lower effective population size of PAVs and fitness advantages due to the purging effect of gene loss. Our results suggest pan-genome shrinkage occurred during wild range expansion. Selection signatures provide evidence that partial or complete gene loss was a key adaptive genetic change in common bean populations with major implications for plant adaptation. The pan-genome is a valuable resource for food legume research and breeding for climate change mitigation and sustainable agriculture.

59 BASIC BIOLOGICAL SCIENCES↗

Seagrass genomes reveal ancient polyploidy and adaptations to the marine environment

Here, we present chromosome-level genome assemblies from representative species of three independently evolved seagrass lineages: Posidonia oceanica, Cymodocea nodosa, Thalassia testudinum and Zostera marina. We also include a draft genome of Potamogeton acutifolius, belonging to a freshwater sister lineage to Zosteraceae. All seagrass species share an ancient whole-genome triplication, while additional whole-genome duplications were uncovered for C. nodosa, Z. marina and P. acutifolius. Comparative analysis of selected gene families suggests that the transition from submerged-freshwater to submerged-marine environments mainly involved fine-tuning of multiple processes (such as osmoregulation, salinity, light capture, carbon acquisition and temperature) that all had to happen in parallel, probably explaining why adaptation to a marine lifestyle has been exceedingly rare. Major gene losses related to stomata, volatiles, defence and lignification are probably a consequence of the return to the sea rather than the cause of it. These new genomes will accelerate functional studies and solutions, as continuing losses of the savannahs of the sea are of major concern in times of climate change and loss of biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗

CheckV assesses the quality and completeness of metagenome-assembled viral genomes

Abstract Millions of new viral sequences have been identified from metagenomes, but the quality and completeness of these sequences vary considerably. Here we present CheckV, an automated pipeline for identifying closed viral genomes, estimating the completeness of genome fragments and removing flanking host regions from integrated proviruses. CheckV estimates completeness by comparing sequences with a large database of complete viral genomes, including 76,262 identified from a systematic search of publicly available metagenomes, metatranscriptomes and metaviromes. After validation on mock datasets and comparison to existing methods, we applied CheckV to large and diverse collections of metagenome-assembled viral sequences, including IMG/VR and the Global Ocean Virome. This revealed 44,652 high-quality viral genomes (that is, >90% complete), although the vast majority of sequences were small fragments, which highlights the challenge of assembling viral genomes from short-read metagenomes. Additionally, we found that removal of host contamination substantially improved the accurate identification of auxiliary metabolic genes and interpretation of viral-encoded functions.

59 BASIC BIOLOGICAL SCIENCES↗

A unified catalog of 204,938 reference genomes from the human gut microbiome

Comprehensive, high-quality reference genomes are required for functional characterization and taxonomic assignment of the human gut microbiota. We present the Unified Human Gastrointestinal Genome (UHGG) collection, comprising 204,938 nonredundant genomes from 4,644 gut prokaryotes. These genomes encode >170 million protein sequences, which we collated in the Unified Human Gastrointestinal Protein (UHGP) catalog. The UHGP more than doubles the number of gut proteins in comparison to those present in the Integrated Gene Catalog. More than 70% of the UHGG species lack cultured representatives, and 40% of the UHGP lack functional annotations. Intraspecies genomic variation analyses revealed a large reservoir of accessory genes and single-nucleotide variants, many of which are specific to individual human populations. The UHGG and UHGP collections will enable studies linking genotypes to phenotypes in the human gut microbiome.

59 BASIC BIOLOGICAL SCIENCES↗

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES↗