Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

ContScout: sensitive detection and removal of contamination from annotated genomes

Contamination of genomes is an increasingly recognized problem affecting several downstream applications, from comparative evolutionary genomics to metagenomics. Here we introduce ContScout, a precise tool for eliminating foreign sequences from annotated genomes. It achieves high specificity and sensitivity on synthetic benchmark data even when the contaminant is a closely related species, outperforms competing tools, and can distinguish horizontal gene transfer from contamination. A screen of 844 eukaryotic genomes for contamination identified bacteria as the most common source, followed by fungi and plants. Furthermore, we show that contaminants in ancestral genome reconstructions lead to erroneous early origins of genes and inflate gene loss rates, leading to a false notion of complex ancestral genomes. Taken together, we offer here a tool for sensitive removal of foreign proteins, identify and remove contaminants from diverse eukaryotic genomes and evaluate their impact on phylogenomic analyses.

59 BASIC BIOLOGICAL SCIENCES↗

High-quality genome of the basidiomycete yeast Dioszegia hungarica PDD-24b-2 isolated from cloud water

The genome of the basidiomycete yeast Dioszegia hungarica strain PDD-24b-2 isolated from cloud water at the summit of puy de $D\hat{o}me$ (France) was sequenced using a hybrid PacBio and Illumina sequencing strategy. The obtained assembled genome of 20.98 Mb and a GC content of 57% is structured in 16 large-scale contigs ranging from 90 kb to 5.56Mb, and another 27.2 kb contig representing the complete circular mitochondrial genome. In total, 8,234 proteins were predicted from the genome sequence. The mitochondrial genome shows 16.2% cgu codon usage for arginine but has no canonical cognate tRNA to translate this codon. Detected transposable element (TE)-related sequences account for about 0.63% of the assembled genome. A dataset of 2,068 hand-picked public environmental metagenomes, representing over 20 Tbp of raw reads, was probed for D. hungarica related ITS sequences, and revealed worldwide distribution of this species, particularly in aerial habitats. Growth experiments suggested a psychrophilic phenotype and the ability to disperse by producing ballistospores. The high-quality assembled genome obtained for this D. hungarica strain will help investigate the behavior and ecological functions of this species in the environment.

59 BASIC BIOLOGICAL SCIENCES↗

Insights into the Ecological Diversification of the Hymenochaetales based on Comparative Genomics and Phylogenomics With an Emphasis on Coltricia

To elucidate the genomic traits of ecological diversification in the Hymenochaetales, we sequenced 15 new genomes, with attention to ectomycorrhizal (EcM) Coltricia species. Together with published data, 32 genomes, including 31 Hymenochaetales and one outgroup, were comparatively analyzed in total. Compared with those of parasitic and saprophytic members, EcM species have significantly reduced number of plant cell wall degrading enzyme genes, and expanded transposable elements, genome sizes, small secreted proteins, and secreted proteases. EcM species still retain some of secreted carbohydrate-active enzymes (CAZymes) and have lost the key secreted CAZymes to degrade lignin and cellulose, while possess a strong capacity to degrade a microbial cell wall containing chitin and peptidoglycan. There were no significant differences in secreted CAZymes between fungi growing on gymnosperms and angiosperms, suggesting that the secreted CAZymes in the Hymenochaetales evolved before differentiation of host trees into gymnosperms and angiosperms. Nevertheless, parasitic and saprophytic species of the Hymenochaetales are very similar in many genome features, which reflect their close phylogenetic relationships both being white rot fungi. Phylogenomic and molecular clock analyses showed that the EcM genus Coltricia formed a clade located at the base of the Hymenochaetaceae and divergence time later than saprophytic species. And Coltricia remains one to two genes of AA2 family. These indicate that the ancestors of Coltricia appear to have originated from saprophytic ancestor with the ability to cause a white rot. This study provides new genomic data for EcM species and insights into the ecological diversification within the Hymenochaetales based on comparative genomics and phylogenomics analyses.

59 BASIC BIOLOGICAL SCIENCES↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Comparative genomic and transcriptomic analyses of trans-kingdom pathogen Fusarium solani species complex reveal degrees of compartmentalization

Background: The Fusarium solani species complex (FSSC) comprises fungal pathogens responsible for mortality in a diverse range of animals and plants, but their genome diversity and transcriptome responses in animal pathogenicity remain to be elucidated. We sequenced, assembled and annotated six chromosome-level FSSC clade 3 genomes of aquatic animal and plant host origins. We established a pathosystem and investigated the expression data of F. falciforme and F. keratoplasticum in Chinese softshell turtle (Pelodiscus sinensis) host. Results: Comparative analyses between the FSSC genomes revealed a spectrum of conservation patterns in chromosomes categorised into three compartments: core, fast-core (FC), and lineage-specific (LS). LS chromosomes contribute to variations in genomes size, with up to 42.2% of variations between F. vanettenii strains. Each chromosome compartment varied in structural architectures, with FC and LS chromosomes contain higher proportions of repetitive elements with genes enriched in functions related to pathogenicity and niche expansion. We identified differences in both selection in the coding sequences and DNA methylation levels between genome features and chromosome compartments which suggest a multi-speed evolution that can be traced back to the last common ancestor of Fusarium. We further demonstrated that F. falciforme and F. keratoplasticum are opportunistic pathogens by inoculating P. sinensis eggs and identified differentially expressed genes also associated with plant pathogenicity. These included the most upregulated genes encoding the CFEM (Common in Fungal Extracellular Membrane) domain. Conclusions: The high-quality genome assemblies provided new insights into the evolution of FSSC chromosomes, which also serve as a resource for studies of fungal genome evolution and pathogenesis. This study also establishes an animal model for fungal pathogens of trans-kingdom hosts.

59 BASIC BIOLOGICAL SCIENCES↗

The Tetracentron genome provides insight into the early evolution of eudicots and the formation of vessel elements

Tetracentron sinense is an endemic and endangered deciduous tree. It belongs to the Trochodendrales, one of four early diverging lineages of eudicots known for having vesselless secondary wood. Sequencing and resequencing of the T. sinense genome will help us understand eudicot evolution, the genetic basis of tracheary element development, and the genetic diversity of this relict species. Here, we report a chromosome-scale assembly of the T. sinense genome. We assemble the 1.07 Gb genome sequence into 24 chromosomes and annotate 32,690 protein-coding genes. Phylogenomic analyses verify that the Trochodendrales and core eudicots are sister lineages and showed that two whole-genome duplications occurred in the Trochodendrales approximately 82 and 59 million years ago. Synteny analyses suggest that the γ event, resulting in paleohexaploidy, may have only happened in core eudicots. Interestingly, we find that vessel elements are present in T. sinense, which has two orthologs of AtVND7, the master regulator of vessel formation. T. sinense also has several key genes regulated by or regulating TsVND7.2 and their regulatory relationship resembles that in Arabidopsis thaliana. Resequencing and population genomics reveals high levels of genetic diversity of T. sinense and identifies four refugia in China. The T. sinense genome provides a unique reference for inferring the early evolution of eudicots and the mechanisms underlying vessel element formation. Population genomics analysis of T. sinense reveals its genetic diversity and geographic structure with implications for conservation.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of twelve genomes of the bacterium Kerstersia gyiorum from brown-throated sloths ( Bradypus variegatus ), the first from a non-human host

Kerstersia gyiorum is a Gram-negative bacterium found in various animals, including humans, where it has been associated with various infections. Knowledge of the basic biology of K. gyiorum is essential to understand the evolutionary strategies of niche adaptation and how this organism contributes to infectious diseases; however, genomic data about K. gyiorum is very limited, especially from non-human hosts. In this work, we sequenced 12 K. gyiorum genomes isolated from healthy free-living brown-throated sloths (Bradypus variegatus) in the Parque Estadual das Fontes do Ipiranga (São Paulo, Brazil), and compared them with genomes from isolates of human origin, in order to gain insights into genomic diversity, phylogeny, and host specialization of this species. Phylogenetic analysis revealed that these K. gyiorum strains are structured according to host. Despite the fact that sloth isolates were sampled from a single geographic location, the intra-sloth K. gyiorum diversity was divided into three clusters, with differences of more than 1,000 single nucleotide polymorphisms between them, suggesting the circulation of various K. gyiorum lineages in sloths. Genes involved in mobilome and defense mechanisms against mobile genetic elements were the main source of gene content variation between isolates from different hosts. Sloth-specific K. gyiorum genome features include an IncN2 plasmid, a phage sequence, and a CRISPR-Cas system. The broad diversity of defense elements in K. gyiorum (14 systems) may prevent further mobile element flow and explain the low amount of mobile genetic elements in K. gyiorum genomes. Gene content variation may be important for the adaptation of K. gyiorum to different host niches. This study furthers our understanding of diversity, host adaptation, and evolution of K. gyiorum, by presenting and analyzing the first genomes of non-human isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Gene and genome duplications have contrasting impacts on biosynthetic and flower developmental pathways in California poppy

Benzylisoquinoline alkaloids (BIAs) represent a vast group of specialized plant metabolites with diverse pharmaceutical applications, synthesized by a variety of gene families. Among the multiple plant lineages that produce BIAs, the most notable is the poppy family (Papaveraceae), with California poppy (Eschscholzia californica) emerging as a model organism. Here, we report a haplotype-resolved genome assembly, in combination with a high-density expression atlas, for California poppy. Genome analyses reveal recent diversification of BIA biosynthesis genes in poppy through localized duplications. Furthermore, we demonstrate that the degree of phylogenetic relatedness among paralogs within BIA biosynthesis-associated gene families correlates with similarities in gene expression. In contrast, gene families involved in carotenoid biosynthesis, which contributes to the intense orange petal pigmentation, are not phylogenetically clustered, and floral developmental regulators exhibit a high degree of retention of gene duplicates associated with ancient polyploidy events. These findings illustrate alternative roles for gene and genome duplications as drivers of trait evolution. Given the position of California poppy in the angiosperm phylogeny, the high-quality genomic resources generated for this work constitute a valuable resource for comparative genomic and transcriptomic analyses for poppies and flowering plants more generally.

Rössner, Le-Han [Justus-Liebig University, Giessen↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata

Viruses are widely recognized as critical members of all microbiomes. Metagenomics enables large-scale exploration of the global virosphere, progressively revealing the extensive genomic diversity of viruses on Earth and highlighting the myriad of ways by which viruses impact biological processes. IMG/VR provides access to the largest collection of viral sequences obtained from (meta)genomes, along with functional annotation and rich metadata. A web interface enables users to efficiently browse and search viruses based on genome features and/or sequence similarity. Here, for this work, we present the fourth version of IMG/VR, composed of >15 million virus genomes and genome fragments, a ≈6-fold increase in size compared to the previous version. These clustered into 8.7 million viral operational taxonomic units, including 231 408 with at least one high-quality representative. Viral sequences in IMG/VR are now systematically identified from genomes, metagenomes, and metatranscriptomes using a new detection approach (geNomad), and IMG standard annotation are complemented with genome quality estimation using CheckV, taxonomic classification reflecting the latest taxonomic standards, and microbial host taxonomy prediction. IMG/VR v4 is available at https://img.jgi.doe.gov/vr, and the underlying data are available to download at https://genome.jgi.doe.gov/portal/IMG_VR.

59 BASIC BIOLOGICAL SCIENCES↗

VIBES: a workflow for annotating and visualizing viral sequences integrated into bacterial genomes

Abstract Bacteriophages are viruses that infect bacteria. Many bacteriophages integrate their genomes into the bacterial chromosome and become prophages. Prophages may substantially burden or benefit host bacteria fitness, acting in some cases as parasites and in others as mutualists. Some prophages have been demonstrated to increase host virulence. The increasing ease of bacterial genome sequencing provides an opportunity to deeply explore prophage prevalence and insertion sites. Here we present VIBES (Viral Integrations in Bacterial genomES), a workflow intended to automate prophage annotation in complete bacterial genome sequences. VIBES provides additional context to prophage annotations by annotating bacterial genes and viral proteins in user-provided bacterial and viral genomes. The VIBES pipeline is implemented as a Nextflow-driven workflow, providing a simple, unified interface for execution on local, cluster and cloud computing environments. For each step of the pipeline, a container including all necessary software dependencies is provided. VIBES produces results in simple tab-separated format and generates intuitive and interactive visualizations for data exploration. Despite VIBES’s primary emphasis on prophage annotation, its generic alignment-based design allows it to be deployed as a general-purpose sequence similarity search manager. We demonstrate the utility of the VIBES prophage annotation workflow by searching for 178 Pf phage genomes across 1072 Pseudomonas spp. genomes.

59 BASIC BIOLOGICAL SCIENCES↗

The final piece of the Triangle of U: Evolution of the tetraploid Brassica carinata genome

Abstract Ethiopian mustard (Brassica carinata) is an ancient crop with remarkable stress resilience and a desirable seed fatty acid profile for biofuel uses. Brassica carinata is one of six Brassica species that share three major genomes from three diploid species (AA, BB, and CC) that spontaneously hybridized in a pairwise manner to form three allotetraploid species (AABB, AACC, and BBCC). Of the genomes of these species, that of B. carinata is the least understood. Here, we report a chromosome scale 1.31-Gbp genome assembly with 156.9-fold sequencing coverage for B. carinata, completing the reference genomes comprising the classic Triangle of U, a classical theory of the evolutionary relationships among these six species. Our assembly provides insights into the hybridization event that led to the current B. carinata genome and the genomic features that gave rise to the superior agronomic traits of B. carinata. Notably, we identified an expansion of transcription factor networks and agronomically important gene families. Completion of the Triangle of U comparative genomics platform has allowed us to examine the dynamics of polyploid evolution and the role of subgenome dominance in the domestication and continuing agronomic improvement of B. carinata and other Brassica species.

Biochemistry & Molecular Biology↗

Genome organization and botanical diversity

Abstract The rich diversity of angiosperms, both the planet's dominant flora and the cornerstone of agriculture, is integrally intertwined with a distinctive evolutionary history. Here, we explore the interplay between angiosperm genome organization and botanical diversity, empowered by genomic approaches ranging from genetic linkage mapping to analysis of gene regulation. Commonality in the genetic hardware of plants has enabled robust comparative genomics that has provided a broad picture of angiosperm evolution and implicated both general processes and specific elements in contributing to botanical diversity. We argue that the hardware of plant genomes—both in content and in dynamics—has been shaped by selection for rather substantial differences in gene regulation between plants and animals such as maize and human, organisms of comparable genome size and gene number. Their distinctive genome content and dynamics may reflect in part the indeterminate development of plants that puts strikingly different demands on gene regulation than in animals. Repeated polyploidization of plant genomes and multiplication of individual genes together with extensive rearrangement and differential retention provide rich raw material for selection of morphological and/or physiological variations conferring fitness in specific niches, whether natural or artificial. These findings exemplify the burgeoning information available to employ in increasing knowledge of plant biology and in modifying selected plants to better meet human needs.

Biochemistry & Molecular Biology↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary transition to the ectomycorrhizal habit in the genomes of a hyperdiverse lineage of mushroom-forming fungi

The ectomycorrhizal (ECM) symbiosis has independently evolved from diverse types of saprotrophic ancestors. In this study, we seek to identify genomic signatures of the transition to the ECM habit within the hyperdiverse Russulaceae. We present comparative analyses of the genomic architecture and the total and secreted gene repertoires of 18 species across the order Russulales, of which 13 are newly sequenced, including a representative of a saprotrophic member of Russulaceae, Gloeopeniophorella convolvens. The genomes of ECM Russulaceae are characterized by a loss of genes for plant cell wall-degrading enzymes (PCWDEs), an expansion of genome size through increased transposable element (TE) content, a reduction in secondary metabolism clusters, and an association of small secreted proteins (SSPs) with TE ‘nests’, or dense aggregations of TEs. Here, some PCWDEs have been retained or even expanded, mostly in a species-specific manner. The genome of G. convolvens possesses some characteristics of ECM genomes (e.g. loss of some PCWDEs, TE expansion, reduction in secondary metabolism clusters). Functional specialization in ECM decomposition may drive diversification. Accelerated gene evolution predates the evolution of the ECM habit, indicating that changes in genome architecture and gene content may be necessary to prime the evolutionary switch.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗