Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

RNAseq-based transcriptome assembly of Clostridium acetobutylicum for functional genome annotation and discovery

Accurate genome annotations are essential in modern biology and biotechnology, yet they are still largely based on genome sequencing and comparative analyses. We show that the Clostridium acetobutylicum genome annotation can be markedly improved by integrating bioinformatic predictions with RNA sequencing (RNAseq) data. Samples were acquired under butanol, butyrate, and unstressed treatments across various growth conditions. Analysis of an initial assembly revealed errors due to background signals and limitations of assembly algorithms. Hurdles for RNAseq transcriptome mapping include optimizing library complexity and sequencing depth, yet most studies report low sequencing depth and ignore the effect of ribosomal RNA abundance. An integrative analysis was developed to combine motif predictions, single-nucleotide resolution sequencing depth, and library complexity to resolve difficulties in assembly curation. This minimized false positive error and determined gene boundaries, in some cases, to the exact base-pair of prior studies. This will be the first strand-specific transcriptome assembly in a Clostridium organism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Four chromosome scale genomes and a pan-genome annotation to accelerate pecan tree breeding

Genome-enabled biotechnologies have the potential to accelerate breeding efforts in long-lived perennial crop species. Despite the transformative potential of molecular tools in pecan and other outcrossing tree species, highly heterozygous genomes, significant presence–absence gene content variation, and histories of interspecific hybridization have constrained breeding efforts. To overcome these challenges, here, we present diploid genome assemblies and annotations of four outbred pecan genotypes, including a PacBio HiFi chromosome-scale assembly of both haplotypes of the ‘Pawnee’ cultivar. Comparative analysis and pan-genome integration reveal substantial and likely adaptive interspecific genomic introgressions, including an over-retained haplotype introgressed from bitternut hickory into pecan breeding pedigrees. Further, by leveraging our pan-genome presence–absence and functional annotation database among genomes and within the two outbred haplotypes of the ‘Lakota’ genome, we identify candidate genes for pest and pathogen resistance. Combined, these analyses and resources highlight significant progress towards functional and quantitative genomics in highly diverse and outbred crops.

54 ENVIRONMENTAL SCIENCES↗

Natural variation and improved genome annotation of the emerging biofuel crop field pennycress ( Thlaspi arvense )

The Brassicaceae family comprises more than 3,700 species with a diversity of phenotypic characteristics, including seed oil content and composition. Recently, the global interest in Thlaspi arvense L. (pennycress) has grown as the seed oil composition makes it a suitable source for biodiesel and aviation fuel production. However, many wild traits of this species need to be domesticated to make pennycress ideal for cultivation. Molecular breeding and engineering efforts require the availability of an accurate genome sequence of the species. Here, we describe pennycress genome annotation improvements, using a combination of long- and short-read transcriptome data obtained from RNA derived from embryos of 22 accessions, in addition to public genome and gene expression information. Our analysis identified 27,213 protein-coding genes, as well as on average 6,188 biallelic SNPs. In addition, we used the identified SNPs to evaluate the population structure of our accessions. The data from this analysis support that the accession Ames 32872, originally from Armenia, is highly divergent from the other accessions, while the accessions originating from Canada and the United States cluster together. When we evaluated the likely signatures of natural selection from alternative SNPs, we found 7 candidate genes under likely recent positive selection. These genes are enriched with functions related to amino acid metabolism and lipid biosynthesis and highlight possible future targets for crop improvement efforts in pennycress.

59 BASIC BIOLOGICAL SCIENCES↗

Annotated Genome Sequence of the High-Biomass-Producing Yellow-Green Alga Tribonema minus

Here, we report the annotated genome sequence for a heterokont alga from the class Xanthophyceae. This high-biomass-producing strain, Tribonema minus UTEX B 3156, was isolated from a wastewater treatment plant in California. It is stable in outdoor raceway ponds and is a promising industrial feedstock for biofuels and bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

ContScout: sensitive detection and removal of contamination from annotated genomes

Contamination of genomes is an increasingly recognized problem affecting several downstream applications, from comparative evolutionary genomics to metagenomics. Here we introduce ContScout, a precise tool for eliminating foreign sequences from annotated genomes. It achieves high specificity and sensitivity on synthetic benchmark data even when the contaminant is a closely related species, outperforms competing tools, and can distinguish horizontal gene transfer from contamination. A screen of 844 eukaryotic genomes for contamination identified bacteria as the most common source, followed by fungi and plants. Furthermore, we show that contaminants in ancestral genome reconstructions lead to erroneous early origins of genes and inflate gene loss rates, leading to a false notion of complex ancestral genomes. Taken together, we offer here a tool for sensitive removal of foreign proteins, identify and remove contaminants from diverse eukaryotic genomes and evaluate their impact on phylogenomic analyses.

59 BASIC BIOLOGICAL SCIENCES↗

Galba: genome annotation with miniprot and AUGUSTUS

The Earth Biogenome Project has rapidly increased the number of available eukaryotic genomes, but most released genomes continue to lack annotation of protein-coding genes. In addition, no transcriptome data is available for some genomes. Various gene annotation tools have been developed but each has its limitations. Here, we introduce GALBA, a fully automated pipeline that utilizes miniprot, a rapid protein-to-genome aligner, in combination with AUGUSTUS to predict genes with high accuracy. Accuracy results indicate that GALBA is particularly strong in the annotation of large vertebrate genomes. We also present use cases in insects, vertebrates, and a land plant. GALBA is fully open source and available as a docker image for easy execution with Singularity in high-performance computing environments. Our pipeline addresses the critical need for accurate gene annotation in newly sequenced genomes, and we believe that GALBA will greatly facilitate genome annotation for diverse organisms.

59 BASIC BIOLOGICAL SCIENCES↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome assembled and annotated genome sequence of Aspergillus flavus NRRL 3357

Abstract Aspergillus flavus is an opportunistic pathogen of crops, including peanuts and maize, and is the second leading cause of aspergillosis in immunocompromised patients. A. flavus is also a major producer of the mycotoxin, aflatoxin, a potent carcinogen, which results in significant crop losses annually. The A. flavus isolate NRRL 3357 was originally isolated from peanut and has been used as a model organism for understanding the regulation and production of secondary metabolites, such as aflatoxin. A draft genome of NRRL 3357 was previously constructed, enabling the development of molecular tools and for understanding population biology of this particular species. Here, we describe an updated, near complete, telomere-to-telomere assembly and re-annotation of the eight chromosomes of A. flavus NRRL 3357 genome, accomplished via long-read PacBio and Oxford Nanopore technologies combined with Illumina short-read sequencing. A total of 13,715 protein-coding genes were predicted. Using RNA-seq data, a significant improvement was achieved in predicted 5’ and 3’ untranslated regions, which were incorporated into the new gene models.

59 BASIC BIOLOGICAL SCIENCES↗

Annotated genome sequence of a fast-growing diploid clone of red alder ( Alnus rubra Bong.)

Abstract Red alder (Alnus rubra Bong.) is an ecologically significant and important fast-growing commercial tree species native to western coastal and riparian regions of North America, having highly desirable wood, pigment, and medicinal properties. We have sequenced the genome of a rapidly growing clone. The assembly is nearly complete, containing the full complement of expected genes. This supports our objectives of identifying and studying genes and pathways involved in nitrogen-fixing symbiosis and those related to secondary metabolites that underlie red alder's many interesting defense, pigmentation, and wood quality traits. We established that this clone is most likely diploid and identified a set of SNPs that will have utility in future breeding and selection endeavors, as well as in ongoing population studies. We have added a well-characterized genome to others from the order Fagales. In particular, it improves significantly upon the only other published alder genome sequence, that of Alnus glutinosa. Our work initiated a detailed comparative analysis of members of the order Fagales and established some similarities with previous reports in this clade, suggesting a biased retention of certain gene functions in the vestiges of an ancient genome duplication when compared with more recent tandem duplications.

59 BASIC BIOLOGICAL SCIENCES↗

kb_DRAM: annotation and metabolic profiling of genomes with DRAM in KBase

Microbial genome annotation is the process of identifying structural and functional elements in DNA sequences and subsequently attaching biological information to those elements. DRAM is a tool developed to annotate bacterial, archaeal, and viral genomes derived from pure cultures or metagenomes. DRAM goes beyond traditional annotation tools by distilling multiple gene annotations to genome level summaries of functional potential. Despite these benefits, a downside of DRAM is the requirement of large computational resources, which limits its accessibility. Further, it did not integrate with downstream metabolic modeling tools that require genome annotation. To alleviate these constraints, DRAM and the viral counterpart, DRAM-v, are now available and integrated with the freely accessible KBase cyberinfrastructure. With kb_DRAM users can generate DRAM annotations and functional summaries from microbial or viral genomes in a point-and-click interface, as well as generate genome-scale metabolic models from DRAM annotations.

59 BASIC BIOLOGICAL SCIENCES↗

Status of genome function annotation in model organisms and crops

Abstract Since the entry into genome‐enabled biology several decades ago, much progress has been made in determining, describing, and disseminating the functions of genes and their products. Yet, this information is still difficult to access for many scientists and for most genomes. To provide easy access and a graphical summary of the status of genome function annotation for model organisms and bioenergy and food crop species, we created a web application ( https://genomeannotation.rheelab.org ) to visualize, search, and download genome annotation data for 28 species. The summary graphics and data tables will be updated semi‐annually, and snapshots will be archived to provide a historical record of the progress of genome function annotation efforts. Clear and simple visualization of up‐to‐date genome function annotation status, including the extent of what is unknown, will help address the grand challenge of elucidating the functions of all genes in organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-genome Phage Annotation Toolkit and Evaluator

Summary: To address the need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of the multiPhATE code, multiPhATE2 includes comparative genomics codes for gene matching among sets of input bacteriophage genomes, and scales well to large input data sets due to incorporation of multiprocessing in the functional annotation and comparative genomics subsystems. Furthermore, additional search algorithms and databases have been added to the functional annotation subsystem. MultiPhATE2 was implemented in Python 3.7, and runs as a command-line code under Linux or MAC-OS.

Kimbrel, JeffreyA.↗

A scaffolded and annotated reference genome of giant kelp (Macrocystis pyrifera)

Abstract Macrocystis pyrifera (giant kelp), is a brown macroalga of great ecological importance as a primary producer and structure-forming foundational species that provides habitat for hundreds of species. It has many commercial uses (e.g. source of alginate, fertilizer, cosmetics, feedstock). One of the limitations to exploiting giant kelp’s economic potential and assisting in giant kelp conservation efforts is a lack of genomic tools like a high quality, contiguous reference genome with accurate gene annotations. Reference genomes attempt to capture the complete genomic sequence of an individual or species, and importantly provide a universal structure for comparison across a multitude of genetic experiments, both within and between species. We assembled the giant kelp genome of a haploid female gametophyte de novo using PacBio reads, then ordered contigs into chromosome level scaffolds using Hi-C. We found the giant kelp genome to be 537 MB, with a total of 35 scaffolds and 188 contigs. The assembly N50 is 13,669,674 with GC content of 50.37%. We assessed the genome completeness using BUSCO, and found giant kelp contained 94% of the BUSCO genes from the stramenopile clade. Annotation of the giant kelp genome revealed 25,919 genes. Additionally, we present genetic variation data based on 48 diploid giant kelp sporophytes from three different Southern California populations that confirms the population structure found in other studies of these populations. This work resulted in a high-quality giant kelp genome that greatly increases the genetic knowledge of this ecologically and economically vital species.

60 APPLIED LIFE SCIENCES↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Accuracy-Based Annotation Quality Score (ABAQS) v1.0

Assessing genome annotation quality is crucial for downstream analyses, but current methods are inadequate for eukaryotes. We present Accuracy-Based Annotation Quality Score (ABAQS), a novel, minimal-data-driven method that comprehensively assesses annotation quality. ABAQS evaluates multiple factors, including genome completeness, gene model validity, and protein profile accuracy, outperforming other metrics like BUSCO and PSAURON. We applied ABAQS to over 2500 eukaryotic genomes and showed its robustness and effectiveness in evaluating genome annotation quality, making it a valuable tool for researchers working with genomic data. ABAQS reveals significant variation in annotation quality and highlights the importance of filtering in improving annotation quality and accuracy.

Haridas, Sajeet [Lawrence Berkeley National Labora↗