Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

KBase Narrative - Genomic and environmental controls on Castellaniella biogeography in an anthropogenically disturbed site

Genome assemblies were imported into KBase using the Batch Import Assembly from Staging Area (v1.0.57) function. All assemblies were annotated using the Annotated Multiple Microbial Assemblies with RASTtk - v1.073 tool. Annotated genomes were grouped into sets using the Add Genomes to GenomeSet - v1.7.6 function. Individual annotated genomes can be found both below and in the Data menu to the left. Taxonomy was assigned using the Classify Microbes with GTDB-Tk-v1.7.0 tool. The results of this analysis are shown below. Analysis of the Castellaniella pangenome was performed using the Compute Pangenome (v0.0.7) tool. Using the same method, we also computed the ORR-specific and non-ORR Castellaniella pangenomes. All pangenome results (including the presence/absence matrix) can be found below.

Szink, Elizabeth↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

Quantitative analysis of bristle number in Drosophila mutants identifies genes involved in neural development

BACKGROUND: The identification of the function of all genes that contribute to specific biological processes and complex traits is one of the major challenges in the postgenomic era. One approach is to employ forward genetic screens in genetically tractable model organisms. In Drosophila melanogaster, P element-mediated insertional mutagenesis is a versatile tool for the dissection of molecular pathways, and there is an ongoing effort to tag every gene with a P element insertion. However, the vast majority of P element insertion lines are viable and fertile as homozygotes and do not exhibit obvious phenotypic defects, perhaps because of the tendency for P elements to insert 5' of transcription units. Quantitative genetic analysis of subtle effects of P element mutations that have been induced in an isogenic background may be a highly efficient method for functional genome annotation. RESULTS: Here, we have tested the efficacy of this strategy by assessing the extent to which screening for quantitative effects of P elements on sensory bristle number can identify genes affecting neural development. We find that such quantitative screens uncover an unusually large number of genes that are known to function in neural development, as well as genes with yet uncharacterized effects on neural development, and novel loci. CONCLUSIONS: Our findings establish the use of quantitative trait analysis for functional genome annotation through forward genetics. Similar analyses of quantitative effects of P element insertions will facilitate our understanding of the genes affecting many other complex traits in Drosophila.

Non-NASA Center↗

The Roseibium album (Labrenzia alba) Genome Possesses Multiple Symbiosis Factors Possibly Underpinning Host-Microbe Relationships in the Marine Benthos

Here, we announce the genomes of eight Roseibium album (synonym Labrenzia alba ) strains that were obtained from the octocoral Eunicella labiata . Genome annotation revealed multiple symbiosis factors common to all genomes, such as eukaryotic-like repeat protein- and multidrug resistance-encoding genes, which likely underpin symbiotic relationships with marine invertebrate hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Knowledge-matching based computational framework for genome-scale metabolic model refinement

Genome-scale metabolic models (GEMs) are mathematically structured knowledge base reconstructed from annotated genome of different organisms. With the advancement of next-generation sequencing technology, many organisms have had their genomes sequenced. However, obtaining a high-quality GEM is highly time-consuming, even with the introduction of several genome-scale reconstruction tools that offer automated draft network generation and gap filling. It has been recognized that the iterative process of manual curation and refinement is the limiting step of GEM development, and how to expedite the GEM refinement is still an open question. As cellular metabolism is a complex system with very high degree of freedom and redundancy, the principles and techniques developed in process systems engineering can be adapted to expedite GEM refinement. In this paper we present a knowledge-matching based computation framework for GEM refinement, and demonstrate the effectiveness of the proposed solution using the refinement of a GEM for Clostridium tyrobutyricum.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

pnnl-predictive-phenomics/csc052cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc040cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc009cyc

Using the genome annotation as input, Pathway-tools generates a database containing all the information that can be inferred from the genome. The Pathway/Genome database (PGDB) can subsequently be curated manually. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii

Tilletia maclaganii is a smut fungal pathogen that causes significant biomass reduction of switchgrass ( Panicum virgatum ) used for animal forage and biofuel production. Here we present the annotated genome of T. maclaganii , strain Tm001-NY21, estimated at 42.79 Mb in size, in 53 assembled contigs and encoding 10,235 predicted genes. This genome will be important for future comparative studies of Ustilaginales across its geographic and host range.

PacBio↗

Assembly, comparative analysis, and utilization of a single haplotype reference genome for soybean

Cultivar Williams 82 has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as Wm82-ISU-01 (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48 387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of Williams 82 reveal large regions of genomic heterogeneity, including regions of differential introgression from the cultivar Kingwa within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (> 20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the ‘Williams’ haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of Williams 82. In addition to the Wm82.a6 assembly, we also assembled the genome of ‘Fiskeby III,’ a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with Fiskeby III revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and Fiskeby III genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.

59 BASIC BIOLOGICAL SCIENCES↗

Near-complete genome sequence of Lipomyces tetrasporous NRRL Y-64009, an oleaginous yeast capable of growing on lignocellulosic hydrolysates

ABSTRACT Lipomyces tetrasporous is an oleaginous yeast that can utilize a variety of plant-based sugars. It accumulates lipids during growth on lignocellulosic biomass hydrolysates. We present the annotated genome sequence of L. tetrasporous NRRL Y-64009 to aid in its development as a platform organism for producing lipids and lipid-based bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗

Dynamic genome evolution in a model fern

The large size and complexity of most fern genomes have hampered efforts to elucidate fundamental aspects of fern biology and land plant evolution through genome-enabled research. Here we present a chromosomal genome assembly and associated methylome, transcriptome and metabolome analyses for the model fern species Ceratopteris richardii. The assembly reveals a history of remarkably dynamic genome evolution including rapid changes in genome content and structure following the most recent whole-genome duplication approximately 60 million years ago. These changes include massive gene loss, rampant tandem duplications and multiple horizontal gene transfers from bacteria, contributing to the diversification of defence-related gene families. The insertion of transposable elements into introns has led to the large size of the Ceratopteris genome and to exceptionally long genes relative to other plants. Gene family analyses indicate that genes directing seed development were co-opted from those controlling the development of fern sporangia, providing insights into seed plant evolution. Our findings and annotated genome assembly extend the utility of Ceratopteris as a model for investigating and teaching plant biology.

59 BASIC BIOLOGICAL SCIENCES↗

Glacier ice archives nearly 15,000-year-old microbes and phages

Background Glacier ice archives information, including microbiology, that helps reveal paleoclimate histories and predict future climate change. Though glacier-ice microbes are studied using culture or amplicon approaches, more challenging metagenomic approaches, which provide access to functional, genome-resolved information and viruses, are under-utilized, partly due to low biomass and potential contamination. Results We expand existing clean sampling procedures using controlled artificial ice-core experiments and adapted previously established low-biomass metagenomic approaches to study glacier-ice viruses. Controlled sampling experiments drastically reduced mock contaminants including bacteria, viruses, and free DNA to background levels. Amplicon sequencing from eight depths of two Tibetan Plateau ice cores revealed common glacier-ice lineages including Janthinobacterium, Polaromonas, Herminiimonas, Flavobacterium, Sphingomonas, and Methylobacterium as the dominant genera, while microbial communities were significantly different between two ice cores, associating with different climate conditions during deposition. Separately, ~355- and ~14,400-year-old ice were subject to viral enrichment and low-input quantitative sequencing, yielding genomic sequences for 33 vOTUs. These were virtually all unique to this study, representing 28 novel genera and not a single species shared with 225 environmentally diverse viromes. Further, 42.4% of the vOTUs were identifiable temperate, which is significantly higher than that in gut, soil, and marine viromes, and indicates that temperate phages are possibly favored in glacier-ice environments before being frozen. In silico host predictions linked 18 vOTUs to co-occurring abundant bacteria (Methylobacterium, Sphingomonas, and Janthinobacterium), indicating that these phages infected ice-abundant bacterial groups before being archived. Functional genome annotation revealed four virus-encoded auxiliary metabolic genes, particularly two motility genes suggest viruses potentially facilitate nutrient acquisition for their hosts. Finally, given their possible importance to methane cycling in ice, we focused on Methylobacterium viruses by contextualizing our ice-observed viruses against 123 viromes and prophages extracted from 131 Methylobacterium genomes, revealing that the archived viruses might originate from soil or plants. Conclusions Together, these efforts further microbial and viral sampling procedures for glacier ice and provide a first window into viral communities and functions in ancient glacier environments. Such methods and datasets can potentially enable researchers to contextualize new discoveries and begin to incorporate glacier-ice microbes and their viruses relative to past and present climate change in geographically diverse regions globally.

59 BASIC BIOLOGICAL SCIENCES↗

Large-scale genomic analyses with machine learning uncover predictive patterns associated with fungal phytopathogenic lifestyles and traits

Abstract Invasive plant pathogenic fungi have a global impact, with devastating economic and environmental effects on crops and forests. Biosurveillance, a critical component of threat mitigation, requires risk prediction based on fungal lifestyles and traits. Recent studies have revealed distinct genomic patterns associated with specific groups of plant pathogenic fungi. We sought to establish whether these phytopathogenic genomic patterns hold across diverse taxonomic and ecological groups from the Ascomycota and Basidiomycota, and furthermore, if those patterns can be used in a predictive capacity for biosurveillance. Using a supervised machine learning approach that integrates phylogenetic and genomic data, we analyzed 387 fungal genomes to test a proof-of-concept for the use of genomic signatures in predicting fungal phytopathogenic lifestyles and traits during biosurveillance activities. Our machine learning feature sets were derived from genome annotation data of carbohydrate-active enzymes (CAZymes), peptidases, secondary metabolite clusters (SMCs), transporters, and transcription factors. We found that machine learning could successfully predict fungal lifestyles and traits across taxonomic groups, with the best predictive performance coming from feature sets comprising CAZyme, peptidase, and SMC data. While phylogeny was an important component in most predictions, the inclusion of genomic data improved prediction performance for every lifestyle and trait tested. Plant pathogenicity was one of the best-predicted traits, showing the promise of predictive genomics for biosurveillance applications. Furthermore, our machine learning approach revealed expansions in the number of genes from specific CAZyme and peptidase families in the genomes of plant pathogens compared to non-phytopathogenic genomes (saprotrophs, endo- and ectomycorrhizal fungi). Such genomic feature profiles give insight into the evolution of fungal phytopathogenicity and could be useful to predict the risks of unknown fungi in future biosurveillance activities.

59 BASIC BIOLOGICAL SCIENCES↗

Genome, transcriptome and secretome analyses of the antagonistic, yeast-like fungus Aureobasidium pullulans to identify potential biocontrol genes

Aureobasidium pullulans is an extremotolerant, cosmopolitan yeast-like fungus that successfully colonises vastly different ecological niches. The species is widely used in biotechnology and successfully applied as a commercial biocontrol agent against postharvest diseases and fireblight. However, the exact mechanisms that are responsible for its antagonistic activity against diverse plant pathogens are not known at the molecular level. Thus, it is difficult to optimise and improve the biocontrol applications of this species. As a foundation for elucidating biocontrol mechanisms, we have de novo assembled a high-quality reference genome of a strongly antagonistic A. pullulans strain, performed dual RNA-seq experiments, and analysed proteins secreted during the interaction with the plant pathogen Fusarium oxysporum. Based on the genome annotation, potential biocontrol genes were predicted to encode secreted hydrolases or to be part of secondary metabolite clusters (e.g., NRPS-like, NRPS, T1PKS, terpene, and β-lactone clusters). Transcriptome and secretome analyses defined a subset of 79 A. pullulans genes (among the 10,925 annotated genes) that were transcriptionally upregulated or exclusively detected at the protein level during the competition with F. oxysporum. These potential biocontrol genes comprised predicted secreted hydrolases such as glycosylases, esterases, and proteases, as well as genes encoding enzymes, which are predicted to be involved in the synthesis of secondary metabolites. This study highlights the value of a sequential approach starting with genome mining and consecutive transcriptome and secretome analyses in order to identify a limited number of potential target genes for detailed, functional analyses.

transcriptome↗