Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Co‑cultivation of the anaerobic fungus Caecomyces churrovis with Methanobacterium bryantii enhances transcription of carbohydrate binding modules, dockerins, and pyruvate formate lyases on specific substrates

Abstract Anaerobic fungi and methanogenic archaea are two classes of microorganisms found in the rumen microbiome that metabolically interact during lignocellulose breakdown. Here, stable synthetic co-cultures of the anaerobic fungus Caecomyces churrovis and the methanogen Methanobacterium bryantii (not native to the rumen) were formed, demonstrating that microbes from different environments can be paired based on metabolic ties. Transcriptional and metabolic changes induced by methanogen co-culture were evaluated in C. churrovis across a variety of substrates to identify mechanisms that impact biomass breakdown and sugar uptake. A high-quality genome of C. churrovis was obtained and annotated, which is the first sequenced genome of a non-rhizoid-forming anaerobic fungus. C. churrovis possess an abundance of CAZymes and carbohydrate binding modules and, in agreement with previous studies of early-diverging fungal lineages, N6-methyldeoxyadenine (6mA) was associated with transcriptionally active genes. Co-culture with the methanogen increased overall transcription of CAZymes, carbohydrate binding modules, and dockerin domains in co-cultures grown on both lignocellulose and cellulose and caused upregulation of genes coding associated enzymatic machinery including carbohydrate binding modules in family 18 and dockerin domains across multiple growth substrates relative to C. churrovis monoculture. Two other fungal strains grown on a reed canary grass substrate in co-culture with the same methanogen also exhibited high log2-fold change values for upregulation of genes encoding carbohydrate binding modules in families 1 and 18. Transcriptional upregulation indicated that co-culture of the C. churrovis strain with a methanogen may enhance pyruvate formate lyase (PFL) function for growth on xylan and fructose and production of bottleneck enzymes in sugar utilization pathways, further supporting the hypothesis that co-culture with a methanogen may enhance certain fungal metabolic functions. Upregulation of CBM18 may play a role in fungal–methanogen physical associations and fungal cell wall development and remodeling.

09 BIOMASS FUELS↗

A roadmap for the functional annotation of protein families: a community perspective

Over the last 25 years, biology has entered the genomic era and is becoming a science of ‘big data’. Most interpretations of genomic analyses rely on accurate functional annotations of the proteins encoded by more than 500 000 genomes sequenced to date. By different estimates, only half the predicted sequenced proteins carry an accurate functional annotation, and this percentage varies drastically between different organismal lineages. Such a large gap in knowledge hampers all aspects of biological enterprise and, thereby, is standing in the way of genomic biology reaching its full potential. A brainstorming meeting to address this issue funded by the National Science Foundation was held during 3–4 February 2022. Bringing together data scientists, biocurators, computational biologists and experimentalists within the same venue allowed for a comprehensive assessment of the current state of functional annotations of protein families. Further, major issues that were obstructing the field were identified and discussed, which ultimately allowed for the proposal of solutions on how to move forward.

59 BASIC BIOLOGICAL SCIENCES↗

Supervised extraction of near-complete genomes from metagenomic samples: A new service in PATRIC

Large amounts of metagenomically-derived data are submitted to PATRIC for analysis. In the future, we expect even more jobs submitted to PATRIC will use metagenomic data. One in-demand use case is the extraction of near-complete draft genomes from assembled contigs of metagenomic origin. The PATRIC metagenome binning service utilizes the PATRIC database to furnish a large, diverse set of reference genomes. We provide a new service for supervised extraction and annotation of high-quality, near-complete genomes from metagenomically-derived contigs. Reference genomes are assigned to putative draft genome bins based on the presence of single-copy universal marker roles in the sample, and contigs are sorted into these bins by their similarity to reference genomes in PATRIC. Each set of binned contigs represents a draft genome that will be annotated by RASTtk in PATRIC. A structured-language binning report is provided containing quality measurements and taxonomic information about the contig bins. The PATRIC metagenome binning service emphasizes extraction of high-quality genomes for downstream analysis using other PATRIC tools and services. Due to its supervised nature, the binning service is not appropriate for mining novel or extremely low-coverage genomes from metagenomic samples.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

6051R & 6051S Assembly and Annotation

We report the draft genomes of two morphologically distinct variants of Bacillus subtilis ATCC 6051 [NCBI3610]. The two isolates exhibit differences in not only morphology but also their genetics, despite identical 16S rRNA sequences. Investigating the genetic differences of colony morphology variation in this model organism can provide valuable insights.

59 BASIC BIOLOGICAL SCIENCES↗

The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes

Abstract For over 10 years, ModelSEED has been a primary resource for the construction of draft genome-scale metabolic models based on annotated microbial or plant genomes. Now being released, the biochemistry database serves as the foundation of biochemical data underlying ModelSEED and KBase. The biochemistry database embodies several properties that, taken together, distinguish it from other published biochemistry resources by: (i) including compartmentalization, transport reactions, charged molecules and proton balancing on reactions; (ii) being extensible by the user community, with all data stored in GitHub; and (iii) design as a biochemical ‘Rosetta Stone’ to facilitate comparison and integration of annotations from many different tools and databases. The database was constructed by combining chemical data from many resources, applying standard transformations, identifying redundancies and computing thermodynamic properties. The ModelSEED biochemistry is continually tested using flux balance analysis to ensure the biochemical network is modeling-ready and capable of simulating diverse phenotypes. Ontologies can be designed to aid in comparing and reconciling metabolic reconstructions that differ in how they represent various metabolic pathways. ModelSEED now includes 33,978 compounds and 36,645 reactions, available as a set of extensible files on GitHub, and available to search at https://modelseed.org and KBase.

59 BASIC BIOLOGICAL SCIENCES↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗

KBase Silver Case Study: Determining Media Formulation Requirements for Isolation of Microbiome Constituents

KBase has powerful tools for extracting microbial genomes from metagenomes and performing phylogenomic analysis and metabolic modeling. These tools can be used to predict key media ingredients for isolating uncultured members of microbiomes. Essential to this process are high-quality genomes extracted from metagenomic assemblies, and Kbase has a tool for assessing the quality of genomes too. Below we identify growth factors for a myxobacteria ("slime bacteria") yet to be isolated from the rhizosphere of Miscanthus xgiganteus (hybrid of "Silvergrass"), cultivated at the Kellogg Biological Station in Michigan. Data was transferred with Globus from JGI-IMG. This narrative and the "KBase Gold Case Study: Can you find Delftia?" make up the Silver and Gold Narrative Set for teaching metagenomics concepts to students in the BIT 477/577 course at North Carolina State University. This tutorial will guide the user through the process of extracting and annotating high-quality genomes from a metagenomic data, performing phylogenomic analysis, building a metabolic model, and using these to predict nutrient requirments for growth and isolation of corresponding microbes.

59 BASIC BIOLOGICAL SCIENCES↗

Autotrophic and mixotrophic metabolism of an anammox bacterium revealed by in vivo 13 C and 2 H metabolic network mapping

Anaerobic ammonium-oxidizing (anammox) bacteria mediate a key step in the biogeochemical nitrogen cycle and have been applied worldwide for the energy-efficient removal of nitrogen from wastewater. However, outside their core energy metabolism, little is known about the metabolic networks driving anammox bacterial anabolism and use of different carbon and energy substrates beyond genome-based predictions. Here, we experimentally resolved the central carbon metabolism of the anammox bacterium Candidatus ‘Kuenenia stuttgartiensis’ using time-series 13 C and 2 H isotope tracing, metabolomics, and isotopically nonstationary metabolic flux analysis. Our findings confirm predicted metabolic pathways used for CO 2 fixation, central metabolism, and amino acid biosynthesis in K. stuttgartiensis , and reveal several instances where genomic predictions are not supported by in vivo metabolic fluxes. This includes the use of the oxidative branch of an incomplete tricarboxylic acid cycle for alpha-ketoglutarate biosynthesis, despite the genome not having an annotated citrate synthase. We also demonstrate that K. stuttgartiensis is able to directly assimilate extracellular formate via the Wood–Ljungdahl pathway instead of oxidizing it completely to CO 2 followed by reassimilation. In contrast, our data suggest that K. stuttgartiensis is not capable of using acetate as a carbon or energy source in situ and that acetate oxidation occurred via the metabolic activity of a low-abundance microorganism in the bioreactor’s side population. Together, these findings provide a foundation for understanding the carbon metabolism of anammox bacteria at a systems-level and will inform future studies aimed at elucidating factors governing their function and niche differentiation in natural and engineered ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Suillus : an emerging model for the study of ectomycorrhizal ecology and evolution

Research on mycorrhizal symbiosis has been slowed by a lack of established study systems. To address this challenge, we have been developing Suillus, a widespread ecologically and economically relevant fungal genus primarily associated with the plant family Pinaceae, into a model system for studying ectomycorrhizal (ECM) associations. Over the last decade, we have compiled extensive genomic resources, culture libraries, a phenotype database, and protocols for manipulating Suillus fungi with and without their tree partners. Our efforts have already resulted in a large number of publicly available genomes, transcriptomes, and respective annotations, as well as advances in our understanding of mycorrhizal partner specificity and host communication, fungal and plant nutrition, environmental adaptation, soil nutrient cycling, interspecific competition, and biological invasions. Here, we highlight the most significant recent findings enabled by Suillus, present a suite of protocols for working with the genus, and discuss how Suillus is emerging as an important model to elucidate the ecology and evolution of ECM interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Thaumarchaea Genome Sequences from a High Arctic Active Layer

The role of archaeal ammonia oxidizers often exceeds that of bacterial ammonia oxidizers in marine and terrestrial environments but has been understudied in permafrost, where thawing has the potential to release ammonia. Here, three thaumarchaea genomes were assembled and annotated from metagenomic data sets from carbon-poor Canadian High Arctic active-layer cryosols.

Sun, Emily Wei-Hsin↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation

Preliminary: This final report corresponds to a grant (DE-SC0021216) that was awarded to the University of Montana. Mid-way through the grant period, I relocated from the University of Montana to the University of Arizona. The grant was ended at University of Montana in late 2022, with all efforts concluding on 08/26/22; the remaining funds supporting the project were relinquished by University of Montana, and were later awarded to University of Arizona under a new grant, with start date 04/01/23. This report focuses on results of research efforts at UMontana through 08/26/22. Results: We made progress in each of the three aims of the proposal. We released software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes. We made substantial progress in developing software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we made notable progress in developing AI methods (specifically: a neural embedding model) for identifying similarities between protein sequences based on amino-wise latent vectors. These efforts were supplemented by development of methods for protein modeling in support of predicting protein-drug binding activity, and by my leadership of a team in the NIH/DOE 2021 Petabyte-Scale Sequence Search hack-a-thon.

59 BASIC BIOLOGICAL SCIENCES↗

Energetics and Kinetics of Syntrophic Aromatic Degradation (Final Technical Report)

This DOE Basic Energy Physical Biosciences project spanned a 28-year period and two project investigators. Numerous significant research discoveries have occurred over this time frame. Much of this work has been reported in peer-reviewed literature, with the publication of results from the last four years forthcoming. The overarching project themes have focused on probing the metabolism of syntrophic bacteria and their environments. Research supported by this project has been central to ten doctoral dissertations and one master’s thesis at the University of Oklahoma. The project has also supported numerous undergraduate research projects throughout the years. Additionally, several microbial genomes were sequenced and annotated in association with this work 12-16. Key findings and publications from this work are highlighted. More detailed results are provided for unpublished and embargoed work. A complete list of publications and research products associated with this project is at the report's end.

59 BASIC BIOLOGICAL SCIENCES↗

Conditional Filamentation of Paraburkholderia elongata 5N - Data Container

Overview of Dataset This narrative contains assemblies for all Paraburkholderia discussed in Karasz et al. 2022, where their phosphate-solubilizing activity was measured. Assemblies were downloaded from NCBI and annotated with Prokka. These genomes were created in many Narratives and collected here to be more accessible Note: The following naming of assemblies does not correspond with current taxonomic names for the following: Paraburkholderia 5N = Paraburkholderia elongata 5NT Paraburkholderia 1N = Paraburkholderia solitsuga 1NT Paraburkholderia vancouverensis = Paraburkholderia madseniana RL16

59 BASIC BIOLOGICAL SCIENCES↗