Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS↗

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete↗

A chromosome-level genome assembly of the varied leaved jewelflower, Streptanthus diversifolius, reveals a recent whole genome duplication

Abstract The Streptanthoid complex, a clade of primarily Streptanthus and Caulanthus species in the Thelypodieae (Brassicaceae) is an emerging model system for ecological and evolutionary studies. This complex spans the full range of the California Floristic Province including desert, foothill, and mountain environments. The ability of these related species to radiate into dramatically different environments makes them a desirable study subject for exploring how plant species expand their ranges and adapt to new environments over time. Ecological and evolutionary studies for this complex have revealed fascinating variation in serpentine soil adaptation, defense compounds, germination, flowering, and life history strategies. Until now a lack of publicly available genome assemblies has hindered the ability to relate these phenotypic observations to their underlying genetic and molecular mechanisms. To help remedy this situation, we present here a chromosome-level genome assembly and annotation of Streptanthus diversifolius, a member of the Streptanthoid Complex, developed using Illumina, Hi-C, and HiFi sequencing technologies. Construction of this assembly also provides further evidence to support the previously reported recent whole genome duplication unique to the Thelypodieae. This whole genome duplication may have provided individuals in the Streptanthoid Complex the genetic arsenal to rapidly radiate throughout the California Floristic Province and to occupy commonly inhospitable environments including serpentine soils.

Genetics & Heredity↗

TranSyT , an innovative framework for identifying transport systems

The importance and rate of development of genome-scale metabolic models have been growing for the last few years, increasing the demand for software solutions that automate several steps of this process. However, since TRIAGE’s release, software development for the automatic integration of transport reactions into models has stalled. Here, in this paper, we present the Transport Systems Tracker (TranSyT). Unlike other transport systems annotation software, TranSyT does not rely on manual curation to expand its internal database, which is derived from highly curated records retrieved from the Transporters Classification Database and complemented with information from other data sources. TranSyT compiles information regarding transporter families and proteins, and derives reactions into its internal database, making it available for rapid annotation of complete genomes. All transport reactions have GPR associations and can be exported with identifiers from four different metabolite databases. TranSyT is currently available as a plugin for merlin v4.0 and an app for KBase.

59 BASIC BIOLOGICAL SCIENCES↗

DRAM example narrative

DRAM example narrative DRAM on KBase let's anyone run annotations using DRAM in the cloud. DRAM is an annotation tool that can annotate bacterial, archaeal and viral genomes and distills those annotatios into represetations of the functional genomic potential of those organisms. If you want to read more about DRAM you can check out the GitHub, wiki and journal article. DRAM annotate assemblies In KBase Assembly objects contain nucleotide sequences from genomes or metagenomes. DRAM can predict genes and annotate their function from KBase Assembly objects which may be microbial isolate genomes, metagenome assembled genomes or metagenomes. This is done with the Annotate and Distill Assemblies with DRAM app. This app can also anntoate AssemblySet objects which contain collection of Assembly objects. It also generates a Genome object and a GenomeSet object which can be used for further analysis with other KBase apps. The full annotations and other DRAM files are also available for download in the app.

59 BASIC BIOLOGICAL SCIENCES↗

Whole-Genome Comparisons of Ergot Fungi Reveals the Divergence and Evolution of Species within the Genus Claviceps Are the Result of Varying Mechanisms Driving Genome Evolution and Host Range Expansion

The genus Claviceps has been known for centuries as an economically important fungal genus for pharmacology and agricultural research. Only recently have researchers begun to unravel the evolutionary history of the genus, with origins in South America and classification of four distinct sections through ecological, morphological, and metabolic features (Claviceps sects. Citrinae, Paspalorum, Pusillae, and Claviceps). The first three sections are additionally characterized by narrow host range, whereas section Claviceps is considered evolutionarily more successful and adaptable as it has the largest host range and biogeographical distribution. However, the reasons for this success and adaptability remain unclear. Our study elucidates factors influencing adaptability by sequencing and annotating 50 Claviceps genomes, representing 21 species, for a comprehensive comparison of genome architecture and plasticity in relation to host range potential. Our results show the trajectory from specialized genomes (sects. Citrinae and Paspalorum) toward adaptive genomes (sects. Pusillae and Claviceps) through colocalization of transposable elements around predicted effectors and a putative loss of repeat-induced point mutation resulting in unconstrained tandem gene duplication coinciding with increased host range potential and speciation. Alterations of genomic architecture and plasticity can substantially influence and shape the evolutionary trajectory of fungal pathogens and their adaptability. Furthermore, our study provides a large increase in available genomic resources to propel future studies of Claviceps in pharmacology and agricultural research, as well as, research into deeper understanding of the evolution of adaptable plant pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Physiological, genomic, and sulfur isotopic characterization of methanol metabolism by Desulfovibrio carbinolicus

Methanol is often considered as a non-competitive substrate for methanogenic archaea, but an increasing number of sulfate-reducing microorganisms (SRMs) have been reported to be capable of respiring with methanol as an electron donor. A better understanding of the fate of methanol in natural or artificial anaerobic systems thus requires knowledge of the methanol dissimilation by SRMs. In this study, we describe the growth kinetics and sulfur isotope effects of Desulfovibrio carbinolicus, a methanol-oxidizing sulfate-reducing deltaproteobacterium, together with its genome sequence and annotation. D. carbinolicus can grow with a series of alcohols from methanol to butanol. Compared to longer-chain alcohols, however, specific growth and respiration rates decrease by several fold with methanol as an electron donor. Larger sulfur isotope fractionation accompanies slowed growth kinetics, indicating low chemical potential at terminal reductive steps of respiration. In a medium containing both ethanol and methanol, D. carbinolicus does not consume methanol even after the cessation of growth on ethanol. Among the two known methanol dissimilatory systems, the genome of D. carbinolicus contains the genes coding for alcohol dehydrogenase but lacks enzymes analogous to methanol methyltransferase. We analyzed the genomes of 52 additional species of sulfate-reducing bacteria that have been tested for methanol oxidation. There is no apparent relationship between phylogeny and methanol metabolizing capacity, but most gram-negative methanol oxidizers grow poorly, and none carry homologs for methyltransferase (mtaB). Although the amount of available data is limited, it is notable that more than half of the known gram-positive methanol oxidizers have both enzymatic systems, showing enhanced growth relative to the SRMs containing only alcohol dehydrogenase genes. Thus, physiological, genomic, and sulfur isotopic results suggest that D. carbinolicus and close relatives have the ability to metabolize methanol but likely play a limited role in methanol degradation in most natural environments.

54 ENVIRONMENTAL SCIENCES↗

The reference genome and abiotic stress responses of the model perennial grass Brachypodium sylvaticum

Abstract Perennial grasses are important forage crops and emerging biomass crops and have the potential to be more sustainable grain crops. However, most perennial grass crops are difficult experimental subjects due to their large size, difficult genetics, and/or their recalcitrance to transformation. Thus, a tractable model perennial grass could be used to rapidly make discoveries that can be translated to perennial grass crops. Brachypodium sylvaticum has the potential to serve as such a model because of its small size, rapid generation time, simple genetics, and transformability. Here, we provide a high-quality genome assembly and annotation for B. sylvaticum, an essential resource for a modern model system. In addition, we conducted transcriptomic studies under 4 abiotic stresses (water, heat, salt, and freezing). Our results indicate that crowns are more responsive to freezing than leaves which may help them overwinter. We observed extensive transcriptional responses with varying temporal dynamics to all abiotic stresses, including classic heat-responsive genes. These results can be used to form testable hypotheses about how perennial grasses respond to these stresses. Taken together, these results will allow B. sylvaticum to serve as a truly tractable perennial model system.

59 BASIC BIOLOGICAL SCIENCES↗

The Ontology of Biological Attributes (OBA)—computational traits for the life sciences

Abstract Existing phenotype ontologies were originally developed to represent phenotypes that manifest as a character state in relation to a wild-type or other reference. However, these do not include the phenotypic trait or attribute categories required for the annotation of genome-wide association studies (GWAS), Quantitative Trait Loci (QTL) mappings or any population-focussed measurable trait data. The integration of trait and biological attribute information with an ever increasing body of chemical, environmental and biological data greatly facilitates computational analyses and it is also highly relevant to biomedical and clinical applications. The Ontology of Biological Attributes (OBA) is a formalised, species-independent collection of interoperable phenotypic trait categories that is intended to fulfil a data integration role. OBA is a standardised representational framework for observable attributes that are characteristics of biological entities, organisms, or parts of organisms. OBA has a modular design which provides several benefits for users and data integrators, including an automated and meaningful classification of trait terms computed on the basis of logical inferences drawn from domain-specific ontologies for cells, anatomical and other relevant entities. The logical axioms in OBA also provide a previously missing bridge that can computationally link Mendelian phenotypes with GWAS and quantitative traits. The term components in OBA provide semantic links and enable knowledge and data integration across specialised research community boundaries, thereby breaking silos.

59 BASIC BIOLOGICAL SCIENCES↗

Expression profiling of MADS-box gene family revealed its role in vegetative development and stem ripening in S. spontaneum

Sugarcane is the most important sugar and biofuel crop. MADS-box genes encode transcription factors that are involved in developmental control and signal transduction in plants. Systematic analyses of MADS-box genes have been reported in many plant species, but its identification and characterization were not possible until a reference genome of autotetraploid wild type sugarcane specie, Saccharum spontaneum is available recently. We identified 182 MADS-box sequences in the S. spontaneum genome, which were annotated into 63 genes, including 6 (9.5%) genes with four alleles, 21 (33.3%) with three, 29 (46%) with two, 7 (11.1%) with one allele. Paralogs (tandem duplication and disperse duplicated) were also identified and characterized. These MADS-box genes were divided into two groups; Type-I (21 Mα, 4 Mβ, 4 Mγ) and Type-II (32 MIKCc, 2 MIKC*) through phylogenetic analysis with orthologs in Arabidopsis and sorghum. Structural diversity and distribution of motifs were studied in detail. Chromosomal localizations revealed that S. spontaneum MADS-box genes were randomly distributed across eight homologous chromosome groups. The expression profiles of these MADS-box genes were analyzed in leaves, roots, stem sections and after hormones treatment. Important alleles based on promoter analysis and expression variations were dissected. qRT-PCR analysis was performed to verify the expression pattern of pivotal S. spontaneum MADS-box genes and suggested that flower timing genes ( SOC1 and SVP ) may regulate vegetative development.

59 BASIC BIOLOGICAL SCIENCES↗

IMG Annotation Pipeline (IMGAP) v5.1.13

The IMG Annotation Pipeline is a collection of Bash and Python scripts to control a workflow for structural and functional annotation of prokaryotic genomes, metagenomes, and metatranscriptomes. The bash scripts in general control the overall workflow and are wrappers around 3rd party executables (not included in repo) that predict features or functions. Whereas the Python scripts do post-processing of raw output in terms of filtering or format transformation and in some cases contain some logic for picking the correct predictions or resolving overlaps. The pipeline is tailored to produce results required by IMG (https://img.jgi.doe.gov/) and is executed on every dataset submitted to IMG via https://img.jgi.doe.gov/submit. These consist of internal genomes, metagenomes and metatranscriptomes sequenced and assembled at the JGI, as well as datasets submitted by external users (non-lab/JGI affiliates).

Huntemann, Marcel↗

KBase Narrative - StRoNG Net Part 2: Genome Analysis - Student - Static

Here we continue our exploration of novel genomes discovered in module 1. We will explore different approaches for examining the genome sequence and annotation data, investigating the best approaches for answering our question: Do benthic microbes have functioning circadian clocks?

Schirmer, Aaron↗

Combinatorial Glycomic Analyses to Direct CAZyme Discovery for the Tailored Degradation of Canola Meal Non-Starch Dietary Polysaccharides

Canola meal (CM), the protein-rich by-product of canola oil extraction, has shown promise as an alternative feedstuff and protein supplement in poultry diets, yet its use has been limited due to the abundance of plant cell wall fibre, specifically non-starch polysaccharides (NSP) and lignin. The addition of exogenous enzymes to promote the digestion of CM NSP in chickens has potential to increase the metabolizable energy of CM. We isolated chicken cecal bacteria from a continuous-flow mini-bioreactor system and selected for those with the ability to metabolize CM NSP. Of 100 isolates identified, Bacteroides spp. and Enterococcus spp. were the most common species with these capabilities. To identify enzymes specifically for the digestion of CM NSP, we used a combination of glycomics techniques, including enzyme-linked immunosorbent assay characterization of the plant cell wall fractions, glycosidic linkage analysis (methylation-GC-MS analysis) of CM NSP and their fractions, bacterial growth profiles using minimal media supplemented with CM NSP, and the sequencing and de novo annotation of bacterial genomes of high-efficiency CM NSP utilizing bacteria. The SACCHARIS pipeline was used to select plant cell wall active enzymes for recombinant production and characterization. This approach represents a multidisciplinary innovation platform to bioprospect endogenous CAZymes from the intestinal microbiota of herbivorous and omnivorous animals which is adaptable to a variety of applications and dietary polysaccharides.

glycome profiling↗

KBase Narrative - kb_DRAM E. coli annotation

Here we are annotating a E. coli K-12 genome using both DRAM and RAST then using those annotations to build models. The genome is from NCBI RefSeq ID NC_000913.

Shaffer, Michael↗

Relics of interspecific hybridization retained in the genome of a drought-adapted peanut cultivar

Peanut (Arachis hypogaea L.) is a globally important oil and food crop frequently grown in arid, semi-arid, or dryland environments. Improving drought tolerance is a key goal for peanut crop improvement efforts. Here, we present the genome assembly and gene model annotation for “Line8,” a peanut genotype bred from drought-tolerant cultivars. Our assembly and annotation are the most contiguous and complete peanut genome resources currently available. The high contiguity of the Line8 assembly allowed us to explore structural variation both between peanut genotypes and subgenomes. We detect several large inversions between Line8 and other peanut genome assemblies, and there is a trend for the inversions between more genetically diverged genotypes to have higher gene content. We also relate patterns of subgenome exchange to structural variation between Line8 homeologous chromosomes. Unexpectedly, we discover that Line8 harbors an introgression from A.cardenasii, a diploid peanut relative and important donor of disease resistance alleles to peanut breeding populations. The fully resolved sequences of both haplotypes in this introgression provide the first in situ characterization of A.cardenasii candidate alleles that can be leveraged for future targeted improvement efforts. The completeness of our genome will support peanut biotechnology and broader research into the evolution of hybridization and polyploidy.

60 APPLIED LIFE SCIENCES↗

AlloSHP: deconvoluting single homeologous polymorphism for phylogenetic analysis of allopolyploids

Background The genomic and evolutionary study of allopolyploid organisms involves multiple copies of homeologous chromosomes, making their assembly, annotation, and phylogenetic analysis challenging. Bioinformatics tools and protocols have been developed to study polyploid genomes, but sometimes require the assembly of their genomes, or at least the genes, limiting their use. Results We have developed AlloSHP, a command-line tool for detecting and extracting single homeologous polymorphisms (SHPs) from the subgenomes of allopolyploid species. This tool integrates three main algorithms, WGA, VCF2ALIGNMENT and VCF2SYNTENY, and allows the detection of SHPs for the study of diploid-polyploid complexes with available diploid progenitor genomes, without assembling and annotating the genomes of the allopolyploids under study. AlloSHP has been validated on three diploid-polyploid plant complexes, Brachypodium, Brassica, and Triticum-Aegilops, and a set of synthetic hybrid yeasts and their progenitors of the genus Saccharomyces. The results and congruent phylogenies obtained from the four datasets demonstrate the potential of AlloSHP for the evolutionary analysis of allopolyploids with a wide range of ploidy and genome sizes. Conclusions AlloSHP combines the strategies of simultaneous mapping against multiple reference genomes and syntenic alignment of these genomes to call SHPs, using as input data a single VCF file and the reference genomes of the known or closest extant diploid progenitor species. This novel approach provides a valuable tool for the evolutionary study of allopolyploid species, both at the interspecific and intraspecific levels, allowing the simultaneous analysis of a large number of accessions and avoiding the complex process of assembling polyploid genomes.

Allopolyploids↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗