Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

MaizeMine: A Data Mining Warehouse for the Maize Genetics and Genomics Database

MaizeMine is the data mining resource of the Maize Genetics and Genome Database (MaizeGDB; http://maizemine.maizegdb.org). It enables researchers to create and export customized annotation datasets that can be merged with their own research data for use in downstream analyses. MaizeMine uses the InterMine data warehousing system to integrate genomic sequences and gene annotations from the Zea mays B73 RefGen_v3 and B73 RefGen_v4 genome assemblies, Gene Ontology annotations, single nucleotide polymorphisms, protein annotations, homologs, pathways, and precomputed gene expression levels based on RNA-seq data from the Z. mays B73 Gene Expression Atlas. MaizeMine also provides database cross references between genes of alternative gene sets from Gramene and NCBI RefSeq. MaizeMine includes several search tools, including a keyword search, built-in template queries with intuitive search menus, and a QueryBuilder tool for creating custom queries. The Genomic Regions search tool executes queries based on lists of genome coordinates, and supports both the B73 RefGen_v3 and B73 RefGen_v4 assemblies. The List tool allows you to upload identifiers to create custom lists, perform set operations such as unions and intersections, and execute template queries with lists. When used with gene identifiers, the List tool automatically provides gene set enrichment for Gene Ontology (GO) and pathways, with a choice of statistical parameters and background gene sets. With the ability to save query outputs as lists that can be input to new queries, MaizeMine provides limitless possibilities for data integration and meta-analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie↗

The state of algal genome quality and diversity

The genomic era of biology has created unprecedented opportunity to study how life works, including understanding evolutionary principles, bioprospecting for novel antibiotics, and genetic manipulation of bioeconomically-relevant species. While genome sequencing and genomic analysis was previously restricted to large-scale projects and consortia, sequencing has become democratized. However, as genomics has become commonplace, the cataloging of sequenced organisms has become challenging, and standardized practices for sequencing, assembly, and annotation have not been adopted. This is equally true in research fields such as algal biology, despite the growing importance of algae in a bio-based economy. Here, in this study, we provide a comprehensive review of the state of eukaryotic algal genomics, explore the quality of algal genome assemblies, and identify the biases and gaps in the current species distribution to inform the development of future genome projects. Overall, we find a trend of declining quality of genomic resources, including a reduction in assembly quality, gene annotation quality, and genome completeness. Potential solutions to improve genome quality include widespread utilization of long read and scaffolding technologies, implementation of standards for assembly quality, evidence-based gene annotation and requisite publication, and support of continued development of gene prediction and genome assessment software.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii V.2

The head smut (Tilletia maclaganii) is a significant pathogen of the bioenergy crop switchgrass. T. maclaganii typically is more prevalent in older stands of switchgrass and can contribute to significant biomass loss. Here, we outline the methods for the sequencing, assembly, and annotation of the first reference genome for Tilletia maclaganii.

Benucci, Gian Maria Niccolò [GLBRC - Michigan Stat↗

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES↗

The Tetracentron genome provides insight into the early evolution of eudicots and the formation of vessel elements

Tetracentron sinense is an endemic and endangered deciduous tree. It belongs to the Trochodendrales, one of four early diverging lineages of eudicots known for having vesselless secondary wood. Sequencing and resequencing of the T. sinense genome will help us understand eudicot evolution, the genetic basis of tracheary element development, and the genetic diversity of this relict species. Here, we report a chromosome-scale assembly of the T. sinense genome. We assemble the 1.07 Gb genome sequence into 24 chromosomes and annotate 32,690 protein-coding genes. Phylogenomic analyses verify that the Trochodendrales and core eudicots are sister lineages and showed that two whole-genome duplications occurred in the Trochodendrales approximately 82 and 59 million years ago. Synteny analyses suggest that the γ event, resulting in paleohexaploidy, may have only happened in core eudicots. Interestingly, we find that vessel elements are present in T. sinense, which has two orthologs of AtVND7, the master regulator of vessel formation. T. sinense also has several key genes regulated by or regulating TsVND7.2 and their regulatory relationship resembles that in Arabidopsis thaliana. Resequencing and population genomics reveals high levels of genetic diversity of T. sinense and identifies four refugia in China. The T. sinense genome provides a unique reference for inferring the early evolution of eudicots and the mechanisms underlying vessel element formation. Population genomics analysis of T. sinense reveals its genetic diversity and geographic structure with implications for conservation.

59 BASIC BIOLOGICAL SCIENCES↗

Omics-guided metabolic pathway discovery in plants: Resources, approaches, and opportunities

Plants produce a vast array of metabolites, the biosynthetic routes of which remain largely undetermined. Genome-scale enzyme and pathway annotations and omics technologies have revolutionized research to decrypt plant metabolism and produced a growing list of functionally characterized metabolic genes and pathways. However, what is known is still a tiny fraction of the metabolic capacity harbored by plants. Here, in this work, we review plant enzyme and pathway annotation resources and cutting-edge omics approaches to guide discovery and characterization of plant metabolic pathways. We also discuss strategies for improving enzyme function prediction by integrating protein 3D structure information and single cell omics. This review aims to serve as a primer for plant biologists to leverage omics datasets to facilitate understanding and engineering plant metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Comparative genomic and transcriptomic analyses of trans-kingdom pathogen Fusarium solani species complex reveal degrees of compartmentalization

Background: The Fusarium solani species complex (FSSC) comprises fungal pathogens responsible for mortality in a diverse range of animals and plants, but their genome diversity and transcriptome responses in animal pathogenicity remain to be elucidated. We sequenced, assembled and annotated six chromosome-level FSSC clade 3 genomes of aquatic animal and plant host origins. We established a pathosystem and investigated the expression data of F. falciforme and F. keratoplasticum in Chinese softshell turtle (Pelodiscus sinensis) host. Results: Comparative analyses between the FSSC genomes revealed a spectrum of conservation patterns in chromosomes categorised into three compartments: core, fast-core (FC), and lineage-specific (LS). LS chromosomes contribute to variations in genomes size, with up to 42.2% of variations between F. vanettenii strains. Each chromosome compartment varied in structural architectures, with FC and LS chromosomes contain higher proportions of repetitive elements with genes enriched in functions related to pathogenicity and niche expansion. We identified differences in both selection in the coding sequences and DNA methylation levels between genome features and chromosome compartments which suggest a multi-speed evolution that can be traced back to the last common ancestor of Fusarium. We further demonstrated that F. falciforme and F. keratoplasticum are opportunistic pathogens by inoculating P. sinensis eggs and identified differentially expressed genes also associated with plant pathogenicity. These included the most upregulated genes encoding the CFEM (Common in Fungal Extracellular Membrane) domain. Conclusions: The high-quality genome assemblies provided new insights into the evolution of FSSC chromosomes, which also serve as a resource for studies of fungal genome evolution and pathogenesis. This study also establishes an animal model for fungal pathogens of trans-kingdom hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Genomic mechanisms of climate adaptation in polyploid bioenergy switchgrass

Long-term climate change and periodic environmental extremes threaten food and fuel security and global crop productivity. Although molecular and adaptive breeding strategies can buffer the effects of climatic stress and improve crop resilience, these approaches require sufficient knowledge of the genes that underlie productivity and adaptation—knowledge that has been limited to a small number of well-studied model systems. Here we present the assembly and annotation of the large and complex genome of the polyploid bioenergy crop switchgrass ( Panicum virgatum ). Analysis of biomass and survival among 732 resequenced genotypes, which were grown across 10 common gardens that span 1,800 km of latitude, jointly revealed extensive genomic evidence of climate adaptation. Climate–gene–biomass associations were abundant but varied considerably among deeply diverged gene pools. Furthermore, we found that gene flow accelerated climate adaptation during the postglacial colonization of northern habitats through introgression of alleles from a pre-adapted northern gene pool. The polyploid nature of switchgrass also enhanced adaptive potential through the fractionation of gene function, as there was an increased level of heritable genetic diversity on the nondominant subgenome. In addition to investigating patterns of climate adaptation, the genome resources and gene–trait associations developed here provide breeders with the necessary tools to increase switchgrass yield for the sustainable production of bioenergy.

09 BIOMASS FUELS↗

Make the Most Out of Genome Announcements with KBase

Genomics is a dynamic field driven by new technology, massive growth in data, and changes in how scientists share knowledge. This brings new challenges for researchers looking to maximize the impact of their science and ensure the longevity of their data. These changes are especially relevant for the practice of publishing genome announcements. Genome announcements are meant to share and promote newly sequenced organisms with the larger scientific community. A few decades ago, genome announcements were highly anticipated hallmarks of novel science. But, what was once a publishing coup is now an every day event. It is harder and harder to keep abreast of all the genomes that have been sequenced and promote your own new genomes to get the attention and recognition of your research community. Fortunately, KBase can help construct genomes from your sequencing data, generate reports for your announcement, and promote your data to your research community.

59 BASIC BIOLOGICAL SCIENCES↗

JGI Plant Gene Atlas: an updateable transcriptome resource to improve functional gene descriptions across the plant kingdom

Abstract Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an updateable data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we analyzed expression profiles, built gene clusters that exhibited tissue/condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles for genes without any previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas (https://plantgeneatlas.jgi.doe.gov) and Phytozome (https://phytozome.jgi.doe.gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentially expressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.

59 BASIC BIOLOGICAL SCIENCES↗

MONet/1000 Soils metagenome pathway modelling narrative w/ auto batch import

This narrative performs metabolic modeling and flux balance analysis (FBA) using metagenome-assembled genomes (MAGs) from the 1000 Soils samples, as described by Song et al. (2026, accepted). The set of MAGs (in FASTA format) is converted into a set of assembly objects compatible with functional annotation via RASTtk, yielding a set of genome objects that undergo metabolic modeling via OMEGGA. These genome objects are then used to conduct FBA, generating tables of metabolite uptake rates across the MAGs under investigation.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

Identification of mobile genetic elements with geNomad

Identifying and characterizing mobile genetic elements in sequencing data is essential for understanding their diversity, ecology, biotechnological applications and impact on public health. Here we introduce geNomad, a classification and annotation framework that combines information from gene content and a deep neural network to identify sequences of plasmids and viruses. geNomad uses a dataset of more than 200,000 marker protein profiles to provide functional gene annotation and taxonomic assignment of viral genomes. Using a conditional random field model, geNomad also detects proviruses integrated into host genomes with high precision. In benchmarks, geNomad achieved high classification performance for diverse plasmids and viruses (Matthews correlation coefficient of 77.8% and 95.3%, respectively), substantially outperforming other tools. Leveraging geNomad’s speed and scalability, we processed over 2.7 trillion base pairs of sequencing data, leading to the discovery of millions of viruses and plasmids that are available through the IMG/VR and IMG/PR databases. geNomad is available at https://portal.nersc.gov/genomad.

59 BASIC BIOLOGICAL SCIENCES↗

Multi‐season analysis reveals hundreds of drought‐responsive genes in sorghum

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3 years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Cole, Benjamin [USDOE Joint Genome Institute (JGI)↗