Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Assembly and comparative genome analysis of a Patagonian Aureobasidium pullulans isolate reveals unexpected intraspecific variation

Aureobasidium pullulans is a yeast-like fungus with remarkable phenotypic plasticity widely studied for its importance for the pharmaceutical and food industries. So far, genomic studies with strains from all over the world suggest they constitute a genetically unstructured population, with no association by habitat. However, the mechanisms by which this genome supports so many phenotypic permutations are still poorly understood. Recent works have shown the importance of sequencing yeast genomes from extreme environments to increase the repertoire of phenotypic diversity of unconventional yeasts. In this study, we present the genomic draft of A. pullulans strain from a Patagonian yeast diversity hotspot, re-evaluate its taxonomic classification based on taxogenomic approaches, and annotate its genome with high-depth transcriptomic data. Here, our analysis suggests this isolate could be considered a novel variant at an early stage of the speciation process. The discovery of divergent strains in a genomically homogeneous group, such as A. pullulans, can be valuable in understanding the evolution of the species. The identification and characterization of new variants will not only allow finding unique traits of biotechnological importance, but also optimize the choice of strains whose phenotypes will be characterized, providing new elements to explore questions about plasticity and adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

CyanoCyc cyanobacterial web portal

CyanoCyc is a web portal that integrates an exceptionally rich database collection of information about cyanobacterial genomes with an extensive suite of bioinformatics tools. It was developed to address the needs of the cyanobacterial research and biotechnology communities. The 277 annotated cyanobacterial genomes currently in CyanoCyc are supplemented with computational inferences including predicted metabolic pathways, operons, protein complexes, and orthologs; and with data imported from external databases, such as protein features and Gene Ontology (GO) terms imported from UniProt. Five of the genome databases have undergone manual curation with input from more than a dozen cyanobacteria experts to correct errors and integrate information from more than 1,765 published articles. CyanoCyc has bioinformatics tools that encompass genome, metabolic pathway and regulatory informatics; omics data analysis; and comparative analyses, including visualizations of multiple genomes aligned at orthologous genes, and comparisons of metabolic networks for multiple organisms. CyanoCyc is a high-quality, reliable knowledgebase that accelerates scientists’ work by enabling users to quickly find accurate information using its powerful set of search tools, to understand gene function through expert mini-reviews with citations, to acquire information quickly using its interactive visualization tools, and to inform better decision-making for fundamental and applied research.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/PR: a database of plasmids from genomes and metagenomes with rich annotations and metadata

Plasmids are mobile genetic elements found in many clades of Archaea and Bacteria. They drive horizontal gene transfer, impacting ecological and evolutionary processes within microbial communities, and hold substantial importance in human health and biotechnology. To support plasmid research and provide scientists with data of an unprecedented diversity of plasmid sequences, we introduce the IMG/PR database, a new resource encompassing 699 973 plasmid sequences derived from genomes, metagenomes and metatranscriptomes. IMG/PR is the first database to provide data of plasmid that were systematically identified from diverse microbiome samples. IMG/PR plasmids are associated with rich metadata that includes geographical and ecosystem information, host taxonomy, similarity to other plasmids, functional annotation, presence of genes involved in conjugation and antibiotic resistance. The database offers diverse methods for exploring its extensive plasmid collection, enabling users to navigate plasmids through metadata-centric queries, plasmid comparisons and BLAST searches. The web interface for IMG/PR is accessible at https://img.jgi.doe.gov/pr. Plasmid metadata and sequences can be downloaded from https://genome.jgi.doe.gov/portal/IMG_PR.

59 BASIC BIOLOGICAL SCIENCES↗

A genomic perspective on fungal diversity and evolution

Originating from aquatic unicellular ancestors, over the course of ~1 billion years, the fungi have evolved to occupy nearly all aerobic environments on the planet, diversified into millions of different ‘species’ and have developed complex multicellular structures. Their relatively small, simple genomes have facilitated massive-scale sequencing and allowed us to explore genome evolution across an ancient eukaryotic kingdom. With thousands of genomes from diverse lineages now available, this Review will discuss insights into fungal biology and evolution gleaned with genomics and other multi-omics approaches. Using published genomes available through GenBank and the Joint Genome Institute’s MycoCosm platform, we generated kingdom-wide phylogenies and used them to highlight how fungal genomes have changed over time. With this phylogeny as a guide, we also discuss major evolutionary transitions that occurred across the fungal kingdom. Although progress has been made, these efforts are hampered by biases in genome representation and limited characterization of gene functions. Here, in this study, we discuss these challenges and possible future directions to address them, including initiatives to characterize conserved genes of unknown function and scale up sequencing towards 10,000 annotated fungal genomes.

Mondo, Stephen J. [USDOE Joint Genome Institute (J↗

Exploring life’s hidden majority: microbial dark matter symposium highlights

The Microbial Dark Matter Symposium held on August 28–29, 2025, in Laguna Beach, Orange County, CA, convened a multidisciplinary group of scientists to address the vast unknowns in microbial life—from uncultured taxa and uncharacterized proteins to elusive viruses and spacefaring microbes. Set against a scenic coastal backdrop, the symposium highlighted advances in single-cell genomics, proximity ligation sequencing, and artificial intelligence-ready bioinformatics, while also probing the limits of microbial persistence, metabolism, and ecological distribution. Sessions explored microbial dark matter from multiple dimensions: cultivability, where new strategies are enabling recovery of elusive microbes; functional ambiguity, where metagenomic dark zones are illuminated by computational annotation; and genomic representation, where single-cell methods bridge gaps left by shotgun community sequencing. Researchers shared breakthroughs in identifying atmospheric microbiomes, “dark oxygen” production in groundwater ecosystems, and microbial survival on the International Space Station. The symposium emphasized integration of methods, disciplines, and ecosystems, advancing a collective push to illuminate the microbial dark matter on Earth and beyond. By highlighting emerging tools, pressing questions, and cross-domain insights, the symposium underscored the need for collaborative, open, and adaptive approaches to study the microbial unknown. The meeting marks a pivotal moment in microbiology, where cultivating knowledge of the uncultivated promises transformative understanding of life, everywhere.

Podar, Mircea [ORNL] (ORCID:0000000327760205)↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

A unified catalog of 204,938 reference genomes from the human gut microbiome

Comprehensive, high-quality reference genomes are required for functional characterization and taxonomic assignment of the human gut microbiota. We present the Unified Human Gastrointestinal Genome (UHGG) collection, comprising 204,938 nonredundant genomes from 4,644 gut prokaryotes. These genomes encode >170 million protein sequences, which we collated in the Unified Human Gastrointestinal Protein (UHGP) catalog. The UHGP more than doubles the number of gut proteins in comparison to those present in the Integrated Gene Catalog. More than 70% of the UHGG species lack cultured representatives, and 40% of the UHGP lack functional annotations. Intraspecies genomic variation analyses revealed a large reservoir of accessory genes and single-nucleotide variants, many of which are specific to individual human populations. The UHGG and UHGP collections will enable studies linking genotypes to phenotypes in the human gut microbiome.

59 BASIC BIOLOGICAL SCIENCES↗

Bridging Place-Based Astrobiology Education with Genomics, Including Descriptions of Three Novel Bacterial Species Isolated from Mars Analog Sites of Cultural Relevance

Democratizing genomic data science, including bioinformatics, can diversify the STEM workforce and may, in turn, bring new perspectives into the space sciences. In this respect, the development of education and research programs that bridge genome science with “place” and world-views specific to a given region are valuable for Indigenous students and educators. Through a multi-institutional collaboration, we developed an ongoing education program and model that includes Illumina and Oxford Nanopore sequencing, free bioinformatic platforms, and teacher training workshops to address our research and education goals through a place-based science education lens. High school students and researchers cultivated, sequenced, assembled, and annotated the genomes of 13 bacteria from Mars analog sites with cultural relevance, 10 of which were novel species. Students, teachers, and community members assisted with the discovery of new, potentially chemolithotrophic bacteria relevant to astrobiology. This joint education-research program also led to the discovery of species from Mars analog sites capable of producing N-acyl homoserine lactones, which are quorum-sensing molecules used in bacterial communication. Whole genome sequencing was completed in high school classrooms, and connected students to funded space research, increased research output, and provided culturally relevant, place-based science education, with participants naming three novel species described here. Students at St. Andrew's School (Honolulu, Hawai‘i) proposed the name Bradyrhizobium prioritasuperba for the type strain, BL16A T , of the new species (DSM 112479 T = NCTC 14602 T ). The nonprofit organization Kauluakalana proposed the name Brenneria ulupoensis for the type strain, K61 T , of the new species (DSM 116657 T = LMG = 33184 T ), and Hawai‘i Baptist Academy students proposed the name Paraflavitalea speifideiaquila for the type strain, BL16E T , of the new species (DSM 112478 T = NCTC 14603 T ).

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Comparative genome analyses suggest a hemibiotrophic lifestyle and virulence differences for the beech bark disease fungal pathogens Neonectria faginata and Neonectria coccinea

Abstract Neonectria faginata and Neonectria coccinea are the causal agents of the insect-fungus disease complex known as beech bark disease (BBD), known to cause mortality in beech forest stands in North America and Europe. These fungal species have been the focus of extensive ecological and disease management studies, yet less progress has been made toward generating genomic resources for both micro- and macro-evolutionary studies. Here, we report a 42.1 and 42.7 mb highly contiguous genome assemblies of N. faginata and N. coccinea, respectively, obtained using Illumina technology. These species share similar gene number counts (12,941 and 12,991) and percentages of predicted genes with assigned functional categories (64 and 65%). Approximately 32% of the predicted proteomes of both species are homologous to proteins involved in pathogenicity, yet N. coccinea shows a higher number of predicted mitogen-activated protein kinase genes, virulence determinants possibly contributing to differences in disease severity between N. faginata and N. coccinea. A wide range of genes encoding for carbohydrate-active enzymes capable of degradation of complex plant polysaccharides and a small number of predicted secretory effector proteins, secondary metabolite biosynthesis clusters and cytochrome oxidase P450 genes were also found. This arsenal of enzymes and effectors correlates with, and reflects, the hemibiotrophic lifestyle of these two fungal pathogens. Phylogenomic analysis and timetree estimations indicated that the N. faginata and N. coccinea species divergence may have occurred at ∼4.1 million years ago. Differences were also observed in the annotated mitochondrial genomes as they were found to be 81.7 kb (N. faginata) and 43.2 kb (N. coccinea) in size. The mitochondrial DNA expansion observed in N. faginata is attributed to the invasion of introns into diverse intra- and intergenic locations. These first draft genomes of N. faginata and N. coccinea serve as valuable tools to increase our understanding of basic genetics, evolutionary mechanisms and molecular physiology of these two nectriaceous plant pathogenic species.

Salgado-Salazar, Catalina↗

Crowdsourcing biocuration: The Community Assessment of Community Annotation with Ontologies (CACAO)

Experimental data about gene functions curated from the primary literature have enormous value for research scientists in understanding biology. Using the Gene Ontology (GO), manual curation by experts has provided an important resource for studying gene function, especially within model organisms. Unprecedented expansion of the scientific literature and validation of the predicted proteins have increased both data value and the challenges of keeping pace. Capturing literature-based functional annotations is limited by the ability of biocurators to handle the massive and rapidly growing scientific literature. Within the community-oriented wiki framework for GO annotation called the Gene Ontology Normal Usage Tracking System (GONUTS), we describe an approach to expand biocuration through crowdsourcing with undergraduates. This multiplies the number of high-quality annotations in international databases, enriches our coverage of the literature on normal gene function, and pushes the field in new directions. From an intercollegiate competition judged by experienced biocurators, Community Assessment of Community Annotation with Ontologies (CACAO), we have contributed nearly 5,000 literature-based annotations. Many of those annotations are to organisms not currently well-represented within GO. Over a 10-year history, our community contributors have spurred changes to the ontology not traditionally covered by professional biocurators. The CACAO principle of relying on community members to participate in and shape the future of biocuration in GO is a powerful and scalable model used to promote the scientific enterprise. It also provides undergraduate students with a unique and enriching introduction to critical reading of primary literature and acquisition of marketable skills.

59 BASIC BIOLOGICAL SCIENCES↗

An optimized ChIP-Seq framework for profiling histone modifications in Chromochloris zofingiensis

The eukaryotic green alga Chromochloris zofingiensis is a reference organism for studying carbon partitioning and a promising candidate for the production of biofuel precursors. Recent transcriptome profiling transformed our understanding of its biology and generally algal biology, but epigenetic regulation remains understudied and represents a fundamental gap in our understanding of algal gene expression. Chromatin immunoprecipitation followed by deep sequencing (ChIP-Seq) is a powerful tool for the discovery of such mechanisms, by identifying genome-wide histone modification patterns and transcription factor-binding sites alike. Here, we established a ChIP-Seq framework for Chr. zofingiensis yielding over 20 million high-quality reads per sample. The most critical steps in a ChIP experiment were optimized, including DNA shearing to obtain an average DNA fragment size of 250 bp and assessment of the recommended formaldehyde concentration for optimal DNA-protein cross-linking. We used this ChIP-Seq framework to generate a genome-wide map of the H3K4me3 distribution pattern and to integrate these data with matching RNA-Seq data. In line with observations from other organisms, H3K4me3 marks predominantly transcription start sites of genes. Our H3K4me3 ChIP-Seq data will pave the way for improved genome structural annotation in the emerging reference alga Chr. zofingiensis.

59 BASIC BIOLOGICAL SCIENCES↗

Barcoded overexpression screens in gut Bacteroidales identify genes with roles in carbon utilization and stress resistance

Abstract A mechanistic understanding of host-microbe interactions in the gut microbiome is hindered by poorly annotated bacterial genomes. While functional genomics can generate large gene-to-phenotype datasets to accelerate functional discovery, their applications to study gut anaerobes have been limited. For instance, most gain-of-function screens of gut-derived genes have been performed in Escherichia coli and assayed in a small number of conditions. To address these challenges, we develop Barcoded Overexpression BActerial shotgun library sequencing (Boba-seq). We demonstrate the power of this approach by assaying genes from diverse gut Bacteroidales overexpressed in Bacteroides thetaiotaomicron . From hundreds of experiments, we identify new functions and phenotypes for 29 genes important for carbohydrate metabolism or tolerance to antibiotics or bile salts. Highlights include the discovery of a d -glucosamine kinase, a raffinose transporter, and several routes that increase tolerance to ceftriaxone and bile salts through lipid biosynthesis. This approach can be readily applied to develop screens in other strains and additional phenotypic assays.

59 BASIC BIOLOGICAL SCIENCES↗

Reekeekee- and roodoodooviruses, two different Microviridae clades constituted by the smallest DNA phages

Small circular single-stranded DNA viruses of the Microviridae family are both prevalent and diverse in all ecosystems. They usually harbor a genome between 4.3 and 6.3 kb, with a microvirus recently isolated from a marine Alphaproteobacteria being the smallest known genome of a DNA phage (4.248 kb). A subfamily, Amoyvirinae, has been proposed to classify this virus and other related small Alphaproteobacteria-infecting phages. Here, we report the discovery, in meta-omics data sets from various aquatic ecosystems, of sixteen complete microvirus genomes significantly smaller (2.991–3.692 kb) than known ones. Phylogenetic analysis reveals that these sixteen genomes represent two related, yet distinct and diverse, novel groups of microviruses—amoyviruses being their closest known relatives. We propose that these small microviruses are members of two tentatively named subfamilies Reekeekeevirinae and Roodoodoovirinae. As known microvirus genomes encode many overlapping and overprinted genes that are not identified by gene prediction software, we developed a new methodology to identify all genes based on protein conservation, amino acid composition, and selection pressure estimations. Surprisingly, only four to five genes could be identified per genome, with the number of overprinted genes lower than that in phiX174. These small genomes thus tend to have both a lower number of genes and a shorter length for each gene, leaving no place for variable gene regions that could harbor overprinted genes. Even more surprisingly, these two Microviridae groups had specific and different gene content, and major differences in their conserved protein sequences, highlighting that these two related groups of small genome microviruses use very different strategies to fulfill their lifecycle with such a small number of genes. The discovery of these genomes and the detailed prediction and annotation of their genome content expand our understanding of ssDNA phages in nature and are further evidence that these viruses have explored a wide range of possibilities during their long evolution.

59 BASIC BIOLOGICAL SCIENCES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗