Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

3′ RNA-seq is superior to standard RNA-seq in cases of sparse data but inferior at identifying toxicity pathways in a model organism

The application of RNA-sequencing has led to numerous breakthroughs related to investigating gene expression levels in complex biological systems. Among these are knowledge of how organisms, such as the vertebrate model organism zebrafish (Danio rerio), respond to toxicant exposure. Recently, the development of 3' RNA-seq has allowed for the determination of gene expression levels with a fraction of the required reads compared to standard RNA-seq. While 3' RNA-seq has many advantages, a comparison to standard RNA-seq has not been performed in the context of whole organism toxicity and sparse data. Here, we examined samples from zebrafish exposed to perfluorobutane sulfonamide (FBSA) with either 3' or standard RNA-seq to determine the advantages of each with regards to the identification of functionally enriched pathways. We found that 3' and standard RNA-seq showed specific advantages when focusing on annotated or unannotated regions of the genome. We also found that standard RNA-seq identified more differentially expressed genes (DEGs), but that this advantage disappeared under conditions of sparse data. We also found that standard RNA-seq had a significant advantage in identifying functionally enriched pathways via analysis of DEG lists but that this advantage was minimal when identifying pathways via gene set enrichment analysis of all genes. These results show that each approach has experimental conditions where they may be advantageous. Our observations can help guide others in the choice of 3' RNA-seq vs standard RNA sequencing to query gene expression levels in a range of biological systems.

3’ RNA-seq↗

Pseudomonas aeruginosa gene PA4880 encodes a Dps-like protein with a Dps fold, bacterioferritin-type ferroxidase centers, and endonuclease activity

We report the biochemical, structural, and functional characterization of the protein coded by gene PA4880 in the P. aeruginosa PAO1 genome. The PA4880 gene had been annotated as coding a probable bacterioferritin. Our structural work shows that the product of gene PA4880 is a protein that adopts the Dps subunit fold, which oligomerizes into a 12-mer quaternary structure. Unlike Dps, however, the ferroxidase di-iron centers and iron coordinating ligands are buried within each subunit, in a manner identical to that observed in the ferroxidase center of P. aeruginosa bacterioferritin. Since these structural characteristics correspond to Dps-like proteins, we term the protein as P. aeruginosa Dps-like, or Pa DpsL. The ferroxidase centers in Pa DpsL catalyze the oxidation of Fe 2+ utilizing O 2 or H 2 O 2 as oxidant, and the resultant Fe 3+ is compartmentalized in the interior cavity. Interestingly, incubating Pa DpsL with plasmid DNA results in efficient nicking of the DNA and at higher concentrations of Pa DpsL the DNA is linearized and eventually degraded. The nickase and endonuclease activities suggest that Pa DpsL, in addition to participating in the defense of P. aeruginosa cells against iron-induced toxicity, may also participate in the innate immune mechanisms consisting of restriction endonucleases and cognate methyl transferases.

59 BASIC BIOLOGICAL SCIENCES↗

Ecological Trait-Based Digital Categorization of Microbial Genomes for Denitrification Potential

Microorganisms encode proteins that function in the transformations of useful and harmful nitrogenous compounds in the global nitrogen cycle. The major transformations in the nitrogen cycle are nitrogen fixation, nitrification, denitrification, anaerobic ammonium oxidation, and ammonification. The focus of this report is the complex biogeochemical process of denitrification, which, in the complete form, consists of a series of four enzyme-catalyzed reduction reactions that transforms nitrate to nitrogen gas. Denitrification is a microbial strain-level ecological trait (characteristic), and denitrification potential (functional performance) can be inferred from trait rules that rely on the presence or absence of genes for denitrifying enzymes in microbial genomes. Despite the global significance of denitrification and associated large-scale genomic and scholarly data sources, there is lack of datasets and interactive computational tools for investigating microbial genomes according to denitrification trait rules. Therefore, our goal is to categorize archaeal and bacterial genomes by denitrification potential based on denitrification traits defined by rules of enzyme involvement in the denitrification reduction steps. We report the integration of datasets on genome, taxonomic lineage, ecosystem, and denitrifying enzymes to provide data investigations context for the denitrification potential of microbial strains. We constructed an ecosystem and taxonomic annotated denitrification potential dataset of 62,624 microbial genomes (866 archaea and 61,758 bacteria) that encode at least one of the twelve denitrifying enzymes in the four-step canonical denitrification pathway. Our four-digit binary-coding scheme categorized the microbial genomes to one of sixteen denitrification traits including complete denitrification traits assigned to 3280 genomes from 260 bacteria genera. The bacterial strains with complete denitrification potential pattern included Arcobacteraceae strains isolated or detected in diverse ecosystems including aquatic, human, plant, and Mollusca (shellfish). The dataset on microbial denitrification potential and associated interactive data investigations tools can serve as research resources for understanding the biochemical, molecular, and physiological aspects of microbial denitrification, among others. The microbial denitrification data resources produced in our research can also be useful for identifying microbial strains for synthetic denitrifying communities.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Genome analyses reveal population structure and a purple stigma color gene candidate in finger millet

Finger millet is a key food security crop widely grown in eastern Africa, India and Nepal. Long considered a ‘poor man’s crop’, finger millet has regained attention over the past decade for its climate resilience and the nutritional qualities of its grain. To bring finger millet breeding into the 21 st century, here we present the assembly and annotation of a chromosome-scale reference genome. We show that this ~1.3 million years old allotetraploid has a high level of homoeologous gene retention and lacks subgenome dominance. Population structure is mainly driven by the differential presence of large wild segments in the pericentromeric regions of several chromosomes. Trait mapping, followed by variant analysis of gene candidates, reveals that loss of purple coloration of anthers and stigma is associated with loss-of-function mutations in the finger millet orthologs of the maize R1/B1 and Arabidopsis GL3/EGL3 anthocyanin regulatory genes. Proanthocyanidin production in seed is not affected by these gene knockouts.

59 BASIC BIOLOGICAL SCIENCES↗

Multiomic atlas with functional stratification and developmental dynamics of zebrafish cis -regulatory elements

Zebrafish, a popular organism for studying embryonic development and for modeling human diseases, has so far lacked a systematic functional annotation program akin to those in other animal models. To address this, we formed the international DANIO-CODE consortium and created a central repository to store and process zebrafish developmental functional genomic data. Our data coordination center (https://danio-code.zfin.org) combines a total of 1,802 sets of unpublished and re-analyzed published genomic data, which we used to improve existing annotations and show its utility in experimental design. We identified over 140,000 cis-regulatory elements throughout development, including classes with distinct features dependent on their activity in time and space. We delineated the distinct distance topology and chromatin features between regulatory elements active during zygotic genome activation and those active during organogenesis. Finally, we matched regulatory elements and epigenomic landscapes between zebrafish and mouse and predicted functional relationships between them beyond sequence similarity, thus extending the utility of zebrafish developmental genomics to mammals.

59 BASIC BIOLOGICAL SCIENCES↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

FeGenie: a comprehensive tool for the identification of iron genes and iron gene neighborhoods in genomes and metagenome assemblies

Iron is a micronutrient for nearly all life on Earth. It can be used as an electron donor and electron acceptor by iron-oxidizing and iron-reducing microorganisms and is used in a variety of biological processes, including photosynthesis and respiration. While it is the fourth most abundant metal in the Earth’s crust, iron is often limiting for growth in oxic environments because it is readily oxidized and precipitated. Much of our understanding of how microorganisms compete for and utilize iron is based on laboratory experiments. However, the advent of next-generation sequencing and surge in publicly available sequence data has made it possible to probe the structure and function of microbial communities in the environment. To bridge the gap between our understanding of iron acquisition, iron redox cycling, iron storage, and magnetosome formation in model microorganisms and the plethora of sequence data available from environmental studies, we have created a comprehensive database of hidden Markov models (HMMs) based on genes related to iron acquisition, storage, and reduction/oxidation in Bacteria and Archaea. Along with this database, we present FeGenie, a bioinformatics tool that accepts genome and metagenome assemblies as input and uses our comprehensive HMM database to annotate provided datasets with respect to iron-related genes and gene neighborhood. An important contribution of this tool is the efficient identification of genes involved in iron oxidation and dissimilatory iron reduction, which have been largely overlooked by standard annotation pipelines. We validated FeGenie against a selected set of 28 isolate genomes and showcase its utility in exploring iron genes present in 27 metagenomes, 4 isolate genomes from human oral biofilms, and 17 genomes from candidate organisms, including members of the candidate phyla radiation. We show that FeGenie accurately identifies iron genes in isolates. Furthermore, analysis of metagenomes using FeGenie demonstrates that the iron gene repertoire and abundance of each environment is correlated with iron richness. While this tool will not replace the reliability of culture-dependent analyses of microbial physiology, it provides reliable predictions derived from the most up-to-date genetic markers. FeGenie’s database will be maintained and continually updated as new genes are discovered.

59 BASIC BIOLOGICAL SCIENCES↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

GapMind: Automated Annotation of Amino Acid Biosynthesis

ABSTRACT GapMind is a Web-based tool for annotating amino acid biosynthesis in bacteria and archaea ( http://papers.genomics.lbl.gov/gaps ). GapMind incorporates many variant pathways and 130 different reactions, and it analyzes a genome in just 15 s. To avoid error-prone transitive annotations, GapMind relies primarily on a database of experimentally characterized proteins. GapMind correctly handles fusion proteins and split proteins, which often cause errors for best-hit approaches. To improve GapMind’s coverage, we examined genetic data from 35 bacteria that grow in defined media without amino acids, and we filled many gaps in amino acid biosynthesis pathways. For example, we identified additional genes for arginine synthesis with succinylated intermediates in Bacteroides thetaiotaomicron , and we propose that Dyella japonica synthesizes tyrosine from phenylalanine. Nevertheless, for many bacteria and archaea that grow in minimal media, genes for some steps still cannot be identified. To help interpret potential gaps, GapMind checks if they match known gaps in related microbes that can grow in minimal media. GapMind should aid the identification of microbial growth requirements. IMPORTANCE Many microbes can make all of the amino acids (the building blocks of proteins). In principle, we should be able to predict which amino acids a microbe can make, and which it requires as nutrients, by checking its genome sequence for all of the necessary genes. However, in practice, it is difficult to check for all of the alternative pathways. Furthermore, new pathways and enzymes are still being discovered. We built an automated tool, GapMind, to annotate amino acid biosynthesis in bacterial and archaeal genomes. We used GapMind to list gaps: cases where a microbe makes an amino acid but a complete pathway cannot be identified in its genome. We used these gaps, together with data from mutants, to identify new pathways and enzymes. However, for most bacteria and archaea, we still do not know how they can make all of the amino acids.

59 BASIC BIOLOGICAL SCIENCES↗

A large sequenced mutant library – valuable reverse genetic resource that covers 98% of sorghum genes

SUMMARY Mutant populations are crucial for functional genomics and discovering novel traits for crop breeding. Sorghum , a drought and heat‐tolerant C4 species, requires a vast, large‐scale, annotated, and sequenced mutant resource to enhance crop improvement through functional genomics research. Here, we report a sorghum large‐scale sequenced mutant population with 9.5 million ethyl methane sulfonate (EMS)‐induced mutations that covered 98% of sorghum's annotated genes using inbred line BTx623. Remarkably, a total of 610 320 mutations within the promoter and enhancer regions of 18 000 and 11 790 genes, respectively, can be leveraged for novel research of cis ‐regulatory elements. A comparison of the distribution of mutations in the large‐scale mutant library and sorghum association panel (SAP) provides insights into the influence of selection. EMS‐induced mutations appeared to be random across different regions of the genome without significant enrichment in different sections of a gene, including the 5′ UTR, gene body, and 3′‐UTR. In contrast, there were low variation density in the coding and UTR regions in the SAP. Based on the K a / K s value, the mutant library (~1) experienced little selection, unlike the SAP (0.40), which has been strongly selected through breeding. All mutation data are publicly searchable through SorbMutDB ( https://www.depts.ttu.edu/igcast/sorbmutdb.php ) and SorghumBase ( https://sorghumbase.org/ ). This current large‐scale sequence‐indexed sorghum mutant population is a crucial resource that enriched the sorghum gene pool with novel diversity and a highly valuable tool for the Poaceae family, that will advance plant biology research and crop breeding.

59 BASIC BIOLOGICAL SCIENCES↗

Combining GWAS and population genomic analyses to characterize coevolution in a legume‐rhizobia symbiosis

Abstract The mutualism between legumes and rhizobia is clearly the product of past coevolution. However, the nature of ongoing evolution between these partners is less clear. To characterize the nature of recent coevolution between legumes and rhizobia, we used population genomic analysis to characterize selection on functionally annotated symbiosis genes as well as on symbiosis gene candidates identified through a two‐species association analysis. For the association analysis, we inoculated each of 202 accessions of the legume host Medicago truncatula with a community of 88 Sinorhizobia (Ensifer) meliloti strains. Multistrain inoculation, which better reflects the ecological reality of rhizobial selection in nature than single‐strain inoculation, allows strains to compete for nodulation opportunities and host resources and for hosts to preferentially form nodules and provide resources to some strains. We found extensive host by symbiont, that is, genotype‐by‐genotype, effects on rhizobial fitness and some annotated rhizobial genes bear signatures of recent positive selection. However, neither genes responsible for this variation nor annotated host symbiosis genes are enriched for signatures of either positive or balancing selection. This result suggests that stabilizing selection dominates selection acting on symbiotic traits and that variation in these traits is under mutation‐selection balance. Consistent with the lack of positive selection acting on host genes, we found that among‐host variation in growth was similar whether plants were grown with rhizobia or N‐fertilizer, suggesting that the symbiosis may not be a major driver of variation in plant growth in multistrain contexts.

59 BASIC BIOLOGICAL SCIENCES↗

High-Performance Deep Learning Toolbox for Genome-Scale Prediction of Protein Structure and Function

Computational biology is one of many scientific disciplines ripe for innovation and acceleration with the advent of high-performance computing (HPC). In recent years, the field of machine learning has also seen significant benefits from adopting HPC practices. In this work, we present a novel HPC pipeline that incorporates various machine-learning approaches for structure-based functional annotation of proteins on the scale of whole genomes. Our pipeline makes extensive use of deep learning and provides computational insights into best practices for training advanced deep-learning models for high-throughput data such as proteomics data. We showcase methodologies our pipeline currently supports and detail future tasks for our pipeline to envelop, including large-scale sequence comparison using SAdLSA and prediction of protein tertiary structures using AlphaFold2.

Gao, Mu↗

DOE JGI Metagenome Workflow

The DOE Joint Genome Institute (JGI) Metagenome Workflow performs metagenome data processing, including assembly; structural, functional, and taxonomic annotation; and binning of metagenomic data sets that are subsequently included into the Integrated Microbial Genomes and Microbiomes (IMG/M) (I.-M. A. Chen, K. Chu, K. Palaniappan, A. Ratner, et al., Nucleic Acids Res, 49:D751–D763, 2021, https://doi.org/10.1093/nar/gkaa939) comparative analysis system and provided for download via the JGI data portal (https://genome.jgi.doe.gov/portal/). This workflow scales to run on thousands of metagenome samples per year, which can vary by the complexity of microbial communities and sequencing depth. Here, we describe the different tools, databases, and parameters used at different steps of the workflow to help with the interpretation of metagenome data available in IMG and to enable researchers to apply this workflow to their own data. We use 20 publicly available sediment metagenomes to illustrate the computing requirements for the different steps and highlight the typical results of data processing. The workflow modules for read filtering and metagenome assembly are available as a workflow description language (WDL) file (https://code.jgi.doe.gov/BFoster/jgi_meta_wdl). The workflow modules for annotation and binning are provided as a service to the user community at https://img.jgi.doe.gov/submit and require filling out the project and associated metadata descriptions in the Genomes OnLine Database (GOLD) (S. Mukherjee, D. Stamatis, J. Bertsch, G. Ovchinnikova, et al., Nucleic Acids Res, 49:D723–D733, 2021, https://doi.org/10.1093/nar/gkaa983).

59 BASIC BIOLOGICAL SCIENCES↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

Predicting variable gene content in Escherichia coli using conserved genes

Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%–90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943–0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins” was accurately predicted (F1 = 0.902 [0.898–0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876–0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

Detecting operons in bacterial genomes via visual representation learning

Contiguous genes in prokaryotes are often arranged into operons. Detecting operons plays a critical role in inferring gene functionality and regulatory networks. Human experts annotate operons by visually inspecting gene neighborhoods across pileups of related genomes. These visual representations capture the inter-genic distance, strand direction, gene size, functional relatedness, and gene neighborhood conservation, which are the most prominent operon features mentioned in the literature. By studying these features, an expert can then decide whether a genomic region is part of an operon. We propose a deep learning based method named Operon Hunter that uses visual representations of genomic fragments to make operon predictions. Using transfer learning and data augmentation techniques facilitates leveraging the powerful neural networks trained on image datasets by re-training them on a more limited dataset of extensively validated operons. Our method outperforms the previously reported state-of-the-art tools, especially when it comes to predicting full operons and their boundaries accurately. Furthermore, our approach makes it possible to visually identify the features influencing the network’s decisions to be subsequently cross-checked by human experts.

59 BASIC BIOLOGICAL SCIENCES↗

KBase Narrative - Porphyromonadaceae sp. W3.11 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗