Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Elemental profiling and genome-wide association studies reveal genomic variants modulating ionomic composition in Populus trichocarpa leaves

The ionome represents elemental composition in plant tissues and can be an indicator of nutrient status as well as overall plant performance. Thus, identifying genetic determinants governing elemental uptake and storage is an important goal for breeding and engineering biomass feedstocks with improved performance. In this study, we coupled high-throughput ionome characterization of leaf tissues with high-resolution genome-wide association studies (GWAS) to uncover genetic loci that modulate ionomic composition in leaves of poplar ( Populus trichocarpa ). Significant agreement was observed across the three ionomic profiling platforms tested: inductively coupled plasma-mass spectrometry (ICP-MS), neutron activation analysis (NAA) and laser-induced breakdown spectroscopy (LIBS). Relative quantification of 20 elements using ICP-MS across a population of 584 genotypes, revealed larger variation in micro-nutrients and trace elements content than for macro-nutrients across genotypes. The GWAS performed using a set of high-density (>8.2 million) single nucleotide polymorphisms, identified over 600 loci significantly associated with variations in these mineral elements, pointing to numerous uncharacterized candidate genes. A significant enrichment for genes related to ion homeostasis and transport was observed, including several members of the cation-proton antiporters (CPA) family and MATE efflux transporters, previously reported to be critical for plant growth and fitness in other species. Our results also included a polymorphic copy of the high-affinity molybdenum transporter MOT1 found directly associated to molybdenum content. For the first time in a perennial plant, our results provide evidence of genetic control of mineral content in a model tree species.

59 BASIC BIOLOGICAL SCIENCES↗

A method for achieving complete microbial genomes and improving bins from metagenomics data

Metagenomics facilitates the study of the genetic information from uncultured microbes and complex microbial communities. Assembling complete genomes from metagenomics data is difficult because most samples have high organismal complexity and strain diversity. Some studies have attempted to extract complete bacterial, archaeal, and viral genomes and often focus on species with circular genomes so they can help confirm completeness with circularity. However, less than 100 circularized bacterial and archaeal genomes have been assembled and published from metagenomics data despite the thousands of datasets that are available. Circularized genomes are important for (1) building a reference collection as scaffolds for future assemblies, (2) providing complete gene content of a genome, (3) confirming little or no contamination of a genome, (4) studying the genomic context and synteny of genes, and (5) linking protein coding genes to ribosomal RNA genes to aid metabolic inference in 16S rRNA gene sequencing studies. We developed a semi-automated method called Jorg to help circularize small bacterial, archaeal, and viral genomes using iterative assembly, binning, and read mapping. In addition, this method exposes potential misassemblies from k-mer based assemblies. We chose species of the Candidate Phyla Radiation (CPR) to focus our initial efforts because they have small genomes and are only known to have one ribosomal RNA operon. In addition to 34 circular CPR genomes, we present one circular Margulisbacteria genome, one circular Chloroflexi genome, and two circular megaphage genomes from 19 public and published datasets. We demonstrate findings that would likely be difficult without circularizing genomes, including that ribosomal genes are likely not operonic in the majority of CPR, and that some CPR harbor diverged forms of RNase P RNA. Code and a tutorial for this method is available at https://github.com/lmlui/Jorg and is available on the DOE Systems Biology KnowledgeBase as a beta app.

59 BASIC BIOLOGICAL SCIENCES↗

A Genome-Based Model to Predict the Virulence of Pseudomonas aeruginosa Isolates

ABSTRACT: Variation in the genome of Pseudomonas aeruginosa , an important pathogen, can have dramatic impacts on the bacterium’s ability to cause disease. We therefore asked whether it was possible to predict the virulence of P. aeruginosa isolates based on their genomic content. We applied a machine learning approach to a genetically and phenotypically diverse collection of 115 clinical P. aeruginosa isolates using genomic information and corresponding virulence phenotypes in a mouse model of bacteremia. We defined the accessory genome of these isolates through the presence or absence of accessory genomic elements (AGEs), sequences present in some strains but not others. Machine learning models trained using AGEs were predictive of virulence, with a mean nested cross-validation accuracy of 75% using the random forest algorithm. However, individual AGEs did not have a large influence on the algorithm’s performance, suggesting instead that virulence predictions are derived from a diffuse genomic signature. These results were validated with an independent test set of 25 P. aeruginosa isolates whose virulence was predicted with 72% accuracy. Machine learning models trained using core genome single-nucleotide variants and whole-genome k-mers also predicted virulence. Our findings are a proof of concept for the use of bacterial genomes to predict pathogenicity in P. aeruginosa and highlight the potential of this approach for predicting patient outcomes. IMPORTANCE Pseudomonas aeruginosa is a clinically important Gram-negative opportunistic pathogen. P. aeruginosa shows a large degree of genomic heterogeneity both through variation in sequences found throughout the species (core genome) and through the presence or absence of sequences in different isolates (accessory genome). P. aeruginosa isolates also differ markedly in their ability to cause disease. In this study, we used machine learning to predict the virulence level of P. aeruginosa isolates in a mouse bacteremia model based on genomic content. We show that both the accessory and core genomes are predictive of virulence. This study provides a machine learning framework to investigate relationships between bacterial genomes and complex phenotypes such as virulence.

59 BASIC BIOLOGICAL SCIENCES↗

METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks

Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial communities. The genome-scale workflow includes annotation of microbial genomes, motif validation of biochemically validated conserved protein residues, metabolic pathway analyses, and calculation of contributions to individual biogeochemical transformations and cycles. The community-scale workflow supplements genome-scale analyses with determination of genome abundance in the microbiome, potential microbial metabolic handoffs and metabolite exchange, reconstruction of functional networks, and determination of microbial contributions to biogeochemical cycles. METABOLIC can take input genomes from isolates, metagenome-assembled genomes, or single-cell genomes. Results are presented in the form of tables for metabolism and a variety of visualizations including biogeochemical cycling potential, representation of sequential metabolic transformations, community-scale microbial functional networks using a newly defined metric “MW-score” (metabolic weight score), and metabolic Sankey diagrams. METABOLIC takes ~ 3 h with 40 CPU threads to process ~ 100 genomes and corresponding metagenomic reads within which the most compute-demanding part of hmmsearch takes ~ 45 min, while it takes ~ 5 h to complete hmmsearch for ~ 3600 genomes. Tests of accuracy, robustness, and consistency suggest METABOLIC provides better performance compared to other software and online servers. To highlight the utility and versatility of METABOLIC, we demonstrate its capabilities on diverse metagenomic datasets from the marine subsurface, terrestrial subsurface, meadow soil, deep sea, freshwater lakes, wastewater, and the human gut. Conclusion METABOLIC enables the consistent and reproducible study of microbial community ecology and biogeochemistry using a foundation of genome-informed microbial metabolism, and will advance the integration of uncultivated organisms into metabolic and biogeochemical models. METABOLIC is written in Perl and R and is freely available under GPLv3 at https://github.com/AnantharamanLab/METABOLIC.

59 BASIC BIOLOGICAL SCIENCES↗

Accurate and complete genomes from metagenomes

Genomes are an integral component of the biological information about an organism; thus, the more complete the genome, the more informative it is. Historically, bacterial and archaeal genomes were reconstructed from pure (monoclonal) cultures, and the first reported sequences were manually curated to completion. However, the bottleneck imposed by the requirement for isolates precluded genomic insights for the vast majority of microbial life. Shotgun sequencing of microbial communities, referred to initially as community genomics and subsequently as genome-resolved metagenomics, can circumvent this limitation by obtaining metagenome-assembled genomes (MAGs); but gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of these genomes. Here, we discuss genome curation to improve and, in some cases, achieve complete (circularized, no gaps) MAGs (CMAGs). To date, few CMAGs have been generated, although notably some are from very complex systems such as soil and sediment. Through analysis of about 7000 published complete bacterial isolate genomes, we verify the value of cumulative GC skew in combination with other metrics to establish bacterial genome sequence accuracy. The analysis of cumulative GC skew identified potential misassemblies in some reference genomes of isolated bacteria and the repeat sequences that likely gave rise to them. We discuss methods that could be implemented in bioinformatic approaches for curation to ensure that metabolic and evolutionary analyses can be based on very high-quality genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Three founding ancestral genomes involved in the origin of sugarcane

Modern sugarcane cultivars (Saccharum spp.) are high polyploids, aneuploids (2n = ~12 = ~120) derived from interspecific hybridizations between the domesticated sweet species Saccharum officinarum and the wild species S. spontaneum. To analyse the architecture and origin of such a complex genome, we analysed the sequences of all 12 hom(oe)ologous haplotypes (BAC clones) from two distinct genomic regions of a typical modern cultivar, as well as the corresponding sequence in Miscanthus sinense and Sorghum bicolor, and monitored their distribution among representatives of the Saccharum genus. The diversity observed among haplotypes suggested the existence of three founding genomes (A, B, C) in modern cultivars, which diverged between 0.8 and 1.3 Mya. Two genomes (A, B) were contributed by S. officinarum; these were also found in its wild presumed ancestor S. robustum, and one genome (C) was contributed by S. spontaneum. These results suggest that S. officinarum and S. robustum are derived from interspecific hybridization between two unknown ancestors (A and B genomes). The A genome contributed most haplotypes (nine or ten) while the B and C genomes contributed one or two haplotypes in the regions analysed of this typical modern cultivar. Interspecific hybridizations likely involved accessions or gametes with distinct ploidy levels and/or were followed by a series of backcrosses with the A genome. The three founding genomes were found in all S. barberi, S. sinense and modern cultivars analysed. None of the analysed accessions contained only the A genome or the B genome, suggesting that representatives of these founding genomes remain to be discovered. This evolutionary model, which combines interspecificity and high polyploidy, can explain the variable chromosome pairing affinity observed in Saccharum. It represents a major revision of the understanding of Saccharum diversity.

54 ENVIRONMENTAL SCIENCES↗

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

Author Correction: Genome-guided isolation of the hyperthermophilic aerobe Fervidibacter sacchari reveals conserved polysaccharide metabolism in the Armatimonadota

Correction to: Nature Communicationshttps://doi.org/10.1038/s41467-024-53784-3, published online 4 November 2024 In the version of this article initially published, Table 1 did not include the properties of the taxa being proposed or refer directly to another location in the main manuscript describing the properties. As such, the original manuscript did not comply with Rule 27 (2)(c) of the ICNP. Also, Table 1 listed the order Fervidibacterales as the nomenclatural type for the class Fervidibacteria, which violates latest emended version of Rule 15 stating that the nomenclatural type for a class must be a genus. Below we provide a modification of Table 1 containing protologues with these errors corrected. We have also changed the order of the taxa in the table to meet the most common ordering. (Table presented.) Taxon names proposed under the ICNP Proposed taxon Etymology Description Genus Fervidibacter Fer.vi.di.bac’ter. L. masc. adj. fervidus, hot, steaming; N.L. masc. n. bacter, a rod; N.L. masc. n. Fervidibacter, a hot rod Thermophilic or hyperthermophilic inhabitants of freshwater thermal environments. All members are likely polysaccharide-degrading chemoheterotrophs with numerous carbohydrate-active enzymes encoded in their genomes. Aerobic, with high-affinity and/or low-affinity terminal oxidases present in the genomes. The oxidative pentose phosphate pathway and the tricarboxylic acid cycle are complete in genomes belonging to the genus. Gram-stain-negative and diderm cell envelope structure. Ovoid- to rod-shaped morphology. Spores are not formed. The genus is a distinct phylogenetic lineage in the family Fervidibacteraceae, the order Fervidibacterales, and the class Fervidibacteria in the phylum Armatimonadota. The type species is Fervidibacter sacchariT. Species Fervidibacter sacchari sac’cha.ri. N.L. gen. n. sacchari, of sugar Hyperthermophilic, microaerophilic, facultatively anaerobic, and grows chemoheterotrophically on monosaccharides and polysaccharides. Cells are ovoid- to rod-shaped, Gram-stain negative, and are 0.9–1.3 µm in width and 1.6–3.6 µm in length. Grows between 65 and 87.5 °C and an optimum temperature of 80 °C, and a pH range of 6.5–8.6 with an optimum pH of 7.5. Grows at an optimum O2 concentration of 5–10%. Grows on D-arabinose, D-galactose, D-glucose, D-rhamnose, D-ribose, D-xylose, chondroitin sulfate, colloidal chitin, galactan, gellan gum, guar gum, karaya gum, locust bean gum, xantham gum, xyloglucan, β-glucan, glycogen, starch, AFEX-pretreated corn stover, miscanthus, sugarcane bagasse, acetate and casamino acids. Grows weakly on xyloglucan under fermentation conditions. The major fatty acids (>10%) are C16:0, C18:0 and/or cyclo-C17:0, and iso-C16:0. The major respiratory quinones (>10%) are MK-8 and MK-9. The isolate and genomes of the species have been recovered from geothermal springs in the Great Basin, Nevada, USA. GC content of genomes range between 51–52%. Subunits for both the high-affinity and low-affinity terminal oxidases are encoded in the genomes. Genomes also encode a Group 3d [NiFe] hydrogenase, which produces hydrogen as an electron sink for NAD+ regeneration. The type strain PD1T (= JCM 39283T = DSM 113467T) was isolated from Great Boiling Spring in Nevada, USA. Family Fervidibacteraceae Fer.vi.di.bac.te.ra’ce.ae. N.L. masc. n. Fervidibacter type genus of the family; L. suff. -aceae ending to denote a family; N.L. fem. pl. n. Fervidibacteraceae the family of the genus Fervidibacter Thermophilic or hyperthermophilic inhabitants of freshwater thermal environments. All members are likely polysaccharide-degrading chemoheterotrophs with numerous carbohydrate-active enzymes encoded in their genomes. Aerobic, with high-affinity and/or low-affinity terminal oxidases present in the genomes. The oxidative pentose phosphate pathway and the tricarboxylic acid cycle are complete in genomes belonging to the family. The family is a distinct phylogenetic lineage in the order Fervidibacterales and the class Fervidibacteria in the phylum Armatimonadota. The type genus is Fervidibacter. Order Fervidibacterales Fer.vi.di.bac.te.ra’les. N.L. masc. n. Fervidibacter type genus of the order; L. suff. -ales ending to denote an order; N.L. fem. pl. n. Fervidibacterales the order of the genus Fervidibacter Thermophilic or hyperthermophilic inhabitants of freshwater thermal environments. All members are likely polysaccharide-degrading chemoheterotrophs with numerous carbohydrate-active enzymes encoded in their genomes. Aerobic or strictly anaerobic. Phylogenomic placement of this lineage within the Fervidibacteria and relative evolutionary divergence supports delineation of this lineage as an order within the class Fervidibacteria and phylum Armatimonadota. The type genus is Fervidibacter. Class Fervidibacteria Fer.vi.di.bac.te’ri.a. N.L. masc. n. Fervidibacter type genus of the type order of the class; L. suff. -ia ending to denote a class; N.L. neut. pl. n. Fervidibacteria the class of the order Fervidibacterales Thermophilic or hyperthermophilic inhabitants of freshwater thermal environments. All members are likely polysaccharide-degrading chemoheterotrophs with numerous carbohydrate-active enzymes encoded in their genomes. Aerobic or strictly anaerobic. Phylogenomic placement of this lineage within the Armatimonadota and relative evolutionary divergence supports delineation of this lineage as a class within the Armatimonadota. The type genus is Fervidibacter. The error has not been corrected in the PDF or HTML versions of the Article.

Nou, Nancy O↗

A method for achieving complete microbial genomes and improving bins from metagenomics data

Metagenomics facilitates the study of the genetic information from uncultured microbes and complex microbial communities. Assembling complete microbial genomes ( i.e ., circular with no misassemblies) from metagenomics data is difficult because most samples have high organismal complexity and strain diversity. Less than 100 circularized bacterial and archaeal genomes have been assembled from metagenomics data despite the thousands of datasets that are available. Circularized genomes are important for (1) building a reference collection as scaffolds for future assemblies, (2) providing complete gene content of a genome, (3) confirming little or no contamination of a genome, (4) studying the genomic context and synteny of genes, and (5) linking protein coding genes to ribosomal RNA genes to aid metabolic inference in 16S rRNA gene sequencing studies. We developed a method to achieve circularized genomes using iterative assembly, binning, and read mapping. In addition, this method exposes potential misassemblies from k-mer based assemblies. We chose species of the Candidate Phyla Radiation (CPR) to focus our initial efforts because they have small genomes and are only known to have one ribosomal RNA operon. We present 34 circular CPR genomes, one circular Margulisbacteria genome, and two circular megaphage genomes from 19 public and published datasets. We demonstrate findings that would likely be difficult without circularizing genomes, including that ribosomal genes are likely not operonic in the majority of CPR, and that some CPR harbor diverged forms of RNase P RNA. Code and a tutorial for this method is available at https://github.com/lmlui/Jorg .

Lui, Lauren↗

Budding yeasts in the subphylum Saccharomycotina Genome sequencing and assembly

Eukaryotic life depends on the functional elements encoded by both the nuclear genome and organellar genomes, such as those contained within the mitochondria. The content, size, and structure of the mitochondrial genome varies across organisms with potentially large implications for phenotypic variance and resulting evolutionary trajectories. Among yeasts in the subphylum Saccharomycotina, extensive differences have been observed in various species relative to the model yeast Saccharomyces cerevisiae, but mitochondrial genome sampling across many groups has been scarce, even as hundreds of nuclear genomes have become available. By extracting mitochondrial reads from existing short-read genome sequence datasets, we have greatly expanded both the number of available genomes and the coverage across sparsely sampled clades. Comparison of 353 yeast mitochondrial genomes revealed that, while size and GC content were fairly consistent across species, those in the genera Metschnikowia and Saccharomyces trended larger, while several species in the order Saccharomycetales exhibited lower GC content. Extreme examples for both size and GC content were scattered throughout the subphylum. All mitochondrial genomes shared a core set of protein-coding genes for Complexes III, IV, and V, but they varied in the presence or absence of mitochondrially-encoded canonical Complex I genes. We traced the loss of Complex I genes to a major event in the ancestor of the orders Saccharomycetales and Saccharomycodales, but we also observed several independent losses in the orders Phaffomycetales, Pichiales, and Dipodascales. In contrast to prior hypotheses based on smaller-scale datasets, comparison of evolutionary rates in protein-coding genes showed no bias towards elevated rates among aerobically fermenting (Crabtree/Warburg-positive) yeasts. Mitochondrial introns were widely distributed, but highly enriched in some groups. The majority of mitochondrial introns were poorly conserved within groups, but several were shared within groups, between groups, and even across taxonomic orders, which is consistent with horizontal gene transfer, likely involving homing endonucleases acting as selfish elements. As the number of available fungal nuclear genomes continues to expand, the methods described here to retrieve mitochondrial genome sequences from these datasets will prove invaluable to ensuring that studies of fungal mitochondrial genomes keep pace with their nuclear counterparts.

diversity↗