Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome assembly”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring

GAL08 are bacteria belonging to an uncultivated phylogenetic cluster within the phylum Acidobacteria . We detected a natural population of the GAL08 clade in sediment from a pH-neutral hot spring located in British Columbia, Canada. To shed light on the abundance and genomic potential of this clade, we collected and analyzed hot spring sediment samples over a temperature range of 24.2–79.8°C. Illumina sequencing of 16S rRNA gene amplicons and qPCR using a primer set developed specifically to detect the GAL08 16S rRNA gene revealed that absolute and relative abundances of GAL08 peaked at 65°C along three temperature gradients. Analysis of sediment collected over multiple years and locations revealed that the GAL08 group was consistently a dominant clade, comprising up to 29.2% of the microbial community based on relative read abundance and up to 4.7 × 10 5 16S rRNA gene copy numbers per gram of sediment based on qPCR. Using a medium quality threshold, 25 single amplified genomes (SAGs) representing these bacteria were generated from samples taken at 65 and 77°C, and seven metagenome-assembled genomes (MAGs) were reconstructed from samples collected at 45–77°C. Based on average nucleotide identity (ANI), these SAGs and MAGs represented three separate species, with an estimated average genome size of 3.17 Mb and GC content of 62.8%. Phylogenetic trees constructed from 16S rRNA gene sequences and a set of 56 concatenated phylogenetic marker genes both placed the three GAL08 bacteria as a distinct subgroup of the phylum Acidobacteria , representing a candidate order ( Ca. Frugalibacteriales) within the class Blastocatellia. Metabolic reconstructions from genome data predicted a heterotrophic metabolism, with potential capability for aerobic respiration, as well as incomplete denitrification and fermentation. In laboratory cultivation efforts, GAL08 counts based on qPCR declined rapidly under atmospheric levels of oxygen but increased slightly at 1% (v/v) O 2 , suggesting a microaerophilic lifestyle.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and Transcriptomic Evidence Supports Methane Metabolism in Archaeoglobi

Euryarchaeal lineages have been believed to have a methanogenic last common ancestor. However, members of euryarchaeal Archaeoglobi have long been considered nonmethanogenic and their evolutionary history remains elusive. Here, three high-quality metagenomic-assembled genomes (MAGs) retrieved from high-temperature oil reservoir and hot springs, together with three newly assembled Archaeoglobi MAGs from previously reported hot spring metagenomes, are demonstrated to represent a novel genus of Archaeoglobaceae, “Candidatus Methanomixophus.” All “Ca. Methanomixophus” MAGs encode an M methyltransferase (MTR) complex and a traditional type of methyl-coenzyme M reductase (MCR) complex, which is different from the divergent MCR complexes found in “Ca. Polytropus marinifundus.” In addition, “Ca. Methanomixophus dualitatem” MAGs preserve the genomic capacity for dissimilatory sulfate reduction. Comparative phylogenetic analysis supports a laterally transferred origin for an MCR complex and vertical heritage of the MTR complex in this lineage. Metatranscriptomic analysis revealed concomitant in situ activity of hydrogen-dependent methylotrophic methanogenesis and heterotrophic fermentation within populations of “Ca. Methanomixophus hydrogenotrophicum” in a high-temperature oil reservoir.

59 BASIC BIOLOGICAL SCIENCES↗

Novel Toxin Biosynthetic Gene Cluster in Harmful Algal Bloom-Causing Heteroscytonema crispum : Insights into the Origins of Paralytic Shellfish Toxins

Caused by both eukaryotic dinoflagellates and prokaryotic cyanobacteria, harmful algal blooms are events of severe ecological, economic, and public health consequence, and their incidence has become more common of late. Despite coordinated research efforts to identify and characterize the genomes of harmful algal bloom-causing organisms, the genomic basis and evolutionary origins of paralytic shellfish toxins produced by harmful algal blooms remain at best incomplete. The paralytic shellfish toxin saxitoxin has an especially complex genomic architecture and enigmatic phylogenetic distribution, spanning dinoflagellates and multiple cyanobacterial genera. Using filtration and extraction techniques to target the desired cyanobacteria from nonaxenic culture, coupled with a combination of short- and long-read sequencing, we generated a reference-quality hybrid genome assembly for Heteroscytonema crispum UTEX LB 1556, a freshwater, paralytic shellfish toxin-producing cyanobacterium thought to have the largest known genome in its phylum. We report a complete, novel biosynthetic gene cluster for the paralytic shellfish toxin saxitoxin. Leveraging this biosynthetic gene cluster, we find support for the hypothesis that paralytic shellfish toxin production has appeared in divergent Cyanobacteria lineages through widespread and repeated horizontal gene transfer. This work demonstrates the utility of long-read sequencing and metagenomic assembly toward advancing our understanding of paralytic shellfish toxin biosynthetic gene cluster diversity and suggests a mechanism for the origin of paralytic shellfish toxin biosynthetic genes.

59 BASIC BIOLOGICAL SCIENCES↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

A phylogenetically novel cyanobacterium most closely related to Gloeobacter

Clues to the evolutionary steps producing innovations in oxygenic photosynthesis may be preserved in the genomes of organisms phylogenetically placed between non-photosynthetic Vampirovibrionia (formerly Melainabacteria) and the thylakoid-containing Cyanobacteria. However, only two species with published genomes are known to occupy this phylogenetic space, both within the genus Gloeobacter. Here, we describe nearly complete, metagenome-assembled genomes (MAGs) of an uncultured organism phylogenetically placed near Gloeobacter, for which we propose the name Candidatus Aurora vandensis {Au’ro.ra. L. fem. n. aurora, the goddess of the dawn in Roman mythology; van.de’nsis. N.L. fem. adj. vandensis of Lake Vanda, Antarctica}. The MAG of A. vandensis contains homologs of most genes necessary for oxygenic photosynthesis including key reaction center proteins. Many accessory subunits associated with the photosystems in other species either are missing from the MAG or are poorly conserved. The MAG also lacks homologs of genes associated with the pigments phycocyanoerethrin, phycoeretherin and several structural parts of the phycobilisome. Additional characterization of this organism is expected to inform models of the evolution of oxygenic photosynthesis.

59 BASIC BIOLOGICAL SCIENCES↗

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES↗

Methylotrophy in the Mire: direct and indirect routes for methane production in thawing permafrost

While wetlands are major sources of biogenic methane (CH 4 ), our understanding of resident microbial metabolism is incomplete, which compromises the prediction of CH 4 emissions under ongoing climate change. Here, we employed genome-resolved multi-omics to expand our understanding of methanogenesis in the thawing permafrost peatland of Stordalen Mire in Arctic Sweden. In quadrupling the genomic representation of the site’s methanogens and examining their encoded metabolism, we revealed that nearly 20% of the metagenome-assembled genomes (MAGs) encoded the potential for methylotrophic methanogenesis. Further, 27% of the transcriptionally active methanogens expressed methylotrophic genes; for Methanosarcinales and Methanobacteriales MAGs, these data indicated the use of methylated oxygen compounds (e.g., methanol), while for Methanomassiliicoccales, they primarily implicated methyl sulfides and methylamines. In addition to methanogenic methylotrophy, >1,700 bacterial MAGs across 19 phyla encoded anaerobic methylotrophic potential, with expression across 12 phyla. Metabolomic analyses revealed the presence of diverse methylated compounds in the Mire, including some known methylotrophic substrates. Active methylotrophy was observed across all stages of a permafrost thaw gradient in Stordalen, with the most frozen non-methanogenic palsa found to host bacterial methylotrophy and the partially thawed bog and fully thawed fen seen to house both methanogenic and bacterial methylotrophic activities. Methanogenesis across increasing permafrost thaw is thus revised from the sole dominance of hydrogenotrophic production and the appearance of acetoclastic at full thaw to consider the co-occurrence of methylotrophy throughout. Collectively, these findings indicate that methanogenic and bacterial methylotrophy may be an important and previously underappreciated component of carbon cycling and emissions in these rapidly changing wetland habitats.

59 BASIC BIOLOGICAL SCIENCES↗

Genome Extraction from Shotgun Metagenome Sequence Data

Uncultivated Bacteria and Archaea comprise the vast majority of species on Earth, but obtaining their genomes directly from the environment, using shotgun sequencing, has only recently become possible. To realize the hope of capturing Earth’s microbial genetic complement, technologies that accelerate recovery of high-quality genomes are necessary. We present a series of analysis steps and data products for the extraction of high quality metagenome-assembled genomes (MAGs) from microbiomes using the U.S. Department of Energy Systems Biology Knowledgebase (KBase) platform (http://www.kbase.us/). In KBase, the process is end-to-end, allowing a user to go from the initial sequencing reads all the way through to MAG genomes, which can then be analyzed with other KBase capabilities such as phylogenetic placement, functional assignment, metabolic modeling, pangenome functional profiling, RNA-Seq, and others. While portions of such capabilities are individually available from other resources, the combination of the intuitive usability, data interoperability, and integration of tools in a freely available compute resource makes KBase a uniquely powerful platform for obtaining MAGs from microbiomes. While this workflow offers tools for each of the key steps in the genome extraction process, it also provides a scaffold that can be easily extended, with additional MAG recovery and analysis tools, via the KBase SDK (Software Development Kit).

Chivian, Dylan↗

Moab Desert Crust - Sample 4E

Uncultivated Bacteria and Archaea comprise the vast majority of species on Earth, but obtaining their genomes directly from the environment, using shotgun sequencing, has only recently become possible. To realize the hope of capturing Earth’s microbial genetic complement, technologies that accelerate recovery of high-quality genomes are necessary. We present a series of analysis steps and data products for the extraction of high quality metagenome-assembled genomes (MAGs) from microbiomes using the U.S. Department of Energy Systems Biology Knowledgebase (KBase) platform (http://www.kbase.us/). In KBase, the process is end-to-end, allowing a user to go from the initial sequencing reads all the way through to MAG genomes, which can then be analyzed with other KBase capabilities such as phylogenetic placement, functional assignment, metabolic modeling, pangenome functional profiling, RNA-Seq, and others. While portions of such capabilities are individually available from other resources, the combination of the intuitive usability, data interoperability, and integration of tools in a freely available compute resource makes KBase a uniquely powerful platform for obtaining MAGs from microbiomes. While this workflow offers tools for each of the key steps in the genome extraction process, it also provides a scaffold that can be easily extended, with additional MAG recovery and analysis tools, via the KBase SDK (Software Development Kit).

Chivian, Dylan↗

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27↗

Seagrass genomes reveal ancient polyploidy and adaptations to the marine environment

Here, we present chromosome-level genome assemblies from representative species of three independently evolved seagrass lineages: Posidonia oceanica, Cymodocea nodosa, Thalassia testudinum and Zostera marina. We also include a draft genome of Potamogeton acutifolius, belonging to a freshwater sister lineage to Zosteraceae. All seagrass species share an ancient whole-genome triplication, while additional whole-genome duplications were uncovered for C. nodosa, Z. marina and P. acutifolius. Comparative analysis of selected gene families suggests that the transition from submerged-freshwater to submerged-marine environments mainly involved fine-tuning of multiple processes (such as osmoregulation, salinity, light capture, carbon acquisition and temperature) that all had to happen in parallel, probably explaining why adaptation to a marine lifestyle has been exceedingly rare. Major gene losses related to stomata, volatiles, defence and lignification are probably a consequence of the return to the sea rather than the cause of it. These new genomes will accelerate functional studies and solutions, as continuing losses of the savannahs of the sea are of major concern in times of climate change and loss of biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome sequences of strains CBS6241 and CBS6242 of the basidiomycetous yeast Filobasidium floriforme

The Tremellomycetes are a species-rich group within the basidiomycete fungi; however, most analyses of this group to date have focused on pathogenic Cryptococcus species within the order Tremellales. Recent genome-assisted studies of other Tremellomycetes have identified interesting features with respect to biotechnological applications as well as the evolution of genes involved in mating and sexual development. Here, we report genome sequences of two strains of Filobasidium floriforme, a species from the order Filobasidiales, which branches basally to the Tremellales, Trichosporonales, and Holtermanniales. The assembled genomes of strains CBS6241 and CBS6242 are 27.4 Mb and 26.4 Mb in size, respectively, with 8314 and 7695 predicted protein-coding genes. Overall sequence identity at nucleic acid level between the strains is 97%. Among the predicted genes are pheromone precursor and pheromone receptor genes as well as two genes encoding homedomain (HD) transcription factors, which are predicted to be part of the mating type (MAT) locus. Sequence analysis indicates that CBS6241 and CBS6242 carry different alleles for both the pheromone/receptor genes as well as the HD transcription factors. Orthology inference identified 1482 orthogroups exclusively found in F. floriforme, some of which were involved in carbohydrate transport and metabolism. Subsequent CAZyme repertoire characterization identified 267 and 247 enzymes for CBS6241 and CBS6242, respectively, the second highest number of CAZymes among the analyzed Tremellomycete species. In addition, F. floriforme contains five CAZymes absent in other species and several plant-cell-wall degrading CAZymes with the highest copy number in Tremellomycota, indicating the biotechnological potential of this species.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic profiles of four novel cyanobacteria MAGs from Lake Vanda, Antarctica: insights into photosynthesis, cold tolerance, and the circadian clock

Cyanobacteria in polar environments face environmental challenges, including cold temperatures and extreme light seasonality with small diurnal variation, which has implications for polar circadian clocks. However, polar cyanobacteria remain underrepresented in available genomic data, and there are limited opportunities to study their genetic adaptations to these challenges. This paper presents four new Antarctic cyanobacteria metagenome-assembled genomes (MAGs) from microbial mats in Lake Vanda in the McMurdo Dry Valleys in Antarctica. The four MAGs were classified as Leptolyngbya sp. BulkMat.35, Pseudanabaenaceae cyanobacterium MP8IB2.15, Microcoleus sp. MP8IB2.171, and Leptolyngbyaceae cyanobacterium MP9P1.79. The MAGs contain 2.76 Mbp – 6.07 Mbp, and the bin completion ranges from 74.2–92.57%. Furthermore, the four cyanobacteria MAGs have average nucleotide identities (ANIs) under 90% with each other and under 77% with six existing polar cyanobacteria MAGs and genomes. This suggests that they are novel cyanobacteria and demonstrates that polar cyanobacteria genomes are underrepresented in reference databases and there is continued need for genome sequencing of polar cyanobacteria. Analyses of the four novel and six existing polar cyanobacteria MAGs and genomes demonstrate they have genes coding for various cold tolerance mechanisms and most standard circadian rhythm genes with the Leptolyngbya sp. BulkMat.35 and Leptolyngbyaceae cyanobacterium MP9P1.79 contained kaiB3, a divergent homolog of kaiB.

59 BASIC BIOLOGICAL SCIENCES↗

An essential role for tungsten in the ecology and evolution of a previously uncultivated lineage of anaerobic, thermophilic Archaea

Trace metals have been an important ingredient for life throughout Earth’s history. Here, we describe the genome-guided cultivation of a member of the elusive archaeal lineage Caldarchaeales (syn. Aigarchaeota), Wolframiiraptor gerlachensis, and its growth dependence on tungsten. A metagenome-assembled genome (MAG) of W. gerlachensis encodes putative tungsten membrane transport systems, as well as pathways for anaerobic oxidation of sugars probably mediated by tungsten-dependent ferredoxin oxidoreductases that are expressed during growth. Catalyzed reporter deposition-fluorescence in-situ hybridization (CARD-FISH) and nanoscale secondary ion mass spectrometry (nanoSIMS) show that W. gerlachensis preferentially assimilates xylose. Phylogenetic analyses of 78 high-quality Wolframiiraptoraceae MAGs from terrestrial and marine hydrothermal systems suggest that tungsten-associated enzymes were present in the last common ancestor of extant Wolframiiraptoraceae. Our observations imply a crucial role for tungsten-dependent metabolism in the origin and evolution of this lineage, and hint at a relic metabolic dependence on this trace metal in early anaerobic thermophiles.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of muntjac deer genome and chromatin architecture reveals rapid karyotype evolution

Abstract Closely related muntjac deer show striking karyotype differences. Here we describe chromosome-scale genome assemblies for Chinese and Indian muntjacs, Muntiacus reevesi (2 n = 46) and Muntiacus muntjak vaginalis (2 n = 6/7), and analyze their evolution and architecture. The genomes show extensive collinearity with each other and with other deer and cattle. We identified numerous fusion events unique to and shared by muntjacs relative to the cervid ancestor, confirming many cytogenetic observations with genome sequence. One of these M. muntjak fusions reversed an earlier fission in the cervid lineage. Comparative Hi-C analysis showed that the chromosome fusions on the M. muntjak lineage altered long-range, three-dimensional chromosome organization relative to M. reevesi in interphase nuclei including A/B compartment structure. This reshaping of multi-megabase contacts occurred without notable change in local chromatin compaction, even near fusion sites. A few genes involved in chromosome maintenance show evidence for rapid evolution, possibly associated with the dramatic changes in karyotype.

59 BASIC BIOLOGICAL SCIENCES↗

Using online tools at the Bovine Genome Database to manually annotate genes in the new reference genome

With the availability of a new highly contiguous Bos taurus reference genome assembly (ARS-UCD1.2), it is the opportune time to upgrade the bovine gene set by seeking input from researchers. Furthermore, advances in graphical genome annotation tools now make it possible for researchers to leverage sequence data generated with the latest technologies to collaboratively curate genes. For many years the Bovine Genome Database (BGD) has provided tools such as the APOLLO genome annotation editor to support manual bovine gene curation. The goal of this paper is to explain the reasoning behind the decisions made in the manual gene curation process while providing examples using the existing BGD tools. We will describe the sources of gene annotation evidence provided at the BGD, including RNAseq and Iso-Seq data. We will also explain how to interpret various data visualizations when curating gene models, and will demonstrate the value of manual gene annotation. The process described here can be applied to manual gene curation for other species with similar tools. With a better understanding of manual gene annotation, researchers will be encouraged to edit gene models and contribute to the enhancement of livestock gene sets.

59 BASIC BIOLOGICAL SCIENCES↗