Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome assembly”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomes and metagenome-assembled genomes from microbial communities in a biological nutrient removal plant operated at Hamptons Road Sanitation District (HRSD) with high and low dissolved oxygen conditions

Aeration is a major cost at biological nutrient removal (BNR) plants. We report on microbial communities in a pilot-scale BNR system before and after a dissolved oxygen transition from 2.5 to 0.2 mg/L implemented over 18 months. Four PacBio metagenomes and 316 metagenome-assembled genomes are announced.

dissolved oxygen↗

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas↗

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Teucrium is well known for making clerodane-type diterpenoids that are produced from the backbone kolavanyl diphosphate. In order to begin to elucidate some of the complex biosynthetic pathways of these medicinal compounds, we identified and functionally characterized several kolavanyl diphosphate synthases from T. chamaedrys . Along the way, we discovered the genome of this species to be one of the largest genomes published from the Lamiaceae family, to which it belongs. This tetraploid, 3 Gbp genome is especially rich in diterpene synthase genes, with 74 putative sequences identified.

biosynthetic gene cluster (BGC)↗

Phylogenetic analyses and reclassification of the oleaginous marine species Nannochloris sp. “desiccata" (Trebouxiophyceae, Chlorophyta), formerly Chlorella desiccata , supported by a high-quality genome assembly

Microalgae are diverse, with many gaps remaining in phylogenetic and physiological understanding. Thus, studying new microalgae species increases our broader comprehension of biological diversity, and evaluation of new candidates as algal production platforms can lead to improved productivity under a variety of cultivation conditions. Chlorella is a genus of fast-growing species often isolated from freshwater habitats and cultivated as a source of nutritional supplements. However, the use of freshwater increases competition with other freshwater needs. We identified Chlorella desiccata to be worthy of further investigation as a potential algae production strain, due to its isolation from a marine environment and its promising growth and biochemical composition properties. Long-read genomic sequencing was conducted for C. desiccata UTEX 2526, resulting in a high-quality, near chromosome level, diploid genome with an assembly length of 21.55 Mbp in only 18 contigs. We also report complete circular mitochondrial and chloroplast genomes. Phylogenomic and phylogenetic analyses using nuclear, chloroplast, 18S rRNA, and actin sequences revealed that this species clades within strains currently identified as Nannochloris (Trebouxiophyceae, Chlorophyta), leading to its reclassification as Nannochloris sp. “desiccata” UTEX 2526. The mode of cell division for this species is autosporulation, differing from the type species N. bacillaris. As has occurred across multiple microalgae genera, there are repeated examples of Nannochloris species reclassification in the literature. This high-quality genome assembly and phylogenetic analysis of the potential algal production strain Nannochloris sp. “desiccata” UTEX 2526 provides an important reference and useful tool for further studying this region of the phylogenetic tree.

59 BASIC BIOLOGICAL SCIENCES↗

A Chromosome-Scale Genome Assembly of the Flax Rust Fungus Reveals the Two Unusually Large Effector Proteins, AvrM3 and AvrN

Rust fungi comprise thousands of species, many of which cause disease on important crop plants. The flax rust fungus Melampsora lini has been a model species for the genetic dissection of plant immunity since the 1940s; however, the highly fragmented and incomplete reference genome has so far hindered progress in effector gene discovery. Here, we generated a fully phased, chromosome-scale assembly of the two nuclear genomes of M. lini strain CH5, resolving an additional 320 Mbp of the sequence. The 482-Mbp dikaryotic genome is at least 79% repetitive, with a large proportion (approximately 40%) of the genome comprising young, highly similar transposable elements. The assembly resolves the known effector gene loci, some of which carry complex duplications that were collapsed in the previous assembly. Using a genetic map followed by manual correction of gene models, we identified the AvrM3 and AvrN genes, which encode unusually large fungal effector proteins and trigger defense responses when co-expressed with the corresponding resistance genes. We located the genes linked to the tetrapolar mating system on chromosomes 4 and 9, but in contrast to the cereal rusts that have one pheromone receptor gene per haplotype, in flax rust, three pheromone receptor genes were found, with two of them closely linked on one haplotype. Taken together, we show that a high-quality assembly is crucial for resolving complex gene loci, and given the increasing number of fungal effectors of large size, the commonly applied criterion for effector candidates of being small proteins needs to be reconsidered.

Melampsora↗

Metagenome‐Assembled Genomes for Oligotrophic Nitrifiers From a Mountainous Gravelbed Floodplain

Riparian floodplains are important regions for biogeochemical cycling, including nitrogen. Here, we present MAGs from nitrifying microorganisms, including ammonia-oxidising archaea (AOA) and comammox bacteria from Slate River (SR) floodplain sediments (Crested Butte, CO, US). Additionally, we explore MAGs from potential nitrite-oxidising bacteria (NOB) from the Nitrospirales. AOA diversity in SR is lower than observed in other western US floodplain sediments and Nitrosotalea-like lineages such as the genus TA-20 are the dominant AOA. No ammonia-oxidising bacteria (AOB) MAGs were recovered. Microorganisms from the Palsa-1315 genus (clade B comammox) are the most abundant ammonia-oxidizers in SR floodplain sediments. Established NOB are conspicuously absent; however, we recovered MAGs from uncultured lineages of the NS-4 family (Nitrospirales) and Nitrospiraceae that we propose as putative NOB. Nitrite oxidation may be carried out by organisms sister to established Nitrospira NOB lineages based on the genomic content of uncultured Nitrospirales clades. Nitrifier MAGs recovered from SR floodplain sediments harbour genes for using alternative sources of ammonia, such as urea, cyanate, biuret, triuret and nitriles. In conclusion, the SR floodplain therefore appears to be a low ammonia flux environment that selects for oligotrophic nitrifiers.

60 APPLIED LIFE SCIENCES↗

Complete Genome Assemblies of four Saccharomycodales species

To investigate the centromere structures of Saccharomycodales we generated three genome sequences spanning from telomere to telomere for three species in the genus Hanseniaspora. Additionally we generated genome sequences for the species Saccharomycodes pseudoludwigii.

centromeres↗

Five draft genome assemblies from Bacillaceae isolated from a degraded wetland environment

Abstract We isolated 5 Bacillaceae from a degraded wetland environment and sequenced their genomes using Illumina NextSeq. Here, we report draft genome sequences of Bacillus velezensus-SC119, Priestia megaterium-SC120, Bacillus zhangzhouensis-SC123, Bacillus pumilis-SC124, and Bacillus idriensis-SC127. The genomes range between 3,657,353 and 5,772,725 base pairs with %GC between 37.62% and 46.38%. Introduction Wetland environments play critical roles in the terrestrial carbon and water cycles and microbial communities are key players in healthy ecosystem function. Endospore forming bacteria in the Bacillaceae family are metabolically and genomically diverse soil heterotrophs that influence plant health, carbon and nitrogen cycling, and often produce diverse natural products that influence other bacterial and non-bacterial species in their environment (1). We collected two soil samples on January 19, 2023 from 42°43'12.7"N 73°45'01.4"W. One was highly hydrated and within a patch of invasive common reeds (Phragmites sp.) and the other was near the base of an Eastern cottonwood tree (Populus deltoides). Bacillus pumilis strain SC124 was isolated from the soil from near the cottonwood tree, while Bacillus velezensis strain SC119, Priestia megaterium strain SC120, Bacillus zhangzhouensis strain SC123, Bacillus idriensis strain SC127 and were isolated from the marshy soil.

isolate, wetlands, genome announcement↗

Haplotype‐resolved genome assembly of Populus tremula × P. alba reveals aspen‐specific megabase satellite DNA

SUMMARY Populus species play a foundational role in diverse ecosystems and are important renewable feedstocks for bioenergy and bioproducts. Hybrid aspen Populus tremula × P. alba INRA 717‐1B4 is a widely used transformation model in tree functional genomics and biotechnology research. As an outcrossing interspecific hybrid, its genome is riddled with sequence polymorphisms which present a challenge for sequence‐sensitive analyses. Here we report a telomere‐to‐telomere genome for this hybrid aspen with two chromosome‐scale, haplotype‐resolved assemblies. We performed a comprehensive analysis of the repetitive landscape and identified both tandem repeat array‐based and array‐less centromeres. Unexpectedly, the most abundant satellite repeats in both haplotypes lie outside of the centromeres, consist of a 147 bp monomer PtaM147, frequently span >1 megabases, and form heterochromatic knobs. PtaM147 repeats are detected exclusively in aspens (section Populus ) but PtaM147‐like sequences occur in LTR‐retrotransposons of closely related species, suggesting their origin from the retrotransposons. The genomic resource generated for this transformation model genotype has greatly improved the design and analysis of genome editing experiments that are highly sensitive to sequence polymorphisms. The work should motivate future hypothesis‐driven research to probe into the function of the abundant and aspen‐specific PtaM147 satellite DNA.

Zhou, Ran↗

KBase Narrative - Mse_Genome_Assembly

Upon investigating the gut microbiome of mice for microbes related to Polycystic Ovary Syndrome (PCOS), our team found 2 metagenomes that could not be classified using alignment based methods. To further investigate microbial species from these samples, we used common MAGs workflow to create draft genomes which were then later taxonomically classified using GTDB-tk app.

Rastegar, Kiarash↗

A genome assembly and the somatic genetic and epigenetic mutation rate in a wild long-lived perennial Populus trichocarpa

Abstract Background Plants can transmit somatic mutations and epimutations to offspring, which in turn can affect fitness. Knowledge of the rate at which these variations arise is necessary to understand how plant development contributes to local adaption in an ecoevolutionary context, particularly in long-lived perennials. Results Here, we generate a new high-quality reference genome from the oldest branch of a wild Populus trichocarpa tree with two dominant stems which have been evolving independently for 330 years. By sampling multiple, age-estimated branches of this tree, we use a multi-omics approach to quantify age-related somatic changes at the genetic, epigenetic, and transcriptional level. We show that the per-year somatic mutation and epimutation rates are lower than in annuals and that transcriptional variation is mainly independent of age divergence and cytosine methylation. Furthermore, a detailed analysis of the somatic epimutation spectrum indicates that transgenerationally heritable epimutations originate mainly from DNA methylation maintenance errors during mitotic rather than during meiotic cell divisions. Conclusion Taken together, our study provides unprecedented insights into the origin of nucleotide and functional variation in a long-lived perennial plant.

59 BASIC BIOLOGICAL SCIENCES↗

Genome assembly PYCC 8100

The yeast Torulaspora delbrueckii is gaining importance for biotechnology due to its ability to increase wine sensorial complexity and for enhancing pre-frozen bread dough leavening. However, little is known about its population structure, and variation in gene content, or possible domestication routes have not been investigated. Here, we address these issues and update the circumscription of T. delbrueckii, which is composed of five major clades. Among the three European clades, a lineage associated with the wild arboreal niche is sister to the two other lineages that are linked with anthropic environments, one to wine fermentations and the other one to diverse sources including dairy products and bread dough (Mix- Anthropic clade). Using 62 genomes we identified 5629 genes in the pangenome of T. delbrueckii and 270 genes in the cloud genome. A pangenome tree analysis showed that wine strains have a genome composition were more similar to European wild arboreal strains than to those of the Mix Anthropic clade, in contradiction with the phylogenetic analysis. An association of gene content and ecology gave further support to the hypothesis that the Mix - Anthropic clade has the most specialized genome content and indicated that some of the exclusive genes were implicated in galactose and maltose utilization. More detailed analyses traced the acquisition of a cluster of GAL genes in strains associated with dairy products and the expansion and functional diversification of MAL genes in strains isolated from bread dough. Contrary to S. cerevisiae, domestication in T. delbrueckii is not primed by alcoholic fermentation and appears to be a recent event.

Sampaio, Jose P.↗

Genome Assembly and Transcriptome of Colletotrichum sublineola CsGL1, a New Resource to Study Anthracnose Disease in Sorghum

Colletotrichum species are globally distributed and well known as members of a destructive phytopathogenic genus, causing the anthracnose disease in a wide variety of crops and fruits. Colletotrichum sublineola is the causal agent of the anthracnose disease in sorghum, causing losses of up to 50% in yield. Here, we used PacBio sequencing combined with RNA-seq to generate a chromosome-level assembly and annotation of the Colletotrichum sublineola strain CsGL1. [Formula: see text] Copyright © 2021 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license .

Biochemistry & Molecular Biology↗

Metagenomes and Metagenome-Assembled Genomes from Microbiomes Metabolizing Thin Stillage from an Ethanol Biorefinery

Here, we report the metagenomes from five anaerobic bioreactors, operated under different conditions, that were fed carbohydrate-rich thin stillage from a corn starch ethanol plant. The putative functions of the abundant taxa identified here will inform future studies of microbial communities involved in valorizing this and other low-value agroindustrial residues.

Fortney, Nathaniel W.↗

HiFiAdapterFilt, a memory efficient read processing pipeline, prevents occurrence of adapter sequence in PacBio HiFi reads and their negative impacts on genome assembly

Abstract Background Pacific Biosciences HiFi read technology is currently the industry standard for high accuracy long-read sequencing that has been widely adopted by large sequencing and assembly initiatives for generation of de novo assemblies in non-model organisms. Though adapter contamination filtering is routine in traditional short-read analysis pipelines, it has not been widely adopted for HiFi workflows. Results Analysis of 55 publicly available HiFi datasets revealed that a read-sanitation step to remove sequence artifacts derived from PacBio library preparation from read pools is necessary as adapter sequences can be erroneously integrated into assemblies. Conclusions Here we describe the nature of adapter contaminated reads, their consequences in assembly, and present HiFiAdapterFilt, a simple and memory efficient solution for removing adapter contaminated reads prior to assembly.

59 BASIC BIOLOGICAL SCIENCES↗