Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome assembly”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Major proliferation of transposable elements shaped the genome of the soybean rust pathogen Phakopsora pachyrhizi

With >7000 species the order of rust fungi has a disproportionately large impact on agriculture, horticulture, forestry and foreign ecosystems. The infectious spores are typically dikaryotic, a feature unique to fungi in which two haploid nuclei reside in the same cell. A key example is Phakopsora pachyrhizi, the causal agent of Asian soybean rust disease, one of the world’s most economically damaging agricultural diseases. Despite P. pachyrhizi’s impact, the exceptional size and complexity of its genome prevented generation of an accurate genome assembly. Here, we sequence three independent P. pachyrhizi genomes and uncover a genome up to 1.25 Gb comprising two haplotypes with a transposable element (TE) content of ~93%. We study the incursion and dominant impact of these TEs on the genome and show how they have a key impact on various processes such as host range adaptation, stress responses and genetic plasticity.

59 BASIC BIOLOGICAL SCIENCES↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

Impact of short-read sequencing on the misassembly of a plant genome

Abstract Background Availability of plant genome sequences has led to significant advances. However, with few exceptions, the great majority of existing genome assemblies are derived from short read sequencing technologies with highly uneven read coverages indicative of sequencing and assembly issues that could significantly impact any downstream analysis of plant genomes. In tomato for example, 0.6% (5.1 Mb) and 9.7% (79.6 Mb) of short-read based assembly had significantly higher and lower coverage compared to background, respectively. Results To understand what the causes may be for such uneven coverage, we first established machine learning models capable of predicting genomic regions with variable coverages and found that high coverage regions tend to have higher simple sequence repeat and tandem gene densities compared to background regions. To determine if the high coverage regions were misassembled, we examined a recently available tomato long-read based assembly and found that 27.8% (1.41 Mb) of high coverage regions were potentially misassembled of duplicate sequences, compared to 1.4% in background regions. In addition, using a predictive model that can distinguish correctly and incorrectly assembled high coverage regions, we found that misassembled, high coverage regions tend to be flanked by simple sequence repeats, pseudogenes, and transposon elements. Conclusions Our study provides insights on the causes of variable coverage regions and a quantitative assessment of factors contributing to plant genome misassembly when using short reads and the generality of these causes and factors should be tested further in other species.

59 BASIC BIOLOGICAL SCIENCES↗

The final piece of the Triangle of U: Evolution of the tetraploid Brassica carinata genome

Abstract Ethiopian mustard (Brassica carinata) is an ancient crop with remarkable stress resilience and a desirable seed fatty acid profile for biofuel uses. Brassica carinata is one of six Brassica species that share three major genomes from three diploid species (AA, BB, and CC) that spontaneously hybridized in a pairwise manner to form three allotetraploid species (AABB, AACC, and BBCC). Of the genomes of these species, that of B. carinata is the least understood. Here, we report a chromosome scale 1.31-Gbp genome assembly with 156.9-fold sequencing coverage for B. carinata, completing the reference genomes comprising the classic Triangle of U, a classical theory of the evolutionary relationships among these six species. Our assembly provides insights into the hybridization event that led to the current B. carinata genome and the genomic features that gave rise to the superior agronomic traits of B. carinata. Notably, we identified an expansion of transcription factor networks and agronomically important gene families. Completion of the Triangle of U comparative genomics platform has allowed us to examine the dynamics of polyploid evolution and the role of subgenome dominance in the domestication and continuing agronomic improvement of B. carinata and other Brassica species.

Biochemistry & Molecular Biology↗

Large-scale protein level comparison of Deltaproteobacteria reveals cohesive metabolic groups

Abstract Deltaproteobacteria, now proposed to be the phyla Desulfobacterota, Myxococcota, and SAR324, are ubiquitous in marine environments and play essential roles in global carbon, sulfur, and nutrient cycling. Despite their importance, our understanding of these bacteria is biased towards cultured organisms. Here we address this gap by compiling a genomic catalog of 1 792 genomes, including 402 newly reconstructed and characterized metagenome-assembled genomes (MAGs) from coastal and deep-sea sediments. Phylogenomic analyses reveal that many of these novel MAGs are uncultured representatives of Myxococcota and Desulfobacterota that are understudied. To better characterize Deltaproteobacteria diversity, metabolism, and ecology, we clustered ~1 500 genomes based on the presence/absence patterns of their protein families. Protein content analysis coupled with large-scale metabolic reconstructions separates eight genomic clusters of Deltaproteobacteria with unique metabolic profiles. While these eight clusters largely correspond to phylogeny, there are exceptions where more distantly related organisms appear to have similar ecological roles and closely related organisms have distinct protein content. Our analyses have identified previously unrecognized roles in the cycling of methylamines and denitrification among uncultured Deltaproteobacteria. This new view of Deltaproteobacteria diversity expands our understanding of these dominant bacteria and highlights metabolic abilities across diverse taxa.

Langwig, Marguerite V. (ORCID:0000000202472816)↗

Genome-resolved biogeography of Phaeocystales, cosmopolitan bloom-forming algae

Phaeocystales, comprising the genus Phaeocystis and an uncharacterized sister lineage, are nanoplanktonic haptophytes widespread in the global ocean. Several species form mucilaginous colonies and influence key biogeochemical cycles, yet their underlying diversity and ecological strategies remain underexplored. Here, we present new genomic data from 13 strains, including three high-quality reference genomes (N50 > 30 kbp), and integrate previous metagenome-assembled genomes to resolve a robust phylogeny. Divergence timing of P. antarctica aligns with Miocene cooling and Southern Ocean isolation. Genomic traits reveal metabolic flexibility, including mixotrophic nitrogen acquisition in temperate waters and gene expansions linked to polar nutrient adaptation. Concordantly, transcriptomic comparisons between temperate and polar Phaeocystis suggest Southern Ocean populations experience iron and B12 limitation. We also identify signatures of horizontal gene transfer and endogenous giant virus/virophage insertions. Together, these findings highlight Phaeocystales as an ecologically versatile and geographically widespread lineage shaped by evolutionary innovation and adaptation to contrasting environmental stressors.

Füssy, Zoltán↗

Distinct and rich assemblages of giant viruses in Arctic and Antarctic lakes

Giant viruses (GVs) are key players in ecosystem functioning, biogeochemistry, and eukaryotic genome evolution. GV diversity and abundance in aquatic systems can exceed that of prokaryotes, but their diversity and ecology in lakes, especially polar ones, remain poorly understood. We conducted a comprehensive survey and meta-analysis of GV diversity across 20 lakes, spanning polar to temperate regions, combining our extensive lake metagenome database from the Canadian Arctic and subarctic with publicly available datasets. Leveraging a novel GV genome identification tool, we identified 3304 GV metagenome-assembled genomes, revealing lakes as untapped GV reservoirs. Phylogenomic analysis highlighted their dispersion across all Nucleocytoviricota orders. Strong GV population endemism emerged between lakes from similar regions and biomes (Antarctic and Arctic), but a polar/temperate barrier in lacustrine GV populations and differences in their gene content could be observed. Our study establishes a robust genomic reference for future investigations into lacustrine GV ecology in fast changing polar environments.

59 BASIC BIOLOGICAL SCIENCES↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Final Technical Report for DE-SC0022206

This project developed foundational genetic, genomic, and epigenetic tools for anaerobic fungi (Neocallimastigomycota), a group of microorganisms with exceptional natural abilities to deconstruct lignocellulosic biomass. Efficient biomass deconstruction remains a major barrier to economical production of renewable fuels, chemicals, and materials from agricultural and forestry residues. The project sought to enable mechanistic studies and future engineering of anaerobic fungi by improving genomic resources, establishing methods for gene expression, and investigating epigenetic regulation of biomass-degrading pathways. Major accomplishments included generation of the first chromosome-scale genome assemblies for multiple anaerobic fungal species, providing publicly available genomic resources that support both engineering and fundamental biological research. The project established the first reproducible system for heterologous gene expression in anaerobic fungi and identified genomic features and mobile genetic elements that may support future development of stable transformation technologies. In parallel, the project demonstrated direct conversion of untreated lignocellulosic biomass into fuels and specialty chemicals through a fungal-yeast bioprocess and identified anaerobic fungal enzymes with utility for metabolic engineering. The research also revealed that epigenetic regulation plays an important role in controlling fungal gene expression and enzyme production, identifying potential strategies for enhancing biomass degradation. Collectively, this work established anaerobic fungi as a tractable emerging platform for bioenergy and biomanufacturing research, generated valuable public resources, trained the next generation of researchers, and advanced DOE-BER goals related to predictive biology, sustainable bioprocessing, and the circular bioeconomy.

Solomon, Kevin [University of Delaware] (ORCID:000↗

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.↗

Dissecting the dominant hot spring microbial populations based on community-wide sampling at single-cell genomic resolution

With advances in DNA sequencing and miniaturized molecular biology workflows, rapid and affordable sequencing of single-cell genomes has become a reality. Compared to 16S rRNA gene surveys and shotgun metagenomics, large-scale application of single-cell genomics to whole microbial communities provides an integrated snapshot of community composition and function, directly links mobile elements to their hosts, and enables analysis of population heterogeneity of the dominant community members. To that end, we sequenced nearly 500 single-cell genomes from a low diversity hot spring sediment sample from Dewar Creek, British Columbia, and compared this approach to 16S rRNA gene amplicon and shotgun metagenomics applied to the same sample. We found that the broad taxonomic profiles were similar across the three sequencing approaches, though several lineages were missing from the 16S rRNA gene amplicon dataset, likely the result of primer mismatches. At the functional level, we detected a large array of mobile genetic elements present in the single-cell genomes but absent from the corresponding same species metagenome-assembled genomes. Moreover, we performed a single-cell population genomic analysis of the three most abundant community members, revealing differences in population structure based on mutation and recombination profiles. While the average pairwise nucleotide identities were similar across the dominant species-level lineages, we observed differences in the extent of recombination between these dominant populations. Most intriguingly, the creek's Hydrogenobacter sp. population appeared to be so recombinogenic that it more closely resembled a sexual species than a clonally evolving microbe. Together, this work demonstrates that a randomized single-cell approach can be useful for the exploration of previously uncultivated microbes from community composition to population structure.

59 BASIC BIOLOGICAL SCIENCES↗

Phenotypic and genomic characterization of Methanothermobacter wolfeii strain BSEL, a CO 2 -capturing archaeon with minimal nutrient requirements

A new variant of Methanothermobacter wolfeii was isolated from an anaerobic digester using enrichment cultivation in anaerobic conditions. Here, the new isolate was taxonomically identified via 16S rRNA gene sequencing and tagged as M. wolfeii BSEL. The whole genome of the new variant was sequenced and de novo assembled. Genomic variations between the BSEL strain and the type strain were discovered, suggesting evolutionary adaptations of the BSEL strain that conferred advantages while growing under a low concentration of nutrients. M. wolfeii BSEL displayed the highest specific growth rate ever reported for the wolfeii species (0.27 ± 0.03 h –1 ) using carbon dioxide (CO 2 ) as unique carbon source and hydrogen (H 2 ) as electron donor. M. wolfeii BSEL grew at this rate in an environment with ammonium (NH 4 + ) as sole nitrogen source. The minerals content required to cultivate the BSEL strain was relatively low and resembled the ionic background of tap water without mineral supplements. Optimum growth rate for the new isolate was observed at 64°C and pH 8.3. In this work, it was shown that wastewater from a wastewater treatment facility can be used as a low-cost alternative medium to cultivate M. wolfeii BSEL. Continuous gas fermentation fed with a synthetic biogas mimic along with H 2 in a bubble column bioreactor using M. wolfeii BSEL as biocatalyst resulted in a CO 2 conversion efficiency of 97% and a final methane (CH 4 ) titer of 98.5%v, demonstrating the ability of the new strain for upgrading biogas to renewable natural gas.

09 BIOMASS FUELS↗

Genome-Resolved Metagenomics of a Photosynthetic Bioreactor Performing Biological Nutrient Removal

Enhanced biological phosphorus removal (EBPR) is an economically and environmentally significant wastewater treatment process for removing excess phosphorus by harnessing the metabolic physiologies of enriched microbial communities. We present a genome-resolved metagenomic data set consisting of 86 metagenome-assembled genome sequences from a photosynthetically operated lab-scale bioreactor simulating EBPR.

McDaniel, Elizabeth A.↗

Unraveling the ecological success of Iodidimonas in a bioreactor treating oil and gas produced water

Iodidimonas sp., a bacterium found in bioreactors treating oil and gas produced water as well as iodide-rich brines, has garnered attention for its unique ability to oxidize iodine. However, little is known about the metabolic capabilities that enable Iodidimonas sp. to thrive in certain unique ecological niches. In this study, we isolated, characterized, and sequenced three strains belonging to the Iodidimonas genus from the sludge of a membrane bioreactor used for produced water treatment. We investigated the genomic features of these isolates and compared them with the four publicly available isolate genomes from this genus, as well as a metagenome-assembled genome from the source bioreactor. Our Iodidimonas isolates had several genes associated with mitigating salinity, heavy metal, and organic compound stress, which likely help these bacteria to survive in produced water. Phenotyping tests revealed that while the isolates could utilize a wide variety of simple carbon substrates, they failed to degrade aliphatic or aromatic hydrocarbons, consistent with the lack of genes associated with common hydrocarbon degradation pathways in their genomes. We hypothesize that these microbes may lead a scavenging lifestyle in the bioreactor and similar iodide-rich brines. IMPORTANCE: Occupying a niche habitat and having few representative isolates, the genus Iodidimonas is a relatively understudied alphaproteobacterial group. Its ability to corrode pipes in iodine production facilities has economic implications, and its ability to generate potentially carcinogenic iodinated organic compounds during treatment of oil and gas produced water may cause environmental and health concerns with the recycling of treated water. Therefore, detailed characterization of the metabolic potential of the Iodidimonas isolates in this study both sheds light on their adaptation to the environmental conditions they inhabit and has environmental and economic significance.

Acharya, Shwetha M↗

MAGs from alcoholic fermentation in sugarcane biorefineries

Genome-resolved metagenomics was applied to recover microbial metagenome-assembled genomes (MAGs) from alcoholic fermentation samples collected at two sugarcane biorefineries in São Paulo state, Brazil, during the 2024 harvest season.

59 BASIC BIOLOGICAL SCIENCES↗

From Reads to Function Workshop - Milano 2026

The Bicocca Sampling Days (BSDs) model offers a reproducible “citizen science” framework integrating research, education, and public engagement through large-scale microbiome sampling, followed by a workshop of data analysis on select samples. We identified 9 bacterial and archaeal metagenome-assembled genomes from six soil samples across three separate sampling days in two approaches with indidivual sample and replicate co-assembly spanning three unique classes, providing genomic insights into microbial nutrient cycling in these systems.

59 BASIC BIOLOGICAL SCIENCES↗

RNA nanotechnology to build a dodecahedral genome of single-stranded RNA virus

The quest for artificial RNA viral complexes with authentic structure while being non-replicative is on its way for the development of viral vaccines. RNA viruses contain capsid proteins that interact with the genome during morphogenesis. The sequence and properties of the protein and genome determine the structure of the virus. For example, the Pariacoto virus ssRNA genome assembles into a dodecahedron. Virus-inspired nanotechnology has progressed remarkably due to the unique structural and functional properties of viruses, which can inspire the design of novel nanomaterials. RNA is a programmable biopolymer able to self-assemble sophisticated 3D structures with rich functionalities. RNA dodecahedrons mimicking the Pariacoto virus quasi-icosahedral genome structures were constructed from both native and 2'-F modified RNA oligos. The RNA dodecahedron easily self-assembled using the stable pRNA three-way junction of bacteriophage phi29 as building blocks. The RNA dodecahedron cage was further characterized by cryo-electron microscopy and atomic force microscopy, confirming the spontaneous and homogenous formation of the RNA cage. The reported RNA dodecahedron cage will likely provide further studies on the mechanisms of interaction of the capsid protein with the viral genome while providing a template for further construction of the viral RNA scaffold to add capsid proteins for the assembly of the viral nucleocapsid as a model. Understanding the self-assembly and RNA folding of this RNA cage may offer new insights into the 3D organization of viral RNA genomes. Finally, the reported RNA cage also has the potential to be explored as a novel virus-inspired nanocarrier.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Genomics Using the Integrated Microbial Genomes and Microbiomes (IMG/M) System: A Deinococcus Use Case

The Integrated Microbial Genomes and Microbiomes (IMG/M) system is a web-based platform that provides access to the wealth of public sequence data arising from diverse environments and enables the user to answer biological questions. In this review, we explore IMG’s tools and features using genome data for genus Deinococcus isolates as well as metagenome-assembled genomes (MAGs). Here, we use various comparative genomic and visualization tools to investigate this genus and address specific research questions.

59 BASIC BIOLOGICAL SCIENCES↗