Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome assembly”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Insights into convergent evolution of cosexuality in liverworts from the Marchantia quadrata genome

Sex chromosomes are expected to coevolve with their respective sex, potentially disfavoring their co-occurrence as cosexuality evolves. This effect is expected to be stronger where sex chromosomes are restricted to one sex, such as in plants expressing sex in their haploid stage. We assess this hypothesis in liverworts with U/V sex chromosomes, ancestral dioicy, and several independent transitions to monoicy (cosexuality). We report the chromosome-level genome assembly of Marchantia quadrata, which recently evolved monoicy, and perform comparative genomic analyses with its dioicous relative M. polymorpha. We find that monoicy evolved via retention of the V chromosome as a small ninth chromosome, complete loss of the U chromosome, and translocation of key U-linked genes to autosomes, among which the major sex-determining gene (Feminizer) acquired environmental/developmental regulation. Our findings parallel recent observations on Ricciocarpos natans, which evolved monoicy independently, suggesting genetic constraints that may make transitions to monoicy predictable in liverworts.

Potente, Giacomo↗

Novel Gloeobacterales spp. from Diverse Environments across the Globe

Photosynthetic Cyanobacteria and their descendants are the only known organisms capable of oxygenic photosynthesis. Their metabolism permanently changed the Earth’s surface and the evolutionary trajectory of life, but little is known about their evolutionary history. Genomes of the Gloeobacterales, an order of deeply divergent photosynthetic Cyanobacteria, may hold clues about the evolutionary process. However, there are only three published genomes within this order, and it is difficult to make broad inferences based on such little data. Here, I describe five species within the Gloeobacterales retrieved from publicly available databases and examine their photosynthetic gene content and the environments in which Gloeobacterales genomes and 16S rRNA gene sequences are found. The Gloeobacterales contain reduced photosystems and inhabit cold, wet-rock, and low-light environments. They are likely present in low abundances due to their low growth rate. Future searches for Gloeobacterales should target these environments, and samples should be deeply sequenced to capture the low-abundance taxa. Publicly available databases contain undescribed taxa within the Gloeobacterales. However, searching through all available data with current methods is computationally expensive. Therefore, new methods must be developed to search for these and other evolutionarily important taxa. Once identified, these novel photosynthetic Cyanobacteria will help illuminate the origin and evolution of oxygenic photosynthesis. Early branching photosynthetic Cyanobacteria such as the Gloeobacterales may provide clues into the evolutionary history of oxygenic photosynthesis, but there are few genomes or cultured taxa from this order. Five new metagenome-assembled genomes suggest that members of the Gloeobacterales all contain reduced photosystems and lack genes associated with thylakoids and circadian rhythms. Their distribution suggests that they may thrive in environments that are marginal for other species, including wet-rock and cold environments. These traits may aid in the discovery and cultivation of novel species in this clade.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

Four chromosome scale genomes and a pan-genome annotation to accelerate pecan tree breeding

Genome-enabled biotechnologies have the potential to accelerate breeding efforts in long-lived perennial crop species. Despite the transformative potential of molecular tools in pecan and other outcrossing tree species, highly heterozygous genomes, significant presence–absence gene content variation, and histories of interspecific hybridization have constrained breeding efforts. To overcome these challenges, here, we present diploid genome assemblies and annotations of four outbred pecan genotypes, including a PacBio HiFi chromosome-scale assembly of both haplotypes of the ‘Pawnee’ cultivar. Comparative analysis and pan-genome integration reveal substantial and likely adaptive interspecific genomic introgressions, including an over-retained haplotype introgressed from bitternut hickory into pecan breeding pedigrees. Further, by leveraging our pan-genome presence–absence and functional annotation database among genomes and within the two outbred haplotypes of the ‘Lakota’ genome, we identify candidate genes for pest and pathogen resistance. Combined, these analyses and resources highlight significant progress towards functional and quantitative genomics in highly diverse and outbred crops.

54 ENVIRONMENTAL SCIENCES↗

Gene and genome duplications have contrasting impacts on biosynthetic and flower developmental pathways in California poppy

Benzylisoquinoline alkaloids (BIAs) represent a vast group of specialized plant metabolites with diverse pharmaceutical applications, synthesized by a variety of gene families. Among the multiple plant lineages that produce BIAs, the most notable is the poppy family (Papaveraceae), with California poppy (Eschscholzia californica) emerging as a model organism. Here, we report a haplotype-resolved genome assembly, in combination with a high-density expression atlas, for California poppy. Genome analyses reveal recent diversification of BIA biosynthesis genes in poppy through localized duplications. Furthermore, we demonstrate that the degree of phylogenetic relatedness among paralogs within BIA biosynthesis-associated gene families correlates with similarities in gene expression. In contrast, gene families involved in carotenoid biosynthesis, which contributes to the intense orange petal pigmentation, are not phylogenetically clustered, and floral developmental regulators exhibit a high degree of retention of gene duplicates associated with ancient polyploidy events. These findings illustrate alternative roles for gene and genome duplications as drivers of trait evolution. Given the position of California poppy in the angiosperm phylogeny, the high-quality genomic resources generated for this work constitute a valuable resource for comparative genomic and transcriptomic analyses for poppies and flowering plants more generally.

Rössner, Le-Han [Justus-Liebig University, Giessen↗

Hijacking a rapid and scalable metagenomic method reveals subgenome dynamics and evolution in polyploid plants

Premise: The genomes of polyploid plants archive the evolutionary events leading to their present forms. However, plant polyploid genomes present numerous hurdles to the genome comparison algorithms for classification of polyploid types and exploring genome dynamics. Methods: Here, the problem of intra- and inter-genome comparison for examining polyploid genomes is reframed as a metagenomic problem, enabling the use of the rapid and scalable MinHashing approach. To determine how types of polyploidy are described by this metagenomic approach, plant genomes were examined from across the polyploid spectrum for both k-mer composition and frequency with a range of k-mer sizes. In this approach, no subgenome-specific k-mers are identified; rather, whole-chromosome k-mer subspaces were utilized. Results: Given chromosome-scale genome assemblies with sufficient subgenome-specific repetitive element content, literature-verified subgenomic and genomic evolutionary relationships were revealed, including distinguishing auto- from allopolyploidy and putative progenitor genome assignment. The sequences responsible were the rapidly evolving landscape of transposable elements. An investigation into the MinHashing parameters revealed that the downsampled k-mer space (genomic signatures) produced excellent approximations of sequence similarity. Furthermore, the clustering approach used for comparison of the genomic signatures is scrutinized to ensure applicability of the metagenomics-based method. Discussion: The easily implementable and highly computationally efficient MinHashing-based sequence comparison strategy enables comparative subgenomics and genomics for large and complex polyploid plant genomes. Such comparisons provide evidence for polyploidy-type subgenomic assignments. In cases where subgenome-specific repeat signal may not be adequate given a chromosomes' global k-mer profile, alternative methods that are more specific but more computationally complex outperform this approach.

59 BASIC BIOLOGICAL SCIENCES↗

A view of the pan‐genome of domesticated Cowpea ( Vigna unguiculata [L.] Walp.)

Abstract Cowpea, Vigna unguiculata L . Walp., is a diploid warm‐season legume of critical importance as both food and fodder in sub‐Saharan Africa. This species is also grown in Northern Africa, Europe, Latin America, North America, and East to Southeast Asia. To capture the genomic diversity of domesticates of this important legume, de novo genome assemblies were produced for representatives of six subpopulations of cultivated cowpea identified previously from genotyping of several hundred diverse accessions. In the most complete assembly (IT97K‐499‐35), 26,026 core and 4963 noncore genes were identified, with 35,436 pan genes when considering all seven accessions. GO terms associated with response to stress and defense response were highly enriched among the noncore genes, while core genes were enriched in terms related to transcription factor activity, and transport and metabolic processes. Over 5 million single nucleotide polymorphisms (SNPs) relative to each assembly and over 40 structural variants >1 Mb in size were identified by comparing genomes. Vu10 was the chromosome with the highest frequency of SNPs, and Vu04 had the most structural variants. Noncore genes harbor a larger proportion of potentially disruptive variants than core genes, including missense, stop gain, and frameshift mutations; this suggests that noncore genes substantially contribute to diversity within domesticated cowpea.

59 BASIC BIOLOGICAL SCIENCES↗

Robust, versatile DNA FISH probes for chromosome-specific repeats in Caenorhabditis elegans and Pristionchus pacificus

Repetitive DNA sequences are useful targets for chromosomal fluorescence in situ hybridization. We analyzed recent genome assemblies of Caenorhabditis elegans and Pristionchus pacificus to identify tandem repeats with a unique genomic localization. Based on these findings, we designed and validated sets of oligonucleotide probes for each species targeting at least 1 locus per chromosome. These probes yielded reliable fluorescent signals in different tissues and can easily be combined with the immunolocalization of cellular proteins. Synthesis and labeling of these probes are highly cost-effective and require no hands-on labor. The methods presented here can be easily applied in other model and nonmodel organisms with a sequenced genome.

59 BASIC BIOLOGICAL SCIENCES↗

High phenotypic and genotypic plasticity among strains of the mushroom-forming fungus Schizophyllum commune

Schizophyllum commune is a mushroom-forming fungus notable for its distinctive fruiting bodies with split gills. It is used as a model organism to study mushroom development, lignocellulose degradation and mating type loci. It is a hypervariable species with considerable genetic and phenotypic diversity between the strains. In this study, we systematically phenotyped 16 dikaryotic strains for aspects of mushroom development and 18 monokaryotic strains for lignocellulose degradation. There was considerable heterogeneity among the strains regarding these phenotypes. The majority of the strains developed mushrooms with varying morphologies, although some strains only grew vegetatively under the tested conditions. Growth on various carbon sources showed strain-specific profiles. The genomes of seven monokaryotic strains were sequenced and analyzed together with six previously published genome sequences. Moreover, the related species Schizophyllum fasciatum was sequenced. Although there was considerable genetic variation between the genome assemblies, the genes related to mushroom formation and lignocellulose degradation were well conserved. These sequenced genomes, in combination with the high phenotypic diversity, will provide a solid basis for functional genomics analyses of the strains of S. commune.

59 BASIC BIOLOGICAL SCIENCES↗

Quality MAGnified

Here in this Genome Watch highlights different tools and strategies used to enhance the quality of metagenome-assembled genomes (MAGs) generated in microbiome studies.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic and behavioral adaptation of Candida parapsilosis to the microbiome of hospitalized infants revealed by in situ genomics, transcriptomics, and proteomics

Background Candida parapsilosis is a common cause of invasive candidiasis, especially in newborn infants, and infections have been increasing over the past two decades. C. parapsilosis has been primarily studied in pure culture, leaving gaps in understanding of its function in a microbiome context. Results. Here, we compare five unique C. parapsilosis genomes assembled from premature infant fecal samples, three of which are newly reconstructed, and analyze their genome structure, population diversity, and in situ activity relative to reference strains in pure culture. All five genomes contain hotspots of single nucleotide variants, some of which are shared by strains from multiple hospitals. A subset of environmental and hospital-derived genomes share variants within these hotspots suggesting derivation of that region from a common ancestor. Four of the newly reconstructed C. parapsilosis genomes have 4 to 16 copies of the gene RTA3, which encodes a lipid translocase and is implicated in antifungal resistance, potentially indicating adaptation to hospital antifungal use. Time course metatranscriptomics and metaproteomics on fecal samples from a premature infant with a C. parapsilosis blood infection revealed highly variable in situ expression patterns that are distinct from those of similar strains in pure cultures. For example, biofilm formation genes were relatively less expressed in situ, whereas genes linked to oxygen utilization were more highly expressed, indicative of growth in a relatively aerobic environment. In gut microbiome samples, C. parapsilosis co-existed with Enterococcus faecalis that shifted in relative abundance over time, accompanied by changes in bacterial and fungal gene expression and proteome composition. Conclusions The results reveal potentially medically relevant differences in Candida function in gut vs. laboratory environments, and constrain evolutionary processes that could contribute to hospital strain persistence and transfer into premature infant microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome evolution and the genetic basis of agronomically important traits in greater yam

The nutrient-rich tubers of the greater yam, Dioscorea alata L., provide food and income security for millions of people around the world. Despite its global importance, however, greater yam remains an orphan crop. Here, we address this resource gap by presenting a highly contiguous chromosome-scale genome assembly of D. alata combined with a dense genetic map derived from African breeding populations. The genome sequence reveals an ancient allotetraploidization in the Dioscorea lineage, followed by extensive genome-wide reorganization. Using the genomic tools, we find quantitative trait loci for resistance to anthracnose, a damaging fungal pathogen of yam, and several tuber quality traits. Genomic analysis of breeding lines reveals both extensive inbreeding as well as regions of extensive heterozygosity that may represent interspecific introgression during domestication. These tools and insights will enable yam breeders to unlock the potential of this staple crop and take full advantage of its adaptability to varied environments.

59 BASIC BIOLOGICAL SCIENCES↗

MONet/1000 Soils metagenome pathway modelling narrative w/ auto batch import

This narrative performs metabolic modeling and flux balance analysis (FBA) using metagenome-assembled genomes (MAGs) from the 1000 Soils samples, as described by Song et al. (2026, accepted). The set of MAGs (in FASTA format) is converted into a set of assembly objects compatible with functional annotation via RASTtk, yielding a set of genome objects that undergo metabolic modeling via OMEGGA. These genome objects are then used to conduct FBA, generating tables of metabolite uptake rates across the MAGs under investigation.

59 BASIC BIOLOGICAL SCIENCES↗

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS↗