Engineering PapersSearch

DOE OSTI · 3015676

YeastWGD2025

Abstract

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David, Kyle (ORCID:000000019907789X), Rokas, Antonis [Department of Biological Sciences, Vanderbilt University, Nashville, TN 37235, USA; Evolutionary Studies Initiative, Vanderbilt University, Nashville, TN 37235, USA]. 2025-01-01. YeastWGD2025. https://doi.org/10.6084/m9.figshare.29852876

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Data for Comparison of Genotyping Assays for Detection of Targeted CRISPR/Cas Mutagenesis in Highly Polyploid Sugarcane

Sugarcane ( Saccharum spp.) is an important biofuel feedstock and a leading source of global table sugar. Saccharum hybrid cultivars are highly polyploid (2n = 100–130), containing large numbers of functionally redundant hom(e)ologs in their genomes. Genome editing with sequence-specific nucleases holds tremendous promise for sugarcane breeding. However, identification of plants with the desired level of co-editing within a pool of primary transformants can be difficult. While DNA sequencing provides direct evidence of targeted mutagenesis, it is cost-prohibitive as a primary screening method in sugarcane and most other methods of identifying mutant lines have not been optimized for use in highly polyploid species. In this study, non-sequencing methods of mutant screening, including capillary electrophoresis (CE), Cas9 RNP assay, and high-resolution melt analysis (HRMA), were compared to assess their potential for CRISPR/Cas9-mediated mutant screening in sugarcane. These assays were used to analyze sugarcane lines containing mutations at one or more of six sgRNA target sites. All three methods distinguished edited lines from wild type, with co-mutation frequencies ranging from 2% to 100%. Cas9 RNP assays were able to identify mutant sugarcane lines with as low as 3.2% co-mutation frequency, and samples could be scored based on undigested band intensity. CE was highlighted as the most comprehensive assay, delivering precise information on both mutagenesis frequency and indel size to a 1 bp resolution across all six targets. This represents an economical and comprehensive alternative to sequencing-based genotyping methods which could be applied in other polyploid species.

Genomics

Data for A Role for Differential Rubisco Activase Isoform Expression in C4 Bioenergy Grasses at High Temperature

Rubisco activase (Rca) facilitates the release of sugar-phosphate inhibitors at Rubisco catalytic sites during CO2 fixation. Most plant species express two Rca isoforms, the larger Rca-α and the shorter Rca-β, either by alternative splicing from a single gene or expression from separate genes. The mechanism of Rubisco activation by Rca isoforms has been intensively studied in C3 plants. However, the functional role of Rca in C4 plants where Rubisco and Rca are located in a much higher [CO2] compartment is less clear. In this study, we selected four C4 bioenergy grasses and the model C4 grass setaria ( Setaria viridis ) to investigate the role of Rca in C4 photosynthesis. All five C4 grass species contained two Rca genes, one encoding Rca-α and the other Rca-β, which were positioned closely together in the genomes. A variety of abiotic stress-related motifs were identified in the Rca-α promoter of each grass, and while the Rca-β gene was constantly highly expressed at ambient temperature, Rca-α isoforms were expressed only at high temperature but never surpassed 30% of Rca-β content. The pattern of Rca-α induction on transition to high temperature and reduction on return to ambient temperature was the same in all five C4 grasses. In sorghum ( Sorghum bicolor ), sugarcane ( Saccharum officinarum ), and setaria, the induction rate of Rca-α was similar to the recovery rate of photosynthesis and Rubisco activation at high temperature. This association between Rca-α isoform expression and maintenance of Rubisco activation at high temperature suggests that Rca-α has a functional thermo-protective role in carbon fixation in C4 grasses by sustaining Rubisco activation at high temperature.

Genomics

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics