Engineering PapersSearch

Engineering topics

Grimwood, Jane

Publications and source records attributed to Grimwood, Jane.

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri

Insights into convergent evolution of cosexuality in liverworts from the Marchantia quadrata genome

Sex chromosomes are expected to coevolve with their respective sex, potentially disfavoring their co-occurrence as cosexuality evolves. This effect is expected to be stronger where sex chromosomes are restricted to one sex, such as in plants expressing sex in their haploid stage. We assess this hypothesis in liverworts with U/V sex chromosomes, ancestral dioicy, and several independent transitions to monoicy (cosexuality). We report the chromosome-level genome assembly of Marchantia quadrata, which recently evolved monoicy, and perform comparative genomic analyses with its dioicous relative M. polymorpha. We find that monoicy evolved via retention of the V chromosome as a small ninth chromosome, complete loss of the U chromosome, and translocation of key U-linked genes to autosomes, among which the major sex-determining gene (Feminizer) acquired environmental/developmental regulation. Our findings parallel recent observations on Ricciocarpos natans, which evolved monoicy independently, suggesting genetic constraints that may make transitions to monoicy predictable in liverworts.

Potente, Giacomo

Scaffolded and annotated nuclear and organelle genomes of the North American brown alga Saccharina latissima

Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.

DeWeese, Kelly

ZW sex chromosome structure in Amborella trichopoda

Sex chromosomes have evolved hundreds of times across the flowering plant tree of life; their recent origins in some members of this clade can shed light on the early consequences of suppressed recombination, a crucial step in sex chromosome evolution. Amborella trichopoda, the sole species of a lineage that is sister to all other extant flowering plants, is dioecious with a young ZW sex determination system. Here we present a haplotype-resolved genome assembly, including highly contiguous assemblies of the Z and W chromosomes. We identify a ~3-megabase sex-determination region (SDR) captured in two strata that includes a ~300-kilobase inversion that is enriched with repetitive sequences and contains a homologue of the Arabidopsis METHYLTHIOADENOSINE NUCLEOSIDASE (MTN1-2) genes, which are known to be involved in fertility. However, the remainder of the SDR does not show patterns typically found in non-recombining SDRs, such as repeat accumulation and gene loss. These findings are consistent with the hypothesis that dioecy is derived in Amborella and the sex chromosome pair has not significantly degenerated.

59 BASIC BIOLOGICAL SCIENCES

Assembly, comparative analysis, and utilization of a single haplotype reference genome for soybean

Cultivar Williams 82 has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as Wm82-ISU-01 (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48 387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of Williams 82 reveal large regions of genomic heterogeneity, including regions of differential introgression from the cultivar Kingwa within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (> 20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the ‘Williams’ haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of Williams 82. In addition to the Wm82.a6 assembly, we also assembled the genome of ‘Fiskeby III,’ a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with Fiskeby III revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and Fiskeby III genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.

59 BASIC BIOLOGICAL SCIENCES

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES

Relics of interspecific hybridization retained in the genome of a drought-adapted peanut cultivar

Peanut (Arachis hypogaea L.) is a globally important oil and food crop frequently grown in arid, semi-arid, or dryland environments. Improving drought tolerance is a key goal for peanut crop improvement efforts. Here, we present the genome assembly and gene model annotation for “Line8,” a peanut genotype bred from drought-tolerant cultivars. Our assembly and annotation are the most contiguous and complete peanut genome resources currently available. The high contiguity of the Line8 assembly allowed us to explore structural variation both between peanut genotypes and subgenomes. We detect several large inversions between Line8 and other peanut genome assemblies, and there is a trend for the inversions between more genetically diverged genotypes to have higher gene content. We also relate patterns of subgenome exchange to structural variation between Line8 homeologous chromosomes. Unexpectedly, we discover that Line8 harbors an introgression from A.cardenasii, a diploid peanut relative and important donor of disease resistance alleles to peanut breeding populations. The fully resolved sequences of both haplotypes in this introgression provide the first in situ characterization of A.cardenasii candidate alleles that can be leveraged for future targeted improvement efforts. The completeness of our genome will support peanut biotechnology and broader research into the evolution of hybridization and polyploidy.

60 APPLIED LIFE SCIENCES

Genome resources for three modern cotton lines guide future breeding efforts

Cotton ( Gossypium hirsutum L.) is the key renewable fibre crop worldwide, yet its yield and fibre quality show high variability due to genotype-specific traits and complex interactions among cultivars, management practices and environmental factors. Modern breeding practices may limit future yield gains due to a narrow founding gene pool. Precision breeding and biotechnological approaches offer potential solutions, contingent on accurate cultivar-specific data. Here we address this need by generating high-quality reference genomes for three modern cotton cultivars (‘UGA230’, ‘UA48’ and ‘CSX8308’) and updating the ‘TM-1’ cotton genetic standard reference. Despite hypothesized genetic uniformity, considerable sequence and structural variation was observed among the four genomes, which overlap with ancient and ongoing genomic introgressions from ‘Pima’ cotton, gene regulatory mechanisms and phenotypic trait divergence. Differentially expressed genes across fibre development correlate with fibre production, potentially contributing to the distinctive fibre quality traits observed in modern cotton cultivars. These genomes and comparative analyses provide a valuable foundation for future genetic endeavours to enhance global cotton yield and sustainability.

59 BASIC BIOLOGICAL SCIENCES