Engineering PapersSearch

SEARCH · Engineering Papers

Results for “breeding population”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

34 records · Page 2

Optimizing resource allocation in Miscanthus breeding via sparse testing designs for genomic prediction

Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.

Miscanthus sacchariflorus (MSA)

Multi-trait multi-environment genomic prediction strategies for Miscanthus sacchariflorus

Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.

genomic prediction (GP)

Which Plant Traits Increase Soil Carbon Sequestration? Empirical Evidence From a Long‐Term Poplar Genetic Diversity Trial

Plants play a key role in mediating soil response to global change, and breeding or engineering crops to increase soil organic carbon (SOC) storage is a potential route to land-based carbon dioxide removal in agricultural systems. However, due to limited observational datasets plus shifting paradigms of SOC stabilization, it is unclear which plant traits are most important for enhancing different types of soil organic matter. Existing long-term common gardens of genetically diverse plant populations may provide an opportunity to evaluate biological controls on SOC, separate from environmental or management variability. Here we report on soil and root chemical data collected for 24 genotypes within a 13-year-old common garden in northwestern Oregon planted with a large natural variant population of Populus trichocarpa. Fractionating surface soil (0–15 cm) revealed substantial variation in stocks of mineral-associated organic matter (MAOM; 18–67 t C/ha) and particulate organic matter (POM; 2–22 t C/ha). Tree genotype explained 24% and 26% of the MAOM and POM stock variability, respectively, after controlling for background variability. We found minimal association between SOC concentration and either aboveground tree productivity or root biomass recalcitrance (C/N ratios and lignin content). In contrast, root elemental content appeared influential for MAOM-C concentration, which showed a strong positive association with root aluminum (Al) and a strong negative association with root boron (B) and magnesium (Mg). Furthermore, root concentrations of these elements were highly heritable (57%–78%) and not simply a reflection of background variation in soil elemental concentrations. We estimate that surface SOC stocks under these 24 genotypes have diverged at rates of up to 1.2–4.3 t C/ha/year. These results suggest that long-term genetic diversity trials have value for elucidating biological controls on soil organic matter dynamics, and that traits associated with root elemental content may be a useful target for enhancing biosequestration.

biomass recalcitrance

Genomic Analysis of the Natural Variation of Fatty Acid Composition in Seed Oils of Camelina sativa

Camelina sativa is an oilseed crop that has shown strong promise as a biofuel feedstock. The profile of fatty acids greatly influences the oil quality; however, genetic mechanisms that determine the natural variation of fatty acid composition in camelina are not fully understood. A genome wide association study (GWAS) was performed to uncover genetic loci that may contribute to the contents of major fatty acids such as oleic and linolenic acids in camelina seed. Two approaches were taken to improve the GWAS efficiency. First, growing a diversity panel of 212 accessions in four locations and two nitrogen fertilization conditions revealed great variation in fatty acid contents in seeds. Second, using an improved reference genome, abundant markers, including 203,320 single nucleotide polymorphisms (SNPs) and 99,067 insertions/deletions (indels), were developed, which refined the population structure of the diversity panel. GWAS resulted in 118 genetic markers across 31 trait/treatment conditions. Closely linked markers were determined based on linkage decay and by comparing secondarily associated markers when highly associated ones were removed. Candidate genes were examined by comparing the pangenomes of 12 high-quality reference genomes. This study provides new resources to understand seed lipid metabolism and improve camelina oils through molecular breeding.

Life Sciences & Biomedicine - Other Topics

Elemental profiling and genome-wide association studies reveal genomic variants modulating ionomic composition in Populus trichocarpa leaves

The ionome represents elemental composition in plant tissues and can be an indicator of nutrient status as well as overall plant performance. Thus, identifying genetic determinants governing elemental uptake and storage is an important goal for breeding and engineering biomass feedstocks with improved performance. In this study, we coupled high-throughput ionome characterization of leaf tissues with high-resolution genome-wide association studies (GWAS) to uncover genetic loci that modulate ionomic composition in leaves of poplar ( Populus trichocarpa ). Significant agreement was observed across the three ionomic profiling platforms tested: inductively coupled plasma-mass spectrometry (ICP-MS), neutron activation analysis (NAA) and laser-induced breakdown spectroscopy (LIBS). Relative quantification of 20 elements using ICP-MS across a population of 584 genotypes, revealed larger variation in micro-nutrients and trace elements content than for macro-nutrients across genotypes. The GWAS performed using a set of high-density (>8.2 million) single nucleotide polymorphisms, identified over 600 loci significantly associated with variations in these mineral elements, pointing to numerous uncharacterized candidate genes. A significant enrichment for genes related to ion homeostasis and transport was observed, including several members of the cation-proton antiporters (CPA) family and MATE efflux transporters, previously reported to be critical for plant growth and fitness in other species. Our results also included a polymorphic copy of the high-affinity molybdenum transporter MOT1 found directly associated to molybdenum content. For the first time in a perennial plant, our results provide evidence of genetic control of mineral content in a model tree species.

59 BASIC BIOLOGICAL SCIENCES

Higher_wood_density_lowers_feedstock_cost_and_has_minimal_impact_on_biomass_conversion_to_biofuels

Poplar and other woody feedstocks have the potential to provide up to 200 million tons of biomass per year that could be converted to liquid fuels. Most forestry strategies that aim at increasing biomass productivity per hectare rely on short rotation plantations of fast-growing varieties. The improvement of wood density as a key trait itself has largely been overlooked. We evaluated natural variation in wood density across a population of genetically diversePopulus trichocarpatrees grown in a common garden. Wood density varies greatly within this population but is heritable higher wood density was not systematically associated with reduced growth, challenging assumptions of a trade-off between wood density and biomass accumulation. Furthermore, denser wood led to significant improvements throughout the supply chain, including, lowering biomass production and transportation costs. Higher density not correlate to changes in biomass composition. Density did not impact bioconversion in the two feedstock-to-fuel pipelines tested (pretreatment by ionic liquids or soaking in aqueous ammonia, and fermentation to ethanol) on a representative subset of poplars. These findings highlight wood density as a promising breeding target for accelerating the development of high-yielding, conversion-efficient bioenergy crops and as an avenue for increasing land-use efficiency and reducing biomass transportation cost. This data set contain three datasets.

CBI

Higher Wood Density Lowers Feedstock Cost and Has Minimal Impact on Biomass Conversion to Biofuels

Poplar and other woody feedstocks have the potential to provide up to 200 million tons of biomass per year that can be converted to liquid fuels. Most forestry strategies that aim to increase biomass productivity per hectare rely on short rotation plantations of fast-growing varieties. The improvement of the wood density as a key trait itself has largely been overlooked. We evaluated natural variation in wood density across a population of genetically diverse Populus trichocarpa trees grown in a common garden. Wood density varies greatly within this population but is heritable; higher wood density was not systematically associated with reduced growth, challenging assumptions of a trade-off between wood density and biomass accumulation. Furthermore, denser wood led to significant improvements throughout the supply chain including lowering biomass production and transportation costs. Higher density did not correlate with changes in biomass composition. Density did not impact bioconversion in the two feedstock-to-fuel pipelines tested (pretreatment by ionic liquids and fermentation to bisabolene or soaking in aqueous ammonia and fermentation to ethanol) on a representative subset of poplars. These findings highlight wood density as a promising breeding target for accelerating the development of high-yielding, conversion-efficient bioenergy crops and as an avenue for increasing landuse efficiency and reducing biomass transportation cost.

09 BIOMASS FUELS

Population Genomics of Pseudocercospora griseola Reveals New Groups in the Middle American Clade and the Presence of the Endophytic Bacterium Achromobacter xylosoxidans

Angular leaf spot (ALS), caused by Pseudocercospora griseola is an important disease of common beans. P. griseola, is highly variable and has co-evolved with its host. In this study, 48 isolates of P. griseola from Puerto Rico, Guatemala, Honduras and Tanzania were sequenced (3RADseq), resulting in the de novo assembly of 42,214 contigs. Phylogenomic, population genetic structure and principal component analyses using 1,260 SNPs divided these isolates into two populations, Andean and Middle American, while the Middle American population was further divided into three sub-populations. There were moderate to high levels of differentiation between P. griseola populations, with pairwise Fst values ranging from 0.11 to 0.95. The Andean population was composed of isolates from Tanzania, and was separated from the Middle American population (Fst = 0.95). The Middle American population was separated into 3 subpopulations including isolates from: 1. Guatemala and Honduras, 2. Tanzania, and 3. Puerto Rico. Pathogenicity testing of 27 isolates from Puerto Rico, using 12 common bean differential lines, identified ten races, but these races were not associated with SNPs found in virulence genes. DNA of an endophytic bacterium (Achromobacter xylosoxidans) was found in seven mildly virulent isolates suggesting a possible role of the bacterium in the observed virulence patterns. To understand the evolution and diversity of P. griseola, further study of the virulence genes and the interactions among the endophytic bacterium, the fungus, and the host plant is required. Such information is critical to inform breeding strategies for the development of resistant germplasm and cultivars.

Serrato-Diaz, Luz M. [U.S. Department of Agricultu

A single genomic region controls primocane fruiting in tetraploid blackberry

The fresh-market blackberry ( Rubus subgenus Rubus ) industry has expanded dramatically in the past 2 decades, driven in part by improved cultivars. Introgression of the primocane-fruiting (PF; annual flowering) trait into elite germplasm has enabled dual cropping in a single year, season extension, and cultivation in tropical and subtropical regions. Despite its economic performance, the genetic basis of PF is not well understood. It has been proposed that the PF trait is controlled by a major recessive locus, but its genomic location is unclear. Here, a genome-wide association study (GWAS) of 365 tetraploid blackberry genotypes identified a single genomic region on chromosome Ra03 (∼33 Mb) strongly associated with PF. Genetic linkage analysis in a biparental population confirmed that the same interval (32–35 Mb) was linked to the PF phenotype. Ten putative candidate genes were identified in this region. Allele mining using whole-genome resequencing of 17 genotypes highlighted 2 high-priority candidates: a CCCH-type zinc finger gene and an ubiquitin-specific protease gene. Use of an improved Rubus argutus “Hillquist” genome annotation (v1.2) enabled refined variant interpretation, including identification of regulatory 3′ UTR polymorphisms in the zinc finger homolog. Two diagnostic KASP markers (PF1 and PF2), designed from the most significant GWAS SNPs, predicted the PF phenotype with over 96% accuracy in a validation panel of 494 tetraploid blackberries from multiple breeding programs. Together, these results provide the first high-resolution mapping of the PF locus in blackberry, identify candidate genes for flowering regulation in Rubus , and deliver diagnostic markers that can be immediately deployed in breeding programs.

GWAS

Identification of a QTL region for tomato brown rugose fruit virus resistance in Solanum pimpinellifolium

Abstract Tomato (Solanum lycopersicumL.), one of the most widely grown vegetables in the world, has been seriously impacted in the past decade by the emerging tomato brown rugose fruit virus (ToBRFV). ToBRFV is a seed-borne tobamovirus, with ability to overcome the commonly usedTm-2 2 resistance gene in tomato. The objective of this study was to conduct quantitative trait locus (QTL) mapping and identify single-nucleotide polymorphism (SNP) markers associated with ToBRFV resistance in tomato. Two F 2 populations were used for QTL mapping: One derived from a cross betweenS. pimpinellifoliumUSVL333 (PI 390718) × USVL332 (PI 390717) and another from ‘Moneymaker’ × USVL332 (PI 390717), with population sizes of 195 and 79 plants, respectively. The resistance trait was derived from theS. pimpinellifoliumaccession USVL332 (PI 390717). A major QTL for ToBRFV resistance was identified on chromosome 11 (SL4.0ch11), with the peak located at approximately 46.84 Mbp. This QTL spans a 22-kb interval between 46,825,788 bp and 46,847,421 bp, as determined through both genome-wide association study (GWAS) and QTL linkage mapping. Three SNP markers, SL4.0ch11_46825788, SL4.0ch11_46847421, and SL4.0ch11_46850215, demonstrated the most significant association with high LOD values (LOD = 13 in the Blink model) in GWAS analysis. In this genomic region, two disease resistance gene analogs, Solyc11g062150 (TIR-NBS-LRR resistance protein, Toll-Interleukin receptor) and Solyc11g062180 (disease resistance protein, leucine-rich repeat), were identified, which may serve as candidates for ToBRFV resistance. The QTL identified in this study could be valuable for plant breeders in facilitating tomato breeding with ToBRFV resistance.

Agriculture

QTL Mapping of Seed Fatty Acid Contents in Camelina sativa Under Heat Stress

Heat stress alters oil quality in oilseed crops, yet its genetic underpinnings in Camelina sativa remain unclear. This study investigated the genetic basis of heat-induced changes in seed fatty acids using a recombinant inbred line (RIL) population derived from a cross between two camelina varieties, Suneson and Pryzeth. Exposure to high temperature during reproductive growth led to increased proportions of saturated (C16:0, C18:0) and monounsaturated (C18:1) fatty acids, whereas polyunsaturated C18:3, total unsaturated fatty acids (UFA) and the PUFA/MUFA ratio were decreased, suggesting an inhibition of the C18:1 → C18:2 → C18:3 desaturation pathway. A high-density linkage map (4981 bins across 20 chromosomes) was built, and 25 QTLs for fatty acids were detected, with hotspots on chromosomes 1, 9, 12, 13, 16, and 20. A major QTL on chromosome 1 (~ 80 cM) explained the largest variance component for PUFA/MUFA under heat. Three desaturase genes (FAD2, FAD7, FAD8) were located within key QTL intervals, nominating them as candidates for modulating unsaturation under elevated temperature. These results provide a genetic basis for fine mapping and functional validation, supporting future molecular and breeding efforts to stabilize oil quality under warming conditions.

Camelina

Genetic analyses of leaf traits in an interspecific Zoysia japonica × Zoysia matrella F2 population

Zoysiagrass (Zoysia spp.) is an important warm-season turfgrass cultivated across tropical, subtropical, and temperate regions of the world. The genus is characterized by the presence of salt-secreting glands on the adaxial leaf surface, which contribute to its high salt tolerance. In this study, we analyzed an interspecific F2 population, derived from selfing an F1 from a cross between Z. japonica acc. Meyer and Z. matrella acc. PI 231146, for variation in adaxial salt gland density, leaf width, and vein count. Using composite interval mapping with a previously constructed genetic map as a framework, we identified three quantitative trait loci (QTL) for leaf width, two QTL for vein count, and two QTL for salt gland density. We complemented the QTL analysis with bulked segregant RNA-seq (BSR-seq) to identify shared genomic regions and candidate genes for leaf width and salt gland density. BSR-seq identified four trait-associated regions, but only a single region identified for leaf width on Chr08 overlapped with a QTL for the same trait. We highlight putative candidate genes underlying the leaf width and salt gland density QTL and discuss their potential roles in leaf development. Together, the QTL and candidate genes provide an important resource for breeding stress-resilient Zoysia germplasm.

Pradhan, Shreena [University of Georgia, Athens]

Genetics of Flooding Tolerance in an F 2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis Population

Miscanthus is a warm-season, perennial grass cultivated as a feedstock for bioenergy and bioproducts. M. sacchariflorus ssp. lutarioriparius has high yield potential and is well-adapted to seasonal flooding, but little is known about the genetics of this adaptation. We conducted a quantitative trait locus (QTL) analysis on a population of 332 diploid Miscanthus ×giganteus (Mxg) F2s derived from an initial cross between diploid M. sacchariflorus ssp. lutarioriparius ‘PF30022’ and diploid M. sinensis ‘PMS-014’, followed by intermating 50 F 1 s. Using tanks in a greenhouse to assess the effects of partial submergence on actively growing plants, we compared an aerobic soil control to a 6-week flood treatment. The study's primary objectives were to (1) identify QTL for flooding tolerance in Miscanthus , (2) identify candidate genes and (3) compare ethylene response factors in Miscanthus with those in rice and Arabidopsis , sorghum and maize for binding site sequence homology and synteny, especially those associated with flooding tolerance. In total, 10 QTL and 66 candidate genes for partial submergence tolerance were identified (including many for ethylene signalling), a first report for Miscanthus . Notably, none of the Miscanthus candidates were orthologs of rice Sub1A, SK1 or SK2 , yet the ‘PF30022’ parent exhibited a snorkeling phenotype, indicating convergent evolution. This study will facilitate breeding of climate-resiliant Mxg.

abiotic stress tolerance

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri