Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Survey of Thirteen Novel Pseudomonas putida Bacteriophages

Bacteriophages have been widely investigated as a promising treatment of food, medical equipment, and humans colonized by antibiotic-resistant bacteria. Phages pose particular interest in combating those bacteria which form biofilms, such as the medically important human pathogen Pseudomonas aeruginosa and several plant pathogens, including P. syringae . In an undergraduate lab course, P. putida was used as the host to isolate novel anti-pseudomonal bacteriophages. Environmental samples of soil and water were collected, and purified phage isolates were obtained. After Illumina sequencing, genomes of these phages were assembled de novo and annotated. Assembled genomes were compared with known genomes in the literature and GenBank to identify taxonomic relations and to refine their functional annotations. The thirteen phages described are sipho-, myo-, and podoviruses in several families of Caudoviricetes , spanning several novel genera, with genomes ranging from 40,000 to 96,000 bp. One phage (DDSR119) is unique and is the first reported P. putida siphovirus. The remaining 12 can be clustered into four distinct groups. Six are highly related to each other and to previously described Autotranscriptaviridae phages: Waldo5, PlaquesPlease, and Laces98 all belong to the Waldovirus genus, whereas Stalingrad, Bosely, and Stamos belong to the Troedvirus genus. Zuri was previously classified as the founding member of a new genus Zurivirus within the family Schitoviridae . Ebordelon and Holyagarpour each represent different species within Zurivirus , whereas Meara is a more distantly related member of the Schitoviridae . Dolphis and Jeremy are similar enough to form a genus but have only a few distant relatives among sequenced phages and are notable for being temperate. We identified the lysis cassettes in all 13 phages, compared tail spike structures, and found auxiliary metabolic genes in several. Studies like these, which isolate and characterize infectious virions, enable the identification of novel proteins and molecular systems and also provide the raw materials for further study, evaluation, and manipulation of phage proteins and their hosts.

Pseudomonas putida

A multi-omic characterization of the physiological responses to salt stress in Scenedesmus obliquus UTEX393

Scenedesmus obliquus UTEX393 is a promising microalgal candidate for sustainable biomanufacturing but its limited halotolerance hinders large-scale cultivation in saline environments. To investigate the molecular basis of salt stress responses, we conducted a comprehensive multi-omic analysis integrating genomics, transcriptomics, proteomics, lipidomics, metabolomics, and DNA affinity purification sequencing (DAP-seq). An improved nuclear genome assembly and annotation yielded 19,017 gene models and a 97% BUSCO completeness score, enabling construction of a genome-scale metabolic model. Comparing 15 ppt salinity stress to 5 ppt control, growth and productivity were significantly reduced, accompanied by widespread transcriptomic and proteomic changes. Transcriptomic analysis revealed downregulation of photosynthetic machinery and energy conservation genes, and upregulation of stress-responsive elements such as expansins, flavodoxins, and osmoprotectants. Lipidomic profiling showed accumulation of triacylglycerols (TAGs) and degradation of galactosyl lipids, consistent with a shift toward lipid biosynthesis to mitigate redox imbalance. Depletion of key polar metabolites and branched-chain amino acids suggested a rerouting of central carbon metabolism under stress. DAP-seq identified key transcription factors, including LHY1 and SPL12, that target central metabolic enzymes involved in redox balancing, such as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and malate dehydrogenase (MDH). These findings establish a regulatory-metabolic framework linking redox stress to lipid accumulation and reveal potential engineering targets to enhance salt tolerance. Overall, the multi-omic analysis supports the “overflow” hypothesis, where impaired photosynthesis results in excess reducing equivalents being diverted into TAG synthesis and highlights transcriptional regulators as candidates for improving algal robustness in brackish environments.

09 BIOMASS FUELS

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES

Exploring life’s hidden majority: microbial dark matter symposium highlights

The Microbial Dark Matter Symposium held on August 28–29, 2025, in Laguna Beach, Orange County, CA, convened a multidisciplinary group of scientists to address the vast unknowns in microbial life—from uncultured taxa and uncharacterized proteins to elusive viruses and spacefaring microbes. Set against a scenic coastal backdrop, the symposium highlighted advances in single-cell genomics, proximity ligation sequencing, and artificial intelligence-ready bioinformatics, while also probing the limits of microbial persistence, metabolism, and ecological distribution. Sessions explored microbial dark matter from multiple dimensions: cultivability, where new strategies are enabling recovery of elusive microbes; functional ambiguity, where metagenomic dark zones are illuminated by computational annotation; and genomic representation, where single-cell methods bridge gaps left by shotgun community sequencing. Researchers shared breakthroughs in identifying atmospheric microbiomes, “dark oxygen” production in groundwater ecosystems, and microbial survival on the International Space Station. The symposium emphasized integration of methods, disciplines, and ecosystems, advancing a collective push to illuminate the microbial dark matter on Earth and beyond. By highlighting emerging tools, pressing questions, and cross-domain insights, the symposium underscored the need for collaborative, open, and adaptive approaches to study the microbial unknown. The meeting marks a pivotal moment in microbiology, where cultivating knowledge of the uncultivated promises transformative understanding of life, everywhere.

Podar, Mircea [ORNL] (ORCID:0000000327760205)

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

Comprehensive SNP Data for 1,323 GWAS Population in Populus trichocarpa and Combined Annotation Files for P. trichocarpa v3.0 and v3.1

The VCF dataset includes genetic variations found in 1,323 Populus trichocarpa genotypes, providing valuable information for scientists studying plant genetics. Researchers have generated this dataset using whole-genome DNA short-read sequencing on the Illumina Genome Analyzer, HiSeq 2000, and HiSeq 2500 platforms. This sequencing effort ensured a minimum expected sequencing depth of 15×. The dataset comprises more than 9.7 million single nucleotide polymorphisms (SNPs) and indel variants. The combined annotation files are derived from P. trichocarpa v3.0 and v3.1. We merged these files to create a comprehensive annotation file used for GWAS analysis. In total, 38,830 genes overlapped between the two versions. For overlapping genes, we defined the start as the smaller and the end as the larger among the two versions to increase the likelihood of locating candidate genetic loci. Additionally, we included 2,505 unique genes from v3.0 and 4,120 unique genes from v3.1, resulting in a total of 45,455 genes in the updated annotation file.

09 BIOMASS FUELS

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON

Luteolibacter sp. strain Populi

Luteolibacter sp. strain Populi is bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6 Mb chromosome was completely sequenced using Oxford Nanopore long-reads and is predicted to encode 5301 proteins and 60 RNAs. The bacteria was isolated from the rhizosphere of a mature Populus trichocarpa from the Tieton riverwatershed of Washington state, USA (Lat: 46°42’9” N, Lon: 120°25 39’36” W). A rhizosphere sample (fine roots and adhering soil) was used to obtain a microbial fraction by centrifugation on Histodenz (12) and stained with 5µM Syto59 (Thermo Fisher Scientific Inc). A Cytopeia Influx cell sorter (BD, Franklin Lakes, NJ) was used to sort and array single cells (100 per plate) based on forward-side scatter and fluorescence intensity on asparagine-glucose nutrient agar (ATCC medium 184). The Luteolibacter sp. Populi genome sequence has been deposited in GenBank under the accession number CP161812. A draft genome annotated with Prokka and DRAM is available in this Narrative as Luteolibacter_sp_Prokka.240711.

59 BASIC BIOLOGICAL SCIENCES

Systematic identification of transcriptional activation domains from non-transcription factor proteins in plants and yeast

Transcription factors can promote gene expression through activation domains. Whole-genome screens have systematically mapped activation domains in transcription factors but not in non-transcription factor proteins (e.g., chromatin regulators and coactivators). To fill this knowledge gap, we employed the activation domain predictor PADDLE to analyze the proteomes of Arabidopsis thaliana and Saccharomyces cerevisiae. We screened 18,000 predicted activation domains from >800 non-transcription factor genes in both species, confirming that 89% of candidate proteins contain active fragments. Our work enables the annotation of hundreds of nuclear proteins as putative coactivators, many of which have never been ascribed any function in plants. Analysis of peptide sequence compositions reveals how the distribution of key amino acids dictates activity. Finally, we validated short, "universal" activation domains with comparable performance to state-of-the-art activation domains used for genome engineering. Our approach enables the genome-wide discovery and annotation of activation domains that can function across diverse eukaryotes.

59 BASIC BIOLOGICAL SCIENCES

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES

A chromosome-level genome assembly of the varied leaved jewelflower, Streptanthus diversifolius, reveals a recent whole genome duplication

Abstract The Streptanthoid complex, a clade of primarily Streptanthus and Caulanthus species in the Thelypodieae (Brassicaceae) is an emerging model system for ecological and evolutionary studies. This complex spans the full range of the California Floristic Province including desert, foothill, and mountain environments. The ability of these related species to radiate into dramatically different environments makes them a desirable study subject for exploring how plant species expand their ranges and adapt to new environments over time. Ecological and evolutionary studies for this complex have revealed fascinating variation in serpentine soil adaptation, defense compounds, germination, flowering, and life history strategies. Until now a lack of publicly available genome assemblies has hindered the ability to relate these phenotypic observations to their underlying genetic and molecular mechanisms. To help remedy this situation, we present here a chromosome-level genome assembly and annotation of Streptanthus diversifolius, a member of the Streptanthoid Complex, developed using Illumina, Hi-C, and HiFi sequencing technologies. Construction of this assembly also provides further evidence to support the previously reported recent whole genome duplication unique to the Thelypodieae. This whole genome duplication may have provided individuals in the Streptanthoid Complex the genetic arsenal to rapidly radiate throughout the California Floristic Province and to occupy commonly inhospitable environments including serpentine soils.

Genetics & Heredity

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri