Engineering PapersSearch

SEARCH · Engineering Papers

Results for “phylogenetic trees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Transcriptomics outputs and phylogenetic trees used for pathway discovery of diterpenoid alkaloids in Delphinium and Aconitum

Transcriptome assemblies, open reading frames in nucleotide and peptide sequences, clustered transcriptomes and corresponding amino acid files, and expression matrices in TPM and raw counts for RNA-seq datasets from Delphinium grandiflorum, Aconitum plicatum, Aconitum lycoctonum, Aconitum carmichaelii, Aconitum japonicum, Aconitum kusnezoffii, and Aconitum vilmorinianum. Also included are phylogenetic trees for terpene synthases and cytochromes P450 mined from these assemblies.

biosynthesis

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES

Sugar Relese Supplementary Text and Figures

Phylogenetic tree of GAUT Protein Family and gene model, RNAi construct, and relative transcript abundance of GAUT4 in switchgrass, rice and poplar knockdown (KD) lines.

bio engineered

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES

Small signaling peptides, phylogenetic analysis

ML Phylogenetic analysis of small signaling peptides in Arabidopsis, Sorghum bicolor, Rice, Wheat, Maize, and Brachypodium. Sorghum only phylogenetic trees were used to name genes. All trees but RALF tree are rooted at midpoint.

Kurtz, Evan [Department of Biochemistry and Biophy

Exploring the Structural, Biochemical, and Functional Diversity of Glycoside Hydrolase Family 12 from Penicillium subrubescens

Glycoside hydrolases (GHs) play an essential role in plant biomass degradation and modification for the sustainable production of biochemicals. The filamentous Ascomycete fungus Penicillium subrubescens contains a higher number of GH12 candidates compared to related species. Therefore, we aimed to compare P. subrubescens GH12s for their ability and substrate specificity for plant cell wall polysaccharide degradation and species’ potential as a source of novel enzymes for plant biomass valorization. Our re-evaluated phylogenetic analysis of fungal GH12 members showed that the P. subrubescens GH12s were located in different (new) clades. Biochemical characterization marked PsEglA as an endoglucanase and four other P. subrubescens GH12s (i.e., PsXegA–D) as xyloglucanases. Interestingly, structural features of PsXegD and PsXegE were more comparable to those of Basidiomycete GH12 xyloglucanases with a unique open substrate-binding cleft. PsUegA displayed dual xyloglucanase and endoglucanase activity and also showed distinct structural features. Comparative transcriptome analysis supported the functional diversity of P. subrubescens GH12s in plant biomass degradation. The gene encoding PsUegA was expressed under diverse conditions, suggesting a scouting role for this enzyme.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology

Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution

A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion)

Khurana, Mridul

Characterization of switchgrass ( Panicum virgatum L.) PvKSL1 as a levopimaradiene/abietadiene‐type diterpene synthase

Abstract The diverse class of plant diterpenoid metabolites serves important functions in mediating growth, chemical defence, and ecological adaptation. In major monocot crops, such as maize (Zea mays), rice (Oryza sativa), and barley (Hordeum vulgare), diterpenoids function as core components of biotic and abiotic stress resilience. Switchgrass (Panicum virgatum) is a perennial grass valued as a stress‐resilient biofuel model crop. Previously we identified an unusually large diterpene synthase family that produces both common and species‐specific diterpenoids, several of which accumulate in response to abiotic stress. Here, we report discovery and functional characterization of a previously unrecognized monofunctional class I diterpene synthase (PvKSL1) viain vivoco‐expression assays with different copalyl pyrophosphate (CPP) isomers, structural and mutagenesis studies, as well as genomic and transcriptomic analyses. In particular, PvKSL1 convertsent‐CPP intoent‐abietadiene,ent‐palustradiene,ent‐levopimaradiene, andent‐neoabietadiene via a 13‐hydroxy‐8(14)‐ent‐abietene intermediate. Notably, although featuring a distinctent‐stereochemistry, this product profile is near‐identical to bifunctional (+)‐levopimaradiene/abietadiene synthases occurring in conifer trees. PvKSL1 has three of four active site residues previously shown to control (+)‐levopimaradiene/abietadiene synthase catalytic specificity. However, mutagenesis studies suggest a distinct catalytic mechanism in PvKSL1. Genome localization ofPvKSL1distant from other diterpene synthases, and its phylogenetic distinctiveness from known abietane‐forming diterpene synthases, support an independent evolution of PvKSL1 activity. Albeit at low levels,PvKSL1gene expression predominantly in roots suggests a role of diterpenoid formation in belowground tissue. Together, these findings expand the known chemical and functional space of diterpenoid metabolism in monocot crops.

Plant Sciences

Ornamental origins and genomic frontiers: a review of big-bracted dogwood research

The big-bracted (Benthamidia) dogwood clade consists of small- to medium-sized deciduous trees within the genus Cornus, known for their showy spring-time floral bract display. Cornus is within the family Cornaceae and order Cornales, and as Cornales is one of the earliest diverging asterids, these taxa have been important for phylogenetic research. Three species within the big-bracted clade, flowering (Cornus florida), kousa (C. kousa), and Pacific (C. nuttallii) dogwoods, are popular ornamental landscape plants in North America, with more than 130 cultivars released. Despite their commercial popularity, numerous research gaps have limited the expansion of fundamental research and dogwood breeding programs. In this present review, we aim to provide a thorough overview of our current understanding of 1) the phylogenetic and biogeographic context, 2) plant biology and major pests and pathogens impacting commercialization, 3) historical commercialization and propagation methods, and 4) genetic and genomic resources and how they have been implemented to understand these species. Research gaps and future directions to advance basic research and breeding of big-bracted ornamental dogwoods are discussed throughout.

Cornus florida

Beyond Solanaceae: incorporation of feruloyltyramine and feruloyloctopamine into Cannabaceae lignins

The ferulic acid amides, feruloyltyramine and feruloyloctopamine, have been widely reported as integral constituents in the lignins in several species of Solanaceae in which they function as authentic lignin monomers. In the present study, we demonstrate that these ferulic acid amides are likewise incorporated into the lignins of species within Cannabaceae, including hemp (Cannabis sativa), hops (Humulus lupulus), and European nettle tree (Celtis australis). Structural analyses using derivatization followed by reductive cleavage (DFRC) and two-dimensional nuclear magnetic resonance (2D-NMR) spectroscopy revealed that these ferulic acid amides are incorporated via 4−O- and 8−O-ether linkages, as well as through 8−5′ linkages forming phenylcoumaran structures. Examination of a broad phylogenetic range of plant families demonstrated the absence of these ferulic acid amides from the lignins of all families studied except Solanaceae and Cannabaceae. Given the distant phylogenetic relationship between Solanaceae and Cannabaceae, the recruitment of these ferulic acid amides as lignin monomers in both lineages likely constitutes a case for convergent evolution at the level of lignin biosynthetic pathways. The significance of these ferulic acid amides lies in their unique role as the sole nitrogen-containing phenolic compounds known to participate in lignin formation.

Cannabaceae

Diversity of Sordariales Fungi: Identification of Seven New Species of Naviculisporaceae Through Morphological Analyses and Genome Sequencing

Thanks to next-generation sequencing (NGS) technologies, the diversity of fungi can now be investigated through the analysis of their genome sequences. Naviculisporaceae is a family within the Sordariales, whose diversity is not well-known, with only one genome sequence published for this family. Here, we report on the isolation and cultivation of 20 new strains of Naviculisporaceae. Their genome sequences, as well as those of the five commercially available strains, were determined, thus providing complete genome sequences for 25 new Naviculisporaceae strains. Species delimitation was conducted using a combination of (1) ITS + LSU phylogenetic analysis of the new isolates along with other known species of the family, (2) comparisons between DNA barcode sequences of the new strains with those of the known species, and (3) average genome-wide nucleotide identity calculation. We built a phylogenomic tree and studied the organization of the mating-type locus. In vitro fruiting was obtained for 16 strains, enabling the definition of seven new species, namely Pseudorhypophila gallica, Pseudorhypophila guyanensis Rhypophila alpibus, Rhypophila brasiliensis, Rhypophila camarguensis, Rhypophila reunionensis and Rhypophila thailandica, as well as two new combinations, namely Pseudorhypophila latipes and Pseudorhypophila oryzae. Eight strains for which in vitro fruiting was not obtained may belong to additional new species. These results expand the known diversity of the Naviculisporaceae and greatly enlarge the genomic data available for the family.

Naviculisporaceae

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni

Phylogenomic insights into the taxonomy, ecology, and mating systems of the lorchel family Discinaceae (Pezizales, Ascomycota)

Lorchels, also known as false morels (Gyromitra sensu lato), are iconic due to their brain-shaped mushrooms and production of gyromitrin, a deadly mycotoxin. Molecular phylogenetic studies have hitherto failed to resolve deep-branching relationships in the lorchel family, Discinaceae, hampering our ability to settle longstanding taxonomic debates and to reconstruct the evolution of toxin production. We generated 75 draft genomes from cultures and ascomata (some collected as early as 1960), conducted phylogenomic analyses using 1542 single-copy orthologs to infer the early evolutionary history of lorchels, and identified genomic signatures of trophic mode and mating-type loci to better understand lorchel ecology and reproductive biology. Our phylogenomic tree was supported by high gene tree concordance, facilitating taxonomic revisions in Discinaceae. We recognized 10 genera across two tribes: tribe Discineae (Discina, Maublancomyces, Neogyromitra, Piscidiscina, and Pseudodiscina) and tribe Gyromitreae (Gyromitra, Hydnotrya, Paragyromitra, Pseudorhizina, and Pseudoverpa); Piscidiscina was newly erected and 26 new combinations were formalized. Paradiscina melaleuca and Marcelleina donadinii formed their own family-level clade sister to Morchellaceae, which merits further taxonomic study. Genome size and CAZyme content were consistent with a mycorrhizal lifestyle for the truffle species (Hydnotrya spp.), whereas the other Discinaceae genera possessed genomic properties of a saprotrophic habit. Lorchels were found to be predominantly heterothallic-either MAT1-1 or MAT1-2-but a single occurrence of colocalized mating-type idiomorphs indicative of homothallism was observed in Gyromitra esculenta strain CBS101906 and requires additional confirmation and follow-up study. Lastly, we confirmed that gyromitrin has a phylogenetically discontinuous distribution, having been detected exclusively in two distantly related genera (Gyromitra and Piscidiscina) belonging to separate tribes. Our genomic dataset will facilitate further investigations into the gyromitrin biosynthesis genes and their evolutionary history. With additional sampling of Geomoriaceae and Helvellaceae-two closely related families with no publicly available genomes-these data will enable comprehensive studies on the independent evolution of truffles and ecological diversification in an economically important group of pezizalean fungi.

Dirks, Alden C

A haplotype-resolved, chromosome-scale genome assembly for the southern live oak, Quercus virginiana

Hybridization is a major force driving diversification, migration, and adaptation in Quercus species. While population genetics and phylogenetics have traditionally been used for studying these processes, advances in sequencing technology now enable us to incorporate comparative and pan-genomic approaches as well. Here, we present a highly contiguous, chromosome-scale and haplotype-resolved genome assembly for the southern live oak, Quercus virginiana, the first reference genome for section Virentes, as part of the American Campus Tree Genomes program. Originating from a clone of Auburn University's historic “Toomer's Oak,” this assembly contributes to the pool of genomic resources for investigating recombination, haplotype variation, and structural genomic changes influencing hybridization potential in this clade and across Quercus. It also provides insights into the architecture of the putative centromeric regions within the genus. Alongside other oak references, the Q. virginiana genome will support research into the evolution and adaptation of the Quercus genus.

Quercus virginiana