Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “phylogenetics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

CasPEDIA Database: a functional classification system for class 2 CRISPR-Cas enzymes

Abstract CRISPR-Cas enzymes enable RNA-guided bacterial immunity and are widely used for biotechnological applications including genome editing. In particular, the Class 2 CRISPR-associated enzymes (Cas9, Cas12 and Cas13 families), have been deployed for numerous research, clinical and agricultural applications. However, the immense genetic and biochemical diversity of these proteins in the public domain poses a barrier for researchers seeking to leverage their activities. We present CasPEDIA (http://caspedia.org), the Cas Protein Effector Database of Information and Assessment, a curated encyclopedia that integrates enzymatic classification for hundreds of different Cas enzymes across 27 phylogenetic groups spanning the Cas9, Cas12 and Cas13 families, as well as evolutionarily related IscB and TnpB proteins. All enzymes in CasPEDIA were annotated with a standard workflow based on their primary nuclease activity, target requirements and guide-RNA design constraints. Our functional classification scheme, CasID, is described alongside current phylogenetic classification, allowing users to search related orthologs by enzymatic function and sequence similarity. CasPEDIA is a comprehensive data portal that summarizes and contextualizes enzymatic properties of widely used Cas enzymes, equipping users with valuable resources to foster biotechnological development. CasPEDIA complements phylogenetic Cas nomenclature and enables researchers to leverage the multi-faceted nucleic-acid targeting rules of diverse Class 2 Cas enzymes.

59 BASIC BIOLOGICAL SCIENCES↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

The Influence of the Number of Tree Searches on Maximum Likelihood Inference in Phylogenomics

Maximum likelihood (ML) phylogenetic inference is widely used in phylogenomics. As heuristic searches most likely find suboptimal trees, it is recommended to conduct multiple (e.g., 10) tree searches in phylogenetic analyses. However, beyond its positive role, how and to what extent multiple tree searches aid ML phylogenetic inference remains poorly explored. Here, we found that a random starting tree was not as effective as the BioNJ and parsimony starting trees in inferring the ML gene tree and that RAxML-NG and PhyML were less sensitive to different starting trees than IQ-TREE. We then examined the effect of the number of tree searches on ML tree inference with IQ-TREE and RAxML-NG, by running 100 tree searches on 19,414 gene alignments from 15 animal, plant, and fungal phylogenomic datasets. We found that the number of tree searches substantially impacted the recovery of the best-of-100 ML gene tree topology among 100 searches for a given ML program. In addition, all of the concatenation-based trees were topologically identical if the number of tree searches was ≥10. Quartet-based ASTRAL trees inferred from 1 to 80 tree searches differed topologically from those inferred from 100 tree searches for 6/15 phylogenomic datasets. Lastly, our simulations showed that gene alignments with lower difficulty scores had a higher chance of finding the best-of-100 gene tree topology and were more likely to yield the correct trees.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic analysis of 1710 surveillance-based Neisseria gonorrhoeae isolates from the USA in 2019 identifies predominant strain types and chromosomal antimicrobial-resistance determinants

This study characterized high-quality whole-genome sequences of a sentinel, surveillance-based collection of 1710 Neisseria gonorrhoeae (GC) isolates from 2019 collected in the USA as part of the Gonococcal Isolate Surveillance Project (GISP). It aims to provide a detailed report of strain diversity, phylogenetic relationships and resistance determinant profiles associated with reduced susceptibilities to antibiotics of concern. The 1710 isolates represented 164 multilocus sequence types and 21 predominant phylogenetic clades. Common genomic determinants defined most strains’ phenotypic, reduced susceptibility to current and historic antibiotics (e.g. bla TEM plasmid for penicillin, tetM plasmid for tetracycline, gyrA for ciprofloxacin, 23S rRNA and/or mosaic mtr operon for azithromycin, and mosaic penA for cefixime and ceftriaxone). The most predominant phylogenetic clade accounted for 21 % of the isolates, included a majority of the isolates with low-level elevated MICs to azithromycin (2.0 µg ml –1 ), carried a mosaic mtr operon and variants in PorB, and showed expansion with respect to data previously reported from 2018. The second largest clade predominantly carried the GyrA S91F variant, was largely ciprofloxacin resistant (MIC ≥1.0 µg ml –1 ), and showed significant expansion with respect to 2018. Overall, a low proportion of isolates had medium- to high-level elevated MIC to azithromycin ((≥4.0 µg ml –1 ), based on C2611T or A2059G 23S rRNA variants). One isolate carried the penA 60.001 allele resulting in elevated MICs to cefixime and ceftriaxone of 1.0 µg ml –1 . This high-resolution snapshot of genetic profiles of 1710 GC sequences, through a comparison with 2018 data (1479 GC sequences) within the sentinel system, highlights change in proportions and expansion of select GC strains and the associated genetic mechanisms of resistance. The knowledge gained through molecular surveillance may support rapid identification of outbreaks of concern. Continued monitoring may inform public health responses to limit the development and spread of antibiotic-resistant gonorrhoea.

(3-6) antimicrobial resistance↗

The 1.3 Å resolution structure of the truncated group Ia type IV pilin from Pseudomonas aeruginosa strain P1

The type IV pilus is a diverse molecular machine capable of conferring a variety of functions and is produced by a wide range of bacterial species. The ability of the pilus to perform host-cell adherence makes it a viable target for the development of vaccines against infection by human pathogens such as Pseudomonas aeruginosa . Here, the 1.3 Å resolution crystal structure of the N-terminally truncated type IV pilin from P. aeruginosa strain P1 (ΔP1) is reported, the first structure of its phylogenetically linked group (group I) to be discussed in the literature. The structure was solved from X-ray diffraction data that were collected 20 years ago with a molecular-replacement search model generated using AlphaFold ; the effectiveness of other search models was analyzed. Examination of the high-resolution ΔP1 structure revealed a solvent network that aids in maintaining the fold of the protein. On comparing the sequence and structure of P1 with a variety of type IV pilins, it was observed that there are cases of higher structural similarities between the phylogenetic groups of P. aeruginosa than there are between the same phylogenetic group, indicating that a structural grouping of pilins may be necessary in developing antivirulence drugs and vaccines. These analyses also identified the α–β loop as the most structurally diverse domain of the pilins, which could allow it to serve a role in pilus recognition. Studies of ΔP1 in vitro polymerization demonstrate that the optimal hydrophobic catalyst for the oligomerization of the pilus from strain K122 is not conducive for pilus formation of ΔP1; a model of a three-start helical assembly using the ΔP1 structure indicates that the α–β loop and the D-loop prevent in vitro polymerization.

Bragagnolo, Nicholas↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Characterization of a novel aromatic substrate-processing microcompartment in Actinobacteria

ABSTRACT We have discovered a new cluster of genes that is found exclusively in the Actinobacteria phylum. This locus includes genes for the 2-aminophenol meta -cleavage pathway and the shell proteins of a bacterial microcompartment (BMC) and has been named aromatics (ARO) for its putative role in the breakdown of aromatic compounds. In this study, we provide details about the distribution and composition of the ARO BMC locus and conduct phylogenetic, structural, and functional analyses of the first two enzymes in the catabolic pathway: a unique 2-aminophenol dioxygenase, which is exclusively found alongside BMC shell genes in Actinobacteria, and a semialdehyde dehydrogenase, which works downstream of the dioxygenase. Genomic analysis reveals variations in the complexity of the ARO loci across different orders. Some loci are simple, containing shell proteins and enzymes for the initial steps of the catabolic pathway, while others are extensive, encompassing all the necessary genes for the complete breakdown of 2-aminophenol into pyruvate and acetyl-CoA. Furthermore, our analysis uncovers two subtypes of ARO BMC that likely degrade either 2-aminophenol or catechol, depending on the presence of a pathway-specific gene within the ARO locus. The precise precursor of 2-aminophenol, which serves as the initial substrate and/or inducer for the ARO pathway, remains unknown, as our model organism Micromonospora rosaria cannot utilize 2-aminophenol as its sole energy source. However, using enzymatic assays, we demonstrate the dioxygenase’s ability to cleave both 2-aminophenol and catechol in vitro , in collaboration with the aldehyde dehydrogenase, to facilitate the rapid conversion of these unstable and toxic intermediates. IMPORTANCE Bacterial microcompartments (BMCs) are proteinaceous organelles that are widespread among bacteria and provide a competitive advantage in specific environmental niches. Studies have shown that the genetic information necessary to form functional BMCs is encoded in loci that contain genes encoding shell proteins and the enzymatic core. This allows the bioinformatic discovery of BMCs with novel functions and expands our understanding of the metabolic diversity of BMCs. ARO loci, found only in Actinobacteria, contain genes encoding for phylogenetically remote shell proteins and homologs of the meta -cleavage degradation pathway enzymes that were shown to convert central aromatic intermediates into pyruvate and acetyl-CoA in gamma Proteobacteria. By analyzing the gene composition of ARO BMC loci and characterizing two core enzymes phylogenetically, structurally, and functionally, we provide an initial functional characterization of the ARO BMC, the most unusual BMC identified to date, distinctive among the repertoire of studied BMCs.

2-AP 1,6-dioxygenase↗

A genome-informed higher rank classification of the biotechnologically important fungal subphylum Saccharomycotina

The subphylum Saccharomycotina is a lineage in the fungal phylum Ascomycota that exhibits levels of genomic diversity similar to those of plants and animals. The Saccharomycotina consist of more than 1 200 known species currently divided into 16 families, one order, and one class. Species in this subphylum are ecologically and metabolically diverse and include important opportunistic human pathogens, as well as species important in biotechnological applications. Many traits of biotechnological interest are found in closely related species and often restricted to single phylogenetic clades. However, the biotechnological potential of most yeast species remains unexplored. Although the subphylum Saccharomycotina has much higher rates of genome sequence evolution than its sister subphylum, Pezizomycotina, it contains only one class compared to the 16 classes in Pezizomycotina. The third subphylum of Ascomycota, the Taphrinomycotina, consists of six classes and has approximately 10 times fewer species than the Saccharomycotina. These data indicate that the current classification of all these yeasts into a single class and a single order is an underappreciation of their diversity. Our previous genome-scale phylogenetic analyses showed that the Saccharomycotina contains 12 major and robustly supported phylogenetic clades; seven of these are current families (Lipomycetaceae, Trigonopsidaceae, Alloascoideaceae, Pichiaceae, Phaffomycetaceae, Saccharomycodaceae, and Saccharomycetaceae), one comprises two current families (Dipodascaceae and Trichomonascaceae), one represents the genus Sporopachydermia, and three represent lineages that differ in their translation of the CUG codon (CUG-Ala, CUG-Ser1, and CUG-Ser2). Using these analyses in combination with relative evolutionary divergence and genome content analyses, we propose an updated classification for the Saccharomycotina, including seven classes and 12 orders that can be diagnosed by genome content. This updated classification is consistent with the high levels of genomic diversity within this subphylum and is necessary to make the higher rank classification of the Saccharomycotina more comparable to that of other fungi, as well as to communicate efficiently on lineages that are not yet formally named.

59 BASIC BIOLOGICAL SCIENCES↗

GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring

GAL08 are bacteria belonging to an uncultivated phylogenetic cluster within the phylum Acidobacteria . We detected a natural population of the GAL08 clade in sediment from a pH-neutral hot spring located in British Columbia, Canada. To shed light on the abundance and genomic potential of this clade, we collected and analyzed hot spring sediment samples over a temperature range of 24.2–79.8°C. Illumina sequencing of 16S rRNA gene amplicons and qPCR using a primer set developed specifically to detect the GAL08 16S rRNA gene revealed that absolute and relative abundances of GAL08 peaked at 65°C along three temperature gradients. Analysis of sediment collected over multiple years and locations revealed that the GAL08 group was consistently a dominant clade, comprising up to 29.2% of the microbial community based on relative read abundance and up to 4.7 × 10 5 16S rRNA gene copy numbers per gram of sediment based on qPCR. Using a medium quality threshold, 25 single amplified genomes (SAGs) representing these bacteria were generated from samples taken at 65 and 77°C, and seven metagenome-assembled genomes (MAGs) were reconstructed from samples collected at 45–77°C. Based on average nucleotide identity (ANI), these SAGs and MAGs represented three separate species, with an estimated average genome size of 3.17 Mb and GC content of 62.8%. Phylogenetic trees constructed from 16S rRNA gene sequences and a set of 56 concatenated phylogenetic marker genes both placed the three GAL08 bacteria as a distinct subgroup of the phylum Acidobacteria , representing a candidate order ( Ca. Frugalibacteriales) within the class Blastocatellia. Metabolic reconstructions from genome data predicted a heterotrophic metabolism, with potential capability for aerobic respiration, as well as incomplete denitrification and fermentation. In laboratory cultivation efforts, GAL08 counts based on qPCR declined rapidly under atmospheric levels of oxygen but increased slightly at 1% (v/v) O 2 , suggesting a microaerophilic lifestyle.

59 BASIC BIOLOGICAL SCIENCES↗

Biochemical characterization of Fsa16295Glu from “Fervidibacter sacchari,” the first hyperthermophilic GH50 with β-1,3-endoglucanase activity and founding member of the subfamily GH50_3

The aerobic hyperthermophile “Fervidibacter sacchari” catabolizes diverse polysaccharides and is the only cultivated member of the class “Fervidibacteria” within the phylum Armatimonadota. It encodes 117 putative glycoside hydrolases (GHs), including two from GH family 50 (GH50). In this study, we expressed, purified, and functionally characterized one of these GH50 enzymes, Fsa16295Glu. We show that Fsa16295Glu is a β-1,3-endoglucanase with optimal activity on carboxymethyl curdlan (CM-curdlan) and only weak agarase activity, despite most GH50 enzymes being described as β-agarases. The purified enzyme has a wide temperature range of 4–95°C (optimal 80°C), making it the first characterized hyperthermophilic representative of GH50. The enzyme is also active at a broad pH range of at least 5.5–11 (optimal 6.5–10). Fsa16295Glu possesses a relatively high k cat /K M of 1.82 × 10 7 s-1 M-1 with CM-curdlan and degrades CM-curdlan nearly completely to sugar monomers, indicating preferential hydrolysis of glucans containing β-1,3 linkages. Finally, a phylogenetic analysis of Fsa16295Glu and all other GH50 enzymes revealed that Fsa16295Glu is distant from other characterized enzymes but phylogenetically related to enzymes from thermophilic archaea that were likely acquired horizontally from “Fervidibacteria.” Given its functional and phylogenetic novelty, we propose that Fsa16295Glu represents a new enzyme subfamily, GH50_3.

59 BASIC BIOLOGICAL SCIENCES↗

Macroevolutionary diversity of traits and genomes in the model yeast genus Saccharomyces

Species is the fundamental unit to quantify biodiversity. In recent years, the model yeast Saccharomyces cerevisiae has seen an increased number of studies related to its geographical distribution, population structure, and phenotypic diversity. However, seven additional species from the same genus have been less thoroughly studied, which has limited our understanding of the macroevolutionary events leading to the diversification of this genus over the last 20 million years. Here, we show the geographies, hosts, substrates, and phylogenetic relationships for approximately 1,800 Saccharomyces strains, covering the complete genus with unprecedented breadth and depth. We generated and analyzed complete genome sequences of 163 strains and phenotyped 128 phylogenetically diverse strains. This dataset provides insights about genetic and phenotypic diversity within and between species and populations, quantifies reticulation and incomplete lineage sorting, and demonstrates how gene flow and selection have affected traits, such as galactose metabolism. These findings elevate the genus Saccharomyces as a model to understand biodiversity and evolution in microbial eukaryotes.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms↗

Ornamental origins and genomic frontiers: a review of big-bracted dogwood research

The big-bracted (Benthamidia) dogwood clade consists of small- to medium-sized deciduous trees within the genus Cornus, known for their showy spring-time floral bract display. Cornus is within the family Cornaceae and order Cornales, and as Cornales is one of the earliest diverging asterids, these taxa have been important for phylogenetic research. Three species within the big-bracted clade, flowering (Cornus florida), kousa (C. kousa), and Pacific (C. nuttallii) dogwoods, are popular ornamental landscape plants in North America, with more than 130 cultivars released. Despite their commercial popularity, numerous research gaps have limited the expansion of fundamental research and dogwood breeding programs. In this present review, we aim to provide a thorough overview of our current understanding of 1) the phylogenetic and biogeographic context, 2) plant biology and major pests and pathogens impacting commercialization, 3) historical commercialization and propagation methods, and 4) genetic and genomic resources and how they have been implemented to understand these species. Research gaps and future directions to advance basic research and breeding of big-bracted ornamental dogwoods are discussed throughout.

Cornus florida↗

Anatomy of a mega‐radiation: Biogeography and niche evolution in Astragalus

Premise Astragalus (Fabaceae), with more than 3000 species, represents a globally successful radiation of morphologically highly similar species predominant across the northern hemisphere. It has attracted attention from systematists and biogeographers, who have asked what factors might be behind the extraordinary diversity of this important arid-adapted clade and what sets it apart from close relatives with far less species richness. Methods Here, for the first time using extensive phylogenetic sampling, we asked whether (1) Astragalus is uniquely characterized by bursts of radiation or whether diversification instead is uniform and no different from closely related taxa. Then we tested whether the species diversity of Astragalus is attributable specifically to its predilection for (2) cold and arid habitats, (3) particular soils, or to (4) chromosome evolution. Finally, we tested (5) whether Astragalus originated in central Asia as proposed and (6) whether niche evolutionary shifts were subsequently associated with the colonization of other continents. Results Our results point to the importance of heterogeneity in the diversification of Astragalus, with upshifts associated with the earliest divergences but not strongly tied to any abiotic factor or biogeographic regionalization tested here. The only potential correlate with diversification we identified was chromosome number. Biogeographic shifts have a strong association with the abiotic environment and highlight the importance of central Asia as a biogeographic gateway. Conclusions Our investigation shows the importance of phylogenetic and evolutionary studies of logistically challenging “mega-radiations.” Our findings reject any simple key innovation behind high diversity and underline the often nuanced, multifactorial processes leading to species-rich clades.

59 BASIC BIOLOGICAL SCIENCES↗

Phylogenomic insights into the taxonomy, ecology, and mating systems of the lorchel family Discinaceae (Pezizales, Ascomycota)

Lorchels, also known as false morels (Gyromitra sensu lato), are iconic due to their brain-shaped mushrooms and production of gyromitrin, a deadly mycotoxin. Molecular phylogenetic studies have hitherto failed to resolve deep-branching relationships in the lorchel family, Discinaceae, hampering our ability to settle longstanding taxonomic debates and to reconstruct the evolution of toxin production. We generated 75 draft genomes from cultures and ascomata (some collected as early as 1960), conducted phylogenomic analyses using 1542 single-copy orthologs to infer the early evolutionary history of lorchels, and identified genomic signatures of trophic mode and mating-type loci to better understand lorchel ecology and reproductive biology. Our phylogenomic tree was supported by high gene tree concordance, facilitating taxonomic revisions in Discinaceae. We recognized 10 genera across two tribes: tribe Discineae (Discina, Maublancomyces, Neogyromitra, Piscidiscina, and Pseudodiscina) and tribe Gyromitreae (Gyromitra, Hydnotrya, Paragyromitra, Pseudorhizina, and Pseudoverpa); Piscidiscina was newly erected and 26 new combinations were formalized. Paradiscina melaleuca and Marcelleina donadinii formed their own family-level clade sister to Morchellaceae, which merits further taxonomic study. Genome size and CAZyme content were consistent with a mycorrhizal lifestyle for the truffle species (Hydnotrya spp.), whereas the other Discinaceae genera possessed genomic properties of a saprotrophic habit. Lorchels were found to be predominantly heterothallic-either MAT1-1 or MAT1-2-but a single occurrence of colocalized mating-type idiomorphs indicative of homothallism was observed in Gyromitra esculenta strain CBS101906 and requires additional confirmation and follow-up study. Lastly, we confirmed that gyromitrin has a phylogenetically discontinuous distribution, having been detected exclusively in two distantly related genera (Gyromitra and Piscidiscina) belonging to separate tribes. Our genomic dataset will facilitate further investigations into the gyromitrin biosynthesis genes and their evolutionary history. With additional sampling of Geomoriaceae and Helvellaceae-two closely related families with no publicly available genomes-these data will enable comprehensive studies on the independent evolution of truffles and ecological diversification in an economically important group of pezizalean fungi.

Dirks, Alden C↗

Diverse ecophysiological adaptations of subsurface Thaumarchaeota in floodplain sediments revealed through genome-resolved metagenomics

Abstract The terrestrial subsurface microbiome contains vastly underexplored phylogenetic diversity and metabolic novelty, with critical implications for global biogeochemical cycling. Among the key microbial inhabitants of subsurface soils and sediments are Thaumarchaeota, an archaeal phylum that encompasses ammonia-oxidizing archaea (AOA) as well as non-ammonia-oxidizing basal lineages. Thaumarchaeal ecology in terrestrial systems has been extensively characterized, particularly in the case of AOA. However, there is little knowledge on the diversity and ecophysiology of Thaumarchaeota in deeper soils, as most lineages, particularly basal groups, remain uncultivated and underexplored. Here we use genome-resolved metagenomics to examine the phylogenetic and metabolic diversity of Thaumarchaeota along a 234 cm depth profile of hydrologically variable riparian floodplain sediments in the Wind River Basin near Riverton, Wyoming. Phylogenomic analysis of the metagenome-assembled genomes (MAGs) indicates a shift in AOA population structure from the dominance of the terrestrial Nitrososphaerales lineage in the well-drained top ~100 cm of the profile to the typically marine Nitrosopumilales in deeper, moister, more energy-limited sediment layers. We also describe two deeply rooting non-AOA MAGs with numerous unexpected metabolic features, including the reductive acetyl-CoA (Wood-Ljungdahl) pathway, tetrathionate respiration, a form III RuBisCO, and the potential for extracellular electron transfer. These MAGs also harbor tungsten-containing aldehyde:ferredoxin oxidoreductase, group 4f [NiFe]-hydrogenases and a canonical heme catalase, typically not found in Thaumarchaeota. Our results suggest that hydrological variables, particularly proximity to the water table, impart a strong control on the ecophysiology of Thaumarchaeota in alluvial sediments.

59 BASIC BIOLOGICAL SCIENCES↗