Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “phylogenetics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Cation and Anion Channelrhodopsins: Sequence Motifs and Taxonomic Distribution

Cation and anion channelrhodopsins (CCRs and ACRs, respectively) primarily from two algal species, Chlamydomonas reinhardtii and Guillardia theta, have become widely used as optogenetic tools to control cell membrane potential with light. We mined algal and other protist polynucleotide sequencing projects and metagenomic samples to identify 75 channelrhodopsin homologs from four channelrhodopsin families, including one revealed in dinoflagellates in this study. We carried out electrophysiological analysis of 33 natural channelrhodopsin variants from different phylogenetic lineages and 10 metagenomic homologs in search of sequence determinants of ion selectivity, photocurrent desensitization, and spectral tuning in channelrhodopsins. Our results show that association of a reduced number of glutamates near the conductance path with anion selectivity depends on a wider protein context, because prasinophyte homologs with a glutamate pattern identical to that in cryptophyte ACRs are cation selective. Desensitization is also broadly context dependent, as in one branch of stramenopile ACRs and their metagenomic homologs, its extent roughly correlates with phylogenetic relationship of their sequences. Regarding spectral tuning, we identified two prasinophyte CCRs with red-shifted spectra to 585 nm. They exhibit a third residue pattern in their retinal-binding pockets distinctly different from those of the only two types of red-shifted channelrhodopsins known (i.e., the CCR Chrimson and RubyACRs). In cryptophyte ACRs we identified three specific residue positions in the retinal-binding pocket that define the wavelength of their spectral maxima. Lastly, we found that dinoflagellate rhodopsins with a TCP motif in the third transmembrane helix and a metagenomic homolog exhibit channel activity. Channelrhodopsins are widely used in neuroscience and cardiology as research tools and are considered prospective therapeutics, but their natural diversity and mechanisms remain poorly characterized. Genomic and metagenomic sequencing projects are producing an ever-increasing wealth of data, whereas biophysical characterization of the encoded proteins lags behind. In this study, we used manual and automated patch clamp recording of representative members of four channelrhodopsin families, including a family in dinoflagellates that we report in this study. Our results contribute to a better understanding of molecular determinants of ionic selectivity, photocurrent desensitization, and spectral tuning in channelrhodopsins.

59 BASIC BIOLOGICAL SCIENCES↗

Survey of Early-Diverging Lineages of Fungi Reveals Abundant and Diverse Mycoviruses

ABSTRACT Mycoviruses are widespread and purportedly common throughout the fungal kingdom, although most are known from hosts in the two most recently diverged phyla, Ascomycota and Basidiomycota, together called Dikarya. To augment our knowledge of mycovirus prevalence and diversity in underexplored fungi, we conducted a large-scale survey of fungi in the earlier-diverging lineages, using both culture-based and transcriptome-mining approaches to search for RNA viruses. In total, 21.6% of 333 isolates were positive for RNA mycoviruses. This is a greater proportion than expected based on previous taxonomically broad mycovirus surveys and is suggestive of a strong phylogenetic component to mycoviral infection. Our newly found viral sequences are diverse, composed of double-stranded RNA, positive-sense single-stranded RNA (ssRNA), and negative-sense ssRNA genomes and include novel lineages lacking representation in the public databases. These identified viruses could be classified into 2 orders, 5 families, and 5 genera; however, half of the viruses remain taxonomically unassigned. Further, we identified a lineage of virus-like sequences in the genomes of members of Phycomycetaceae and Mortierellales that appear to be novel genes derived from integration of a viral RNA-dependent RNA polymerase gene. The two screening methods largely agreed in their detection of viruses; thus, we suggest that the culture-based assay is a cost-effective means to quickly assess whether a laboratory culture is virally infected. This study used culture collections and publicly available transcriptomes to demonstrate that mycoviruses are abundant in laboratory cultures of early-diverging fungal lineages. The function and diversity of mycoviruses found here will help guide future studies into mycovirus origins and ecological functions. IMPORTANCE Viruses are key drivers of evolution and ecosystem function and are increasingly recognized as symbionts of fungi. Fungi in early-diverging lineages are widespread, ecologically important, and comprise the majority of the phylogenetic diversity of the kingdom. Viruses infecting early-diverging lineages of fungi have been almost entirely unstudied. In this study, we screened fungi for viruses by two alternative approaches: a classic culture-based method and by transcriptome-mining. The results of our large-scale survey demonstrate that early-diverging lineages have higher infection rates than have been previously reported in other fungal taxa and that laboratory strains worldwide are host to infections, the implications of which are unknown. The function and diversity of mycoviruses found in these basal fungal lineages will help guide future studies into mycovirus origins and their evolutionary ramifications and ecological impacts.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic and Antigenic Characterization of an Expanding H3 Influenza A Virus Clade in U.S. Swine Visualized by Nextstrain

Defining factors that influence spatial and temporal patterns of influenza A virus (IAV) is essential to inform vaccine strain selection and strategies to reduce the spread of potentially zoonotic swine-origin IAV. The relative frequency of detection of the H3 phylogenetic clade 1990.4.a (colloquially known as C-IVA) in U.S. swine declined to 7% in 2017 but increased to 32% in 2019. We conducted phylogenetic and phenotypic analyses to determine putative mechanisms associated with increased detection. We created an implementation of Nextstrain to visualize the emergence, spatial spread, and genetic evolution of H3 IAV in swine, identifying two C-IVA clades that emerged in 2017 and cocirculated in multiple U.S. states. Phylodynamic analysis of the hemagglutinin (HA) gene documented low relative genetic diversity from 2017 to 2019, suggesting clonal expansion. The major H3 C-IVA clade contained an N156H amino acid substitution, but hemagglutination inhibition (HI) assays demonstrated no significant antigenic drift. The minor HA clade was paired with the neuraminidase (NA) clade N2-2002B prior to 2016 but acquired and maintained an N2-2002A in 2016, resulting in a loss of antigenic cross-reactivity between N2-2002B- and -2002A-containing H3N2 strains. The major C-IVA clade viruses acquired a nucleoprotein (NP) of the H1N1pdm09 lineage through reassortment in the replacement of the North American swine-lineage NP. Instead of genetic or antigenic diversity within the C-IVA HA, our data suggest that population immunity to H3 2010.1 along with the antigenic diversity of the NA and the acquisition of the H1N1pdm09 NP gene likely explain the reemergence and transmission of C-IVA H3N2 in swine.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of Effector Metabolites Using Exometabolite Profiling of Diverse Microalgae

Dissolved exometabolites mediate algal interactions in aquatic ecosystems, but microalgal exometabolomes remain understudied. We conducted an untargeted metabolomic analysis of nonpolar exometabolites exuded from four phylogenetically and ecologically diverse eukaryotic microalgal strains grown in the laboratory, freshwater Chlamydomonas reinhardtii, brackish Desmodesmus sp., marine Phaeodactylum tricornutum, and marine Microchloropsis salina, to identify released metabolites based on relative enrichment in the exometabolomes compared to cell pellet metabolomes. Exudates from the different taxa were distinct, but we did not observe clear phylogenetic patterns. We used feature-based molecular networking to explore the identities of these metabolites, revealing several distinct di- and tripeptides secreted by each of the algae, lumichrome, a compound that is known to be involved in plant growth and bacterial quorum sensing, and novel prostaglandin-like compounds. We further investigated the impacts of exogenous additions of eight compounds selected based on exometabolome enrichment on algal growth. Of these compounds, five (lumichrome, 5'-S-methyl-5'-thioadenosine, 17-phenyl trinor prostaglandin A2, dodecanedioic acid, and aleuritic acid) impacted growth in at least one of the algal cultures. Two of these compounds (dodecanedioic acid and aleuritic acid) produced contrasting results, increasing growth in some algae and decreasing growth in others. Together, our results reveal new groups of microalgal exometabolites, some of which could alter algal growth when provided exogenously, suggesting potential roles in allelopathy and algal interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Refinement of the “ Candidatus Accumulibacter” genus based on metagenomic analysis of biological nutrient removal (BNR) pilot-scale plants operated with reduced aeration

Members of the “Candidatus Accumulibacter” genus are widely studied as key polyphosphate-accumulating organisms (PAOs) in biological nutrient removal (BNR) facilities performing enhanced biological phosphorus removal (EBPR). This diverse lineage includes 18 “Ca. Accumulibacter” species, which have been proposed based on the phylogenetic divergence of the polyphosphate kinase 1 (ppk1) gene and genome-scale comparisons of metagenome-assembled genomes (MAGs). Phylogenetic classification based on the 16S rRNA genetic marker has been difficult to attain because most “Ca. Accumulibacter” MAGs are incomplete and often do not include the rRNA operon. Here, we investigate the “Ca. Accumulibacter” diversity in pilot-scale treatment trains performing BNR under low dissolved oxygen (DO) conditions using genome-resolved metagenomics. Using long-read sequencing, we recovered medium- and high-quality MAGs for 5 of the 18 “Ca. Accumulibacter” species, all with rRNA operons assembled, which allowed a reassessment of the 16S rRNA-based phylogeny of this genus and an analysis of phylogeny based on the 23S rRNA gene.

59 BASIC BIOLOGICAL SCIENCES↗

Distinguishing Leptothrix and Sphaerotilus genera by an integrated genomic-phenotypic analysis supported by new Leptothrix genomes

The Sphaerotilus-Leptothrix group of bacteria includes one of the first described microorganisms, Leptothrix ochracea, an uncultured type strain, plus isolates of Leptothrix and Sphaerotilus. This group is unified by the ability to form sheaths and oxidize metals, although L. ochracea exhibits obvious ecological, morphological, and functional differences from the rest of Sphaerotilus-Leptothrix. Recently, there have been calls to combine the group into one genus, Sphaerotilus; however, these studies lacked adequate genomic representation of L. ochracea. Here, we present a comprehensive comparative genomic analysis of the Sphaerotilus-Leptothrix group, including expanded representation of L. ochracea, a closely related novel species, Leptothrix toolikensis, and two new isolates (Leptothrix mechoopdaensis). Analysis of 38 genomes resolves three phylogenetic and functional groups: the ochracea-type Leptothrix (Group 1), the mobilis-type Leptothrix (Group 2), and Sphaerotilus (Group 3). Group 1 genomes form a separate genus based on average nucleotide identity and alignment fraction. The genomes clearly diverge from the rest of Sphaerotilus-Leptothrix in phylogeny, size, and metabolic potential. Group 1 genomes are much smaller (2.59–3.04 Mb) than those of Groups 2 (4.55–6.06 Mb) and 3 (3.94–5.07 Mb), while encoding more metal oxidases and fewer carbohydrate-active enzymes. Group 2 clusters with Group 3 phylogenetically and is similar in organic carbon metabolisms but maintains more metal oxidation genes. Group 2 members lack homogeneity in phenotype and genotype, suggesting that additional isolates and genomes are needed for confident classification. However, Group 1 genomes (L. ochracea and L. toolikensis) show clear divergence, precluding their inclusion in Sphaerotilus and supporting the retention of the genus Leptothrix.

Leptothrix↗

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga↗

A new crab of the genus Nanoplax from the Gulf of Mexico, and assignment of Micropanope pusilla to a new genus (Crustacea, Brachyura, Pseudorhombilidae)

Nanoplax thomai n. sp. is described from the Gulf of Mexico, representing the second species of the genus. The description is based upon a number of specimens previously misidentified as Micropanope truncatifrons Rathbun, 1898, including one so represented in recent molecular phylogenetic analyses. As restricted, Micropanope truncatifrons remains known with certainty from only the limited type series, which does not include a mature male, and sequence quality tissues are not available for molecular phylogenetic analyses. Its generic placement remains questionable following morphological study of its type materials and comparisons to specimens representing other present and former members of Micropanope Stimpson, 1871. Those comparisons underscore that morphological and molecular distinctions warrant assignment of Micropanope pusilla A. Milne-Edwards, 1880 to a new genus, herein designated as Pseudopanopeus n. gen.

Zoology↗

Major Revisions in Arthropod Phylogeny Through Improved Supermatrix, With Support for Two Possible Waves of Land Invasion by Chelicerates

Deep phylogeny involving arthropod lineages is difficult to recover because the erosion of phylogenetic signals over time leads to unreliable multiple sequence alignment (MSA) and subsequent phylogenetic reconstruction. One way to alleviate the problem is to assemble a large number of gene sequences to compensate for the weakness in each individual gene. Such an approach has led to many robustly supported but contradictory phylogenies. A close examination shows that the supermatrix approach often suffers from two shortcomings. The first is that MSA is rarely checked for reliability and, as will be illustrated, can be poor. The second is that, to alleviate the problem of homoplasy at the third codon position of protein-coding genes due to convergent evolution of nucleotide frequencies, phylogeneticists may remove or degenerate the third codon position but may do it improperly and introduce new biases. We performed extensive reanalysis of one of such “big data” sets to highlight these two problems, and demonstrated the power and benefits of correcting or alleviating these problems. Our results support a new group with Xiphosura and Arachnopulmonata (Tetrapulmonata + Scorpiones) as sister taxa. This favors a new hypothesis in which the ancestor of Xiphosura and the extinct Eurypterida (sea scorpions, of which many later forms lived in brackish or freshwater) returned to the sea after the initial chelicerate invasion of land. Our phylogeny is supported even with the original data but processed with a new “principled” codon degeneration. We also show that removing the 1673 codon sites with both AGN and UCN codons (encoding serine) in our alignment can partially reconcile discrepancies between nucleotide-based and AA-based tree, partly because two sequences, one with AGN and the other with UCN, would be identical at the amino acid level but quite different at the nucleotide level.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome-level genome assemblies and genetic maps reveal heterochiasmy and macrosynteny in endangered Atlantic Acropora

Abstract Background Over their evolutionary history, corals have adapted to sea level rise and increasing ocean temperatures, however, it is unclear how quickly they may respond to rapid change. Genome structure and genetic diversity contained within may highlight their adaptive potential. Results We present chromosome-scale genome assemblies and linkage maps of the critically endangered Atlantic acroporids,Acropora palmataandA. cervicornis. Both assemblies and linkage maps were resolved into 14 chromosomes with their gene content and colinearity. Repeats and chromosome arrangements were largely preserved between the species. The family Acroporidae and the genusAcroporaexhibited many phylogenetically significant gene family expansions. Macrosynteny decreased with phylogenetic distance. Nevertheless, scleractinians shared six of the 21 cnidarian ancestral linkage groups as well as numerous fission and fusion events compared to other distantly related cnidarians. Genetic linkage maps were constructed from oneA. palmatafamily and 16A. cervicornisfamilies using a genotyping array. The consensus maps span 1,013.42 cM and 927.36 cM forA. palmataandA. cervicornis, respectively. Both species exhibited high genome-wide recombination rates (3.04 to 3.53 cM/Mb) and pronounced sex-based differences, known as heterochiasmy, with 2 to 2.5X higher recombination rates estimated in the female maps. Conclusions Together, the chromosome-scale assemblies and genetic maps we present here are the first detailed look at the genomic landscapes of the critically endangered Atlantic acroporids. These data sets revealed that adaptive capacity of Atlantic acroporids is not limited by their recombination rates. The sister species maintain macrosynteny with few genes with high sequence divergence that may act as reproductive barriers between them. In the AtlanticAcropora, hybridization between the two sister species yields an F1 hybrid with limited fertility despite the high levels of macrosynteny and gene colinearity of their genomes. Together, these resources now enable genome-wide association studies and discovery of quantitative trait loci, two tools that can aid in the conservation of these species.

Biotechnology & Applied Microbiology↗

Enrichment of root-associated Streptomyces strains in response to drought is driven by diverse functional traits and does not predict beneficial effects on plant growth

The genus Streptomyces has consistently been found enriched in drought-stressed plant root microbiomes, yet the ecological basis and functional variation underlying this enrichment at the strain and isolate level remain unclear. Using two 16S rRNA sequencing methods with different levels of taxonomic resolution, we confirmed drought-associated enrichment (DE) of Streptomyces in field-grown sorghum roots and identified five closely related but distinct amplicon sequence variants (ASVs) belonging to the genus with variable drought enrichment patterns. From a culture collection of sorghum root endophytes, we selected 12 Streptomyces isolates representing these ASVs for phenotypic and genomic characterization. Whole-genome sequencing revealed substantial variation in gene content, even among closely related isolates, and exometabolomic profiling showed distinct metabolic responses to media supplemented with drought- versus well-watered root tissue. Traits linked to drought survival, including osmotic stress tolerance, siderophore production, and carbon utilization, varied widely among isolates and were not phylogenetically conserved. Using a broader panel of 48 Streptomyces, we demonstrate that DE scores, determined through mono-association experiments in gnotobiotic sorghum systems, showed high variability and lacked correlation with plant growth promotion. Pangenome-wide association identified orthogroups involved in osmolyte transport (e.g., proP) and membrane biosynthesis (e.g., fabG) as positively associated with DE, though most associations lacked phylogenetic signal. Collectively, these results demonstrate that Streptomyces DE is not a conserved genus-level trait but is instead strain-specific and functionally heterogeneous. Furthermore, DE in the root microbiome was shown not to predict beneficial effects on plant growth. This work underscores the need to resolve functional traits at the strain level and highlights the complexity of microbe-host-environment interactions under abiotic stress.

Fonseca-Garcia, Citlali↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

An Abundant and Diverse New Family of Electron Bifurcating Enzymes With a Non-canonical Catalytic Mechanism

Microorganisms utilize electron bifurcating enzymes in metabolic pathways to carry out thermodynamically unfavorable reactions. Bifurcating FeFe-hydrogenases (HydABC) reversibly oxidize NADH (E′∼−280 mV, under physiological conditions) and reduce protons to H 2 gas (E°′−414 mV) by coupling this endergonic reaction to the exergonic reduction of protons by reduced ferredoxin (Fd) (E′∼−500 mV). We show here that HydABC homologs are surprisingly ubiquitous in the microbial world and are represented by 57 phylogenetically distinct clades but only about half are FeFe-hydrogenases. The others have replaced the hydrogenase domain with another oxidoreductase domain or they contain additional subunits, both of which enable various third reactions to be reversibly coupled to NAD + and Fd reduction. We hypothesize that all of these enzymes carry out electron bifurcation and that their third substrates can include hydrogen peroxide, pyruvate, carbon monoxide, aldehydes, aryl-CoA thioesters, NADP + , cofactor F 420 , formate, and quinones, as well as many yet to be discovered. Some of the enzymes are proposed to be integral membrane-bound proton-translocating complexes. These different functionalities are associated with phylogenetically distinct clades and in many cases with specific microbial phyla. We propose that this new and abundant class of electron bifurcating enzyme be referred to as the Bfu family whose defining feature is a conserved bifurcating BfuBC core. This core contains FMN and six iron sulfur clusters and it interacts directly with ferredoxin (Fd) and NAD(H). Electrons to or from the third substrate are fed into the BfuBC core via BfuA. The other three known families of electron bifurcating enzyme (abbreviated as Nfn, EtfAB, and HdrA) contain a special FAD that bifurcates electrons to high and low potential pathways. The Bfu family are proposed to use a different electron bifurcation mechanism that involves a combination of FMN and three adjacent iron sulfur clusters, including a novel [2Fe-2S] cluster with pentacoordinate and partial non-Cys coordination. The absolute conservation of the redox cofactors of BfuBC in all members of the Bfu enzyme family indicate they have the same non-canonical mechanism to bifurcate electrons. A hypothetical catalytic mechanism is proposed as a basis for future spectroscopic analyses of Bfu family members.

59 BASIC BIOLOGICAL SCIENCES↗

Characterization of Senecavirus A Isolates Collected From the Environment of U.S. Sow Slaughter Plants

Vesicular disease caused by Senecavirus A (SVA) is clinically indistinguishable from foot-and-mouth disease (FMD) and other vesicular diseases of swine. When a vesicle is observed in FMD-free countries, a costly and time-consuming foreign animal disease investigation (FADI) is performed to rule out FMD. Recently, there has been an increase in the number of FADIs and SVA positive samples at slaughter plants in the U.S. The objectives of this investigation were to: (1) describe the environmental burden of SVA in sow slaughter plants; (2) determine whether there was a correlation between PCR diagnostics, virus isolation (VI), and swine bioassay results; and (3) phylogenetically characterize the genetic diversity of contemporary SVA isolates. Environmental swabs were collected from three sow slaughter plants (Plants 1-3) and one market-weight slaughter plant (Plant 4) between June to December 2020. Of the 426 samples taken from Plants 1-3, 304 samples were PCR positive and 107 were VI positive. There was no detection of SVA by PCR or VI at Plant 4. SVA positive samples were most frequently found in the summer (78.3% June-September, vs. 59.4% October-December), with a peak at 85% in August. Eighteen PCR positive environmental samples with a range of C t values were selected for a swine bioassay: a single sample infected piglets ( n = 2). A random subset of the PCR positive samples was sequenced; and phylogenetic analysis demonstrated co-circulation and divergence of two genetically distinct groups of SVA. These data demonstrate that SVA was frequently found in the environment of sow slaughter plants, but environmental persistence and diagnostic detection was not indicative of whether a sampled was infectious to swine. Consequently, a more detailed understanding of the epidemiology of SVA and its environmental persistence in the marketing chain is necessary to reduce the number of FADIs and aide in the development of control measures to reduce the spread of SVA.

59 BASIC BIOLOGICAL SCIENCES↗

classLog: Logistic regression for the classification of genetic sequences

Introduction Sequencing and phylogenetic classification have become a common task in human and animal diagnostic laboratories. It is routine to sequence pathogens to identify genetic variations of diagnostic significance and to use these data in realtime genomic contact tracing and surveillance. Under this paradigm, unprecedented volumes of data are generated that require rapid analysis to provide meaningful inference. Methods We present a machine learning logistic regression pipeline that can assign classifications to genetic sequence data. The pipeline implements an intuitive and customizable approach to developing a trained prediction model that runs in linear time complexity, generating accurate output rapidly, even with incomplete data. Our approach was benchmarked against porcine respiratory and reproductive syndrome virus (PRRSv) and swine H1 influenza A virus (IAV) datasets. Trained classifiers were tested against sequences and simulated datasets that artificially degraded sequence quality at 0, 10, 20, 30, and 40%. Results When applied to a poor-quality sequence data, the classifier achieved between >85% to 95% accuracy for the PRRSv and the swine H1 IAV HA dataset and this increased to near perfect accuracy when using the full dataset. The model also identifies amino acid positions used to determine genetic clade identity through a feature selection ranking within the model. These positions can be mapped onto a maximum-likelihood phylogenetic tree, allowing for the inference of clade defining mutations. Discussion Our approach is implemented as a python package with code available at https://github.com/flu-crew/classLog .

Zeller, Michael A.↗

Optimization of Molecular Methods for Detecting Duckweed-Associated Bacteria

The bacterial colonization dynamics of plants can differ between phylogenetically similar bacterial strains and in the context of complex bacterial communities. Quantitative methods that can resolve closely related bacteria within complex communities can lead to a better understanding of plant–microbe interactions. However, current methods often lack the specificity to differentiate phylogenetically similar bacterial strains. In this study, we describe molecular strategies to study duckweed–associated bacteria. We first systematically optimized a bead-beating protocol to co-isolate nucleic acids simultaneously from duckweed and bacteria. We then developed a generic fingerprinting assay to detect bacteria present in duckweed samples. To detect specific duckweed–bacterium associations, we developed a genomics-based computational pipeline to generate bacterial strain-specific primers. These strain-specific primers differentiated bacterial strains from the same genus and enabled the detection of specific duckweed–bacterium associations present in a community context. Moreover, we used these strain-specific primers to quantify the bacterial colonization of duckweed by normalization to a plant reference gene and revealed differences in colonization levels between strains from the same genus. Lastly, confocal microscopy of inoculated duckweed further supported our PCR results and showed bacterial colonization of the duckweed root–frond interface and root interior. The molecular methods introduced in this work should enable the tracking and quantification of specific plant-microbe associations within plant-microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

Phylogenomics and diversification of Sordariomycetes

The Sordariomycetes is a specious, morphologically diverse, and widely distributed class of the phylum Ascomycota that forms a well-supported clade diverged from Leotiomycetes. Aside from their ecological significance as plant and human pathogens, saprobes, endophytes, and fungicolous taxa, species of Sordariomycetes produces a wide range of chemically novel and diverse metabolites used in important fields. Recent phylogenetic analyses derived from a small number of genes have considerably increased our understanding of the family, order, and subclass relationships within Sordariomycetes, but several important groups have not been resolved well. In addition, there are various paraphyletic or polyphyletic groups. Moreover, the criteria used to establish higher ranks remain highly variable across different studies. Therefore, the taxonomy of Sordariomycetes is in constant flux, remains poorly understood, and is subject to much controversy. Here, for the first time, we have assembled a phylogenetic dataset containing 638 genomes representing the 156 genera, 50 families, and 17 orders and 5 subclasses of Sordariomycetes. This data set is based on 1124 genes and results in a well-resolved phylogenomic tree. We further constructed an evolutionary timeline of Sordariomycetes diversification based on the genomic data sets. Our divergence time estimate results are inconsistent with previous studies, suggesting estimates of node ages are less precise and varied. Based on these results, we discuss the higher ranks of Sordariomycetes and empirically propose an unprecedented taxonomic framework for the class.

59 BASIC BIOLOGICAL SCIENCES↗

Evolution of a plant gene cluster in Solanaceae and emergence of metabolic diversity

Plants produce phylogenetically and spatially restricted, as well as structurally diverse specialized metabolites via multistep metabolic pathways. Hallmarks of specialized metabolic evolution include enzymatic promiscuity and recruitment of primary metabolic enzymes and examples of genomic clustering of pathway genes. Solanaceae glandular trichomes produce defensive acylsugars, with sidechains that vary in length across the family. We describe a tomato gene cluster on chromosome 7 involved in medium chain acylsugar accumulation due to trichome specific acyl-CoA synthetase and enoyl-CoA hydratase genes. This cluster co-localizes with a tomato steroidal alkaloid gene cluster and is syntenic to a chromosome 12 region containing another acylsugar pathway gene. We reconstructed the evolutionary events leading to this gene cluster and found that its phylogenetic distribution correlates with medium chain acylsugar accumulation across the Solanaceae. This work reveals insights into the dynamics behind gene cluster evolution and cell-type specific metabolite diversity.

59 BASIC BIOLOGICAL SCIENCES↗