Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “phylogenetics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Recombination smooths the time-signal disrupted by latency in within-host HIV phylogenies

Within-host HIV evolution involves several features that may disrupt standard phylogenetic reconstruction. One important feature is re-activation of latently integrated provirus, which has the potential to disrupt the temporal signal, leading to variation in the branch lengths and apparent evolutionary rates in a tree. Yet, real within-host HIV phylogenies tend to show clear, ladder-like trees structured by the time of sampling. Another important feature is recombination, which violates the fundamental assumption that evolutionary history can be represented by a single bifurcating tree. Thus, recombination complicates the within-host HIV dynamic by mixing genomes and creating evolutionary loop structures that cannot be represented in bifurcating trees. In this paper, we develop a coalescent-based simulator of within-host HIV evolution that includes latency, recombination, and effective population size dynamics that allows us to study the relationship between the true, complex genealogy of within-host HIV evolution, encoded as an Ancestral Recombination Graph (ARG), and the observed phylogenetic tree. To compare our ARG results to the familiar phylogeny format, we calculate the expected bifurcating tree after decomposing the ARG into all unique site trees, their combined distance matrix, and the overall corresponding bifurcating tree. While latency and recombination separately disrupt the phylogenetic signal, remarkably, we find that recombination recovers the temporal signal of within-host HIV evolution caused by latency by mixing fragments of old, latent genomes into the contemporary population. In effect, recombination averages over extant heterogeneity, whether it stems from mixed time-signals or population bottlenecks. Further, we establish that the signals of latency and recombination can be observed in phylogenetic trees despite being an incorrect representation of the true evolutionary history. Using an Approximate Bayesian Computation method, we develop a set of statistical probes to tune our simulation model to nine longitudinally-sampled within-host HIV phylogenies. Because ARGs are exceedingly difficult to infer from real HIV data, our simulation system allows investigating effects of latency, recombination, and population size bottlenecks by matching decomposed ARGs to real data as observed in standard phylogenies.

59 BASIC BIOLOGICAL SCIENCES↗

Patterns of Element Incorporation in Calcium Carbonate Biominerals Recapitulate Phylogeny for a Diverse Range of Marine Calcifiers

Elemental ratios in biogenic marine calcium carbonates are widely used in geobiology, environmental science, and paleoenvironmental reconstructions. It is generally accepted that the elemental abundance of biogenic marine carbonates reflects a combination of the abundance of that ion in seawater, the physical properties of seawater, the mineralogy of the biomineral, and the pathways and mechanisms of biomineralization. Here we report measurements of a suite of nine elemental ratios (Li/Ca, B/Ca, Na/Ca, Mg/Ca, Zn/Ca, Sr/Ca, Cd/Ca, Ba/Ca, and U/Ca) in 18 species of benthic marine invertebrates spanning a range of biogenic carbonate polymorph mineralogies (low-Mg calcite, high-Mg calcite, aragonite, mixed mineralogy) and of phyla (including Mollusca, Echinodermata, Arthropoda, Annelida, Cnidaria, Chlorophyta, and Rhodophyta) cultured at a single temperature (25°C) and a range of p CO 2 treatments (ca. 409, 606, 903, and 2856 ppm). This dataset was used to explore various controls over elemental partitioning in biogenic marine carbonates, including species-level and biomineralization-pathway-level controls, the influence of internal pH regulation compared to external pH changes, and biocalcification responses to changes in seawater carbonate chemistry. The dataset also enables exploration of broad scale phylogenetic patterns of elemental partitioning across calcifying species, exhibiting high phylogenetic signals estimated from both uni- and multivariate analyses of the elemental ratio data (univariate: λ = 0–0.889; multivariate: λ = 0.895–0.99). Comparing partial R 2 values returned from non-phylogenetic and phylogenetic regression analyses echo the importance of and show that phylogeny explains the elemental ratio data 1.4–59 times better than mineralogy in five out of nine of the elements analyzed. Therefore, the strong associations between biomineral elemental chemistry and species relatedness suggests mechanistic controls over element incorporation rooted in the evolution of biomineralization mechanisms.

58 GEOSCIENCES↗

New Data Define the Molecular Phylogeny and Taxonomy of Four Freshwater Suctorian Ciliates With Redefinition of Two Families Heliophryidae and Cyclophryidae (Ciliophora, Phyllopharyngea, Suctoria)

Four suctorian ciliates, Cyclophrya magna Gönnert, 1935, Peridiscophrya florea (Kormos & Kormos, 1958) Dovgal, 2002, Heliophrya rotunda (Hentschel, 1916) Matthes, 1954 and Dendrosoma radians Ehrenberg, 1838, were collected from a freshwater lake in Ningbo, China. The morphological redescription and molecular phylogenetic analyses of these ciliates were investigated. Phylogenetic analyses inferred from SSU rDNA sequences show that all three suctorian orders, Endogenida, Evaginogenida, and Exogenida, are monophyletic and that the latter two clusters as sister clades. The newly sequenced P. florea forms sister branches with C. magna , while sequences of D. radians group with those from H. rotunda within Endogenida. The family Heliophryidae, which is comprised of only two genera, Heliophrya and Cyclophrya , was previously assigned to Evaginogenida. There is now sufficient evidence, however, that the type genus Heliophrya reproduces by endogenous budding, which corresponds to the definitive feature of Endogenida. In line with this and with the support of molecular phylogenetic analyses, we therefore transfer the family Heliophryidae with the type genus Heliophrya to Endogenida. The other genus, Cyclophrya , still remains in Evaginogenida because of its evaginative budding. Therefore, combined with morphological and phylogenetic analysis, Cyclophyidae are reactivated, and it belongs to Evaginogenida.

Ma, Mingzhen↗

An expanded role for the transcription factor WRINKLED1 in the biosynthesis of triacylglycerols during seed development

The transcription factor WRINKLED1 ( WRI1 ) is known as a master regulator of fatty acid synthesis in developing oilseeds of Arabidopsis thaliana and other species. WRI1 is known to directly stimulate the expression of many fatty acid biosynthetic enzymes and a few targets in the lower part of the glycolytic pathway. However, it remains unclear to what extent and how the conversion of sugars into fatty acid biosynthetic precursors is controlled by WRI 1. To shortlist possible gene targets for future in-planta experimental validation, here we present a strategy that combines phylogenetic foot printing of cis-regulatory elements with additional layers of evidence. Upstream regions of protein-encoding genes in A. thaliana were searched for the previously described DNA-binding consensus for WRI1, the ASML1/WRI1 (AW)-box. For about 900 genes, AW-box sites were found to be conserved across orthologous upstream regions in 11 related species of the crucifer family. For 145 select potential target genes identified this way, affinity of upstream AW-box sequences to WRI1 was assayed by Microscale Thermophoresis. This allowed definition of a refined WRI1 DNA-binding consensus. We find that known WRI1 gene targets are predictable with good confidence when upstream AW-sites are phylogenetically conserved, specifically binding WRI1 in the in vitro assay, positioned in proximity to the transcriptional start site, and if the gene is co-expressed with WRI1 during seed development. When targets predicted in this way are mapped to central metabolism, a conserved regulatory blueprint emerges that infers concerted control of contiguous pathway sections in glycolysis and fatty acid biosynthesis by WRI1. Several of the newly predicted targets are in the upper glycolysis pathway and the pentose phosphate pathway. Of these, plastidic isoforms of fructokinase ( FRK 3) and of phosphoglucose isomerase ( PGI 1) are particularly corroborated by previously reported seed phenotypes of respective null mutations.

59 BASIC BIOLOGICAL SCIENCES↗

Fire and Drought Affect Multiple Aspects of Diversity in a Migratory Bird Stopover Community

Drought and high-severity, stand-replacing wildfires can have substantial impacts on the composition of avian communities, including stop-over communities during migration. An inextricable link exists between drought and wildfire, each operating and impacting across different timescales. Many studies have found nonlinear avian abundance trends in breeding community time series data that include pre- and post-fire observations, describing an initial decrease in abundance followed by rapid increases that can attenuate over time. Here, we use a fall bird-banding dataset to evaluate shifts in a drought-impacted avian community following wildfire from taxonomic, functional, and phylogenetic perspectives. We looked at the community as a whole and also categorized birds as residents, migrants, and breeders to assess potential varying responses at the study site. We observed post-fire shifts in functional and phylogenetic diversity that corresponded to changes in vegetation. An influx of migratory insectivores post-fire drove much of the variation between pre- and post-fire avian communities and toward a more related, less phylogenetically dispersed community. A concurrent monsoon season drought was also associated with functional and phylogenetic diversity, highlighting the intertwined pulse press effects on avian communities. Overall, our results suggest that, although bird communities are immediately impacted by fire-driven resource changes, they can rebound over time, it is unclear how long-term drought may continue to shape the composition of these avian communities.

59 BASIC BIOLOGICAL SCIENCES↗

Higher-order structure of rRNA

A comparative search for phylogenetically covarying basepair replacements within potential helices has been the only reliable method to determine the correct secondary structure of the 3 rRNAs, 5S, 16S, and 23S. The analysis of 16S from a wide phylogenetic spectrum, that includes various branches of the eubacteria, archaebacteria, eucaryotes, in addition to the mitochondria and chloroplast, is beginning to reveal the constraints on the secondary structures of these rRNAs. Based on the success of this analysis, and the assumption that higher order structure will also be phylogenetically conserved, a comparative search was initiated for positions that show co-variation not involved in secondary structure helices. From a list of potential higher order interactions within 16S rRNA, two higher-order interactions are presented. The first of these interactions involves positions 570 and 866. Based on the extent of phylogenetic covariation between these positions while maintaining Watson-Crick pairing, this higher-order interaction is considered proven. The other interaction involves a minimum of six positions between the 1400 and 1500 regions of the 16S rRNA. Although these patterns of covariation are not as striking as the 570/866 interaction, the fact that they all exist in an anti-parallel fashion and that experimental methods previously implicated these two regions of the molecule in tRNA function suggests that these interactions be given serious consideration.

Gutell, R. R.↗

A detailed phylogeny for the Methanomicrobiales

The small subunit rRNA sequence of twenty archaea, members of the Methanomicrobiales, permits a detailed phylogenetic tree to be inferred for the group. The tree confirms earlier studies, based on far fewer sequences, in showing the group to be divided into two major clusters, temporarily designated the "methanosarcina" group and the "methanogenium" group. The tree also defines phylogenetic relationships within these two groups, which in some cases do not agree with the phylogenetic relationships implied by current taxonomic names--a problem most acute for the genus Methanogenium and its relatives. The present phylogenetic characterization provides the basis for a consistent taxonomic restructuring of this major methanogenic taxon.

NASA Discipline Number 52-30↗

The Molecular Ecology of Guerrero Negro: Justifying the Need for Environmental Genomics

The record of life on the only planet where it is known to exist is contained in the biogeochemical processes that organisms catalyze for their survival, in the compounds that they produce, and in their phylogenetic (evolutionary) relationships to each other. We manipulated sulfate and nutrient concentrations in intact microbial mats over periods of time up to a year. The objectives of the manipulations were: 1) characterize the diversity of process-associated functional genes; 2) understand environmental conditions leading to shifts in microbial guilds; 3) monitor/identify competitive responses of organisms sharing a metabolic niche. Characterization of functional genes associated with carbon (mcrA), nitrogen (nifH, nirK) and sulfur (dsrkB) cycling performed to date provided insight into the diversity and metabolic potential of the system; however, we only identified broad scale correlations between gene abundances and changes in mat physiology. For instance, increases in methane production by mats subjected to lowered sulfate and salinity concentrations were correlated with an observed increase in abundance of hydrogenotroph-like mcrA genes. However, due to low sequence similarity to any cultured isolates, phylogenetic associations only allow order level taxonomic commentary, preventing any associations being made on the cellular level. In each of the genes characterized from these experiments, a significant portion of sequences recovered show minimal phylogenetic affiliation to cultured organisms, preventing any understanding of inter-community dynamics and the functional capacities of these unknown organisms. Environmental genomics may provide insight into mat systems by allowing the correlation of functional genes with phylogenetic markers.

Smith, Jason M.↗

Methods for determining the genetic affinity of microorganisms and viruses

Selecting which sub-sequences in a database of nucleic acid such as 16S rRNA are highly characteristic of particular groupings of bacteria, microorganisms, fungi, etc. on a substantially phylogenetic tree. Also applicable to viruses comprising viral genomic RNA or DNA. A catalogue of highly characteristic sequences identified by this method is assembled to establish the genetic identity of an unknown organism. The characteristic sequences are used to design nucleic acid hybridization probes that include the characteristic sequence or its complement, or are derived from one or more characteristic sequences. A plurality of these characteristic sequences is used in hybridization to determine the phylogenetic tree position of the organism(s) in a sample. Those target organisms represented in the original sequence database and sufficient characteristic sequences can identify to the species or subspecies level. Oligonucleotide arrays of many probes are especially preferred. A hybridization signal can comprise fluorescence, chemiluminescence, or isotopic labeling, etc.; or sequences in a sample can be detected by direct means, e.g. mass spectrometry. The method's characteristic sequences can also be used to design specific PCR primers. The method uniquely identifies the phylogenetic affinity of an unknown organism without requiring prior knowledge of what is present in the sample. Even if the organism has not been previously encountered, the method still provides useful information about which phylogenetic tree bifurcation nodes encompass the organism.

Fox, George E.↗

Origin and evolution of HIV-1 subtype A6

Background: HIV outbreaks in the Former Soviet Union (FSU) countries were characterized by repeated transmission of the HIV variant AFSU, which is now classified as a distinct subtype A sub-subtype called A6. The current study used phylogenetic/phylodynamic and signature mutation analyses to determine likely evolutionary relationship between subtype A6 and other subtype A sub-subtypes. Methods: For this study, an initial Maximum Likelihood phylogenetic analysis was performed using a total of 553 full-length, publicly available, reverse transcriptase sequences, from A1, A2, A3, A4, A5, and A6 sub-subtypes of subtype A. For phylogenetic clustering and signature mutation analysis, a total of 5961 and 3959 pol and env sequences, respectively, were used. Results: Phylogenetic and signature mutation analysis showed that HIV-1 sub-subtype A6 likely originated from sub-subtype A1 of African origin. A6 and A1 pol and env genes shared several signature mutations that indicate genetic similarity between the two subtypes. For A6, tMRCA dated to 1975, 15 years later than that of A1. Conclusion: The current study provides insights into the evolution and diversification of A6 in the backdrop of FSU countries and indicates that A6 in FSU countries evolved from A1 of African origin and is getting bridged outside the FSU region.

60 APPLIED LIFE SCIENCES↗

Modest functional diversity decline and pronounced composition shifts of microbial communities in a mixed waste-contaminated aquifer

Background: Microbial taxonomic diversity declines with increased environmental stress. Yet, few studies have explored whether phylogenetic and functional diversities track taxonomic diversity along the stress gradient. Here, we investigated microbial communities within an aquifer in Oak Ridge, Tennessee, USA, which is characterized by a broad spectrum of stressors, including extremely high levels of nitrate, heavy metals like cadmium and chromium, radionuclides such as uranium, and extremely low pH (< 3). Results: Both taxonomic and phylogenetic α-diversities were reduced in the most impacted wells, while the decline in functional α-diversity was modest and statistically insignificant, indicating a more robust buffering capacity to environmental stress. Differences in functional gene composition (i.e., functional β-diversity) were pronounced in highly contaminated wells, while convergent functional gene composition was observed in uncontaminated wells. The relative abundances of most carbon degradation genes were decreased in contaminated wells, but genes associated with denitrification, adenylylsulfate reduction, and sulfite reduction were increased. Compared to taxonomic and phylogenetic compositions, environmental variables played a more significant role in shaping functional gene composition, suggesting that niche selection could be more closely related to microbial functionality than taxonomy. Conclusions: Overall, we demonstrated that despite a reduced taxonomic α-diversity, microbial communities under stress maintained functionality underpinned by environmental selection.

59 BASIC BIOLOGICAL SCIENCES↗

Convergent reductive evolution and host adaptation in Mycoavidus bacterial endosymbionts of Mortierellaceae fungi

Intimate associations between fungi and intracellular bacterial endosymbionts are becoming increasingly well understood. Phylogenetic analyses demonstrate that bacterial endosymbionts of Mucoromycota fungi are related either to free-living Burkholderia or Mollicutes species. The so-called Burkholderia-related endosymbionts or BRE comprise Mycoavidus, Mycetohabitans and Candidatus Glomeribacter gigasporarum. These endosymbionts are marked by genome contraction thought to be associated with intracellular selection. However, the conclusions drawn thus far are based on a very small subset of endosymbiont genomes, and the mechanisms leading to genome streamlining are not well understood. The purpose of this study was to better understand how intracellular existence shapes Mycoavidus and BRE functionally at the genome level. To this end we generated and analyzed 14 novel draft genomes for Mycoavidus living within the hyphae of Mortierellomycotina fungi. We found that our novel Mycoavidus genomes were significantly reduced compared to free-living Burkholderiales relatives. Using a genome-scale phylogenetic approach including the novel and available existing genomes of Mycoavidus, we show that the genus is an assemblage composed of two independently derived lineages including three well supported clades of Mycoavidus. Using a comparative genomic approach, we shed light on the functional implications of genome reduction, documenting shared and unique gene loss patterns between the three Mycoavidus clades. We found that many endosymbiont isolates demonstrate patterns of vertical transmission and host-specificity, but others are present in phylogenetically disparate hosts. We discuss how reductive evolution and host specificity reflect convergent adaptation to the intrahyphal selective landscape, and commonalities of eukaryotic endosymbiont genome evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

Codon Optimization Improves the Prediction of Xylose Metabolism from Gene Content in Budding Yeasts

Xylose is the second most abundant monomeric sugar in plant biomass. Consequently, xylose catabolism is an ecologically important trait for saprotrophic organisms, as well as a fundamentally important trait for industries that hope to convert plant mass to renewable fuels and other bioproducts using microbial metabolism. Although common across fungi, xylose catabolism is rare within Saccharomycotina, the subphylum that contains most industrially relevant fermentative yeast species. The genomes of several yeasts unable to consume xylose have been previously reported to contain the full set of genes in the XYL pathway, suggesting the absence of a gene–trait correlation for xylose metabolism. Here, we measured growth on xylose and systematically identified XYL pathway orthologs across the genomes of 332 budding yeast species. Although the XYL pathway coevolved with xylose metabolism, we found that pathway presence only predicted xylose catabolism about half of the time, demonstrating that a complete XYL pathway is necessary, but not sufficient, for xylose catabolism. We also found that XYL1 copy number was positively correlated, after phylogenetic correction, with xylose utilization. We then quantified codon usage bias of XYL genes and found that XYL3 codon optimization was significantly higher, after phylogenetic correction, in species able to consume xylose. Finally, we showed that codon optimization of XYL2 was positively correlated, after phylogenetic correction, with growth rates in xylose medium. We conclude that gene content alone is a weak predictor of xylose metabolism and that using codon optimization enhances the prediction of xylose metabolism from yeast genome sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

Substitution Models of Protein Evolution with Selection on Enzymatic Activity

Abstract Substitution models of evolution are necessary for diverse evolutionary analyses including phylogenetic tree and ancestral sequence reconstructions. At the protein level, empirical substitution models are traditionally used due to their simplicity, but they ignore the variability of substitution patterns among protein sites. Next, in order to improve the realism of the modeling of protein evolution, a series of structurally constrained substitution models were presented, but still they usually ignore constraints on the protein activity. Here, we present a substitution model of protein evolution with selection on both protein structure and enzymatic activity, and that can be applied to phylogenetics. In particular, the model considers the binding affinity of the enzyme–substrate complex as well as structural constraints that include the flexibility of structural flaps, hydrogen bonds, amino acids backbone radius of gyration, and solvent-accessible surface area that are quantified through molecular dynamics simulations. We applied the model to the HIV-1 protease and evaluated it by phylogenetic likelihood in comparison with the best-fitting empirical substitution model and a structurally constrained substitution model that ignores the enzymatic activity. We found that accounting for selection on the protein activity improves the fitting of the modeled functional regions with the real observations, especially in data with high molecular identity, which recommends considering constraints on the protein activity in the development of substitution models of evolution.

Ferreiro, David↗

CasPEDIA Database: a functional classification system for class 2 CRISPR-Cas enzymes

Abstract CRISPR-Cas enzymes enable RNA-guided bacterial immunity and are widely used for biotechnological applications including genome editing. In particular, the Class 2 CRISPR-associated enzymes (Cas9, Cas12 and Cas13 families), have been deployed for numerous research, clinical and agricultural applications. However, the immense genetic and biochemical diversity of these proteins in the public domain poses a barrier for researchers seeking to leverage their activities. We present CasPEDIA (http://caspedia.org), the Cas Protein Effector Database of Information and Assessment, a curated encyclopedia that integrates enzymatic classification for hundreds of different Cas enzymes across 27 phylogenetic groups spanning the Cas9, Cas12 and Cas13 families, as well as evolutionarily related IscB and TnpB proteins. All enzymes in CasPEDIA were annotated with a standard workflow based on their primary nuclease activity, target requirements and guide-RNA design constraints. Our functional classification scheme, CasID, is described alongside current phylogenetic classification, allowing users to search related orthologs by enzymatic function and sequence similarity. CasPEDIA is a comprehensive data portal that summarizes and contextualizes enzymatic properties of widely used Cas enzymes, equipping users with valuable resources to foster biotechnological development. CasPEDIA complements phylogenetic Cas nomenclature and enables researchers to leverage the multi-faceted nucleic-acid targeting rules of diverse Class 2 Cas enzymes.

59 BASIC BIOLOGICAL SCIENCES↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

The Influence of the Number of Tree Searches on Maximum Likelihood Inference in Phylogenomics

Maximum likelihood (ML) phylogenetic inference is widely used in phylogenomics. As heuristic searches most likely find suboptimal trees, it is recommended to conduct multiple (e.g., 10) tree searches in phylogenetic analyses. However, beyond its positive role, how and to what extent multiple tree searches aid ML phylogenetic inference remains poorly explored. Here, we found that a random starting tree was not as effective as the BioNJ and parsimony starting trees in inferring the ML gene tree and that RAxML-NG and PhyML were less sensitive to different starting trees than IQ-TREE. We then examined the effect of the number of tree searches on ML tree inference with IQ-TREE and RAxML-NG, by running 100 tree searches on 19,414 gene alignments from 15 animal, plant, and fungal phylogenomic datasets. We found that the number of tree searches substantially impacted the recovery of the best-of-100 ML gene tree topology among 100 searches for a given ML program. In addition, all of the concatenation-based trees were topologically identical if the number of tree searches was ≥10. Quartet-based ASTRAL trees inferred from 1 to 80 tree searches differed topologically from those inferred from 100 tree searches for 6/15 phylogenomic datasets. Lastly, our simulations showed that gene alignments with lower difficulty scores had a higher chance of finding the best-of-100 gene tree topology and were more likely to yield the correct trees.

59 BASIC BIOLOGICAL SCIENCES↗