Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Radiation-induced genomic instability: radiation quality and dose response

Genomic instability is a term used to describe a phenomenon that results in the accumulation of multiple changes required to convert a stable genome of a normal cell to an unstable genome characteristic of a tumor. There has been considerable recent debate concerning the importance of genomic instability in human cancer and its temporal occurrence in the carcinogenic process. Radiation is capable of inducing genomic instability in mammalian cells and instability is thought to be the driving force responsible for radiation carcinogenesis. Genomic instability is characterized by a large collection of diverse endpoints that include large-scale chromosomal rearrangements and aberrations, amplification of genetic material, aneuploidy, micronucleus formation, microsatellite instability, and gene mutation. The capacity of radiation to induce genomic instability depends to a large extent on radiation quality or linear energy transfer (LET) and dose. There appears to be a low dose threshold effect with low LET, beyond which no additional genomic instability is induced. Low doses of both high and low LET radiation are capable of inducing this phenomenon. This report reviews data concerning dose rate effects of high and low LET radiation and their capacity to induce genomic instability assayed by chromosomal aberrations, delayed lethal mutations, micronuclei and apoptosis.

Review↗

Improved Mobilome Delineation in Fragmented Genomes

The mobilome of a microbe, i.e., its set of mobile elements, has major effects on its ecology, and is important to delineate properly in each genome. This becomes more challenging for incomplete genomes, and even more so for metagenome-assembled genomes (MAGs), where misbinning of scaffolds and other losses can occur. Genomic islands (GIs), which integrate into the host chromosome, are a major component of the mobilome. Our GI-detection software TIGER, unique in its precise mapping of GI termini, was applied to 74,561 genomes from 2,473 microbial species, each species containing at least one MAG and one isolate genome. A species-normalized deficit of ∼1.6 GIs/genome was measured for MAGs relative to isolates. To test whether this undercount was due to the higher fragmentation of MAG genomes, TIGER was updated to enable detection of split GIs whose termini are on separate scaffolds or that wrap around the origin of a circular replicon. This doubled GI yields, and the new split GIs matched the quality of single-scaffold GIs, except that highly fragmented GIs may lack central portions. Cross-scaffold search is an important upgrade to GI detection as fragmented genomes increasingly dominate public databases. TIGER2 better captures MAG microdiversity, recovering niche-defining GIs and supporting microbiome research aims such as virus-host linking and ecological assessment.

54 ENVIRONMENTAL SCIENCES↗

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES↗

What Is in Umbilicaria pustulata? A Metagenomic Approach to Reconstruct the Holo-Genome of a Lichen

Lichens are valuable models in symbiosis research and promising sources of biosynthetic genes for biotechnological applications. Most lichenized fungi grow slowly, resist aposymbiotic cultivation, and are poor candidates for experimentation. Obtaining contiguous, high-quality genomes for such symbiotic communities is technically challenging. Here, we present the first assembly of a lichen holo-genome from metagenomic whole-genome shotgun data comprising both PacBio long reads and Illumina short reads. The nuclear genomes of the two primary components of the lichen symbiosis—the fungus Umbilicaria pustulata (33 Mb) and the green alga Trebouxia sp. (53 Mb)—were assembled at contiguities comparable to single-species assemblies. The analysis of the read coverage pattern revealed a relative abundance of fungal to algal nuclei of ~20:1. Gap-free, circular sequences for all organellar genomes were obtained. The bacterial community is dominated by Acidobacteriaceae and encompasses strains closely related to bacteria isolated from other lichens. Gene set analyses showed no evidence of horizontal gene transfer from algae or bacteria into the fungal genome. Our data suggest a lineage-specific loss of a putative gibberellin-20-oxidase in the fungus, a gene fusion in the fungal mitochondrion, and a relocation of an algal chloroplast gene to the algal nucleus. Major technical obstacles during reconstruction of the holo-genome were coverage differences among individual genomes surpassing three orders of magnitude. Moreover, we show that GC-rich inverted repeats paired with nonrandom sequencing error in PacBio data can result in missing gene predictions. This likely poses a general problem for genome assemblies based on long reads.

54 ENVIRONMENTAL SCIENCES↗

Horizontal Gene Transfer to a Defensive Symbiont with a Reduced Genome in a Multipartite Beetle Microbiome

Symbiotic mutualisms of bacteria and animals are ubiquitous in nature, running a continuum from facultative to obligate from the perspectives of both partners. The loss of functions required for living independently but not within a host gives rise to reduced genomes in many symbionts. Although the phenomenon of genome reduction can be explained by existing evolutionary models, the initiation of the process is not well understood. Here, we describe the microbiome associated with the eggs of the beetle Lagria villosa, consisting of multiple bacterial symbionts related to Burkholderia gladioli, including a reduced-genome symbiont thought to be the exclusive producer of the defensive compound lagriamide. We show that the putative lagriamide-producing symbiont is the only member of the microbiome undergoing genome reduction and that it has already lost the majority of its primary metabolism and DNA repair pathways. The key step preceding genome reduction in the symbiont was likely the horizontal acquisition of the putative lagriamide lga biosynthetic gene cluster. Unexpectedly, we uncovered evidence of additional horizontal transfers to the symbiont’s genome while genome reduction was occurring and despite a current lack of genes needed for homologous recombination. These gene gains may have given the genome-reduced symbiont a selective advantage in the microbiome, especially given the maintenance of the large lga gene cluster despite ongoing genome reduction.

59 BASIC BIOLOGICAL SCIENCES↗

Supervised extraction of near-complete genomes from metagenomic samples: A new service in PATRIC

Large amounts of metagenomically-derived data are submitted to PATRIC for analysis. In the future, we expect even more jobs submitted to PATRIC will use metagenomic data. One in-demand use case is the extraction of near-complete draft genomes from assembled contigs of metagenomic origin. The PATRIC metagenome binning service utilizes the PATRIC database to furnish a large, diverse set of reference genomes. We provide a new service for supervised extraction and annotation of high-quality, near-complete genomes from metagenomically-derived contigs. Reference genomes are assigned to putative draft genome bins based on the presence of single-copy universal marker roles in the sample, and contigs are sorted into these bins by their similarity to reference genomes in PATRIC. Each set of binned contigs represents a draft genome that will be annotated by RASTtk in PATRIC. A structured-language binning report is provided containing quality measurements and taxonomic information about the contig bins. The PATRIC metagenome binning service emphasizes extraction of high-quality genomes for downstream analysis using other PATRIC tools and services. Due to its supervised nature, the binning service is not appropriate for mining novel or extremely low-coverage genomes from metagenomic samples.

59 BASIC BIOLOGICAL SCIENCES↗

microTrait: A Toolset for a Trait-Based Representation of Microbial Genomes

Remote sensing approaches have revolutionized the study of macroorganisms, allowing theories of population and community ecology to be tested across increasingly larger scales without much compromise in resolution of biological complexity. In microbial ecology, our remote window into the ecology of microorganisms is through the lens of genome sequencing. For microbial organisms, recent evidence from genomes recovered from metagenomic samples corroborate a highly complex view of their metabolic diversity and other associated traits which map into high physiological complexity. Regardless, during the first decades of this omics era, microbial ecological research has primarily focused on taxa and functional genes as ecological units, favoring breadth of coverage over resolution of biological complexity manifested as physiological diversity. Recently, the rate at which provisional draft genomes are generated has increased substantially, giving new insights into ecological processes and interactions. From a genotype perspective, the wide availability of genome-centric data requires new data synthesis approaches that place organismal genomes center stage in the study of environmental roles and functional performance. Extraction of ecologically relevant traits from microbial genomes will be essential to the future of microbial ecological research. Here, we present microTrait , a computational pipeline that infers and distills ecologically relevant traits from microbial genome sequences. microTrait maps a genome sequence into a trait space, including discrete and continuous traits, as well as simple and composite. Traits are inferred from genes and pathways representing energetic, resource acquisition, and stress tolerance mechanisms, while genome-wide signatures are used to infer composite, or life history, traits of microorganisms. This approach is extensible to any microbial habitat, although we provide initial examples of this approach with reference to soil microbiomes.

Karaoz, Ulas↗

Genomic Diversity Evaluation of Populus trichocarpa Germplasm for Rare Variant Genetic Association Studies

Genome-wide association studies are powerful tools to elucidate the genome-to-phenome relationship. In order to explain most of the observed heritability of a phenotypic trait, a sufficient number of individuals and a large set of genetic variants must be examined. The development of high-throughput technologies and cost-efficient resequencing of complete genomes have enabled the genome-wide identification of genetic variation at large scale. As such, almost all existing genetic variation becomes available, and it is now possible to identify rare genetic variants in a population sample. Rare genetic variants that were usually filtered out in most genetic association studies are the most numerous genetic variations across genomes and hold great potential to explain a significant part of the missing heritability observed in association studies. Rare genetic variants must be identified with high confidence, as they can easily be confounded with sequencing errors. In this study, we used a pre-filtered data set of 1,014 pure Populus trichocarpa entire genomes to identify rare and common small genetic variants across individual genomes. We compared variant calls between Platypus and HaplotypeCaller pipelines, and we further applied strict quality filters for improved genetic variant identification. Finally, we only retained genetic variants that were identified by both variant callers increasing calling confidence. Based on these shared variants and after stringent quality filtering, we found high genomic diversity in P. trichocarpa germplasm, with 7.4 million small genetic variants. Importantly, 377k non-synonymous variants (5% of the total) were uncovered. We highlight the importance of genomic diversity and the potential of rare defective genetic variants in explaining a significant portion of P. trichocarpa's phenotypic variability in association genetics. The ultimate goal is to associate both rare and common alleles with poplar's wood quality traits to support selective breeding for an improved bioenergy feedstock.

Genetics & Heredity↗

Whole genome sequencing of Mycobacterium bovis directly from clinical tissue samples without culture

Advancement in next generation sequencing offers the possibility of routine use of whole genome sequencing (WGS) for Mycobacterium bovis (M. bovis) genomes in clinical reference laboratories. To date, the M. bovis genome could only be sequenced if the mycobacteria were cultured from tissue. This requirement for culture has been due to the overwhelmingly large amount of host DNA present when DNA is prepared directly from a granuloma. To overcome this formidable hurdle, we evaluated the usefulness of an RNA-based targeted enrichment method to sequence M. bovis DNA directly from tissue samples without culture. Initial spiking experiments for method development were established by spiking DNA extracted from tissue samples with serially diluted M. bovis BCG DNA at the following concentration range: 0.1 ng/μl to 0.1 pg/μl (10 –1 to 10 –4 ). Library preparation, hybridization and enrichment was performed using SureSelect custom capture library RNA baits and the SureSelect XT HS2 target enrichment system for Illumina paired-end sequencing. The method validation was then assessed using direct WGS of M. bovis DNA extracted from tissue samples from naturally (n = 6) and experimentally (n = 6) infected animals with variable Ct values. Direct WGS of spiked DNA samples achieved 99.1% mean genome coverage (mean depth of coverage: 108×) and 98.8% mean genome coverage (mean depth of coverage: 26.4×) for tissue samples spiked with BCG DNA at 10 –1 (mean Ct value: 20.3) and 10 –2 (mean Ct value: 23.4), respectively. The M. bovis genome from the experimentally and naturally infected tissue samples was successfully sequenced with a mean genome coverage of 99.56% and depth of genome coverage ranging from 9.2× to 72.1×. The spoligoyping and M. bovis group assignment derived from sequencing DNA directly from the infected tissue samples matched that of the cultured isolates from the same sample. Our results show that direct sequencing of M. bovis DNA from tissue samples has the potential to provide accurate sequencing of M. bovis genomes significantly faster than WGS from cultures in research and diagnostic settings.

59 BASIC BIOLOGICAL SCIENCES↗

A chromosome-level genome assembly of the Chinese cork oak (Quercus variabilis)

Quercus variabilis (Fagaceae) is an ecologically and economically important deciduous broadleaved tree species native to and widespread in East Asia. It is a valuable woody species and an indicator of local forest health, and occupies a dominant position in forest ecosystems in East Asia. However, genomic resources from Q. variabilis are still lacking. Here, we present a high-quality Q. variabilis genome generated by PacBio HiFi and Hi-C sequencing. The assembled genome size is 787 Mb, with a contig N50 of 26.04 Mb and scaffold N50 of 64.86 Mb, comprising 12 pseudo-chromosomes. The repetitive sequences constitute 67.6% of the genome, of which the majority are long terminal repeats, accounting for 46.62% of the genome. We used ab initio , RNA sequence-based and homology-based predictions to identify protein-coding genes. A total of 32,466 protein-coding genes were identified, of which 95.11% could be functionally annotated. Evolutionary analysis showed that Q. variabilis was more closely related to Q. suber than to Q. lobata or Q. robur. We found no evidence for species-specific whole genome duplications in Quercus after the species had diverged. This study provides the first genome assembly and the first gene annotation data for Q. variabilis. These resources will inform the design of further breeding strategies, and will be valuable in the study of genome editing and comparative genomics in oak species.

Han, Biao↗

Lake Arrowhead Microbial Genomics Conference

The Lake Arrowhead Microbial Genomics Conference (formerly the E. coli and Small Genomes Conference, and the Microbial Genomes Conference) was held September 11-15, 2022, at the UCLA Conference Center at Lake Arrowhead, California. This conference is part of a yearly meeting initiated in 1991 to bring together genome sequencers, bioinformatics specialists, biologists, and geneticists, to forge interactions that would result in meaningful functional genomics. The goal of the meeting was to translate the influx of new genome sequencing information into useful biological studies. The Lake Arrowhead 2022 meeting had a major focus on microbial communities, the human microbiome, pathogens, phage therapy, bioinformatic methods of analysis, bioenergenics, and environmental microbiology. The field of genomics has reached the point where deriving the sequence of an organism’s entire genome is now seen as a beginning rather than an endpoint. Moreover, understanding the sequences of whole communities, and particularly those that make up different microbiomes, is now a central point of the field. The 2022 meeting covered microorganisms for which extensive analyses exist, and those for which new biological and technical strategies are being developed. The focus on biodiversity, the human microbiome, pathogenic organisms and methods of countering them, and bioenergetics added special significance to this meeting. This meeting was designed to have a mix of 36 invited presentations and poster sessions with more than 80 posters, and had 147 participants. This meeting has become established as a major annual microbial genomics conference during the last 29 years and has assured status and quality. It offers a major opportunity for young investigators and students to attend, present posters, and give talks. In fact, 50% of the invited speakers were early career stage scientists. An important impact of the conference was the forging of new collaborations between scientists and particularly young scientists with the array of multidisciplinary researchers at the meeting.

Miller, Jeffrey H.↗

An evaluation of methodology to determine algal genome completeness

The advancement of sequencing technologies has resulted in a rapid expansion of genome sequencing. One challenge in processing and analyzing this abundance of genomics data is development and application of tools to accurately assess the quality of novel genomes. The quality of a given genome assembly is commonly assessed using the presence of conserved orthologs. However, the applicability and accuracy of identifying conserved orthologs applied to algal genome assemblies, which are uniquely diverse, is unclear. Here, in this study, we analyze the utility of a genome content assessment tool, Benchmarking Universal Single-Copy Orthologs (BUSCO), to evaluate algal genome assembly completeness. We find that Chlorophyta and Stramenopile BUSCO databases are effective tools to analyze genome completeness of these lineages. For other lineages, the Eukaryota database should be utilized until additional lineage-specific databases are developed. Based on these results, suggested best practices for current usage and necessary future improvements are provided.

59 BASIC BIOLOGICAL SCIENCES↗

The complex polyploid genome architecture of sugarcane

Sugarcane, the world’s most harvested crop by tonnage, has shaped global history, trade and geopolitics, and is currently responsible for 80% of sugar production worldwide. While traditional sugarcane breeding methods have effectively generated cultivars adapted to new environments and pathogens, sugar yield improvements have recently plateaued. The cessation of yield gains may be due to limited genetic diversity within breeding populations, long breeding cycles and the complexity of its genome, the latter preventing breeders from taking advantage of the recent explosion of whole-genome sequencing that has benefited many other crops. Thus, modern sugarcane hybrids are the last remaining major crop without a reference-quality genome. Here we take a major step towards advancing sugarcane biotechnology by generating a polyploid reference genome for R570, a typical modern cultivar derived from interspecific hybridization between the domesticated species (Saccharum officinarum) and the wild species (Saccharum spontaneum). In contrast to the existing single haplotype (‘monoploid’) representation of R570, our 8.7 billion base assembly contains a complete representation of unique DNA sequences across the approximately 12 chromosome copies in this polyploid genome. Using this highly contiguous genome assembly, we filled a previously unsized gap within an R570 physical genetic map to describe the likely causal genes underlying the single-copy Bru1 brown rust resistance locus. This polyploid genome assembly with fine-grain descriptions of genome architecture and molecular targets for biotechnology will help accelerate molecular and transgenic breeding and adaptation of sugarcane to future environmental conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

CoreCruncher : Fast and Robust Construction of Core Genomes in Large Prokaryotic Data Sets

The core genome represents the set of genes shared by all, or nearly all, strains of a given population or species of prokaryotes. Inferring the core genome is integral to many genomic analyses, however, most methods rely on the comparison of all the pairs of genomes; a step that is becoming increasingly difficult given the massive accumulation of genomic data. Here, we present CoreCruncher; a program that robustly and rapidly constructs core genomes across hundreds or thousands of genomes. CoreCruncher does not compute all pairwise genome comparisons and uses a heuristic based on the distributions of identity scores to classify sequences as orthologs or paralogs/xenologs. Although it is much faster than current methods, our results indicate that our approach is more conservative than other tools and less sensitive to the presence of paralogs and xenologs. CoreCruncher is freely available from: https://github.com/lbobay/CoreCruncher. CoreCruncher is written in Python 3.7 and can also run on Python 2.7 without modification. It requires the python library Numpy and either Usearch or Blast. Certain options require the programs muscle or mafft.

59 BASIC BIOLOGICAL SCIENCES↗

A Rhodopseudomonas strain with a substantially smaller genome retains the core metabolic versatility of its genus

ABSTRACT Rhodopseudomonas are a group of phototrophic microbes with a marked metabolic versatility and flexibility that underpins their potential use in the production of value-added products, bioremediation, and plant growth promotion. Members of this group have an average genome size of about 5.5 Mb, but two closely related strains have genome sizes of about 4.0 Mb. To identify the types of genes missing in a reduced genome strain, we compared strain DSM127 with other Rhodopseudomonas isolates at the genomic and phenotypic levels. We found that DSM127 can grow as well as other members of the Rhodopseudomonas genus and retains most of their metabolic versatility, but it has many fewer genes associated with high-affinity transport of nutrients, iron uptake, nitrogen metabolism, and biodegradation of aromatic compounds. This analysis indicates genes that can be deleted in genome reduction campaigns and suggests that DSM127 could be a favorable choice for biotechnology applications using Rhodopseudomonas or as a strain that can be engineered further to reside in a specialized natural environment. IMPORTANCE Rhodopseudomonas are a cohort of phototrophic bacteria with broad metabolic versatility. Members of this group are present in diverse soil and water environments, and some strains are found associated with plants and have plant growth-promoting activity. Motivated by the idea that it may be possible to design bacteria with reduced genomes that can survive well only in a specific environment or that may be more metabolically efficient, we compared Rhodopseudomonas strains with typical genome sizes of about 5.5 Mb to a strain with a reduced genome size of 4.0 Mb. From this, we concluded that metabolic versatility is part of the identity of the Rhodopseudomonas group, but high-affinity transport genes and genes of apparent redundant function can be dispensed with.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying intragenic functional modules of genomic variations associated with cancer phenotypes by learning representation of association networks

Background Genome-wide Association Studies (GWAS) aims to uncover the link between genomic variation and phenotype. They have been actively applied in cancer biology to investigate associations between variations and cancer phenotypes, such as susceptibility to certain types of cancer and predisposed responsiveness to specific treatments. Since GWAS primarily focuses on finding associations between individual genomic variations and cancer phenotypes, there are limitations in understanding the mechanisms by which cancer phenotypes are cooperatively affected by more than one genomic variation. Results This paper proposes a network representation learning approach to learn associations among genomic variations using a prostate cancer cohort. The learned associations are encoded into representations that can be used to identify functional modules of genomic variations within genes associated with early- and late-onset prostate cancer. The proposed method was applied to a prostate cancer cohort provided by the Veterans Administration’s Million Veteran Program to identify candidates for functional modules associated with early-onset prostate cancer. The cohort included 33,159 prostate cancer patients, 3181 early-onset patients, and 29,978 late-onset patients. The reproducibility of the proposed approach clearly showed that the proposed approach can improve the model performance in terms of robustness. Conclusions To our knowledge, this is the first attempt to use a network representation learning approach to learn associations among genomic variations within genes. Associations learned in this way can lead to an understanding of the underlying mechanisms of how genomic variations cooperatively affect each cancer phenotype. This method can reveal unknown knowledge in the field of cancer biology and can be utilized to design more advanced cancer-targeted therapies.

60 APPLIED LIFE SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗