Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “phylogenetics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Genomic analysis of 1710 surveillance-based Neisseria gonorrhoeae isolates from the USA in 2019 identifies predominant strain types and chromosomal antimicrobial-resistance determinants

This study characterized high-quality whole-genome sequences of a sentinel, surveillance-based collection of 1710 Neisseria gonorrhoeae (GC) isolates from 2019 collected in the USA as part of the Gonococcal Isolate Surveillance Project (GISP). It aims to provide a detailed report of strain diversity, phylogenetic relationships and resistance determinant profiles associated with reduced susceptibilities to antibiotics of concern. The 1710 isolates represented 164 multilocus sequence types and 21 predominant phylogenetic clades. Common genomic determinants defined most strains’ phenotypic, reduced susceptibility to current and historic antibiotics (e.g. bla TEM plasmid for penicillin, tetM plasmid for tetracycline, gyrA for ciprofloxacin, 23S rRNA and/or mosaic mtr operon for azithromycin, and mosaic penA for cefixime and ceftriaxone). The most predominant phylogenetic clade accounted for 21 % of the isolates, included a majority of the isolates with low-level elevated MICs to azithromycin (2.0 µg ml –1 ), carried a mosaic mtr operon and variants in PorB, and showed expansion with respect to data previously reported from 2018. The second largest clade predominantly carried the GyrA S91F variant, was largely ciprofloxacin resistant (MIC ≥1.0 µg ml –1 ), and showed significant expansion with respect to 2018. Overall, a low proportion of isolates had medium- to high-level elevated MIC to azithromycin ((≥4.0 µg ml –1 ), based on C2611T or A2059G 23S rRNA variants). One isolate carried the penA 60.001 allele resulting in elevated MICs to cefixime and ceftriaxone of 1.0 µg ml –1 . This high-resolution snapshot of genetic profiles of 1710 GC sequences, through a comparison with 2018 data (1479 GC sequences) within the sentinel system, highlights change in proportions and expansion of select GC strains and the associated genetic mechanisms of resistance. The knowledge gained through molecular surveillance may support rapid identification of outbreaks of concern. Continued monitoring may inform public health responses to limit the development and spread of antibiotic-resistant gonorrhoea.

(3-6) antimicrobial resistance↗

The 1.3 Å resolution structure of the truncated group Ia type IV pilin from Pseudomonas aeruginosa strain P1

The type IV pilus is a diverse molecular machine capable of conferring a variety of functions and is produced by a wide range of bacterial species. The ability of the pilus to perform host-cell adherence makes it a viable target for the development of vaccines against infection by human pathogens such as Pseudomonas aeruginosa . Here, the 1.3 Å resolution crystal structure of the N-terminally truncated type IV pilin from P. aeruginosa strain P1 (ΔP1) is reported, the first structure of its phylogenetically linked group (group I) to be discussed in the literature. The structure was solved from X-ray diffraction data that were collected 20 years ago with a molecular-replacement search model generated using AlphaFold ; the effectiveness of other search models was analyzed. Examination of the high-resolution ΔP1 structure revealed a solvent network that aids in maintaining the fold of the protein. On comparing the sequence and structure of P1 with a variety of type IV pilins, it was observed that there are cases of higher structural similarities between the phylogenetic groups of P. aeruginosa than there are between the same phylogenetic group, indicating that a structural grouping of pilins may be necessary in developing antivirulence drugs and vaccines. These analyses also identified the α–β loop as the most structurally diverse domain of the pilins, which could allow it to serve a role in pilus recognition. Studies of ΔP1 in vitro polymerization demonstrate that the optimal hydrophobic catalyst for the oligomerization of the pilus from strain K122 is not conducive for pilus formation of ΔP1; a model of a three-start helical assembly using the ΔP1 structure indicates that the α–β loop and the D-loop prevent in vitro polymerization.

Bragagnolo, Nicholas↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Characterization of a novel aromatic substrate-processing microcompartment in Actinobacteria

ABSTRACT We have discovered a new cluster of genes that is found exclusively in the Actinobacteria phylum. This locus includes genes for the 2-aminophenol meta -cleavage pathway and the shell proteins of a bacterial microcompartment (BMC) and has been named aromatics (ARO) for its putative role in the breakdown of aromatic compounds. In this study, we provide details about the distribution and composition of the ARO BMC locus and conduct phylogenetic, structural, and functional analyses of the first two enzymes in the catabolic pathway: a unique 2-aminophenol dioxygenase, which is exclusively found alongside BMC shell genes in Actinobacteria, and a semialdehyde dehydrogenase, which works downstream of the dioxygenase. Genomic analysis reveals variations in the complexity of the ARO loci across different orders. Some loci are simple, containing shell proteins and enzymes for the initial steps of the catabolic pathway, while others are extensive, encompassing all the necessary genes for the complete breakdown of 2-aminophenol into pyruvate and acetyl-CoA. Furthermore, our analysis uncovers two subtypes of ARO BMC that likely degrade either 2-aminophenol or catechol, depending on the presence of a pathway-specific gene within the ARO locus. The precise precursor of 2-aminophenol, which serves as the initial substrate and/or inducer for the ARO pathway, remains unknown, as our model organism Micromonospora rosaria cannot utilize 2-aminophenol as its sole energy source. However, using enzymatic assays, we demonstrate the dioxygenase’s ability to cleave both 2-aminophenol and catechol in vitro , in collaboration with the aldehyde dehydrogenase, to facilitate the rapid conversion of these unstable and toxic intermediates. IMPORTANCE Bacterial microcompartments (BMCs) are proteinaceous organelles that are widespread among bacteria and provide a competitive advantage in specific environmental niches. Studies have shown that the genetic information necessary to form functional BMCs is encoded in loci that contain genes encoding shell proteins and the enzymatic core. This allows the bioinformatic discovery of BMCs with novel functions and expands our understanding of the metabolic diversity of BMCs. ARO loci, found only in Actinobacteria, contain genes encoding for phylogenetically remote shell proteins and homologs of the meta -cleavage degradation pathway enzymes that were shown to convert central aromatic intermediates into pyruvate and acetyl-CoA in gamma Proteobacteria. By analyzing the gene composition of ARO BMC loci and characterizing two core enzymes phylogenetically, structurally, and functionally, we provide an initial functional characterization of the ARO BMC, the most unusual BMC identified to date, distinctive among the repertoire of studied BMCs.

2-AP 1,6-dioxygenase↗

Genomic characterization of three marine fungi, including Emericellopsis atlantica sp. nov. with signatures of a generalist lifestyle and marine biomass degradation

ABSTRACT Marine fungi remain poorly covered in global genome sequencing campaigns; the 1000 fungal genomes (1KFG) project attempts to shed light on the diversity, ecology and potential industrial use of overlooked and poorly resolved fungal taxa. This study characterizes the genomes of three marine fungi: Emericellopsis sp. TS7, wood-associated Amylocarpus encephaloides and algae-associated Calycina marina. These species were genome sequenced to study their genomic features, biosynthetic potential and phylogenetic placement using multilocus data. Amylocarpus encephaloides and C. marina were placed in the Helotiaceae and Pezizellaceae (Helotiales) , respectively, based on a 15-gene phylogenetic analysis. These two genomes had fewer biosynthetic gene clusters (BGCs) and carbohydrate active enzymes (CAZymes) than Emericellopsis sp. TS7 isolate. Emericellopsis sp. TS7 ( Hypocreales , Ascomycota ) was isolated from the sponge Stelletta normani . A six-gene phylogenetic analysis placed the isolate in the marine Emericellopsis clade and morphological examination confirmed that the isolate represents a new species, which is described here as E. atlantica . Analysis of its CAZyme repertoire and a culturing experiment on three marine and one terrestrial substrates indicated that E. atlantica is a psychrotrophic generalist fungus that is able to degrade several types of marine biomass. FungiSMASH analysis revealed the presence of 35 BGCs including, eight non-ribosomal peptide synthases (NRPSs), six NRPS-like, six polyketide synthases, nine terpenes and six hybrid, mixed or other clusters. Of these BGCs, only five were homologous with characterized BGCs. The presence of unknown BGCs sets and large CAZyme repertoire set stage for further investigations of E. atlantica . The Pezizellaceae genome and the genome of the monotypic Amylocarpus genus represent the first published genomes of filamentous fungi that are restricted in their occurrence to the marine habitat and form thus a valuable resource for the community that can be used in studying ecological adaptions of fungi using comparative genomics.

59 BASIC BIOLOGICAL SCIENCES↗

A genome-informed higher rank classification of the biotechnologically important fungal subphylum Saccharomycotina

The subphylum Saccharomycotina is a lineage in the fungal phylum Ascomycota that exhibits levels of genomic diversity similar to those of plants and animals. The Saccharomycotina consist of more than 1 200 known species currently divided into 16 families, one order, and one class. Species in this subphylum are ecologically and metabolically diverse and include important opportunistic human pathogens, as well as species important in biotechnological applications. Many traits of biotechnological interest are found in closely related species and often restricted to single phylogenetic clades. However, the biotechnological potential of most yeast species remains unexplored. Although the subphylum Saccharomycotina has much higher rates of genome sequence evolution than its sister subphylum, Pezizomycotina, it contains only one class compared to the 16 classes in Pezizomycotina. The third subphylum of Ascomycota, the Taphrinomycotina, consists of six classes and has approximately 10 times fewer species than the Saccharomycotina. These data indicate that the current classification of all these yeasts into a single class and a single order is an underappreciation of their diversity. Our previous genome-scale phylogenetic analyses showed that the Saccharomycotina contains 12 major and robustly supported phylogenetic clades; seven of these are current families (Lipomycetaceae, Trigonopsidaceae, Alloascoideaceae, Pichiaceae, Phaffomycetaceae, Saccharomycodaceae, and Saccharomycetaceae), one comprises two current families (Dipodascaceae and Trichomonascaceae), one represents the genus Sporopachydermia, and three represent lineages that differ in their translation of the CUG codon (CUG-Ala, CUG-Ser1, and CUG-Ser2). Using these analyses in combination with relative evolutionary divergence and genome content analyses, we propose an updated classification for the Saccharomycotina, including seven classes and 12 orders that can be diagnosed by genome content. This updated classification is consistent with the high levels of genomic diversity within this subphylum and is necessary to make the higher rank classification of the Saccharomycotina more comparable to that of other fungi, as well as to communicate efficiently on lineages that are not yet formally named.

59 BASIC BIOLOGICAL SCIENCES↗

GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring

GAL08 are bacteria belonging to an uncultivated phylogenetic cluster within the phylum Acidobacteria . We detected a natural population of the GAL08 clade in sediment from a pH-neutral hot spring located in British Columbia, Canada. To shed light on the abundance and genomic potential of this clade, we collected and analyzed hot spring sediment samples over a temperature range of 24.2–79.8°C. Illumina sequencing of 16S rRNA gene amplicons and qPCR using a primer set developed specifically to detect the GAL08 16S rRNA gene revealed that absolute and relative abundances of GAL08 peaked at 65°C along three temperature gradients. Analysis of sediment collected over multiple years and locations revealed that the GAL08 group was consistently a dominant clade, comprising up to 29.2% of the microbial community based on relative read abundance and up to 4.7 × 10 5 16S rRNA gene copy numbers per gram of sediment based on qPCR. Using a medium quality threshold, 25 single amplified genomes (SAGs) representing these bacteria were generated from samples taken at 65 and 77°C, and seven metagenome-assembled genomes (MAGs) were reconstructed from samples collected at 45–77°C. Based on average nucleotide identity (ANI), these SAGs and MAGs represented three separate species, with an estimated average genome size of 3.17 Mb and GC content of 62.8%. Phylogenetic trees constructed from 16S rRNA gene sequences and a set of 56 concatenated phylogenetic marker genes both placed the three GAL08 bacteria as a distinct subgroup of the phylum Acidobacteria , representing a candidate order ( Ca. Frugalibacteriales) within the class Blastocatellia. Metabolic reconstructions from genome data predicted a heterotrophic metabolism, with potential capability for aerobic respiration, as well as incomplete denitrification and fermentation. In laboratory cultivation efforts, GAL08 counts based on qPCR declined rapidly under atmospheric levels of oxygen but increased slightly at 1% (v/v) O 2 , suggesting a microaerophilic lifestyle.

59 BASIC BIOLOGICAL SCIENCES↗

Biochemical characterization of Fsa16295Glu from “Fervidibacter sacchari,” the first hyperthermophilic GH50 with β-1,3-endoglucanase activity and founding member of the subfamily GH50_3

The aerobic hyperthermophile “Fervidibacter sacchari” catabolizes diverse polysaccharides and is the only cultivated member of the class “Fervidibacteria” within the phylum Armatimonadota. It encodes 117 putative glycoside hydrolases (GHs), including two from GH family 50 (GH50). In this study, we expressed, purified, and functionally characterized one of these GH50 enzymes, Fsa16295Glu. We show that Fsa16295Glu is a β-1,3-endoglucanase with optimal activity on carboxymethyl curdlan (CM-curdlan) and only weak agarase activity, despite most GH50 enzymes being described as β-agarases. The purified enzyme has a wide temperature range of 4–95°C (optimal 80°C), making it the first characterized hyperthermophilic representative of GH50. The enzyme is also active at a broad pH range of at least 5.5–11 (optimal 6.5–10). Fsa16295Glu possesses a relatively high k cat /K M of 1.82 × 10 7 s-1 M-1 with CM-curdlan and degrades CM-curdlan nearly completely to sugar monomers, indicating preferential hydrolysis of glucans containing β-1,3 linkages. Finally, a phylogenetic analysis of Fsa16295Glu and all other GH50 enzymes revealed that Fsa16295Glu is distant from other characterized enzymes but phylogenetically related to enzymes from thermophilic archaea that were likely acquired horizontally from “Fervidibacteria.” Given its functional and phylogenetic novelty, we propose that Fsa16295Glu represents a new enzyme subfamily, GH50_3.

59 BASIC BIOLOGICAL SCIENCES↗

Evolution of early life inferred from protein and ribonucleic acid sequences

The chemical structures of ferredoxin, 5S ribosomal RNA, and c-type cytochrome sequences have been employed to construct a phylogenetic tree which connects all major photosynthesizing organisms: the three types of bacteria, blue-green algae, and chloroplasts. Anaerobic and aerobic bacteria, eukaryotic cytoplasmic components and mitochondria are also included in the phylogenetic tree. Anaerobic nonphotosynthesizing bacteria similar to Clostridium were the earliest organisms, arising more than 3.2 billion years ago. Bacterial photosynthesis evolved nearly 3.0 billion years ago, while oxygen-evolving photosynthesis, originating in the blue-green algal line, came into being about 2.0 billion years ago. The phylogenetic tree supports the symbiotic theory of the origin of eukaryotes.

Dayhoff, M. O.↗

Characterization of Two Microbial Isolates from Andean Lakes in Bolivia

We are currently investigating the biological population present in the highest and least explored perennial lakes on earth in the Bolivian and Chilean Andes, including several volcanic crater lakes of more than 6000 m elevation, in combination of microbiological and molecular biological methods. Our samples were collected in saline lakes of the Laguna Blanca Laguna Verde area in the Bolivian Altiplano and in the Licancabur volcano crater (27 deg. 47 min S/67 deg. 47 min. W) in the ongoing project studying high altitude lakes. The main goal of the project is to look for analogies with Martian paleolakes. These Bolivian lakes can be described as Andean lakes following the classification of Chong. We have attempted to isolate pure cultures and phylogenetically characterize prokaryotes that grew under laboratory conditions. Sediment samples taken from the Licancabur crater lake (LC), Laguna Verde (LV), and Laguna Blanca (LB) were analyzed and cultured using enriched liquid media under both aerobic and anaerobic conditions. All cultures were incubated at room temperature (15 to 20 C) and under light exposure. For the reported isolates, 36 hours incubation were necessary for reaching optimal optical densities to consider them viable cultures. Ten serial dilutions starting from 1% inoculum were required to obtain a suitable enriched cell culture to transfer into solid media. Cultures on solid medium were necessary to verify the formation of colonies in order to isolate pure cultures. Different solid media were prepared using several combinations of both trace minerals and carbohydrates sources in order to fit their nutrient requirements. The microorganisms formed individual colonies on solid media enriched with tryptone, yeast extract and sodium chloride. Cells morphology was studied by optical and electronic microscopy. Rodshape morphologies were observed in most cases. Total bacterial genomic DNA was isolated from 50 ml late-exponential phase culture by using the CTAB miniprep protocol. The 16S rRNA genes were amplified by PCR using both Bacteria- and Archaeauniversal primer sets: 27f and 1492r, 21f and 1492r respectively. Sequences of 16S rRNA gene were determined and initially compared with reference sequences contained in the EMBL nucleotide sequence database by using the BLAST program and were subsequently aligned with 16S rRNA reference sequences in the ARB package (http://www.mikro.biologie.tu-muenchen.de). Aligned sequences were inserted within a stable phylogenetic tree by using the ARB parsimony tool. In this work we report the morphology and phylogenetic characterization of two isolates belonged to Laguna Blanca sediments.

Demergasso, C.↗

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

NASA Discipline Life Sciences Technologies↗

Evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus

The extension of ecological tolerance limits may be an important mechanism by which microorganisms adapt to novel environments, but it may come at the evolutionary cost of reduced performance under ancestral conditions. We combined a comparative physiological approach with phylogenetic analyses to study the evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus. Among the 20 laboratory clones of Synechococcus isolated from collections made along an Oregon hot spring thermal gradient, four different 16S rRNA gene sequences were identified. Phylogenies constructed by using the sequence data indicated that the clones were polyphyletic but that three of the four sequence groups formed a clade. Differences in thermotolerance were observed for clones with different 16S rRNA gene sequences, and comparison of these physiological differences within a phylogenetic framework provided evidence that more thermotolerant lineages of Synechococcus evolved from less thermotolerant ancestors. The extension of the thermal limit in these bacteria was correlated with a reduction in the breadth of the temperature range for growth, which provides evidence that enhanced thermotolerance has come at the evolutionary cost of increased thermal specialization. This study illustrates the utility of using phylogenetic comparative methods to investigate how evolutionary processes have shaped historical patterns of ecological diversification in microorganisms.

Synechococcus Group/classification/growth & develo↗

Does Collection Time Bias the Ecology of Cleanroom Air Samples?

Microbial monitoring of astromaterials collections has taken on increased importance with the return of biologically sensitive samples from the asteroids Ryugu and Bennu and the initiation of the Mars Sample Return Program. Terrestrial bacteria and fungi can alter the mineralogy and organic composition of our collections causing irreversible contamination of pristine samples and increasing the risk of false positives for life detection measurements. NASA has conducted routine microbial monitoring of its existing collections since 20181. Initial monitoring focused on surface samples collected with foam swabs. Although, airborne microbiology is often decoupled from surface microbiology in the built environment2 culture-based air sampling techniques like impactors were not compliant with existing contamination control requirements. Bringing organic rich media, gelatin or liquids into curation cleanrooms presents an unacceptable risk to pristine samples. In 2022 NASA purchased a materials complaint air sampler and began collecting air samples from the cleanrooms in addition to surface samples3. The new instrument uses an electret filter to collect samples that are suitable for cultivating organisms or for direct DNA sequencing. Preliminary DNA sequencing results appeared to indicate that longer sampling times biased the microbial community in favor of hearty, spore-forming bacteria3. We present the results of a study comparing overnight sampling (17 hours) to short (1 hour) sampling of unoccupied curation cleanrooms. The results will help us optimize our monitoring protocols and develop a more detailed inventory of the ecology of astromaterials curation cleanrooms. Methods: We analyzed 72 paired air samples from six different cleanrooms including the meteorite processing lab (ISO 7 equivalent, 16 samples), the lunar lab (ISO 6 equivalent, 10 samples), the stardust lab (ISO 5 equivalent 14 samples), the OSIRIS-REx lab (ISO 5 equivalent, 12 samples), the Hayabusa2 lab (ISO 5 equivalent, 14 samples), and the Genesis lab (ISO 4 equivalent, 6 samples). All the samples were collected with an InnovaPrep Bobcat air sampler operating at a sampling rate of 200 L/min. The sampler operates for 5 minutes out of every 20 minute period. Half of the samples were collected by filtering 3,000L (15 min. of active sampling) of air across an electret filter for one hour. The rest of the samples were collected by filtering approximately 51,000 L air across the filter overnight (~17 hours, 255 min. of active sampling). Cells were eluted from the filter using 6-7 ml of pressurized 0.15% tween 20 in PBS (phosphate buffered saline). This liquid was used to cultivate bacteria according to previously published methods1,4,5 and for DNA extraction and next generation sequencing. DNA was extracted with a Qiagen MagAttract PowerMicrobiome kit6. To identify bacteria and archaea, the 16S rRNA gene was amplified using Earth Microbiome primers for the V4 region 7. The amplified DNA was sequenced on an Illumina MiSeq using a V3 reagent kit. The resulting sequences were processed using DADA2 and QIIME2 as implemented on the EDGE bioinformatics platform8–10. Results: Only two of the 72 samples had no amplifiable DNA. Amplified DNA concentrations ranged from 2.67 – 0.272 ng/µl. The median concentration of amplified DNA for the 1 hour samples was 0.770 ± 0.368 ng/µl. The median concentration of amplified DNA for the overnight samples was 0.877 ± 0.434 ng/µl. On average the overnight samples had slightly more sequences (58,960 vs. 59,456) and ASV’s (amplicon sequence variants) (60 vs 64.5) than the one hour samples, but these differences are not statistically significant. The most abundant ASV in every sample mapped to the genus Cupravidus. ASV’s mapping to the genuses Bacillus, Schlegelella, Thermus, and Staphylococcus were also common. Discussion and Future Work: Alpha diversity statistics like Shannon Entropy and Faith Phylogenetic Diversity are used to describe the diversity of organisms in a single sample. If a longer sampling time was biasing the data, we would expect to see a change in these diversity statistics vs. sample time. However, we did not observe this in our data. The median Shannon entropy was slightly higher for the overnight samples (3.773 vs 3.611) as was the Faith Phylogenetic Diversity (4.042 vs 3.596), but both values were within a standard deviation of each other for the two sampling times (Fig. 1). It is unlikely, that the longer sampling time is introducing bias into our data. We do observe a significant decrease in diversity when comparing the air samples by lab. The Genesis lab (ISO 4 equivalent) has a lower median number of ASV’s (45.5) than the other labs (62). Median values for Shannon Entropy (3.717 vs. 3.430) and Faith Phylogenetic Diversity (3.796 vs. 3.548) are also lower for Genesis, but those values are with one standard deviation of each other for the different sampling times. This is consistent with previous culture-based results suggesting that the environment in cleanrooms tends to select for a core group of organisms capable of surviving under dry, low nutrient, conditions. The presence of the ASV’s mapping to Cupravidus and Thermus in our sequencing blanks and controls suggests that several of the most common organisms in our samples represent contaminants from the reagents used to perform the DNA extractions and sequencing. Further work is needed to identify these contaminants, remove them from our data and recalculate the diversity statistics. This is a systematic error. Therefore, we do not expect removing the sequencing contaminants to change our conclusions. Longer air sample collection times appear to result in slightly higher diversity and do not bias the results towards “hardy” bacteria like spore-formers. Based on these preliminary results we conclude that sampling at least 3,000 liters of air is sufficient to capture the microbial diversity of cleanrooms, and that air samples can also be collected overnight without negatively impacting diversity. These results allow us to be flexible when designing microbial monitoring plans so that they do not interfere with routine lab activity. References: 1. Regberg, A. B. et al. 49th Lunar and Planetary Science Conference (2018). 2. The United States Pharmacopeial Convention. USP General Chapter <1116> (2013). 3. Regberg, A. B., et al. 54th Lunar and Planetary Science Conference (2023). 4. Regberg, A. B. et al. 53rd Lunar and Planetary Science Conference ( 2022). 5. Davis, R. E.,et al. 50th Lunar and Planetary Science Conference (2019). 6. Qiagen. MagAttract® PowerMicrobiome® DNA/RNA EP Kit Handbook. (2018). 7. Walters, W. et al. mSystems 1, (2015). 8. Callahan, B. J. et al. Nat. Methods 13, 581–583 (2016). 9. Hall, M. & Beiko, R. G. Microbiome Analysis: Methods and Protocols113–129 (Springer, 2018). 10. Philipson, C. et al. Bio-Protoc. 7, e2622 (2017).

A. B. Regberg↗

Macroevolutionary diversity of traits and genomes in the model yeast genus Saccharomyces

Species is the fundamental unit to quantify biodiversity. In recent years, the model yeast Saccharomyces cerevisiae has seen an increased number of studies related to its geographical distribution, population structure, and phenotypic diversity. However, seven additional species from the same genus have been less thoroughly studied, which has limited our understanding of the macroevolutionary events leading to the diversification of this genus over the last 20 million years. Here, we show the geographies, hosts, substrates, and phylogenetic relationships for approximately 1,800 Saccharomyces strains, covering the complete genus with unprecedented breadth and depth. We generated and analyzed complete genome sequences of 163 strains and phenotyped 128 phylogenetically diverse strains. This dataset provides insights about genetic and phenotypic diversity within and between species and populations, quantifies reticulation and incomplete lineage sorting, and demonstrates how gene flow and selection have affected traits, such as galactose metabolism. These findings elevate the genus Saccharomyces as a model to understand biodiversity and evolution in microbial eukaryotes.

59 BASIC BIOLOGICAL SCIENCES↗

Common origin of ornithine–urea cycle in opisthokonts and stramenopiles

Eukaryotic complex phototrophs exhibit a colorful evolutionary history. At least three independent endosymbiotic events accompanied by the gene transfer from the endosymbiont to host assembled a complex genomic mosaic. Resulting patchwork may give rise to unique metabolic capabilities; on the other hand, it can also blur the reconstruction of phylogenetic relationships. The ornithine–urea cycle (OUC) belongs to the cornerstone of the metabolism of metazoans and, as found recently, also photosynthetic stramenopiles. We have analyzed the distribution and phylogenetic positions of genes encoding enzymes of the urea synthesis pathway in eukaryotes. We show here that metazoan and stramenopile OUC enzymes share common origins and that enzymes of the OUC found in primary algae (including plants) display different origins. The impact of this fact on the evolution of stramenopiles is discussed here.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms↗