Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

Abstract The distribution of ecotypic variation in natural populations is influenced by neutral and adaptive evolutionary forces that are challenging to disentangle. This study provides a high‐resolution portrait of genomic variation in Chinook salmon ( Oncorhynchus tshawytscha ) with emphasis on a region of major effect for ecotypic variation in migration timing. With a filtered data set of ~13 million single nucleotide polymorphisms (SNPs) from low‐coverage whole genome resequencing of 53 populations (3566 barcoded individuals), we contrasted patterns of genomic structure within and among major lineages and examined the extent of a selective sweep at a major effect region underlying migration timing (GREB1L/ROCK1). Neutral variation provided support for fine‐scale structure of populations, while allele frequency variation in GREB1L/ROCK1 was highly correlated with mean return timing for early and late migrating populations within each of the lineages ( r 2 = .58–.95; p < .001). However, the extent of selection within the genomic region controlling migration timing was much narrower in one lineage (interior stream‐type) compared to the other two major lineages, which corresponded to the breadth of phenotypic variation in migration timing observed among lineages. Evidence of a duplicated block within GREB1L/ROCK1 may be responsible for reduced recombination in this portion of the genome and contributes to phenotypic variation within and across lineages. Lastly, SNP positions across GREB1L/ROCK1 were assessed for their utility in discriminating migration timing among lineages, and we recommend multiple markers nearest the duplication to provide highest accuracy in conservation applications such as those that aim to protect early migrating Chinook salmon. These results highlight the need to investigate variation throughout the genome and the effects of structural variants on ecologically relevant phenotypic variation in natural species.

Horn, Rebekah L.↗

Predicting Antimicrobial Resistance Using Partial Genome Alignments

Antimicrobial resistance (AMR) is an important global health threat that impacts millions of people worldwide each year. Developing methods that can detect and predict AMR phenotypes can help to mitigate the spread of AMR by informing clinical decision making and appropriate mitigation strategies. Many bioinformatic methods have been developed for predicting AMR phenotypes from whole-genome sequences and AMR genes, but recent studies have indicated that predictions can be made from incomplete genome sequence data. In order to more systematically understand this, we built random forest-based machine learning classifiers for predicting susceptible and resistant phenotypes for Klebsiella pneumoniae (1,640 strains), Mycobacterium tuberculosis (2,497 strains), and Salmonella enterica (1,981 strains). We started by building models from alignments that were based on a reference chromosome for each species. We then subsampled each chromosomal alignment and built models for the resulting subalignments, finding that very small regions, representing approximately 0.1 to 0.2% of the chromosome, are predictive. In K. pneumoniae, M. tuberculosis, and S. enterica, the subalignments are able to predict multiple AMR phenotypes with at least 70% accuracy, even though most do not encode an AMR-related function. We used these models to identify regions of the chromosome with high and low predictive signals. Finally, subalignments that retain high accuracy across larger phylogenetic distances were examined in greater detail, revealing genes and intergenic regions with potential links to AMR, virulence, transport, and survival under stress conditions. IMPORTANCE Antimicrobial resistance causes thousands of deaths annually worldwide. Understanding the regions of the genome that are involved in antimicrobial resistance is important for developing mitigation strategies and preventing transmission. Machine learning models are capable of predicting antimicrobial resistance phenotypes from bacterial genome sequence data by identifying resistance genes, mutations, and other correlated features. They are also capable of implicating regions of the genome that have not been previously characterized as being involved in resistance. In this study, we generated global chromosomal alignments for Klebsiella pneumoniae, Mycobacterium tuberculosis, and Salmonella enterica and systematically searched them for small conserved regions of the genome that enable the prediction of antimicrobial resistance phenotypes. In addition to known antimicrobial resistance genes, this analysis identified genes involved in virulence and transport functions, as well as many genes with no previous implication in antimicrobial resistance.

59 BASIC BIOLOGICAL SCIENCES↗

Heterotrophic Thaumarchaea with Small Genomes Are Widespread in the Dark Ocean

The Thaumarchaeota is a diverse archaeal phylum comprising numerous lineages that play key roles in global biogeochemical cycling, particularly in the ocean. To date, all genomically characterized marine thaumarchaea are reported to be chemolithoautotrophic ammonia oxidizers. In this study, we report a group of putatively heterotrophic marine thaumarchaea (HMT) with small genome sizes that is globally abundant in the mesopelagic, apparently lacking the ability to oxidize ammonia. We assembled five HMT genomes from metagenomic data and show that they form a deeply branching sister lineage to the ammonia-oxidizing archaea (AOA). We identify this group in metagenomes from mesopelagic waters in all major ocean basins, with abundances reaching up to 6% of that of AOA. Surprisingly, we predict the HMT have small genomes of ~1 Mbp, and our ancestral state reconstruction indicates this lineage has undergone substantial genome reduction compared to other related archaea. The genomic repertoire of HMT indicates a versatile metabolism for aerobic chemoorganoheterotrophy that includes a divergent form III-a RuBisCO, a 2M respiratory complex I that has been hypothesized to increase energetic efficiency, and a three-subunit heme-copper oxidase complex IV that is absent from AOA. We also identify 21 pyrroloquinoline quinone (PQQ)-dependent dehydrogenases that are predicted to supply reducing equivalents to the electron transport chain and are among the most highly expressed HMT genes, suggesting these enzymes play an important role in the physiology of this group. Our results suggest that heterotrophic members of the Thaumarchaeota are widespread in the ocean and potentially play key roles in global chemical transformations.

59 BASIC BIOLOGICAL SCIENCES↗

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology↗

Chromosome-level genome assemblies and genetic maps reveal heterochiasmy and macrosynteny in endangered Atlantic Acropora

Abstract Background Over their evolutionary history, corals have adapted to sea level rise and increasing ocean temperatures, however, it is unclear how quickly they may respond to rapid change. Genome structure and genetic diversity contained within may highlight their adaptive potential. Results We present chromosome-scale genome assemblies and linkage maps of the critically endangered Atlantic acroporids,Acropora palmataandA. cervicornis. Both assemblies and linkage maps were resolved into 14 chromosomes with their gene content and colinearity. Repeats and chromosome arrangements were largely preserved between the species. The family Acroporidae and the genusAcroporaexhibited many phylogenetically significant gene family expansions. Macrosynteny decreased with phylogenetic distance. Nevertheless, scleractinians shared six of the 21 cnidarian ancestral linkage groups as well as numerous fission and fusion events compared to other distantly related cnidarians. Genetic linkage maps were constructed from oneA. palmatafamily and 16A. cervicornisfamilies using a genotyping array. The consensus maps span 1,013.42 cM and 927.36 cM forA. palmataandA. cervicornis, respectively. Both species exhibited high genome-wide recombination rates (3.04 to 3.53 cM/Mb) and pronounced sex-based differences, known as heterochiasmy, with 2 to 2.5X higher recombination rates estimated in the female maps. Conclusions Together, the chromosome-scale assemblies and genetic maps we present here are the first detailed look at the genomic landscapes of the critically endangered Atlantic acroporids. These data sets revealed that adaptive capacity of Atlantic acroporids is not limited by their recombination rates. The sister species maintain macrosynteny with few genes with high sequence divergence that may act as reproductive barriers between them. In the AtlanticAcropora, hybridization between the two sister species yields an F1 hybrid with limited fertility despite the high levels of macrosynteny and gene colinearity of their genomes. Together, these resources now enable genome-wide association studies and discovery of quantitative trait loci, two tools that can aid in the conservation of these species.

Biotechnology & Applied Microbiology↗

Exposing new taxonomic variation with inflammation — a murine model-specific genome database for gut microbiome researchers

The murine CBA/J mouse model widely supports immunology and enteric pathogen research. This model has illuminated Salmonella interactions with the gut microbiome since pathogen proliferation does not require disruptive pretreatment of the native microbiota, nor does it become systemic, thereby representing an analog to gastroenteritis disease progression in humans. Despite the value to broad research communities, microbiota in CBA/J mice are not represented in current murine microbiome genome catalogs. Here we present the first microbial and viral genomic catalog of the CBA/J murine gut microbiome. Using fecal microbial communities from untreated and Salmonella-infected, highly inflamed mice, we performed genomic reconstruction to determine the impacts on gut microbiome membership and functional potential. From high depth whole community sequencing (~ 42.4 Gbps/sample), we reconstructed 2281 bacterial and 4516 viral draft genomes. Salmonella challenge significantly altered gut membership in CBA/J mice, revealing 30 genera and 98 species that were conditionally rare and unsampled in non-inflamed mice. Additionally, inflamed communities were depleted in microbial genes that modulate host anti-inflammatory pathways and enriched in genes for respiratory energy generation. Our findings suggest decreases in butyrate concentrations during Salmonella infection corresponded to reductions in the relative abundance in members of the Alistipes. Strain-level comparison of CBA/J microbial genomes to prominent murine gut microbiome databases identified newly sampled lineages in this resource, while comparisons to human gut microbiomes extended the host relevance of dominant CBA/J inflammation-resistant strains. This CBA/J microbiome database provides the first genomic sampling of relevant, uncultivated microorganisms within the gut from this widely used laboratory model. Using this resource, we curated a functional, strain-resolved view on how Salmonella remodels intact murine gut communities, advancing pathobiome understanding beyond inferences from prior amplicon-based approaches. Salmonella-induced inflammation suppressed Alistipes and other dominant members, while rarer commensals like Lactobacillus and Enterococcus endure. The rare and novel species sampled across this inflammation gradient advance the utility of this microbiome resource to benefit the broad research needs of the CBA/J scientific community, and those using murine models for understanding the impact of inflammation on the gut microbiome more generally.

59 BASIC BIOLOGICAL SCIENCES↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Model Soil Consortium 2 (MSC-2) Bacterial Isolate Genomes

Model Soil Consortium 2 (MSC-2) bacterial isolate genome collections are derived from the soil consortium MSC-1 multi-species cultured isolate genome collections from a previously reported WA-IsoC_MSC1.1.0 collection of a naturally evolved, model soil consortia (10.25584/WAIsoCMSC1/1635272). This collection contains 8 different isolate species cultivated under variable carbon and nitrogen sources, originating from the IAREC grassland soil field site located in Prosser, WA, USA. Genomic sequencing of one or more organisms, or genomes in general, such as meta-information on genomes, genome projects, gene names of a given organism within a natural environment. The version described here is the first version.

Soil microbiome, defined community, chitin↗

KBase Narrative - Porphyromonadaceae sp. W3.11 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

KBase Narrative - Lachnospiraceae sp. C1.1 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

A genome-informed higher rank classification of the biotechnologically important fungal subphylum Saccharomycotina

The subphylum Saccharomycotina is a lineage in the fungal phylum Ascomycota that exhibits levels of genomic diversity similar to those of plants and animals. The Saccharomycotina consist of more than 1 200 known species currently divided into 16 families, one order, and one class. Species in this subphylum are ecologically and metabolically diverse and include important opportunistic human pathogens, as well as species important in biotechnological applications. Many traits of biotechnological interest are found in closely related species and often restricted to single phylogenetic clades. However, the biotechnological potential of most yeast species remains unexplored. Although the subphylum Saccharomycotina has much higher rates of genome sequence evolution than its sister subphylum, Pezizomycotina, it contains only one class compared to the 16 classes in Pezizomycotina. The third subphylum of Ascomycota, the Taphrinomycotina, consists of six classes and has approximately 10 times fewer species than the Saccharomycotina. These data indicate that the current classification of all these yeasts into a single class and a single order is an underappreciation of their diversity. Our previous genome-scale phylogenetic analyses showed that the Saccharomycotina contains 12 major and robustly supported phylogenetic clades; seven of these are current families (Lipomycetaceae, Trigonopsidaceae, Alloascoideaceae, Pichiaceae, Phaffomycetaceae, Saccharomycodaceae, and Saccharomycetaceae), one comprises two current families (Dipodascaceae and Trichomonascaceae), one represents the genus Sporopachydermia, and three represent lineages that differ in their translation of the CUG codon (CUG-Ala, CUG-Ser1, and CUG-Ser2). Using these analyses in combination with relative evolutionary divergence and genome content analyses, we propose an updated classification for the Saccharomycotina, including seven classes and 12 orders that can be diagnosed by genome content. This updated classification is consistent with the high levels of genomic diversity within this subphylum and is necessary to make the higher rank classification of the Saccharomycotina more comparable to that of other fungi, as well as to communicate efficiently on lineages that are not yet formally named.

59 BASIC BIOLOGICAL SCIENCES↗

Diversity of Sordariales Fungi: Identification of Seven New Species of Naviculisporaceae Through Morphological Analyses and Genome Sequencing

Thanks to next-generation sequencing (NGS) technologies, the diversity of fungi can now be investigated through the analysis of their genome sequences. Naviculisporaceae is a family within the Sordariales, whose diversity is not well-known, with only one genome sequence published for this family. Here, we report on the isolation and cultivation of 20 new strains of Naviculisporaceae. Their genome sequences, as well as those of the five commercially available strains, were determined, thus providing complete genome sequences for 25 new Naviculisporaceae strains. Species delimitation was conducted using a combination of (1) ITS + LSU phylogenetic analysis of the new isolates along with other known species of the family, (2) comparisons between DNA barcode sequences of the new strains with those of the known species, and (3) average genome-wide nucleotide identity calculation. We built a phylogenomic tree and studied the organization of the mating-type locus. In vitro fruiting was obtained for 16 strains, enabling the definition of seven new species, namely Pseudorhypophila gallica, Pseudorhypophila guyanensis Rhypophila alpibus, Rhypophila brasiliensis, Rhypophila camarguensis, Rhypophila reunionensis and Rhypophila thailandica, as well as two new combinations, namely Pseudorhypophila latipes and Pseudorhypophila oryzae. Eight strains for which in vitro fruiting was not obtained may belong to additional new species. These results expand the known diversity of the Naviculisporaceae and greatly enlarge the genomic data available for the family.

Naviculisporaceae↗

Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation

Viruses influence global patterns of microbial diversity and nutrient cycles. Though viral metagenomics (viromics), specifically targeting dsDNA viruses, has been critical for revealing viral roles across diverse ecosystems, its analyses differ in many ways from those used for microbes. To date, viromics benchmarking has covered read pre-processing, assembly, relative abundance, read mapping thresholds and diversity estimation, but other steps would benefit from benchmarking and standardization. Here we use in silico-generated datasets and an extensive literature survey to evaluate and highlight how dataset composition (i.e., viromes vs bulk metagenomes) and assembly fragmentation impact (i) viral contig identification tool, (ii) virus taxonomic classification, and (iii) identification and curation of auxiliary metabolic genes (AMGs). The in silico benchmarking of five commonly used virus identification tools show that gene-content-based tools consistently performed well for long (≥3 kbp) contigs, while k -mer- and blast-based tools were uniquely able to detect viruses from short (≤3 kbp) contigs. Notably, however, the performance increase of k -mer- and blast-based tools for short contigs was obtained at the cost of increased false positives (sometimes up to ~5% for virome and ~75% bulk samples), particularly when eukaryotic or mobile genetic element sequences were included in the test datasets. Furthermore, for viral classification, variously sized genome fragments were assessed using gene-sharing network analytics to quantify drop-offs in taxonomic assignments, which revealed correct assignations ranging from ~95% (whole genomes) down to ~80% (3 kbp sized genome fragments). A similar trend was also observed for other viral classification tools such as VPF-class, ViPTree and VIRIDIC, suggesting that caution is warranted when classifying short genome fragments and not full genomes. Finally, we highlight how fragmented assemblies can lead to erroneous identification of AMGs and outline a best-practices workflow to curate candidate AMGs in viral genomes assembled from metagenomes. Together, these benchmarking experiments and annotation guidelines should aid researchers seeking to best detect, classify, and characterize the myriad viruses ‘hidden’ in diverse sequence datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES↗

The Ancient Salicoid Genome Duplication Event: A Platform for Reconstruction of De Novo Gene Evolution in Populus trichocarpa

Orphan genes are characteristic genomic features that have no detectable homology to genes in any other species and represent an important attribute of genome evolution as sources of novel genetic functions. Here, we identified 445 genes specific to Populus trichocarpa. Of these, we performed deeper reconstruction of 13 orphan genes to provide evidence of de novo gene evolution. Populus and its sister genera Salix are particularly well suited for the study of orphan gene evolution because of the Salicoid whole-genome duplication event which resulted in highly syntenic sister chromosomal segments across the Salicaceae. We leveraged this genomic feature to reconstruct de novo gene evolution from intergenera, interspecies, and intragenomic perspectives by comparing the syntenic regions within the P. trichocarpa reference, then P. deltoides, and finally Salix purpurea. Furthermore, we demonstrated that 86.5% of the putative orphan genes had evidence of transcription. Additionally, we also utilized the Populus genome-wide association mapping panel, a collection of 1,084 undomesticated P. trichocarpa genotypes to further determine putative regulatory networks of orphan genes using expression quantitative trait loci (eQTL) mapping. Functional enrichment of these eQTL subnetworks identified common biological themes associated with orphan genes such as response to stress and defense response. We also identify a putative cis-element for a de novo gene and leverage conserved synteny to describe evolution of a putative transcription factor binding site. Overall, 45% of orphan genes were captured in trans-eQTL networks.

59 BASIC BIOLOGICAL SCIENCES↗

scMicrobe PTA: near complete genomes from single bacterial cells

Microbial genomes produced by standard single-cell amplification methods are largely incomplete. Here, we show that primary template-directed amplification (PTA), a novel single-cell amplification technique, generated nearly complete genomes from three bacterial isolate species. Furthermore, taxonomically diverse genomes recovered from aquatic and soil microbiomes using PTA had a median completeness of 81%, whereas genomes from standard multiple displacement amplification-based approaches were usually <30% complete. PTA-derived genomes also included more associated viruses and biosynthetic gene clusters.

59 BASIC BIOLOGICAL SCIENCES↗

Population genomics of the pathogenic yeast Candida tropicalis identifies hybrid isolates in environmental samples

Candida tropicalis is a human pathogen that primarily infects the immunocompromised. Whereas the genome of one isolate, C. tropicalis MYA-3404, was originally sequenced in 2009, there have been no large-scale, multi-isolate studies of the genetic and phenotypic diversity of this species. Here, we used whole genome sequencing and phenotyping to characterize 77 isolates of C. tropicalis from clinical and environmental sources from a variety of locations. We show that most C. tropicalis isolates are diploids with approximately 2–6 heterozygous variants per kilobase. The genomes are relatively stable, with few aneuploidies. However, we identified one highly homozygous isolate and six isolates of C. tropicalis with much higher heterozygosity levels ranging from 36–49 heterozygous variants per kilobase. Our analyses show that the heterozygous isolates represent two different hybrid lineages, where the hybrids share one parent (A) with most other C. tropicalis isolates, but the second parent (B or C) differs by at least 4% at the genome level. Four of the sequenced isolates descend from an AB hybridization, and two from an AC hybridization. The hybrids are MTLa/α heterozygotes. Hybridization, or mating, between different parents is therefore common in the evolutionary history of C. tropicalis. The new hybrids were predominantly found in environmental niches, including from soil. Hybridization is therefore unlikely to be associated with virulence. In addition, we used genotype-phenotype correlation and CRISPR-Cas9 editing to identify a genome variant that results in the inability of one isolate to utilize certain branched-chain amino acids as a sole nitrogen source.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Virulence and Genomic Analysis of Streptococcus suis Isolates

Streptococcus suis is a zoonotic bacterial swine pathogen causing substantial economic and health burdens to the pork industry. Mechanisms used by S. suis to colonize and cause disease remain unknown and vaccines and/or intervention strategies currently do not exist. Studies addressing virulence mechanisms used by S. suis have been complicated because different isolates can cause a spectrum of disease outcomes ranging from lethal systemic disease to asymptomatic carriage. The objectives of this study were to evaluate the virulence capacity of nine United States S. suis isolates following intranasal challenge in swine and then perform comparative genomic analyses to identify genomic attributes associated with swine-virulent phenotypes. No correlation was found between the capacity to cause disease in swine and the functional characteristics of genome size, serotype, sequence type (ST), or in vitro virulence-associated phenotypes. A search for orthologs found in highly virulent isolates and not found in non-virulent isolates revealed numerous predicted protein coding sequences specific to each category. While none of these predicted protein coding sequences have been previously characterized as potential virulence factors, this analysis does provide a reliable one-to-one assignment of specific genes of interest that could prove useful in future allelic replacement and/or functional genomic studies. Collectively, this report provides a framework for future allelic replacement and/or functional genomic studies investigating genetic characteristics underlying the spectrum of disease outcomes caused by S. suis isolates.

59 BASIC BIOLOGICAL SCIENCES↗