Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome assembly”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Rapid identification of enteric bacteria from whole genome sequences using average nucleotide identity metrics

Identification of enteric bacteria species by whole genome sequence (WGS) analysis requires a rapid and an easily standardized approach. We leveraged the principles of average nucleotide identity using MUMmer (ANIm) software, which calculates the percent bases aligned between two bacterial genomes and their corresponding ANI values, to set threshold values for determining species consistent with the conventional identification methods of known species. The performance of species identification was evaluated using two datasets: the Reference Genome Dataset v2 (RGDv2), consisting of 43 enteric genome assemblies representing 32 species, and the Test Genome Dataset (TGDv1), comprising 454 genome assemblies which is designed to represent all species needed to query for identification, as well as rare and closely related species. The RGDv2 contains six Campylobacter spp., three Escherichia/Shigella spp., one Grimontia hollisae, six Listeria spp., one Photobacterium damselae, two Salmonella spp., and thirteen Vibrio spp., while the TGDv1 contains 454 enteric bacterial genomes representing 42 different species. The analysis showed that, when a standard minimum of 70% genome bases alignment existed, the ANI threshold values determined for these species were ≥95 for Escherichia/Shigella and Vibrio species, ≥93% for Salmonella species, and ≥92% for Campylobacter and Listeria species. Using these metrics, the RGDv2 accurately classified all validation strains in TGDv1 at the species level, which is consistent with the classification based on previous gold standard methods.

59 BASIC BIOLOGICAL SCIENCES↗

High-quality genome of the basidiomycete yeast Dioszegia hungarica PDD-24b-2 isolated from cloud water

The genome of the basidiomycete yeast Dioszegia hungarica strain PDD-24b-2 isolated from cloud water at the summit of puy de $D\hat{o}me$ (France) was sequenced using a hybrid PacBio and Illumina sequencing strategy. The obtained assembled genome of 20.98 Mb and a GC content of 57% is structured in 16 large-scale contigs ranging from 90 kb to 5.56Mb, and another 27.2 kb contig representing the complete circular mitochondrial genome. In total, 8,234 proteins were predicted from the genome sequence. The mitochondrial genome shows 16.2% cgu codon usage for arginine but has no canonical cognate tRNA to translate this codon. Detected transposable element (TE)-related sequences account for about 0.63% of the assembled genome. A dataset of 2,068 hand-picked public environmental metagenomes, representing over 20 Tbp of raw reads, was probed for D. hungarica related ITS sequences, and revealed worldwide distribution of this species, particularly in aerial habitats. Growth experiments suggested a psychrophilic phenotype and the ability to disperse by producing ballistospores. The high-quality assembled genome obtained for this D. hungarica strain will help investigate the behavior and ecological functions of this species in the environment.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic genome evolution in a model fern

The large size and complexity of most fern genomes have hampered efforts to elucidate fundamental aspects of fern biology and land plant evolution through genome-enabled research. Here we present a chromosomal genome assembly and associated methylome, transcriptome and metabolome analyses for the model fern species Ceratopteris richardii. The assembly reveals a history of remarkably dynamic genome evolution including rapid changes in genome content and structure following the most recent whole-genome duplication approximately 60 million years ago. These changes include massive gene loss, rampant tandem duplications and multiple horizontal gene transfers from bacteria, contributing to the diversification of defence-related gene families. The insertion of transposable elements into introns has led to the large size of the Ceratopteris genome and to exceptionally long genes relative to other plants. Gene family analyses indicate that genes directing seed development were co-opted from those controlling the development of fern sporangia, providing insights into seed plant evolution. Our findings and annotated genome assembly extend the utility of Ceratopteris as a model for investigating and teaching plant biology.

59 BASIC BIOLOGICAL SCIENCES↗

Tunturi virus isolates and metagenome-assembled viral genomes provide insights into the virome of Acidobacteriota in Arctic tundra soils

Arctic soils are climate-critical areas, where microorganisms play crucial roles in nutrient cycling processes. Acidobacteriota are phylogenetically and physiologically diverse bacteria that are abundant and active in Arctic tundra soils. Still, surprisingly little is known about acidobacterial viruses in general and those residing in the Arctic in particular. Here, we applied both culture-dependent and -independent methods to study the virome of Acidobacteriota in Arctic soils. Five virus isolates, Tunturi 1–5, were obtained from Arctic tundra soils, Kilpisjärvi, Finland (69°N), using Tunturiibacter spp. strains originating from the same area as hosts. The new virus isolates have tailed particles with podo- (Tunturi 1, 2, 3), sipho- (Tunturi 4), or myovirus-like (Tunturi 5) morphologies. The dsDNA genomes of the viral isolates are 63–98 kbp long, except Tunturi 5, which is a jumbo phage with a 309-kbp genome. Tunturi 1 and Tunturi 2 share 88% overall nucleotide identity, while the other three are not related to one another. For over half of the open reading frames in Tunturi genomes, no functions could be predicted. To further assess the Acidobacteriota-associated viral diversity in Kilpisjärvi soils, bulk metagenomes from the same soils were explored and a total of 1881 viral operational taxonomic units (vOTUs) were bioinformatically predicted. Almost all vOTUs (98%) were assigned to the class Caudoviricetes. For 125 vOTUs, including five (near-)complete ones, Acidobacteriota hosts were predicted. Acidobacteriota-linked vOTUs were abundant across sites, especially in fens. Terriglobia-associated proviruses were observed in Kilpisjärvi soils, being related to proviruses from distant soils and other biomes. Approximately genus- or higher-level similarities were found between the Tunturi viruses, Kilpisjärvi vOTUs, and other soil vOTUs, suggesting some shared groups of Acidobacteriota viruses across soils. This study provides acidobacterial virus isolates as laboratory models for future research and adds insights into the diversity of viral communities associated with Acidobacteriota in tundra soils. Predicted virus-host links and viral gene functions suggest various interactions between viruses and their host microorganisms. Largely unknown sequences in the isolates and metagenome-assembled viral genomes highlight a need for more extensive sampling of Arctic soils to better understand viral functions and contributions to ecosystem-wide cycling processes in the Arctic.

54 ENVIRONMENTAL SCIENCES↗

Atmospheric methane consumption in arid ecosystems acts as a reverse chimney and is accelerated by plant-methanotroph biomes

Drylands cover one-third of the Earth’s surface and are one of the largest terrestrial sinks for methane. Understanding the structure–function interplay between members of arid biomes can provide critical insights into mechanisms of resilience toward anthropogenic and climate-change-driven environmental stressors—water scarcity, heatwaves, and increased atmospheric greenhouse gases. This study integrates in situ measurements with culture-independent and enrichment-based investigations of methane-consuming microbiomes inhabiting soil in the Anza-Borrego Desert, a model arid ecosystem in Southern California, United States. The atmospheric methane consumption ranged between 2.26 and 12.73 μmol m 2 h −1 , peaking during the daytime at vegetated sites. Metagenomic studies revealed similar soil-microbiome compositions at vegetated and unvegetated sites, with Methylocaldum being the major methanotrophic clade. Eighty-four metagenome-assembled genomes were recovered, six represented by methanotrophic bacteria (three Methylocaldum , two Methylobacter , and uncultivated Methylococcaceae ). The prevalence of copper-containing methane monooxygenases in metagenomic datasets suggests a diverse potential for methane oxidation in canonical methanotrophs and uncultivated Gammaproteobacteria. Five pure cultures of methanotrophic bacteria were obtained, including four Methylocaldum . Genomic analysis of Methylocaldum isolates and metagenome-assembled genomes revealed the presence of multiple stand-alone methane monooxygenase subunit C paralogs, which may have functions beyond methane oxidation. Furthermore, these methanotrophs have genetic signatures typically linked to symbiotic interactions with plants, including tryptophan synthesis and indole-3-acetic acid production. Based on in situ fluxes and soil microbiome compositions, we propose the existence of arid-soil reverse chimneys, an empowered methane sink represented by yet-to-be-defined cooperation between desert vegetation and methane-consuming microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗

Microbial life in 25-m-deep boreholes in ancient permafrost illuminated by metagenomics

This study describes the composition and potential metabolic adaptation of microbial communities in northeastern Siberia, a repository of the oldest permafrost in the Northern Hemisphere. Samples of contrasting depth (1.75 to 25.1 m below surface), age (from ~ 10 kyr to 1.1 Myr) and salinity (from low 0.1–0.2 ppt and brackish 0.3–1.3 ppt to saline 6.1 ppt) were collected from freshwater permafrost (FP) of borehole AL1_15 on the Alazeya River, and coastal brackish permafrost (BP) overlying marine permafrost (MP) of borehole CH1_17 on the East Siberian Sea coast. To avoid the limited view provided with culturing work, we used 16S rRNA gene sequencing to show that the biodiversity decreased dramatically with permafrost age. Nonmetric multidimensional scaling (NMDS) analysis placed the samples into three groups: FP and BP together (10–100 kyr old), MP (105–120 kyr old), and FP (> 900 kyr old). Younger FP/BP deposits were distinguished by the presence of Acidobacteriota, Bacteroidota, Chloroflexota_A, and Gemmatimonadota, older FP deposits had a higher proportion of Gammaproteobacteria, and older MP deposits had much more uncultured groups within Asgardarchaeota, Crenarchaeota, Chloroflexota, Patescibacteria, and unassigned archaea. The 60 recovered metagenome-assembled genomes and un-binned metagenomic assemblies suggested that despite the large taxonomic differences between samples, they all had a wide range of taxa capable of fermentation coupled to nitrate utilization, with the exception of sulfur reduction present only in old MP deposits.

54 ENVIRONMENTAL SCIENCES↗

CheckV assesses the quality and completeness of metagenome-assembled viral genomes

Abstract Millions of new viral sequences have been identified from metagenomes, but the quality and completeness of these sequences vary considerably. Here we present CheckV, an automated pipeline for identifying closed viral genomes, estimating the completeness of genome fragments and removing flanking host regions from integrated proviruses. CheckV estimates completeness by comparing sequences with a large database of complete viral genomes, including 76,262 identified from a systematic search of publicly available metagenomes, metatranscriptomes and metaviromes. After validation on mock datasets and comparison to existing methods, we applied CheckV to large and diverse collections of metagenome-assembled viral sequences, including IMG/VR and the Global Ocean Virome. This revealed 44,652 high-quality viral genomes (that is, >90% complete), although the vast majority of sequences were small fragments, which highlights the challenge of assembling viral genomes from short-read metagenomes. Additionally, we found that removal of host contamination substantially improved the accurate identification of auxiliary metabolic genes and interpretation of viral-encoded functions.

59 BASIC BIOLOGICAL SCIENCES↗

JGI QC impact on assembly, binning, phylogenomics, and functional analysis

Background Investigators using metagenomic sequencing to study their microbiomes are often provided data that has been trimmed and decontaminated or do it themselves without knowing the effect these procedures can have on their downstream analyses. Here we evaluated the impact that JGI trimming and decontamination procedures had on assembly and binning metrics, placement of metagenome assembled genomes into species trees, and functional profiles of metagenome-assembled genomes (MAGs) extracted from twenty three complex rhizosphere metagenomes. We also investigated how more aggressive trimming impacts these binning metrics. Results We found that JGI trimmed and decontamination of input reads had some significant impacts in assembly and binning metrics compared to raw reads, and that differences in placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. More aggressive trimming beyond those used by JGI were found to reduce MAG counts. Conclusions Mild trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing? However, mild trimming and decontamination of metagenomic reads with high quality scores is recommended for those who elect to do so.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of BBDuk metagenomic read trimming and decontamination

Background Investigators using metagenomic sequencing to study their microbiomes are often provided data that has been trimmed and decontaminated or do it themselves without knowing the effect these procedures can have on their downstream analyses. Here we evaluated the impact that JGI trimming and decontamination procedures had on assembly and binning metrics, placement of metagenome assembled genomes into species trees, and functional profiles of metagenome-assembled genomes (MAGs) extracted from twenty three complex rhizosphere metagenomes. We also investigated how more aggressive trimming impacts these binning metrics. Results We found that JGI trimmed and decontamination of input reads had some significant impacts in assembly and binning metrics compared to raw reads, and that differences in placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. More aggressive trimming beyond those used by JGI were found to reduce MAG counts. Conclusions Mild trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing? However, mild trimming and decontamination of metagenomic reads with high quality scores is recommended for those who elect to do so.

59 BASIC BIOLOGICAL SCIENCES↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton

ABSTRACT Metagenomics is a powerful method for interpreting the ecological roles and physiological capabilities of mixed microbial communities. Yet, many tools for processing metagenomic data are neither designed to consider eukaryotes nor are they built for an increasing amount of sequence data. EukHeist is an automated pipeline to retrieve eukaryotic and prokaryotic metagenome-assembled genomes (MAGs) from large-scale metagenomic sequence data sets. We developed the EukHeist workflow to specifically process large amounts of both metagenomic and/or metatranscriptomic sequence data in an automated and reproducible fashion. Here, we applied EukHeist to the large-size fraction data (0.8–2,000 µm) from Tara Oceans to recover both eukaryotic and prokaryotic MAGs, which we refer to as TOPAZ (Tara Oceans Particle-Associated MAGs). The TOPAZ MAGs consisted of >900 environmentally relevant eukaryotic MAGs and >4,000 bacterial and archaeal MAGs. The bacterial and archaeal TOPAZ MAGs expand upon the phylogenetic diversity of likely particle- and host-associated taxa. We use these MAGs to demonstrate an approach to infer the putative trophic mode of the recovered eukaryotic MAGs. We also identify ecological cohorts of co-occurring MAGs, which are driven by specific environmental factors and putative host-microbe associations. These data together add to a number of growing resources of environmentally relevant eukaryotic genomic information. Complementary and expanded databases of MAGs, such as those provided through scalable pipelines like EukHeist, stand to advance our understanding of eukaryotic diversity through increased coverage of genomic representatives across the tree of life. IMPORTANCE Single-celled eukaryotes play ecologically significant roles in the marine environment, yet fundamental questions about their biodiversity, ecological function, and interactions remain. Environmental sequencing enables researchers to document naturally occurring protistan communities, without culturing bias, yet metagenomic and metatranscriptomic sequencing approaches cannot separate individual species from communities. To more completely capture the genomic content of mixed protistan populations, we can create bins of sequences that represent the same organism (metagenome-assembled genomes [MAGs]). We developed the EukHeist pipeline, which automates the binning of population-level eukaryotic and prokaryotic genomes from metagenomic reads. We show exciting insight into what protistan communities are present and their trophic roles in the ocean. Scalable computational tools, like EukHeist, may accelerate the identification of meaningful genetic signatures from large data sets and complement researchers’ efforts to leverage MAG databases for addressing ecological questions, resolving evolutionary relationships, and discovering potentially novel biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic factors shape carbon and nitrogen metabolic niche breadth across Saccharomycotina yeasts

Organisms exhibit extensive variation in ecological niche breadth, from very narrow (specialists) to very broad (generalists). Two general paradigms have been proposed to explain this variation: (i) trade-offs between performance efficiency and breadth and (ii) the joint influence of extrinsic (environmental) and intrinsic (genomic) factors. We assembled genomic, metabolic, and ecological data from nearly all known species of the ancient fungal subphylum Saccharomycotina (1154 yeast strains from 1051 species), grown in 24 different environmental conditions, to examine niche breadth evolution. We found that large differences in the breadth of carbon utilization traits between yeasts stem from intrinsic differences in genes encoding specific metabolic pathways, but we found limited evidence for trade-offs. Furthermore, these comprehensive data argue that intrinsic factors shape niche breadth variation in microbes.

59 BASIC BIOLOGICAL SCIENCES↗

Commentary: Duckweeds as model organisms for metabolic studies

Duckweeds have many practical applications, for example in human nutrition, as animal feed, in the production of bioplastics or vaccines, and phytoremediation (Acosta et al., 2021). Under most conditions, they reproduce asexually which provides genetically uniform material with predictable patterns of growth that make them ideal as sentinel organisms for phytotoxicity testing (Park et al., 2021). Asexual growth also results in high biomass production which makes duckweeds promising candidates as biofuel feedstocks (Acosta et al., 2021; Liang et al., 2023). In addition, duckweed species like Lemna minor and Spirodela polyrhiza are also reemerging as model organisms in plant biology as high-quality full genome assemblies and other genomic resources become available (Chang et al., 2016; Acosta et al., 2021). We argue that duckweed species are particularly of interest for the study of primary plant metabolism. Primary metabolism concerns the part of metabolism that is directly involved in the growth and development of plants, and which tends to be highly conserved among plant species. What makes duckweeds particularly attractive is that when grown on liquid media more precise control of physiological conditions can be attained relative to growth of plants in soil. Also, due to their relatively simple anatomical structure and asexual reproduction of fronds by budding, precise characterization of the physiological state under study is possible through one simple metric, i.e., the specific growth rate (rate of dry weight increase per existing dry weight), which can be incorporated relatively easily into metabolic models. This is not possible for land plants, such as Arabidopsis, because over the course of their life cycle, they go through multiple growth stages and phases of anatomical differentiation, which are much more complex to quantify. Furthermore, duckweeds can grow on organic substrates under heterotrophic or photomixotrophic conditions that facilitate isotope tracer studies. For example, in a previous study on duckweed by one of the authors, Lemna gibba (L). was grown on glucose with a position-specific 13 C-label that can be detected and resolved by Mass Spectrometry or Nuclear Magnetic Resonance spectrometry. Using this approach, the 13 C-label was traced into biomass compounds formed from glucose, particularly isoprenoid compounds. Some of the resulting labeling patterns were in apparent disagreement with predictions based on known metabolic pathways for the biosynthesis of isopentenyl pyrophosphate, the universal building block for isoprenoids, (Lichtenthaler et al., 1997). From this data it was deduced that isoprenoid compounds such as carotenoids and isoprenoid chains of phytol and plastoquinone, synthesized in the chloroplast, are produced via a previously unreported plant metabolic pathway, now known as the methylerythitol/deoxyxylulose-5-phosphate pathway (Lichtenthaler et al., 1997).

59 BASIC BIOLOGICAL SCIENCES↗

METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks

Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial communities. The genome-scale workflow includes annotation of microbial genomes, motif validation of biochemically validated conserved protein residues, metabolic pathway analyses, and calculation of contributions to individual biogeochemical transformations and cycles. The community-scale workflow supplements genome-scale analyses with determination of genome abundance in the microbiome, potential microbial metabolic handoffs and metabolite exchange, reconstruction of functional networks, and determination of microbial contributions to biogeochemical cycles. METABOLIC can take input genomes from isolates, metagenome-assembled genomes, or single-cell genomes. Results are presented in the form of tables for metabolism and a variety of visualizations including biogeochemical cycling potential, representation of sequential metabolic transformations, community-scale microbial functional networks using a newly defined metric “MW-score” (metabolic weight score), and metabolic Sankey diagrams. METABOLIC takes ~ 3 h with 40 CPU threads to process ~ 100 genomes and corresponding metagenomic reads within which the most compute-demanding part of hmmsearch takes ~ 45 min, while it takes ~ 5 h to complete hmmsearch for ~ 3600 genomes. Tests of accuracy, robustness, and consistency suggest METABOLIC provides better performance compared to other software and online servers. To highlight the utility and versatility of METABOLIC, we demonstrate its capabilities on diverse metagenomic datasets from the marine subsurface, terrestrial subsurface, meadow soil, deep sea, freshwater lakes, wastewater, and the human gut. Conclusion METABOLIC enables the consistent and reproducible study of microbial community ecology and biogeochemistry using a foundation of genome-informed microbial metabolism, and will advance the integration of uncultivated organisms into metabolic and biogeochemical models. METABOLIC is written in Perl and R and is freely available under GPLv3 at https://github.com/AnantharamanLab/METABOLIC.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-Resolved Metagenomics of Nitrogen Transformations in the Switchgrass Rhizosphere Microbiome on Marginal Lands

Switchgrass (Panicum virgatum L.) remains the preeminent American perennial (C4) bioenergy crop for cellulosic ethanol, that could help displace over a quarter of the US current petroleum consumption. Intriguingly, there is often little response to nitrogen fertilizer once stands are established. The rhizosphere microbiome plays a critical role in nitrogen cycling and overall plant nutrient uptake. We used high-throughput metagenomic sequencing to characterize the switchgrass rhizosphere microbial community before and after a nitrogen fertilization event for established stands on marginal land. We examined community structure and bulk metabolic potential, and resolved 29 individual bacteria genomes via metagenomic de novo assembly. Community structure and diversity were not significantly different before and after fertilization; however, the bulk metabolic potential of carbohydrate-active enzymes was depleted after fertilization. We resolved 29 metagenomic assembled genomes, including some from the ‘most wanted’ soil taxa such as Verrucomicrobia, Candidate phyla UBA10199, Acidobacteria (rare subgroup 23), Dormibacterota, and the very rare Candidatus Eisenbacteria. The Dormibacterota (formally candidate division AD3) we identified have the potential for autotrophic CO utilization, which may impact carbon partitioning and storage. Our study also suggests that the rhizosphere microbiome may be involved in providing associative nitrogen fixation (ANF) via the novel diazotroph Janthinobacterium to switchgrass.

60 APPLIED LIFE SCIENCES↗

Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases

Whole genome sequencing datasets, involving more than 600 groundwater samples, from nine countries, were analyzed to identify metagenome assembled genomes (MAGs) containing full operons for propane monooxygenase, soluble methane monooxygease, toluene monooxygenase and particulate ammonia/methane monooxygenase. The enzymes encoded by these genes are a focus of interest because of their ability to degrade common groundwater contaminants. Due to the large amount of data, sequence analyses involved more than 80 individual KBase narratives. The approach followed the KBase tutorial called "Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment" The generated MAGs were exported from each individual narrative into separate summary KBase narratives for each monooxygenase. Three KBase narratives were generated for particulate ammonia/methane monooxygenase, due to the large number of MAGs identified.

59 BASIC BIOLOGICAL SCIENCES↗

Draft assemblies for three geographically distinct isolates of the human-infecting microsporidium Encephalitozoon intestinalis

Encephalitozoon intestinalis is an obligate intracellular parasite that causes enteritis, bronchitis, conjunctivitis, and/or encephalitis in immunosuppressed and immunocompetent individuals. To better assess its genetic diversity, here we report near-chromosome-level draft genome assemblies for three geographically distinct isolates, quadrupling the number of genome assemblies available for the enigmatic fungal pathogen.

Biological and medical sciences↗

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES↗