Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The state of algal genome quality and diversity

The genomic era of biology has created unprecedented opportunity to study how life works, including understanding evolutionary principles, bioprospecting for novel antibiotics, and genetic manipulation of bioeconomically-relevant species. While genome sequencing and genomic analysis was previously restricted to large-scale projects and consortia, sequencing has become democratized. However, as genomics has become commonplace, the cataloging of sequenced organisms has become challenging, and standardized practices for sequencing, assembly, and annotation have not been adopted. This is equally true in research fields such as algal biology, despite the growing importance of algae in a bio-based economy. Here, in this study, we provide a comprehensive review of the state of eukaryotic algal genomics, explore the quality of algal genome assemblies, and identify the biases and gaps in the current species distribution to inform the development of future genome projects. Overall, we find a trend of declining quality of genomic resources, including a reduction in assembly quality, gene annotation quality, and genome completeness. Potential solutions to improve genome quality include widespread utilization of long read and scaffolding technologies, implementation of standards for assembly quality, evidence-based gene annotation and requisite publication, and support of continued development of gene prediction and genome assessment software.

59 BASIC BIOLOGICAL SCIENCES↗

Assembly, comparative analysis, and utilization of a single haplotype reference genome for soybean

Cultivar Williams 82 has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as Wm82-ISU-01 (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48 387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of Williams 82 reveal large regions of genomic heterogeneity, including regions of differential introgression from the cultivar Kingwa within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (> 20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the ‘Williams’ haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of Williams 82. In addition to the Wm82.a6 assembly, we also assembled the genome of ‘Fiskeby III,’ a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with Fiskeby III revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and Fiskeby III genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.

59 BASIC BIOLOGICAL SCIENCES↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

Large-scale genomic analyses with machine learning uncover predictive patterns associated with fungal phytopathogenic lifestyles and traits

Abstract Invasive plant pathogenic fungi have a global impact, with devastating economic and environmental effects on crops and forests. Biosurveillance, a critical component of threat mitigation, requires risk prediction based on fungal lifestyles and traits. Recent studies have revealed distinct genomic patterns associated with specific groups of plant pathogenic fungi. We sought to establish whether these phytopathogenic genomic patterns hold across diverse taxonomic and ecological groups from the Ascomycota and Basidiomycota, and furthermore, if those patterns can be used in a predictive capacity for biosurveillance. Using a supervised machine learning approach that integrates phylogenetic and genomic data, we analyzed 387 fungal genomes to test a proof-of-concept for the use of genomic signatures in predicting fungal phytopathogenic lifestyles and traits during biosurveillance activities. Our machine learning feature sets were derived from genome annotation data of carbohydrate-active enzymes (CAZymes), peptidases, secondary metabolite clusters (SMCs), transporters, and transcription factors. We found that machine learning could successfully predict fungal lifestyles and traits across taxonomic groups, with the best predictive performance coming from feature sets comprising CAZyme, peptidase, and SMC data. While phylogeny was an important component in most predictions, the inclusion of genomic data improved prediction performance for every lifestyle and trait tested. Plant pathogenicity was one of the best-predicted traits, showing the promise of predictive genomics for biosurveillance applications. Furthermore, our machine learning approach revealed expansions in the number of genes from specific CAZyme and peptidase families in the genomes of plant pathogens compared to non-phytopathogenic genomes (saprotrophs, endo- and ectomycorrhizal fungi). Such genomic feature profiles give insight into the evolution of fungal phytopathogenicity and could be useful to predict the risks of unknown fungi in future biosurveillance activities.

59 BASIC BIOLOGICAL SCIENCES↗

Scaffolded and annotated nuclear and organelle genomes of the North American brown alga Saccharina latissima

Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.

DeWeese, Kelly↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

$\mathrm{CROPSR}$: an automated platform for complex genome-wide $\mathrm{CRISPR}$ g$\mathrm{RNA}$ design and validation

CRISPR/Cas9 technology has become an important tool to generate targeted, highly specific genome mutations. The technology has great potential for crop improvement, as crop genomes are tailored to optimize specific traits over generations of breeding. Many crops have highly complex and polyploid genomes, particularly those used for bioenergy or bioproducts. The majority of tools currently available for designing and evaluating gRNAs for CRISPR experiments were developed based on mammalian genomes that do not share the characteristics or design criteria for crop genomes. We have developed an open source tool for genome-wide design and evaluation of gRNA sequences for CRISPR experiments, CROPSR. The genome-wide approach provides a significant decrease in the time required to design a CRISPR experiment, including validation through PCR, at the expense of an overhead compute time required once per genome, at the first run. To better cater to the needs of crop geneticists, restrictions imposed by other packages on design and evaluation of gRNA sequences were lifted. A new machine learning model was developed to provide scores while avoiding situations in which the currently available tools sometimes failed to provide guides for repetitive, A/T-rich genomic regions. We show that our gRNA scoring model provides a significant increase in prediction accuracy over existing tools, even in non-crop genomes. CROPSR provides the scientific community with new methods and a new workflow for performing CRISPR/Cas9 knockout experiments. CROPSR reduces the challenges of working in crops, and helps speed gRNA sequence design, evaluation and validation. We hope that the new software will accelerate discovery and reduce the number of failed experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genomic characterization of three marine fungi, including Emericellopsis atlantica sp. nov. with signatures of a generalist lifestyle and marine biomass degradation

ABSTRACT Marine fungi remain poorly covered in global genome sequencing campaigns; the 1000 fungal genomes (1KFG) project attempts to shed light on the diversity, ecology and potential industrial use of overlooked and poorly resolved fungal taxa. This study characterizes the genomes of three marine fungi: Emericellopsis sp. TS7, wood-associated Amylocarpus encephaloides and algae-associated Calycina marina. These species were genome sequenced to study their genomic features, biosynthetic potential and phylogenetic placement using multilocus data. Amylocarpus encephaloides and C. marina were placed in the Helotiaceae and Pezizellaceae (Helotiales) , respectively, based on a 15-gene phylogenetic analysis. These two genomes had fewer biosynthetic gene clusters (BGCs) and carbohydrate active enzymes (CAZymes) than Emericellopsis sp. TS7 isolate. Emericellopsis sp. TS7 ( Hypocreales , Ascomycota ) was isolated from the sponge Stelletta normani . A six-gene phylogenetic analysis placed the isolate in the marine Emericellopsis clade and morphological examination confirmed that the isolate represents a new species, which is described here as E. atlantica . Analysis of its CAZyme repertoire and a culturing experiment on three marine and one terrestrial substrates indicated that E. atlantica is a psychrotrophic generalist fungus that is able to degrade several types of marine biomass. FungiSMASH analysis revealed the presence of 35 BGCs including, eight non-ribosomal peptide synthases (NRPSs), six NRPS-like, six polyketide synthases, nine terpenes and six hybrid, mixed or other clusters. Of these BGCs, only five were homologous with characterized BGCs. The presence of unknown BGCs sets and large CAZyme repertoire set stage for further investigations of E. atlantica . The Pezizellaceae genome and the genome of the monotypic Amylocarpus genus represent the first published genomes of filamentous fungi that are restricted in their occurrence to the marine habitat and form thus a valuable resource for the community that can be used in studying ecological adaptions of fungi using comparative genomics.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018

Integrated virus genomes (prophages) are commonly found in sequenced bacterial genomes but have rarely been described in detail for rhizobial genomes. Cupriavidus taiwanensis STM 6018 is a rhizobial Betaproteobacteria strain that was isolated in 2006 from a root nodule of a Mimosa pudica host in French Guiana, South America. Here we describe features of the genome of STM 6018, focusing on the characterization of two different types of prophages that have been identified in its genome. The draft genome of STM 6018 is 6,553,639 bp, and consists of 80 scaffolds, containing 5,864 protein-coding genes and 61 RNA genes. STM 6018 contains all the nodulation and nitrogen fixation gene clusters common to symbiotic Cupriavidus species; sharing >99.97% bp identity homology to the nod / nif / noeM gene clusters from C. taiwanensis LMG19424 T and “ Cupriavidus neocalidonicus” STM 6070. The STM 6018 genome contains the genomes of two prophages: one complete Mu-like capsular phage and one filamentous phage, which integrates into a putative dif site. This is the first characterization of a filamentous phage found within the genome of a rhizobial strain. Further examination of sequenced rhizobial genomes identified filamentous prophage sequences in several Beta-rhizobial strains but not in any Alphaproteobacterial rhizobia.

59 BASIC BIOLOGICAL SCIENCES↗

Sequencing the Genomes of the First Terrestrial Fungal Lineages: What Have We Learned?

The first genome sequenced of a eukaryotic organism was for Saccharomyces cerevisiae, as reported in 1996, but it was more than 10 years before any of the zygomycete fungi, which are the early-diverging terrestrial fungi currently placed in the phyla Mucoromycota and Zoopagomycota, were sequenced. The genome for Rhizopus delemar was completed in 2008; currently, more than 1000 zygomycete genomes have been sequenced. Genomic data from these early-diverging terrestrial fungi revealed deep phylogenetic separation of the two major clades—primarily plant—associated saprotrophic and mycorrhizal Mucoromycota versus the primarily mycoparasitic or animal-associated parasites and commensals in the Zoopagomycota. Genomic studies provide many valuable insights into how these fungi evolved in response to the challenges of living on land, including adaptations to sensing light and gravity, development of hyphal growth, and co-existence with the first terrestrial plants. Genome sequence data have facilitated studies of genome architecture, including a history of genome duplications and horizontal gene transfer events, distribution and organization of mating type loci, rDNA genes and transposable elements, methylation processes, and genes useful for various industrial applications. Pathogenicity genes and specialized secondary metabolites have also been detected in soil saprobes and pathogenic fungi. Novel endosymbiotic bacteria and viruses have been discovered during several zygomycete genome projects. Overall, genomic information has helped to resolve a plethora of research questions, from the placement of zygomycetes on the evolutionary tree of life and in natural ecosystems, to the applied biotechnological and medical questions.

59 BASIC BIOLOGICAL SCIENCES↗

A genomic perspective on fungal diversity and evolution

Originating from aquatic unicellular ancestors, over the course of ~1 billion years, the fungi have evolved to occupy nearly all aerobic environments on the planet, diversified into millions of different ‘species’ and have developed complex multicellular structures. Their relatively small, simple genomes have facilitated massive-scale sequencing and allowed us to explore genome evolution across an ancient eukaryotic kingdom. With thousands of genomes from diverse lineages now available, this Review will discuss insights into fungal biology and evolution gleaned with genomics and other multi-omics approaches. Using published genomes available through GenBank and the Joint Genome Institute’s MycoCosm platform, we generated kingdom-wide phylogenies and used them to highlight how fungal genomes have changed over time. With this phylogeny as a guide, we also discuss major evolutionary transitions that occurred across the fungal kingdom. Although progress has been made, these efforts are hampered by biases in genome representation and limited characterization of gene functions. Here, in this study, we discuss these challenges and possible future directions to address them, including initiatives to characterize conserved genes of unknown function and scale up sequencing towards 10,000 annotated fungal genomes.

Mondo, Stephen J. [USDOE Joint Genome Institute (J↗

Chromosome-scale Genome Assembly of the Most Abundant Ectomycorrhizal Fungus Cenococcum Geophilum Reveals Massive TE Expansion and RIP Defense Mechanism

Transposable elements (TEs) play crucial roles in genome evolution and ecological adaptation in fungi, yet their dynamics in ectomycorrhizal species remain poorly understood. Cenococcum geophilum, the most widespread ectomycorrhizal fungus in boreal and temperate forests with its large, repeat-rich genome, represents an ideal system to investigate TE-mediated adaptation to the physical environment and symbiotic lifestyle. However, previous studies have been limited by fragmented genome assemblies that prevented the resolution of repeat-rich regions. We assembled a telomere-to-telomere reference genome of C. geophilum strain 1.58 using PacBio HiFi and Hi-C datasets, resulting in a 178.54 Mbp genome with seven contiguous chromosomes. We identified 14,145 genes and over 78% of the genome consists of transposable elements (TEs). Of these, 94% are affected by repeat-induced point mutations (RIP), a genome defense mechanism that acts during the sexual reproduction phase, indicating cryptic or ancient sexual reproduction in this putatively asexual fungus. Long terminal repeat retrotransposons, LINEs, and DNA transposons dominate, with three TE families (Ty3, Ty1, and Tad1) contributing over 60% of the genome size, indicating recent transposition bursts. Screening of 15 additional C. geophilum strains revealed recent and lineage-specific TE expansions, implying that several TEs escaped the RIP machinery and retained potential activity. Supporting TE activity in the context of symbiosis, we found 56 TEs differentially transcribed between ectomycorrhizal and free-living mycelium tissues. An even higher number (n = 66) of TEs were differentially expressed between stress resistance morphology (i.e. sclerotia) and free-living mycelium. This supports that TEs are differentially regulated as a response to symbiotic and stress-related conditions. Our results demonstrate that the C. geophilum genome expansion was driven by a few lineage-specific TE families in recent history, with high RIP activity attesting to sexual reproduction. We also provide insights how TEs could respond to lifestyle transitions and traits associated with desiccation resistance.

Cenococcum geophilum↗

Maize tissue culture, transformation, and genome editing

The importance of maize (Zea mays L.) to global agriculture, world economy, and food security is widely known and increasing. Current maize breeding programs are deeply integrated with recent and rapid technological advances in genome sequencing, computational biology, and new genotyping and phenotyping technologies. Transformation and genome editing capabilities are a central hub to an array of advanced molecular and breeding approaches to crop improvement. Tissue culture and somatic embryogenesis play essential and central roles in maize transformation biology. Synergistic applications of maize transformation, advanced genomics, and genome editing provide a potent interdependent triad for functional genomics research and advanced molecular breeding. Implementation of advanced capabilities to transform maize and conduct genome editing will profoundly influence the dynamics of global agriculture ushering in a new era of varietal development and molecular breeding. With over 60.9 Mha planted in 2019 alone, biotech maize accounts for 31% of the world’s maize production. Up to 10% higher yields are achieved using new varieties generated using genetic modification technologies compared to similar conventional varieties. By extension, the impact new varietal releases developed through genome editing will likely be more significant. Further, advances in transformation and genome editing technologies will facilitate an even wider applicability for the development of new varieties with increasingly complex traits. The introduction of biochemical pathways and the use of synthetic biology have become increasingly more attainable. The future is of genotype independent maize transformation and genome editing, as the working platform will impact world agriculture, global food security, and plant science well into the future.

59 BASIC BIOLOGICAL SCIENCES↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

A scaffolded and annotated reference genome of giant kelp (Macrocystis pyrifera)

Abstract Macrocystis pyrifera (giant kelp), is a brown macroalga of great ecological importance as a primary producer and structure-forming foundational species that provides habitat for hundreds of species. It has many commercial uses (e.g. source of alginate, fertilizer, cosmetics, feedstock). One of the limitations to exploiting giant kelp’s economic potential and assisting in giant kelp conservation efforts is a lack of genomic tools like a high quality, contiguous reference genome with accurate gene annotations. Reference genomes attempt to capture the complete genomic sequence of an individual or species, and importantly provide a universal structure for comparison across a multitude of genetic experiments, both within and between species. We assembled the giant kelp genome of a haploid female gametophyte de novo using PacBio reads, then ordered contigs into chromosome level scaffolds using Hi-C. We found the giant kelp genome to be 537 MB, with a total of 35 scaffolds and 188 contigs. The assembly N50 is 13,669,674 with GC content of 50.37%. We assessed the genome completeness using BUSCO, and found giant kelp contained 94% of the BUSCO genes from the stramenopile clade. Annotation of the giant kelp genome revealed 25,919 genes. Additionally, we present genetic variation data based on 48 diploid giant kelp sporophytes from three different Southern California populations that confirms the population structure found in other studies of these populations. This work resulted in a high-quality giant kelp genome that greatly increases the genetic knowledge of this ecologically and economically vital species.

60 APPLIED LIFE SCIENCES↗

Metagenome-assembled genomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of LBNL (Lawrence Berkeley National Laboratory) TES (Terrestrial Ecosystem Science) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from soil depth profiles collected from 2014 to 2021 from three paired control and warming plots. We collected soil samples across a range of depth profiles (spanning surface to 90 cm deep) from three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. 101 soil metagenomes were sequenced at JGI (Joint Genome Institute) and UCSF (University of California San Francisco) Center for Advanced Technology and can be found under the JGI (Joint Genome Institute) GOLD (Genomes Online Database) Sequencing project Gs0151586 and NCBI (National Center for Biotechnology Information) Projects PRJNA1225762 and PRJEB39497. Metagenomes were assembled using JGI (Joint Genome Institute) Metagenome Workflow (10.1128/mSystems.00804-20). For each metagenome, the assembled contigs were binned into genomes using 3 binning algorithms (cocacola, metabat, and maxbin) and the resulting bins were consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>50%) and contamination (<25%), and dereplicated at 99% ANI (average nucleotide identity) using dRep (https://github.com/MrOlm/drep).The dataset includes a zip file of 2321 MAG (Metagenome Assembled Genome) fasta files, the accession numbers for the underlying metagenomes, and a csv file with MAG (Metagenome Assembled Genome) quality metrics and taxonomic classification (GTDB -Genome Taxonomy Database-RS220). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES↗

Phylogenomics and Comparative Genomics Highlight Specific Genetic Features in Ganoderma Species

The Ganoderma species in Polyporales are ecologically and economically relevant wood decayers used in traditional medicine, but their genomic traits are still poorly documented. In the present study, we carried out a phylogenomic and comparative genomic analyses to better understand the genetic blueprint of this fungal lineage. We investigated seven Ganoderma genomes, including three new genomes, G. australe, G. leucocontextum, and G. lingzhi. The size of the newly sequenced genomes ranged from 60.34 to 84.27 Mb and they encoded 15,007 to 20,460 genes. A total of 58 species, including 40 white-rot fungi, 11 brown-rot fungi, four ectomycorrhizal fungi, one endophyte fungus, and two pathogens in Basidiomycota, were used for phylogenomic analyses based on 143 single-copy genes. It confirmed that Ganoderma species belong to the core polyporoid clade. Comparing to the other selected species, the genomes of the Ganoderma species encoded a larger set of genes involved in terpene metabolism and coding for secreted proteins (CAZymes, lipases, proteases and SSPs). Of note, G. australe has the largest genome size with no obvious genome wide duplication, but showed transposable elements (TEs) expansion and the largest set of terpene gene clusters, suggesting a high ability to produce terpenoids for medicinal treatment. G. australe also encoded the largest set of proteins containing domains for cytochrome P450s, heterokaryon incompatibility and major facilitator families. Besides, the size of G. australe secretome is the largest, including CAZymes (AA9, GH18, A01A), proteases G01, and lipases GGGX, which may enhance the catabolism of cell wall carbohydrates, proteins, and fats during hosts colonization. The current genomic resource will be used to develop further biotechnology and medicinal applications, together with ecological studies of the Ganoderma species.

59 BASIC BIOLOGICAL SCIENCES↗

Radiation-induced genomic instability: radiation quality and dose response

Genomic instability is a term used to describe a phenomenon that results in the accumulation of multiple changes required to convert a stable genome of a normal cell to an unstable genome characteristic of a tumor. There has been considerable recent debate concerning the importance of genomic instability in human cancer and its temporal occurrence in the carcinogenic process. Radiation is capable of inducing genomic instability in mammalian cells and instability is thought to be the driving force responsible for radiation carcinogenesis. Genomic instability is characterized by a large collection of diverse endpoints that include large-scale chromosomal rearrangements and aberrations, amplification of genetic material, aneuploidy, micronucleus formation, microsatellite instability, and gene mutation. The capacity of radiation to induce genomic instability depends to a large extent on radiation quality or linear energy transfer (LET) and dose. There appears to be a low dose threshold effect with low LET, beyond which no additional genomic instability is induced. Low doses of both high and low LET radiation are capable of inducing this phenomenon. This report reviews data concerning dose rate effects of high and low LET radiation and their capacity to induce genomic instability assayed by chromosomal aberrations, delayed lethal mutations, micronuclei and apoptosis.

Review↗