Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “comparative genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Diversity, taxonomy, and evolution of archaeal viruses of the class Caudoviricetes

The archaeal tailed viruses (arTV), evolutionarily related to tailed double-stranded DNA (dsDNA) bacteriophages of the class Caudoviricetes , represent the most common isolates infecting halophilic archaea. Only a handful of these viruses have been genomically characterized, limiting our appreciation of their ecological impacts and evolution. Here, we present 37 new genomes of haloarchaeal tailed virus isolates, more than doubling the current number of sequenced arTVs. Analysis of all 63 available complete genomes of arTVs, which we propose to classify into 14 new families and 3 orders, suggests ancient divergence of archaeal and bacterial tailed viruses and points to an extensive sharing of genes involved in DNA metabolism and counterdefense mechanisms, illuminating common strategies of virus–host interactions with tailed bacteriophages. Coupling of the comparative genomics with the host range analysis on a broad panel of haloarchaeal species uncovered 4 distinct groups of viral tail fiber adhesins controlling the host range expansion. The survey of metagenomes using viral hallmark genes suggests that the global architecture of the arTV community is shaped through recurrent transfers between different biomes, including hypersaline, marine, and anoxic environments.

59 BASIC BIOLOGICAL SCIENCES↗

Diverse and unconventional methanogens, methanotrophs, and methylotrophs in metagenome-assembled genomes from subsurface sediments of the Slate River floodplain, Crested Butte, CO, USA

We use metagenome-assembled genomes (MAGs) to understand single-carbon (C1) compound-cycling—particularly methane-cycling—microorganisms in montane riparian floodplain sediments. We generated 1,233 MAGs (>50% completeness and <10% contamination) from 50- to 150-cm depth below the sediment surface capturing the transition between oxic, unsaturated sediments and anoxic, saturated sediments in the Slate River (SR) floodplain (Crested Butte, CO, USA). We recovered genomes of putative methanogens, methanotrophs, and methylotrophs (n = 57). Methanogens, found only in deep, anoxic depths at SR, originate from three different clades (Methanoregulaceae, Methanotrichaceae, and Methanomassiliicoccales), each with a different methanogenesis pathway; putative methanotrophic MAGs originate from within the Archaea (Candidatus Methanoperedens) in anoxic depths and uncultured bacteria (Ca. Binatia) in oxic depths. Genomes for canonical aerobic methanotrophs were not recovered. Ca. Methanoperedens were exceptionally abundant (~1,400× coverage, >50% abundance in the MAG library) in one sample that also contained aceticlastic methanogens, indicating a potential C1/methane-cycling hotspot. Ca. Methylomirabilis MAGs from SR encode pathways for methylotrophy but do not harbor methane monooxygenase or nitrogen reduction genes. Comparative genomic analysis supports that one clade within the Ca. Methylomirabilis genus is not methanotrophic. The genetic potential for methylotrophy was widespread, with over 10% and 19% of SR MAGs encoding a methanol dehydrogenase or substrate-specific methyltransferase, respectively. MAGs from uncultured Thermoplasmata archaea in the Ca. Gimiplasmatales (UBA10834) contain pathways that may allow for anaerobic methylotrophic acetogenesis. Overall, MAGs from SR floodplain sediments reveal a potential for methane production and consumption in the system and a robust potential for methylotrophy.

58 GEOSCIENCES↗

Leveraging computational genomics to understand the molecular basis of metal homeostasis

Genome-based data is helping to reveal the diverse strategies plants and algae use to maintain metal homeostasis. In addition to acquisition, distribution and storage of metals, acclimating to feast or famine can involve a wealth of genes that we are just now starting to understand. The fast-paced acquisition of genome-based data, however, is far outpacing our ability to experimentally characterize protein function. Computational genomic approaches are needed to fill the gap between what is known and unknown. To avoid misconstruing bioinformatically derived data, which is the root cause of the inaccurate functional annotations that plague databases, functional inferences from diverse sources and contextualization of that evidence with a robust understanding of protein family evolution is needed. Phylogenomic- and comparative-genomic-based studies can aid in the interpretation of experimental data or provide a spark for the discovery of a new function. These analyses not only lead to novel insight into a target protein's function but can generate thought-provoking insights across protein families.

59 BASIC BIOLOGICAL SCIENCES↗

Genomics, Exometabolomics, and Metabolic Probing Reveal Conserved Proteolytic Metabolism of Thermoflexus hugenholtzii and Three Candidate Species From China and Japan

Thermoflexus hugenholtzii JAD2 T , the only cultured representative of the Chloroflexota order Thermoflexales, is abundant in Great Boiling Spring (GBS), NV, United States, and close relatives inhabit geothermal systems globally. However, no defined medium exists for T. hugenholtzii JAD2 T and no single carbon source is known to support its growth, leaving key knowledge gaps in its metabolism and nutritional needs. Here, we report comparative genomic analysis of the draft genome of T. hugenholtzii JAD2 T and eight closely related metagenome-assembled genomes (MAGs) from geothermal sites in China, Japan, and the United States, representing “Candidatus Thermoflexus japonica,” “Candidatus Thermoflexus tengchongensis,” and “Candidatus Thermoflexus sinensis.” Genomics was integrated with targeted exometabolomics and 13 C metabolic probing of T. hugenholtzii. The Thermoflexus genomes each code for complete central carbon metabolic pathways and an unusually high abundance and diversity of peptidases, particularly Metallo- and Serine peptidase families, along with ABC transporters for peptides and some amino acids. The T. hugenholtzii JAD2 T exometabolome provided evidence of extracellular proteolytic activity based on the accumulation of free amino acids. However, several neutral and polar amino acids appear not to be utilized, based on their accumulation in the medium and the lack of annotated transporters. Adenine and adenosine were scavenged, and thymine and nicotinic acid were released, suggesting interdependency with other organisms in situ. Metabolic probing of T. hugenholtzii JAD2 T using 13 C-labeled compounds provided evidence of oxidation of glucose, pyruvate, cysteine, and citrate, and functioning glycolytic, tricarboxylic acid (TCA), and oxidative pentose-phosphate pathways (PPPs). However, differential use of position-specific 13 C-labeled compounds showed that glycolysis and the TCA cycle were uncoupled. Thus, despite the high abundance of Thermoflexus in sediments of some geothermal systems, they appear to be highly focused on chemoorganotrophy, particularly protein degradation, and may interact extensively with other microorganisms in situ.

59 BASIC BIOLOGICAL SCIENCES↗

Whither the genus Caldicellulosiruptor and the order Thermoanaerobacterales: phylogeny, taxonomy, ecology, and phenotype

The order Thermoanaerobacterales currently consists of fermentative anaerobic bacteria, including the genus Caldicellulosiruptor. Caldicellulosiruptor are represented by thirteen species; all, but one, have closed genome sequences. Interest in these extreme thermophiles has been motivated not only by their high optimal growth temperatures (≥70°C), but also by their ability to hydrolyze polysaccharides including, for some species, both xylan and microcrystalline cellulose. Caldicellulosiruptor species have been isolated from geographically diverse thermal terrestrial environments located in New Zealand, China, Russia, Iceland and North America. Evidence of their presence in other terrestrial locations is apparent from metagenomic signatures, including volcanic ash in permafrost. Here, phylogeny and taxonomy of the genus Caldicellulosiruptor was re-examined in light of new genome sequences. Based on genome analysis of 15 strains, a new order, Caldicellulosiruptorales, is proposed containing the family Caldicellulosiruptoraceae, consisting of two genera, Caldicellulosiruptor and Anaerocellum. Furthermore, the order Thermoanaerobacterales also was re-assessed, using 91 genome-sequenced strains, and should now include the family Thermoanaerobacteraceae containing the genera Thermoanaerobacter, Thermoanaerobacterium, Caldanaerobacter, the family Caldanaerobiaceae containing the genus Caldanaerobius, and the family Calorimonaceae containing the genus Calorimonas. A main outcome of ANI/AAI analysis indicates the need to reclassify several previously designated species in the Thermoanaerobacterales and Caldicellulosiruptorales by condensing them into strains of single species. Comparative genomics of carbohydrate-active enzyme inventories suggested differentiating phenotypic features, even among strains of the same species, reflecting available nutrients and ecological roles in their native biotopes.

59 BASIC BIOLOGICAL SCIENCES↗

Post-Fragmentation Whole Genome Amplification-Based Method

This innovation is derived from a proprietary amplification scheme that is based upon random fragmentation of the genome into a series of short, overlapping templates. The resulting shorter DNA strands (<400 bp) constitute a library of DNA fragments with defined 3 and 5 termini. Specific primers to these termini are then used to isothermally amplify this library into potentially unlimited quantities that can be used immediately for multiple downstream applications including gel eletrophoresis, quantitative polymerase chain reaction (QPCR), comparative genomic hybridization microarray, SNP analysis, and sequencing. The standard reaction can be performed with minimal hands-on time, and can produce amplified DNA in as little as three hours. Post-fragmentation whole genome amplification-based technology provides a robust and accurate method of amplifying femtogram levels of starting material into microgram yields with no detectable allele bias. The amplified DNA also facilitates the preservation of samples (spacecraft samples) by amplifying scarce amounts of template DNA into microgram concentrations in just a few hours. Based on further optimization of this technology, this could be a feasible technology to use in sample preservation for potential future sample return missions. The research and technology development described here can be pivotal in dealing with backward/forward biological contamination from planetary missions. Such efforts rely heavily on an increasing understanding of the burden and diversity of microorganisms present on spacecraft surfaces throughout assembly and testing. The development and implementation of these technologies could significantly improve the comprehensiveness and resolving power of spacecraft-associated microbial population censuses, and are important to the continued evolution and advancement of planetary protection capabilities. Current molecular procedures for assaying spacecraft-associated microbial burden and diversity have inherent sample loss issues at practically every step, particularly nucleic acid extraction. In engineering a molecular means of amplifying nucleic acids directly from single cells in their native state within the sample matrix, this innovation has circumvented entirely the need for DNA extraction regimes in the sample processing scheme.

Benardini, James↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Genome and proteome analyses show the gaseous alkane degrader Desulfosarcina sp. strain BuS5 as an extreme metabolic specialist

The metabolic potential of the sulfate-reducing bacterium Desulfosarcina sp. strain BuS5, currently the only pure culture able to oxidize the volatile alkanes propane and butane without oxygen, was investigated via genomics, proteomics and physiology assays. Complete genome sequencing revealed that strain BuS5 encodes a single alkyl-succinate synthase, an enzyme which apparently initiates oxidation of both propane and butane. The formed alkyl-succinates are oxidized to CO 2 via beta oxidation and the oxidative Wood–Ljungdahl pathways as shown by proteogenomics analyses. Strain BuS5 conserves energy via the canonical sulfate reduction pathway and electron bifurcation. An ability to utilize long-chain fatty acids, mannose and oligopeptides, suggested by automated annotation pipelines, was not supported by physiology assays and in-depth analyses of the corresponding genetic systems. Consistently, comparative genomics revealed a streamlined BuS5 genome with a remarkable paucity of catabolic modules. These results establish strain BuS5 as an exceptional metabolic specialist, able to grow only with propane and butane, for which we propose the name Desulfosarcina aeriophaga BuS5. This highly restrictive lifestyle, most likely the result of habitat-driven evolutionary gene loss, may provide D. aeriophaga BuS5 a competitive edge in sediments impacted by natural gas seeps.

59 BASIC BIOLOGICAL SCIENCES↗

Nonphotochemical quenching kinetics GWAS in sorghum identifies genes that may play conserved roles in maize and Arabidopsis thaliana photoprotection

SUMMARY Photosynthetic organisms must cope with rapid fluctuations in light intensity. Nonphotochemical quenching (NPQ) enables the dissipation of excess light energy as heat under high light conditions, whereas its relaxation under low light maximizes photosynthetic productivity. We quantified variation in NPQ kinetics across a large sorghum ( Sorghum bicolor ) association panel in four environments, uncovering significant genetic control for NPQ. A genome‐wide association study (GWAS) confidently identified three unique regions in the sorghum genome associated with NPQ and suggestive associations in an additional 61 regions. We detected strong signals from the sorghum ortholog of Arabidopsis thaliana Suppressor Of Variegation 3 ( SVR3 ) involved in plastid–nucleus signaling. By integrating GWAS results for NPQ across maize ( Zea mays ) and sorghum‐association panels, we identified a second gene, Non‐yellowing 1 ( NYE1 ), originally studied by Gregor Mendel in pea ( Pisum sativum ) and involved in the degradation of photosynthetic pigments in light‐harvesting complexes. Analysis of nye1 insertion alleles in A. thaliana confirmed the effect of this gene on NPQ kinetics in eudicots. We extended our comparative genomics GWAS framework across the entire maize and sorghum genomes, identifying four additional loci involved in NPQ kinetics. These results provide a baseline for increasing the accuracy and speed of candidate gene identification for GWAS in species with high linkage disequilibrium.

Plant Sciences↗

Cultivation of novel Atribacterota from oil well provides new insight into their diversity, ecology, and evolution in anoxic, carbon-rich environments

Background: The Atribacterota are widely distributed in the subsurface biosphere. Recently, the first Atribacterota isolate was described and the number of Atribacterota genome sequences retrieved from environmental samples has increased significantly; however, their diversity, physiology, ecology, and evolution remain poorly understood. Results: We report the isolation of the second member of Atribacterota, Thermatribacter velox gen. nov., sp. nov., within a new family Thermatribacteraceae fam. nov., and the short-term laboratory cultivation of a member of the JS1 lineage, Phoenicimicrobium oleiphilum HX-OS.bin.34 TS , both from a terrestrial oil reservoir. Physiological and metatranscriptomics analyses showed that Thermatribacter velox B11 T and Phoenicimicrobium oleiphilum HX-OS.bin.34 TS ferment sugars and n-alkanes, respectively, producing H 2 , CO 2 , and acetate as common products. Comparative genomics showed that all members of the Atribacterota lack a complete Wood-Ljungdahl Pathway (WLP), but that the Reductive Glycine Pathway (RGP) is widespread, indicating that the RGP, rather than WLP, is a central hub in Atribacterota metabolism. Ancestral character state reconstructions and phylogenetic analyses showed that key genes encoding the RGP (fdhA, fhs, folD, glyA, gcvT, gcvPAB, pdhD) and other central functions were gained independently in the two classes, Atribacteria (OP9) and Phoenicimicrobiia (JS1), after which they were inherited vertically; these genes included fumarate-adding enzymes (faeA; Phoenicimicrobiia only), the CODH/ACS complex (acsABCDE), and diverse hydrogenases (NiFe group 3b, 4b and FeFe group A3, C). Finally, we present genome-resolved community metabolic models showing the central roles of Atribacteria (OP9) and Phoenicimicrobiia (JS1) in acetate- and hydrocarbon-rich environments. Conclusion: Our findings expand the knowledge of the diversity, physiology, ecology, and evolution of the phylum Atribacterota. This study is a starting point for promoting more incisive studies of their syntrophic biology and may guide the rational design of strategies to cultivate them in the laboratory.

59 BASIC BIOLOGICAL SCIENCES↗

A view of the pan‐genome of domesticated Cowpea ( Vigna unguiculata [L.] Walp.)

Abstract Cowpea, Vigna unguiculata L . Walp., is a diploid warm‐season legume of critical importance as both food and fodder in sub‐Saharan Africa. This species is also grown in Northern Africa, Europe, Latin America, North America, and East to Southeast Asia. To capture the genomic diversity of domesticates of this important legume, de novo genome assemblies were produced for representatives of six subpopulations of cultivated cowpea identified previously from genotyping of several hundred diverse accessions. In the most complete assembly (IT97K‐499‐35), 26,026 core and 4963 noncore genes were identified, with 35,436 pan genes when considering all seven accessions. GO terms associated with response to stress and defense response were highly enriched among the noncore genes, while core genes were enriched in terms related to transcription factor activity, and transport and metabolic processes. Over 5 million single nucleotide polymorphisms (SNPs) relative to each assembly and over 40 structural variants >1 Mb in size were identified by comparing genomes. Vu10 was the chromosome with the highest frequency of SNPs, and Vu04 had the most structural variants. Noncore genes harbor a larger proportion of potentially disruptive variants than core genes, including missense, stop gain, and frameshift mutations; this suggests that noncore genes substantially contribute to diversity within domesticated cowpea.

59 BASIC BIOLOGICAL SCIENCES↗

Progress and challenges in sorghum biotechnology, a multipurpose feedstock for the bioeconomy

Abstract Sorghum [Sorghum bicolor (L.) Moench] is the fifth most important cereal crop globally by harvested area and production. Its drought and heat tolerance allow high yields with minimal input. It is a promising biomass crop for the production of biofuels and bioproducts. In addition, as an annual diploid with a relatively small genome compared with other C4 grasses, and excellent germplasm diversity, sorghum is an excellent research species for other C4 crops such as maize. As a result, an increasing number of researchers are looking to test the transferability of findings from other organisms such as Arabidopsis thaliana and Brachypodium distachyon to sorghum, as well as to engineer new biomass sorghum varieties. Here, we provide an overview of sorghum as a multipurpose feedstock crop which can support the growing bioeconomy, and as a monocot research model system. We review what makes sorghum such a successful crop and identify some key traits for future improvement. We assess recent progress in sorghum transformation and highlight how transformation limitations still restrict its widespread adoption. Finally, we summarize available sorghum genetic, genomic, and bioinformatics resources. This review is intended for researchers new to sorghum research, as well as those wishing to include non-food and forage applications in their research.

59 BASIC BIOLOGICAL SCIENCES↗

Novosphingobium aromaticivorans LigR coordinates transcription of genes involved in metabolism of multiple types of aromatics

Aromatic compounds are a ubiquitous and diverse family of chemicals with functions as biomolecules, natural products, industrial chemicals, and pollutants. Novosphingobium aromaticivorans DSM 12444 uses multiple inducible pathways to catabolize H-, G-, and S-type aromatics that contain zero, one, or two methoxy groups, respectively. Here, we obtain a systems-level view of the transcriptional control of its aromatic metabolic pathways. Several in vitro analyses found that a N. aromaticivorans homolog of the Sphingobium lignivorans SYK-6 transcription factor LigR bound genomic DNA upstream of genes involved in metabolism of multiple aromatic types. We found that a ΔLigR mutant had growth defects on all three types of aromatics as sole carbon sources. Transcriptomic analysis revealed that LigR was required to increase expression of gene products that function in metabolism of all three aromatic types. We also found that, in media containing both glucose and an aromatic carbon source, the ΔLigR mutant directed intermediates through alternative aromatic metabolic pathways. Protein-DNA binding assays showed that N. aromaticivorans LigR binds immediately upstream of promoters of genes involved in aromatic metabolism. We found that N. aromaticivorans LigR coordinates the expression of enzymes that function in the catabolism of H-, G-, and S-type aromatics, and that there are differences in the role of LigR in N. aromaticivorans and S. lignivorans. A comparative genomic analysis predicted that LigR homologs and the aromatic-metabolizing genes that it directly regulates are often co-localized in the genomes of Sphingomonadales, but often not found in this arrangement in many other known aromatic metabolizing bacteria.

Aromatic Compound Degradation↗

Catabolism of β-5 linked aromatics by Novosphingobium aromaticivorans

ABSTRACT Aromatic compounds are an important source of commodity chemicals traditionally produced from fossil fuels. Aromatics derived from plant lignin can potentially be converted into commodity chemicals through depolymerization followed by microbial funneling of monomers and low molecular weight oligomers. This study investigates the catabolism of the β-5 linked aromatic dimer dehydrodiconiferyl alcohol (DC-A) by the bacterium Novosphingobium aromaticivorans . We used genome-wide screens to identify candidate genes involved in DC-A catabolism. Subsequent in vivo and in vitro analyses of these candidate genes elucidated a catabolic pathway composed of four required gene products and several partially redundant dehydrogenases that convert DC-A to aromatic monomers that can be funneled into the central aromatic metabolic pathway of N. aromaticivorans . Specifically, a newly identified γ-formaldehyde lyase, PcfL, opens the phenylcoumaran ring to form a stilbene and formaldehyde. A lignostilbene dioxygenase, LsdD, then cleaves the stilbene to generate the aromatic monomers vanillin and 5-formylferulate (5-FF). We also showed that the aldehyde dehydrogenase FerD oxidizes 5-FF before it is decarboxylated by LigW, yielding ferulic acid. We found that some enzymes involved in the β-5 catabolism pathway can act on multiple substrates and that some steps in the pathway can be mediated by multiple enzymes, providing new insights into the robust flexibility of aromatic catabolism in N. aromaticivorans . A comparative genomic analysis predicted that the newly discovered β-5 aromatic catabolic pathway is common within the order Sphingomonadales. IMPORTANCE In the transition to a circular bioeconomy, the plant polymer lignin holds promise as a renewable source of industrially important aromatic chemicals. However, since lignin contains aromatic subunits joined by various chemical linkages, producing single chemical products from this polymer can be challenging. One strategy to overcome this challenge is using microbes to funnel a mixture of lignin-derived aromatics into target chemical products. This approach requires strategies to cleave the major inter-unit linkages of lignin to release monomers for funneling into valuable products. In this study, we report newly discovered aspects of a pathway by which the Novosphingobium aromaticivorans DSM12444 catabolizes aromatics joined by the second most common inter-unit linkage in lignin, the β-5 linkage. This work advances our knowledge of aromatic catabolic pathways, laying the groundwork for future metabolic engineering of this and other microbes for optimized conversion of lignin into products.

59 BASIC BIOLOGICAL SCIENCES↗

The Distinctive Evolution of orfX Clostridium parabotulinum Strains and Their Botulinum Neurotoxin Type A and F Gene Clusters Is Influenced by Environmental Factors and Gene Interactions via Mobile Genetic Elements

Of the seven currently known botulinum neurotoxin-producing species of Clostridium, C. parabotulinum, or C. botulinum Group I, is the species associated with the majority of human botulism cases worldwide. Phylogenetic analysis of these bacteria reveals a diverse species with multiple genomic clades. The neurotoxins they produce are also diverse, with over 20 subtypes currently represented. The existence of different bont genes within very similar genomes and of the same bont genes/gene clusters within different bacterial variants/species indicates that they have evolved independently. The neurotoxin genes are associated with one of two toxin gene cluster types containing either hemagglutinin (ha) genes or orfX genes. These genes may be located within the chromosome or extrachromosomal elements such as large plasmids. Although BoNT-producing C parabotulinum bacteria are distributed globally, they are more ubiquitous in certain specific geographic regions. Notably, northern hemisphere strains primarily contain ha gene clusters while southern hemisphere strains have a preponderance of orfX gene clusters. OrfX C. parabotulinum strains constitute a subset of this species that contain highly conserved bont gene clusters having a diverse range of bont genes. While much has been written about strains with ha gene clusters, less attention has been devoted to those with orfX gene clusters. The recent sequencing of 28 orfX C. parabotulinum strains and the availability of an additional 91 strains for analysis provides an opportunity to compare genomic relationships and identify unique toxin gene cluster characteristics and locations within this species subset in depth. The mechanisms behind the independent processes of bacteria evolution and generation of toxin diversity are explored through the examination of bacterial relationships relating to source locations and evidence of horizontal transfer of genetic material among different bacterial variants, particularly concerning bont gene clusters. Analysis of the content and locations of the bont gene clusters offers insights into common mechanisms of genetic transfer, chromosomal integration, and development of diversity among these genes.

59 BASIC BIOLOGICAL SCIENCES↗

The Moderately (D)efficient Enzyme: Catalysis-Related Damage In Vivo and Its Repair

Enzymes have in vivo life spans. Analysis of life spans, i.e., lifetime totals of catalytic turnovers, suggests that nonsurvivable collateral chemical damage from the very reactions that enzymes catalyze is a common but underdiagnosed cause of enzyme death. Analysis also implies that many enzymes are moderately deficient in that their active-site regions are not naturally as hardened against such collateral damage as they could be, leaving room for improvement by rational design or directed evolution. Enzyme life span might also be improved by engineering systems that repair otherwise fatal active-site damage, of which a handful are known and more are inferred to exist. Unfortunately, the data needed to design and execute such improvements are lacking: there are too few measurements of in vivo life span, and existing information about the extent, nature, and mechanisms of active-site damage and repair during normal enzyme operation is too scarce, anecdotal, and speculative to act on. Fortunately, advances in proteomics, metabolomics, cheminformatics, comparative genomics, and structural biochemistry now empower a systematic, data-driven approach for identifying, predicting, and validating instances of active-site damage and its repair. These capabilities would be practically useful in enzyme redesign and improvement of in-use stability and could change our thinking about which enzymes die young in vivo, and why.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Insights into the dynamics between viruses and their hosts in a hot spring microbial mat

Abstract Our current knowledge of host–virus interactions in biofilms is limited to computational predictions based on laboratory experiments with a small number of cultured bacteria. However, natural biofilms are diverse and chiefly composed of uncultured bacteria and archaea with no viral infection patterns and lifestyle predictions described to date. Herein, we predict the first DNA sequence-based host–virus interactions in a natural biofilm. Using single-cell genomics and metagenomics applied to a hot spring mat of the Cone Pool in Mono County, California, we provide insights into virus–host range, lifestyle and distribution across different mat layers. Thirty-four out of 130 single cells contained at least one viral contig (26%), which, together with the metagenome-assembled genomes, resulted in detection of 59 viruses linked to 34 host species. Analysis of single-cell amplification kinetics revealed a lack of active viral replication on the single-cell level. These findings were further supported by mapping metagenomic reads from different mat layers to the obtained host–virus pairs, which indicated a low copy number of viral genomes compared to their hosts. Lastly, the metagenomic data revealed high layer specificity of viruses, suggesting limited diffusion to other mat layers. Taken together, these observations indicate that in low mobility environments with high microbial abundance, lysogeny is the predominant viral lifestyle, in line with the previously proposed “Piggyback-the-Winner” theory.

59 BASIC BIOLOGICAL SCIENCES↗

Diversification of mandarin citrus by hybrid speciation and apomixis

The origin and dispersal of cultivated and wild mandarin and related citrus are poorly understood. Here, comparative genome analysis of 69 new east Asian genomes and other mainland Asian citrus reveals a previously unrecognized wild sexual species native to the Ryukyu Islands: C. ryukyuensis sp. nov. The taxonomic complexity of east Asian mandarins then collapses to a satisfying simplicity, accounting for tachibana, shiikuwasha, and other traditional Ryukyuan mandarin types as homoploid hybrid species formed by combining C. ryukyuensis with various mainland mandarins. These hybrid species reproduce clonally by apomictic seed, a trait shared with oranges, grapefruits, lemons and many cultivated mandarins. We trace the origin of apomixis alleles in citrus to mangshanyeju wild mandarins, which played a central role in citrus domestication via adaptive wild introgression. Our results provide a coherent biogeographic framework for understanding the diversity and domestication of mandarin-type citrus through speciation, admixture, and rapid diffusion of apomictic reproduction.

59 BASIC BIOLOGICAL SCIENCES↗