Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Horizontal gene transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

SourceFinder: a Machine-Learning-Based Tool for Identification of Chromosomal, Plasmid, and Bacteriophage Sequences from Assemblies

High-throughput genome sequencing technologies enable the investigation of complex genetic interactions, including the horizontal gene transfer of plasmids and bacteriophages. However, identifying these elements from assembled reads remains challenging due to genome sequence plasticity and the difficulty in assembling complete sequences. In this study, we developed a classifier, using random forest, to identify whether sequences originated from bacterial chromosomes, plasmids, or bacteriophages. The classifier was trained on a diverse collection of 23,211 chromosomal, plasmid, and bacteriophage sequences from hundreds of bacterial species. In order to adapt the classifier to incomplete sequences, each complete sequence was subsampled into 5,000 nucleotide fragments and further subdivided into k-mers. This three-class classifier succeeded in identifying chromosomes, plasmids, and bacteriophages using k-mer distributions of complete and partial genome sequences, including simulated metagenomic scaffolds with minimum performance of 0.939 area under the receiver operating characteristic curve (AUC). This classifier, implemented as SourceFinder, has been made available as an online web service to help the community with predicting the chromosomal, plasmid, and bacteriophage sources of assembled bacterial sequence data (https://cge.food.dtu.dk/services/SourceFinder/).

59 BASIC BIOLOGICAL SCIENCES↗

Identifying Genomic Islands with Deep Neural Networks

Background Horizontal gene transfer is the main source of adaptability for bacteria, through which genes are obtained from different sources including bacteria, archaea, viruses, and eukaryotes. This process promotes the rapid spread of genetic information across lineages, typically in the form of clusters of genes referred to as genomic islands (GIs). Different types of GIs exist, and are often classified by the content of their cargo genes or their means of integration and mobility. While various computational methods have been devised to detect different types of GIs, no single method is capable of detecting all types. Results We propose a method, which we call Shutter Island, that uses a deep learning model (Inception V3, widely used in computer vision) to detect genomic islands. The intrinsic value of deep learning methods lies in their ability to generalize. Via a technique called transfer learning, the model is pre-trained on a large generic dataset and then re-trained on images that we generate to represent genomic fragments. We demonstrate that this image-based approach generalizes better than the existing tools. Conclusions We used a deep neural network and an image-based approach to detect the most out of the correct GI predictions made by other tools, in addition to making novel GI predictions. The fact that the deep neural network was re-trained on only a limited number of GI datasets and then successfully generalized indicates that this approach could be applied to other problems in the field where data is still lacking or hard to curate.

Computer Vision↗

Diverse signatures of convergent evolution in cactus-associated yeasts

Many distantly related organisms have convergently evolved traits and lifestyles that enable them to live in similar ecological environments. However, the extent of phenotypic convergence evolving through the same or distinct genetic trajectories remains an open question. Here, we leverage a comprehensive dataset of genomic and phenotypic data from 1,049 yeast species in the subphylum Saccharomycotina (Kingdom Fungi, Phylum Ascomycota) to explore signatures of convergent evolution in cactophilic yeasts, ecological specialists associated with cacti. We inferred that the ecological association of yeasts with cacti arose independently approximately 17 times. Using a machine learning–based approach, we further found that cactophily can be predicted with 76% accuracy from both functional genomic and phenotypic data. The most informative feature for predicting cactophily was thermotolerance, which we found to be likely associated with altered evolutionary rates of genes impacting the cell envelope in several cactophilic lineages. We also identified horizontal gene transfer and duplication events of plant cell wall–degrading enzymes in distantly related cactophilic clades, suggesting that putatively adaptive traits evolved independently through disparate molecular mechanisms. Notably, we found that multiple cactophilic species and their close relatives have been reported as emerging human opportunistic pathogens, suggesting that the cactophilic lifestyle—and perhaps more generally lifestyles favoring thermotolerance—might preadapt yeasts to cause human disease. This work underscores the potential of a multifaceted approach involving high-throughput genomic and phenotypic data to shed light onto ecological adaptation and highlights how convergent evolution to wild environments could facilitate the transition to human pathogenicity.

59 BASIC BIOLOGICAL SCIENCES↗

Dissemination of blaNDM-5 and mcr-8.1 in carbapenem-resistant Klebsiella pneumoniae and Klebsiella quasipneumoniae in an animal breeding area in Eastern China

Animal farms have become one of the most important reservoirs of carbapenem-resistant Klebsiella spp. (CRK) owing to the wide usage of veterinary antibiotics. “One Health”-studies observing animals, the environment, and humans are necessary to understand the dissemination of CRK in animal breeding areas. Based on the concept of “One-Health,” 263 samples of animal feces, wastewater, well water, and human feces from 60 livestock and poultry farms in Shandong province, China were screened for CRK. Five carbapenem-resistant Klebsiella pneumoniae (CRKP) and three carbapenem-resistant Klebsiella quasipneumoniae (CRKQ) strains were isolated from animal feces, human feces, and well water. The eight strains were characterized by antimicrobial susceptibility testing, plasmid conjugation assays, whole-genome sequencing, and bioinformatics analysis. All strains carried the carbapenemase-encoding gene bla NDM-5 , which was flanked by the same core genetic structure (IS 5 - bla NDM-5 - ble MBL - trpF - dsbD -IS 26 -IS Kox3 ) and was located on highly related conjugative IncX3 plasmids. The colistin resistance gene mcr-8.1 was carried by three CRKP and located on self-transmissible IncFII(K)/IncFIA(HI1) and IncFII(pKP91)/IncFIA(HI1) plasmids. The genetic context of mcr-8.1 consisted of IS 903 - orf - mcr-8.1-copR-baeS-dgkA - orf -IS 903 in three strains. Single nucleotide polymorphism (SNP) analysis confirmed the clonal spread of CRKP carrying- bla NDM-5 and mcr-8.1 between two human workers in the same chicken farm. Additionally, the SNP analysis showed clonal expansion of CRKP and CRKQ strains from well water in different farms, and the clonal CRKP was clonally related to isolates from animal farms and a wastewater treatment plant collected in other studies in the same province. These findings suggest that CRKP and CRKQ are capable of disseminating via horizontal gene transfer and clonal expansion and may pose a significant threat to public health unless preventative measures are taken.

Yang, Chengxia↗

Molecular diversity and evolution of far-red light-acclimated photosystem I

The need to acclimate to different environmental conditions is central to the evolution of cyanobacteria. Far-red light (FRL) photoacclimation, or FaRLiP, is an acclimation mechanism that enables certain cyanobacteria to use FRL to drive photosynthesis. During this process, a well-defined gene cluster is upregulated, resulting in changes to the photosystems that allow them to absorb FRL to perform photochemistry. Because FaRLiP is widespread, and because it exemplifies cyanobacterial adaptation mechanisms in nature, it is of interest to understand its molecular evolution. Here, we performed a phylogenetic analysis of the photosystem I subunits encoded in the FaRLiP gene cluster and analyzed the available structural data to predict ancestral characteristics of FRL-absorbing photosystem I. The analysis suggests that FRL-specific photosystem I subunits arose relatively late during the evolution of cyanobacteria when compared with some of the FRL-specific subunits of photosystem II, and that the order Nodosilineales, which include strains like Halomicronema hongdechloris and Synechococcus sp. PCC 7335, could have obtained FaRLiP via horizontal gene transfer. We show that the ancestral form of FRL-absorbing photosystem I contained three chlorophyll f-binding sites in the PsaB2 subunit, and a rotated chlorophyll a molecule in the A0B site of the electron transfer chain. Along with our previous study of photosystem II expressed during FaRLiP, these studies describe the molecular evolution of the photosystem complexes encoded by the FaRLiP gene cluster.

ancestral sequence reconstruction↗

Sequencing the Genomes of the First Terrestrial Fungal Lineages: What Have We Learned?

The first genome sequenced of a eukaryotic organism was for Saccharomyces cerevisiae, as reported in 1996, but it was more than 10 years before any of the zygomycete fungi, which are the early-diverging terrestrial fungi currently placed in the phyla Mucoromycota and Zoopagomycota, were sequenced. The genome for Rhizopus delemar was completed in 2008; currently, more than 1000 zygomycete genomes have been sequenced. Genomic data from these early-diverging terrestrial fungi revealed deep phylogenetic separation of the two major clades—primarily plant—associated saprotrophic and mycorrhizal Mucoromycota versus the primarily mycoparasitic or animal-associated parasites and commensals in the Zoopagomycota. Genomic studies provide many valuable insights into how these fungi evolved in response to the challenges of living on land, including adaptations to sensing light and gravity, development of hyphal growth, and co-existence with the first terrestrial plants. Genome sequence data have facilitated studies of genome architecture, including a history of genome duplications and horizontal gene transfer events, distribution and organization of mating type loci, rDNA genes and transposable elements, methylation processes, and genes useful for various industrial applications. Pathogenicity genes and specialized secondary metabolites have also been detected in soil saprobes and pathogenic fungi. Novel endosymbiotic bacteria and viruses have been discovered during several zygomycete genome projects. Overall, genomic information has helped to resolve a plethora of research questions, from the placement of zygomycetes on the evolutionary tree of life and in natural ecosystems, to the applied biotechnological and medical questions.

59 BASIC BIOLOGICAL SCIENCES↗

Budding yeasts in the subphylum Saccharomycotina Genome sequencing and assembly

Eukaryotic life depends on the functional elements encoded by both the nuclear genome and organellar genomes, such as those contained within the mitochondria. The content, size, and structure of the mitochondrial genome varies across organisms with potentially large implications for phenotypic variance and resulting evolutionary trajectories. Among yeasts in the subphylum Saccharomycotina, extensive differences have been observed in various species relative to the model yeast Saccharomyces cerevisiae, but mitochondrial genome sampling across many groups has been scarce, even as hundreds of nuclear genomes have become available. By extracting mitochondrial reads from existing short-read genome sequence datasets, we have greatly expanded both the number of available genomes and the coverage across sparsely sampled clades. Comparison of 353 yeast mitochondrial genomes revealed that, while size and GC content were fairly consistent across species, those in the genera Metschnikowia and Saccharomyces trended larger, while several species in the order Saccharomycetales exhibited lower GC content. Extreme examples for both size and GC content were scattered throughout the subphylum. All mitochondrial genomes shared a core set of protein-coding genes for Complexes III, IV, and V, but they varied in the presence or absence of mitochondrially-encoded canonical Complex I genes. We traced the loss of Complex I genes to a major event in the ancestor of the orders Saccharomycetales and Saccharomycodales, but we also observed several independent losses in the orders Phaffomycetales, Pichiales, and Dipodascales. In contrast to prior hypotheses based on smaller-scale datasets, comparison of evolutionary rates in protein-coding genes showed no bias towards elevated rates among aerobically fermenting (Crabtree/Warburg-positive) yeasts. Mitochondrial introns were widely distributed, but highly enriched in some groups. The majority of mitochondrial introns were poorly conserved within groups, but several were shared within groups, between groups, and even across taxonomic orders, which is consistent with horizontal gene transfer, likely involving homing endonucleases acting as selfish elements. As the number of available fungal nuclear genomes continues to expand, the methods described here to retrieve mitochondrial genome sequences from these datasets will prove invaluable to ensuring that studies of fungal mitochondrial genomes keep pace with their nuclear counterparts.

diversity↗

An archaeal genomic signature

Comparisons of complete genome sequences allow the most objective and comprehensive descriptions possible of a lineage's evolution. This communication uses the completed genomes from four major euryarchaeal taxa to define a genomic signature for the Euryarchaeota and, by extension, the Archaea as a whole. The signature is defined in terms of the set of protein-encoding genes found in at least two diverse members of the euryarchaeal taxa that function uniquely within the Archaea; most signature proteins have no recognizable bacterial or eukaryal homologs. By this definition, 351 clusters of signature proteins have been identified. Functions of most proteins in this signature set are currently unknown. At least 70% of the clusters that contain proteins from all the euryarchaeal genomes also have crenarchaeal homologs. This conservative set, which appears refractory to horizontal gene transfer to the Bacteria or the Eukarya, would seem to reflect the significant innovations that were unique and fundamental to the archaeal "design fabric." Genomic protein signature analysis methods may be extended to characterize the evolution of any phylogenetically defined lineage. The complete set of protein clusters for the archaeal genomic signature is presented as supplementary material (see the PNAS web site, www.pnas.org).

Non-NASA Center↗

A genomic timescale of prokaryote evolution: insights into the origin of methanogenesis, phototrophy, and the colonization of land

BACKGROUND: The timescale of prokaryote evolution has been difficult to reconstruct because of a limited fossil record and complexities associated with molecular clocks and deep divergences. However, the relatively large number of genome sequences currently available has provided a better opportunity to control for potential biases such as horizontal gene transfer and rate differences among lineages. We assembled a data set of sequences from 32 proteins (approximately 7600 amino acids) common to 72 species and estimated phylogenetic relationships and divergence times with a local clock method. RESULTS: Our phylogenetic results support most of the currently recognized higher-level groupings of prokaryotes. Of particular interest is a well-supported group of three major lineages of eubacteria (Actinobacteria, Deinococcus, and Cyanobacteria) that we call Terrabacteria and associate with an early colonization of land. Divergence time estimates for the major groups of eubacteria are between 2.5-3.2 billion years ago (Ga) while those for archaebacteria are mostly between 3.1-4.1 Ga. The time estimates suggest a Hadean origin of life (prior to 4.1 Ga), an early origin of methanogenesis (3.8-4.1 Ga), an origin of anaerobic methanotrophy after 3.1 Ga, an origin of phototrophy prior to 3.2 Ga, an early colonization of land 2.8-3.1 Ga, and an origin of aerobic methanotrophy 2.5-2.8 Ga. CONCLUSIONS: Our early time estimates for methanogenesis support the consideration of methane, in addition to carbon dioxide, as a greenhouse gas responsible for the early warming of the Earths' surface. Our divergence times for the origin of anaerobic methanotrophy are compatible with highly depleted carbon isotopic values found in rocks dated 2.8-2.6 Ga. An early origin of phototrophy is consistent with the earliest bacterial mats and structures identified as stromatolites, but a 2.6 Ga origin of cyanobacteria suggests that those Archean structures, if biologically produced, were made by anoxygenic photosynthesizers. The resistance to desiccation of Terrabacteria and their elaboration of photoprotective compounds suggests that the common ancestor of this group inhabited land. If true, then oxygenic photosynthesis may owe its origin to terrestrial adaptations.

Methane/metabolism↗

Population dynamics of transgenic strain Escherichia coli Z905/pPHL7 in freshwater and saline lake water microcosms with differing microbial community structures

Populations of Escherichia coli Z905/pPHL7, a transgenic microorganism, were heterogenic in the expression of plasmid genes when adapting to the conditions of water microcosms of various mineralization levels and structure of microbial community. This TM has formed two subpopulations (ampicillin-resistant and ampicillin-sensitive) in every microcosm. Irrespective of mineralization level of a microcosm, when E. coli Z905/pPHL7 alone was introduced, the ampicillin-resistant subpopulation prevailed, while introduction of the TM together with indigenous bacteria led to the dominance of the ampicillin-sensitive subpopulation. A high level of lux gene expression maintained longer in the freshwater microcosms than in sterile saline lake water microcosms. A horizontal gene transfer has been revealed between the jointly introduced TM and Micrococcus sp. 9/pSH1 in microcosms with the Lake Shira sterile water. c2005 COSPAR. Published by Elsevier Ltd. All rights reserved.

Ecosystem↗

Pioneering a Biobased UAS

With the exponential growth of interest in unmanned aerial vehicles (UAVs) and their vast array of applications in both space exploration and terrestrial uses such as the delivery of medicine and monitoring the environment, the 2014 Stanford-Brown-Spelman iGEM team is pioneering the development of a fully biological UAV for scientific and humanitarian missions. The prospect of a biologically-produced UAV presents numerous advantages over the current manufacturing paradigm. First, a foundational architecture built by cells allows for construction or repair in locations where it would be difficult to bring traditional tools of production. Second, a major limitation of current research with UAVs is the size and high power consumption of analytical instruments, which require bulky electrical components and large fuselages to support their weight. By moving these functions into cells with biosensing capabilities - for example, a series of cells engineered to report GFP, green fluorescent protein, when conditions exceed a certain threshold concentration of a compound of interest, enabling their detection post-flight - these problems of scale can be avoided. To this end, we are working to engineer cells to synthesize cellulose acetate as a novel bioplastic, characterize biological methods of waterproofing the material, and program this material's systemic biodegradation. In addition, we aim to use an "amberless" system to prevent horizontal gene transfer from live cells on the material to microorganisms in the flight environment. So far, we have: successfully transformed Gluconacetobacter hansenii, a cellulose-producing bacterium, with a series of promoters to test transformation efficiency before adding the acetylation genes; isolated protein bands present in the wasp nest material; transformed the cellulose-degrading genes into Escherichia coli; and we have confirmed that the amberless construct prevents protein expression in wild-type cells. In addition, as part of our human outreach project, we have been in touch with leaders in the fields of UAVs, synthetic biology, and earth sciences, and it is clear that biodegradable UAVs could have a significant impact on the industry.

Escherichia↗

Towards a Biosynthetic UAV

We are currently working on a series of projects towards the construction of a fully biological unmanned aerial vehicle (UAV) for use in scientific and humanitarian missions. The prospect of a biologically-produced UAV presents numerous advantages over the current manufacturing paradigm. First, a foundational architecture built by cells allows for construction or repair in locations where it would be difficult to bring traditional tools of production. Second, a major limitation of current research with UAVs is the size and high power consumption of analytical instruments, which require bulky electrical components and large fuselages to support their weight. By moving these functions into cells with biosensing capabilities - for example, a series of cells engineered to report GFP, green fluorescent protein, when conditions exceed a certain threshold concentration of a compound of interest, enabling their detection post-flight - these problems of scale can be avoided. To this end, we are working to engineer cells to synthesize cellulose acetate as a novel bioplastic, characterize biological methods of waterproofing the material, and program this material's systemic biodegradation. In addition, we aim to use an "amberless" system to prevent horizontal gene transfer from live cells on the material to microorganisms in the flight environment.

Biological↗

Pioneering a Biobased UAS

With the exponential growth of interest in unmanned aerial vehicles (UAVs) and their vast array of applications in both space exploration and terrestrial uses such as the delivery of medicine and monitoring the environment, the 2014 Stanford-Brown-Spelman iGEM team is pioneering the development of a fully biological UAV for scientific and humanitarian missions. The prospect of a biologically-produced UAV presents numerous advantages over the current manufacturing paradigm. First, a foundational architecture built by cells allows for construction or repair in locations where it would be difficult to bring traditional tools of production. Second, a major limitation of current research with UAVs is the size and high power consumption of analytical instruments, which require bulky electrical components and large fuselages to support their weight. By moving these functions into cells with biosensing capabilities – for example, a series of cells engineered to report GFP, green fluorescent protein, when conditions exceed a certain threshold concentration of a compound of interest, enabling their detection post-flight – these problems of scale can be avoided. To this end, we are working to engineer cells to synthesize cellulose acetate as a novel bioplastic, characterize biological methods of waterproofing the material, and program this material’s systemic biodegradation. In addition, we aim to use an “amberless” system to prevent horizontal gene transfer from live cells on the material to microorganisms in the flight environment. So far, we have: successfully transformed Gluconacetobacter hansenii, a cellulose-producing bacterium, with a series of promoters to test transformation efficiency before adding the acetylation genes; isolated protein bands present in the wasp nest material; transformed the cellulose-degrading genes into Escherichia coli; and we have confirmed that the amberless construct prevents protein expression in wild-type cells. In addition, as part of our human outreach project, we have been in touch with leaders in the fields of UAVs, synthetic biology, and earth sciences, and it is clear that biodegradable UAVs could have a significant impact on the industry.

Escherichia↗

Comparison of Auxenochlorella protothecoides and Chlorella spp. Chloroplast Genomes: Evidence for Endosymbiosis and Horizontal Virus-like Gene Transfer

Resequencing of the chloroplast genome (cpDNA) of Auxenochlorella protothecoides UTEX 25 was completed (GenBank Accession no. KC631634.1), revealing a genome size of 84,576 base pairs and 30.8% GC content, consistent with features reported for the previously sequenced A. protothecoides 0710, (GenBank Accession no. KC843975). The A. protothecoides UTEX 25 cpDNA encoded 78 predicted open reading frames, 32 tRNAs, and 4 rRNAs, making it smaller and more compact than the cpDNA genome of C. variabilis (124,579 bp) and C. vulgaris (150,613 bp). By comparison, the compact genome size of A. protothecoides was attributable primarily to a lower intergenic sequence content. The cpDNA coding regions of all known Chlorella species were found to be organized in conserved colinear blocks, with some rearrangements. The Auxenochlorella and Chlorella species genome structure and composition were similar, and of particular interest were genes influencing photosynthetic efficiency, i.e., chlorophyll synthesis and photosystem subunit I and II genes, consistent with other biofuel species of interest. Phylogenetic analysis revealed that Prototheca cutis is the closest known A. protothecoides relative, followed by members of the genus Chlorella. The cpDNA of A. protothecoides encodes 37 genes that are highly homologous to representative cyanobacteria species, including rrn16, rrn23, and psbA, corroborating a well-recognized symbiosis. Several putative coding regions were identified that shared high nucleotide sequence identity with virus-like sequences, suggestive of horizontal gene transfer. Despite these predictions, no corresponding transcripts were obtained by RT-PCR amplification, indicating they are unlikely to be expressed in the extant lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting variable gene content in Escherichia coli using conserved genes

Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%–90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943–0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins” was accurately predicted (F1 = 0.902 [0.898–0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876–0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

Giant Starship Elements Mobilize Accessory Genes in Fungal Genomes

Accessory genes are variably present among members of a species and are a reservoir of adaptive functions. In bacteria, differences in gene distributions among individuals largely result from mobile elements that acquire and disperse accessory genes as cargo. In contrast, the impact of cargo-carrying elements on eukaryotic evolution remains largely unknown. Here, we show that variation in genome content within multiple fungal species is facilitated by Starships, a newly discovered group of massive mobile elements that are 110 kb long on average, share conserved components, and carry diverse arrays of accessory genes. We identified hundreds of Starship-like regions across every major class of filamentous Ascomycetes, including 28 distinct Starships that range from 27 to 393 kb and last shared a common ancestor ca. 400 Ma. Using new long-read assemblies of the plant pathogen Macrophomina phaseolina, we characterize four additional Starships whose activities contribute to standing variation in genome structure and content. One of these elements, Voyager, inserts into 5S rDNA and contains a candidate virulence factor whose increasing copy number has contrasting associations with pathogenic and saprophytic growth, suggesting Voyager’s activity underlies an ecological trade-off. We propose that Starships are eukaryotic analogs of bacterial integrative and conjugative elements based on parallels between their conserved components and may therefore represent the first dedicated agents of active gene transfer in eukaryotes. Our results suggest that Starships have shaped the content and structure of fungal genomes for millions of years and reveal a new concerted route for evolution throughout an entire eukaryotic phylum.

59 BASIC BIOLOGICAL SCIENCES↗

An orphan gene BOOSTER enhances photosynthetic efficiency and plant productivity

Organelle-to-nucleus DNA transfer is an ongoing process playing an important role in the evolution of eukaryotic life. Here, genome-wide association studies (GWAS) of non-photochemical quenching parameters in 743 Populus trichocarpa accessions identified a nuclear-encoded genomic region associated with variation in photosynthesis under fluctuating light. The identified gene, BOOSTER (BSTR), comprises three exons, two with apparent endophytic origin and the third containing a large fragment of plastid-encoded Rubisco large subunit. Higher expression of BSTR facilitated anterograde signaling between nucleus and plastid, which corresponded to enhanced expression of Rubisco, increased photosynthesis, and up to 35% greater plant height and 88% biomass in poplar accessions under field conditions. Overexpression of BSTR in Populus tremula × P. alba achieved up to a 200% in plant height. Similarly, Arabidopsis plants heterologously expressing BSTR gained up to 200% in biomass and up to 50% increase in seed.

60 APPLIED LIFE SCIENCES↗