Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genomic selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Genomic Features and Pervasive Negative Selection in Rhodanobacter Strains Isolated from Nitrate and Heavy Metal Contaminated Aquifer

Despite the dominance of Rhodanobacter species in the subsurface of the contaminated Oak Ridge Reservation (ORR) site, very little is known about the mechanisms underlying their adaptions to the various stressors present at ORR. Recently, multiple Rhodanobacter strains have been isolated from the ORR groundwater samples from several wells with varying geochemical properties.

59 BASIC BIOLOGICAL SCIENCES↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Simulation of sugar kelp ( Saccharina latissima ) breeding guided by practices to accelerate genetic gains

Abstract Though Saccharina japonica cultivation has been established for many decades in East Asian countries, the domestication process of sugar kelp (Saccharina latissima) in the Northeast United States is still at its infancy. In this study, by using data from our breeding experience, we will demonstrate how obstacles for accelerated genetic gain can be assessed using simulation approaches that inform resource allocation decisions. Thus far, we have used 140 wild sporophytes that were sampled in 2018 from the northern Gulf of Maine to southern New England. From these sporophytes, we sampled gametophytes and made and evaluated over 600 progeny sporophytes from crosses among the gametophytes in 2019 and 2020. The biphasic life cycle of kelp gives a great advantage in selective breeding as we can potentially select both on the sporophytes and gametophytes. However, several obstacles exist, such as the amount of time it takes to complete a breeding cycle, the number of gametophytes that can be maintained in the laboratory, and whether positive selection can be conducted on farm-tested sporophytes. Using the Gulf of Maine population characteristics for heritability and effective population size, we simulated a founder population of 1,000 individuals and evaluated the impact of overcoming these obstacles on rate of genetic gain. Our results showed that key factors to improve current genetic gain rely mainly on our ability to induce reproduction of the best farm-tested sporophytes, and to accelerate the clonal vegetative growth of released gametophytes so that enough gametophyte biomass is ready for making crosses by the next growing season. Overcoming these challenges could improve rates of genetic gain more than 2-fold. Future research should focus on conditions favorable for inducing spring reproduction, and on increasing the amount of gametophyte tissue available in time to make fall crosses in the same year.

59 BASIC BIOLOGICAL SCIENCES↗

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS↗

Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

Abstract The distribution of ecotypic variation in natural populations is influenced by neutral and adaptive evolutionary forces that are challenging to disentangle. This study provides a high‐resolution portrait of genomic variation in Chinook salmon ( Oncorhynchus tshawytscha ) with emphasis on a region of major effect for ecotypic variation in migration timing. With a filtered data set of ~13 million single nucleotide polymorphisms (SNPs) from low‐coverage whole genome resequencing of 53 populations (3566 barcoded individuals), we contrasted patterns of genomic structure within and among major lineages and examined the extent of a selective sweep at a major effect region underlying migration timing (GREB1L/ROCK1). Neutral variation provided support for fine‐scale structure of populations, while allele frequency variation in GREB1L/ROCK1 was highly correlated with mean return timing for early and late migrating populations within each of the lineages ( r 2 = .58–.95; p < .001). However, the extent of selection within the genomic region controlling migration timing was much narrower in one lineage (interior stream‐type) compared to the other two major lineages, which corresponded to the breadth of phenotypic variation in migration timing observed among lineages. Evidence of a duplicated block within GREB1L/ROCK1 may be responsible for reduced recombination in this portion of the genome and contributes to phenotypic variation within and across lineages. Lastly, SNP positions across GREB1L/ROCK1 were assessed for their utility in discriminating migration timing among lineages, and we recommend multiple markers nearest the duplication to provide highest accuracy in conservation applications such as those that aim to protect early migrating Chinook salmon. These results highlight the need to investigate variation throughout the genome and the effects of structural variants on ecologically relevant phenotypic variation in natural species.

Horn, Rebekah L.↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

p53 deficiency alters the yield and spectrum of radiation-induced lacZ mutants in the brain of transgenic mice

Exposure to heavy particle radiation in the galacto-cosmic environment poses a significant risk in space exploration and the evaluation of radiation-induced genetic damage in tissues, especially in the central nervous system, is an important consideration in long-term manned space missions. We used a plasmid-based transgenic mouse model system, with the pUR288 lacZ transgene integrated in the genome of every cell of C57Bl/6(lacZ) mice, to evaluate the genetic damage induced by iron particle radiation. In order to examine the importance of genetic background on the radiation sensitivity of individuals, we cross-bred p53 wild-type lacZ transgenic mice with p53 nullizygous mice, producing lacZ transgenic mice that were either hemizygous or nullizygous for the p53 tumor suppressor gene. Animals were exposed to an acute dose of 1 Gy of iron particles and the lacZ mutation frequency (MF) in the brain was measured at time intervals from 1 to 16 weeks post-irradiation. Our results suggest that iron particles induced an increase in lacZ MF (2.4-fold increase in p53+/+ mice, 1.3-fold increase in p53+/- mice and 2.1-fold increase in p53-/- mice) and that this induction is both temporally regulated and p53 genotype dependent. Characterization of mutants based on their restriction patterns showed that the majority of the mutants arising spontaneously are derived from point mutations or small deletions in all three genotypes. Radiation induced alterations in the spectrum of deletion mutants and reorganization of the genome, as evidenced by the selection of mutants containing mouse genomic DNA. These observations are unique in that mutations in brain tissue after particle radiation exposure have never before been reported owing to technical limitations in most other mutation assays.

Non-NASA Center↗

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST↗

Developing a Cas9-Based Tool to Engineer Native Plasmids in Synechocystis sp. PCC 6803

The oxygenic photosynthetic bacterium Synechocystis sp. PCC 6803 (S6803) is a model cyanobacterium widely used for fundamental research and biotechnology applications. Due to its polyploidy, existing methods for genome engineering of S6803 require multiple rounds of selection to modify all genome copies, which is time consuming and inefficient. In this study, we engineered the Cas9 tool for onestep, segregationfree genome engineering. We further used our Cas9 tool to delete three of seven S6803 native plasmids. Our results show that all three smallsize native plasmids, but not the largesize native plasmids, can be deleted with this tool. To further facilitate heterologous gene expression in S6803, a shuttle vector based on the native plasmid pCC5.2 was created. The shuttle vector can be introduced into Cas9containing S6803 in one step without requiring segregation and can be stably maintained without antibiotic pressure for at least 30 days. Moreover, genes encoded on the shuttle vector remain functional after 30 days of continuous cultivation without selective pressure. Thus, this study provides a set of new tools for rapid modification of the S6803 genome and for stable expression of heterologous genes, potentially facilitating both fundamental research and biotechnology applications using S6803.

biotechnology↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

eYGFPuv-Assisted Transgenic Selection in Populus deltoides WV94 and Multiplex Genome Editing in Protoplasts of P. trichocarpa × P. deltoides Clone ‘52-225’

Although CRISPR/Cas-based genome editing has been widely used for plant genetic engineering, its application in the genetic improvement of trees has been limited, partly because of challenges in Agrobacterium-mediated transformation. As an important model for poplar genomics and biotechnology research, eastern cottonwood (Populus deltoides) clone WV94 can be transformed by A. tumefaciens, but several challenges remain unresolved, including the relatively low transformation efficiency and the relatively high rate of false positives from antibiotic-based selection of transgenic events. Moreover, the efficacy of CRISPR-Cas system has not been explored in P. deltoides yet. Here, we first optimized the protocol for Agrobacterium-mediated stable transformation in P. deltoides WV94 and applied a UV-visible reporter called eYGFPuv in transformation. Our results showed that the transgenic events in the early stage of transformation could be easily recognized and counted in a non-invasive manner to narrow down the number of regenerated shoots for further molecular characterization (at the DNA or mRNA level) using PCR. We found that approximately 8.7% of explants regenerated transgenic shoots with green fluorescence within two months. Next, we examined the efficacy of multiplex CRISPR-based genome editing in the protoplasts derived from P. deltoides WV94 and hybrid poplar clone ‘52-225’ (P. trichocarpa × P. deltoides clone ‘52-225’). The two constructs expressing the Trex2-Cas9 system resulted in mutation efficiency ranging from 31% to 57% in hybrid poplar clone 52-225, but no editing events were observed in P. deltoides WV94 transient assay. The eYGFPuv-assisted plant transformation and genome editing approach demonstrated in this study has great potential for accelerating the genome editing-based breeding process in poplar and other non-model plants species and point to the need for additional CRISPR work in P. deltoides.

59 BASIC BIOLOGICAL SCIENCES↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

In vivo human T cell engineering with enveloped delivery vehicles

Viruses and virally derived particles have the intrinsic capacity to deliver molecules to cells, but the difficulty of readily altering cell-type selectivity has hindered their use for therapeutic delivery. Here, we show that cell surface marker recognition by antibody fragments displayed on membrane-derived particles encapsulating CRISPR–Cas9 protein and guide RNA can deliver genome editing tools to specific cells. Compared to conventional vectors like adeno-associated virus that rely on evolved capsid tropisms to deliver virally encoded cargo, these Cas9-packaging enveloped delivery vehicles (Cas9-EDVs) leverage predictable antibody–antigen interactions to transiently deliver genome editing machinery selectively to cells of interest. Antibody-targeted Cas9-EDVs preferentially confer genome editing in cognate target cells over bystander cells in mixed populations, both ex vivo and in vivo. By using multiplexed targeting molecules to direct delivery to human T cells, Cas9-EDVs enable the generation of genome-edited chimeric antigen receptor T cells in humanized mice, establishing a programmable delivery modality with the potential for widespread therapeutic utility.

42 ENGINEERING↗

Wireworm (Coleoptera: Elateridae) genomic analysis reveals putative cryptic species, population structure, and adaptation to pest control

The larvae of click beetles (Coleoptera: Elateridae), known as “wireworms,” are agricultural pests that pose a substantial economic threat worldwide. We produced one of the first wireworm genome assemblies (Limonius californicus), and investigated population structure and phylogenetic relationships of three species (L. californicus, L. infuscatus, L. canus) across the northwest US and southwest Canada using genome-wide markers (RADseq) and genome skimming. We found two species (L. californicus and L. infuscatus) are comprised of multiple genetically distinct groups that diverged in the Pleistocene but have no known distinguishing morphological characters, and therefore could be considered cryptic species complexes. We also found within-species population structure across relatively short geographic distances. Genome scans for selection provided preliminary evidence for signatures of adaptation associated with different pesticide treatments in an agricultural field trial for L. canus. We demonstrate that genomic tools can be a strong asset in developing effective wireworm control strategies.

59 BASIC BIOLOGICAL SCIENCES↗

Developing Non-Food Grade Brassica Biofuel Feedstock Cultivars with High Yield, Oil Content, and Oil Quality that are Suitable for Low Input Production Dryland Systems (Final Report)

The U.S. uses a substantial amount of fossil fuel as an energy source for a wide range for functions including home heating, agriculture and transportation. In the transportation sector, diesel and jet fuel are consumed at a rapid rate, and alternative liquid energy is being investigated globally and nationally to reduce our dependence on fossil fuel and reduce the impact of our carbon footprint on global climate change. Non-food Brassica crops have the potential of producing high oil yield (over 250 gal acre-1) and have oil quality highly desirable for use as biodiesel or bio jet fuel. Developing oilseed feedstock Brassica cultivars with higher seed and oil yield, with high oil quality and with resistance to pathogens, that can be grown with few chemical inputs will helping break our dependence on fossil fuels and reduce importation of fossil fuels. While some oilseed Brassicas have been grown on a small scale for many years in the Pacific Northwest (PNW), adoption has been limited, and the potential of the crops have not been realized or even fully investigated. This report summarizes the results of a study to develop superior non-food grade winter (B. napus) and spring (B. napus and B. juncea) oilseed cultivars suitable for a range of PNW, and other US environments with high resistance to the biotic and abiotic stresses suitable for high-quality biofuel feedstocks. In conducting this work, genome-wide association selection was used to dissect the genetic architecture of industrial Brassica oilseed germplasm for yield, quality, and resistance to biotic and abiotic stresses. A genome-wide bioinformatics approach was used to identify putative PRR (pattern recognition receptor) - type resistance genes that confer durable resistance to blackleg. A novel transgenic approach was developed to generate resistant non-food oilseed lines using PPR genes Br1033 and Br8486. These genes were inserted into regionally adapted oilseed cultivars.

09 BIOMASS FUELS↗

Population genomics provides insights into the genetic basis of adaptive evolution in the mushroom-forming fungus Lentinula edodes

Introduction: Mushroom-forming fungi comprise diverse species that develop complex multicellular structures. In cultivated species, both ecological adaptation and artificial selection have driven genome evolution. However, little is known about the connections among genotype, phenotype and adaptation in mushroom-forming fungi. Objectives: This study aimed to (1) uncover the population structure and demographic history of Lentinula edodes, (2) dissect the genetic basis of adaptive evolution in L. edodes, and (3) determine if genes related to fruiting body development are involved in adaptive evolution. Methods: We analyzed genomes and fruiting body-related traits (FBRTs) in 133 L. edodes strains and conducted RNA-seq analysis of fruiting body development in the YS69 strain. Combined methods of genomic scan for divergence, genome-wide association studies (GWAS), and RNA-seq were used to dissect the genetic basis of adaptive evolution. Results: We detected three distinct subgroups of L. edodes via single nucleotide polymorphisms, which showed robust phenotypic and temperature response differentiation and correlation with geographical distribution. Demographic history inference suggests that the subgroups diverged 36,871 generations ago. Moreover, L. edodes cultivars in China may have originated from the vicinity of Northeast China. A total of 942 genes were found to be related to genetic divergence by genomic scan, and 719 genes were identified to be candidates underlying FBRTs by GWAS. Integrating results of genomic scan and GWAS, 80 genes were detected to be related to phenotypic differentiation. A total of 364 genes related to fruiting body development were involved in genetic divergence and phenotypic differentiation. Conclusion: Adaptation to the local environment, especially temperature, triggered genetic divergence and phenotypic differentiation of L. edodes. A general model for genetic divergence and phenotypic differentiation during adaptive evolution in L. edodes, which involves in signal perception and transduction, transcriptional regulation, and fruiting body morphogenesis, was also integrated here.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic insights into the origin, domestication and diversification of Brassica juncea

Despite early domestication around 3000 BC, the evolutionary history of the ancient allotetraploid species Brassica juncea (L.) Czern & Coss remains uncertain. Here, we report a chromosome-scale de novo assembly of a yellow-seeded B. juncea genome by integrating long-read and short-read sequencing, optical mapping and Hi-C technologies. Nuclear and organelle phylogenies of 480 accessions worldwide supported that B. juncea is most likely a single origin in West Asia, 8,000–14,000 years ago, via natural interspecific hybridization. Subsequently, new crop types evolved through spontaneous gene mutations and introgressions along three independent routes of eastward expansion. Selective sweeps, genome-wide trait associations and tissue-specific RNA-sequencing analysis shed light on the domestication history of flowering time and seed weight, and on human selection for morphological diversification in this versatile species. Our data provide a comprehensive insight into the origin and domestication and a foundation for genomics-based breeding of B. juncea.

59 BASIC BIOLOGICAL SCIENCES↗