Engineering PapersSearch

SEARCH · Engineering Papers

Results for “genomic selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

2024 IUFRO Tree Biotechnology Conference (Aug 4-8, 2024)

The 2024 IUFRO Tree Biotechnology Conference is the biennial meeting on genomics, molecular biology, and biotechnology of forest trees, associated with the IUFRO Working Party 2.04.06. This year's meeting was held in Annapolis, MD, USA from August 4th to 8th and was hosted by Yiping Qi (University of Maryland), Edward Eisenstein (University of Maryland), Gary Coleman (University of Maryland), and Heather Coleman (Syracuse University). The conference covered seven topics over the course of five days: 1) Biological and ecological insights from OMICS, 2) Advancing technologies for targeted trait manipulation and acceptability to diverse tree species, 3) Genes, development, and physiology, 4) Translating genomics and biotechnology to practice, 5) Trees in a changing world, 6) Genetic and phenotypic diversity for breeding and genomic selection, and 7) Biotechnology for biomaterials and bioeconomy. In addition to the sessions, there were two plenary sessions, provided by John Ralph (University of Wisconsin) and Tanja Pyrhäjärvi (University of Helsinki). The meeting celebrated the second awardees of the newly created IUFRO WG 2.04.06 Award: Excellence in Forest Molecular Biology and Genomics, which was presented to Chung-Jui (C.J.) Tsai (University of Georgia). Greg Goralogia (Oregon State University) was the recipient of the associated Early Career Award. The scientific presentations at the conference highlighted cutting-edge advancements in many facets of forest biotechnology research, including applications of genomic selection in forest genetics and breeding, the use of genetic editing, tree physiology, stress response, molecular breeding, wood development, "omics" technologies, and the social and economic impacts of genetically modified (GM) trees. Scientific take homes from the meeting include the power of NMR to dissect the composition of lignin, the genomic diversity of forest trees that has enormous potential for tree improvement and the integration of systems biology with climate and geographical data. The conference attracted a mix of students (25), postdoctoral fellows (32), and scientists from academia (66) and industry (18). In all, the conference was attended by 141 registered participants, representing 20 countries that participated in 23 invited lectures (including 6 'early-career' keynotes), 27 voluntary talks and 61 poster presentations. Support for the conference was drawn from a wide variety of Academia, Industry, and Government sources, and included financial support from several tree improvement companies. Overall, the conference was a great success, providing an exceptional mix of science and social activities in a relaxed and collegial atmosphere. More information about the meeting can be found at treebiotech.org. The next meeting will be held in Stellenbosch, South Africa, in 2026, hosted jointly by Zander Myburg, Dave Drew (University of Stellenbosch,) and Sanushka Naidoo (University of Pretoria, FABI).

59 BASIC BIOLOGICAL SCIENCES

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES

Scaffolded and annotated nuclear and organelle genomes of the North American brown alga Saccharina latissima

Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.

DeWeese, Kelly

Apomixis in Farmers’ Fields: Overview, Case Studies from Forage Grasses and Considerations for Future Apomictic Crops

Apomixis occurs naturally in several commercially important species from diverse plant families. While in some of these species apomixis is yet to be exploited in breeding schemes aimed at fixing heterosis, genetic progress and cultivar development, in other species apomixis has been integrated at different stages of breeding. Some of the most relevant examples come from the subfamily Panicoideae, the second largest subfamily of the Poaceae, and are the main focus of this review. The subfamily encompasses many tropical and sub-tropical grasses and grains of worldwide economic importance. Apomictic tropical forages are prime examples of how apomixis can be used and exploited in the development of marketable cultivars, which are essential to the meat and milk production industries globally. The main commercial forages used as grass pastures covering millions of hectares in tropical and sub-tropical regions are polyploids exhibiting gametophytic apomixis that belong to the genus Urochloa spp. (brachiariagrasses) and to the species Megathyrsus maximus (guineagrass). Buffel grass (Cenchrus ciliaris) and Paspalum spp. are other important apomictic forages bred and used in these regions. Breeding involves large germplasm collections from the centers of origin of the species, and for most of them, sexually reproducing diploid plants have been found. Chromosomically duplicated plants that maintain sexual reproduction are used in crosses with apomictic genotypes for the development and selection of cultivars to be marketed or used as progenitors in subsequent breeding cycles. The peculiarities of each genus/species breeding programs, the cultivars obtained from these programs, and the impact of use of marker assisted selection in cultivar development are presented. In addition, the test or implementation of new technologies such as high throughput phenotyping, and the use of machine learning methods for trait prediction and genomic selection are positively impacting the selection and speed of development of new polyploid apomictic cultivars. Furthermore, genetic transformation techniques, including genome editing, provide an additional layer for design of tailor-made, customer-oriented cultivars.

Cenchrus

Multi-trait multi-environment genomic prediction strategies for Miscanthus sacchariflorus

Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.

genomic prediction (GP)

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology

Plastome evolution in annual Brachypodium species reveals widespread heteroplasmy and chloroplast capture, lineage-specific codon usage bias, and low positive selection

Comparative genomics and plastome phylogenomics have advanced significantly in recent years, highlighting the diversity, possible admixture, and non-neutral evolution of the predominantly considered non-recombinant chloroplast genomes in angiosperms. The grass genus Brachypodium serves as a powerful model for studying evolutionary processes in monocots. We analyzed 287 plastomes across the native circum-Mediterranean range of the three annual Brachypodium species ( B. distachyon, B. stacei, B. hybridum ), focusing on their structural variation, selection patterns and phylogenomic relationships. Our analyses confirmed the differentiation of the S and D plastomes, inherited respectively from the diploid progenitor species B. stacei and B. distachyon . We identified novel structural rearrangements and indels, and unique repeat motifs, along with widespread heteroplasmy, particularly in ancestral B. hybridum -D plastotypes. SNP diversity varied among plastotypes, reflecting population dynamics and evolutionary histories, with B. hybridum -D plastotypes showing the highest normalized diversity and B. hybridum -S the lowest. Positive selection was detected in 29 plastid genes by Tajima’s neutrality test, and in nine genes by site and branch-site evolutionary models, including matK, ndhF, rbcL, and rpoC2. Phylogenomic analyses revealed well-supported clades corresponding to the S and D plastome lineages, with frequent chloroplast capture events and long-distance dispersals shaping their evolutionary trajectories.

allopolyploidy

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio

Phage-based delivery of CRISPR-associated transposases for targeted bacterial editing

Phage λ, a well-characterized temperate phage, has been recently leveraged for bacterial genome editing by selectively delivering base editors into targeted bacterial species. We extend this concept by engineering phage λ to deliver CRISPR-guided transposases, accomplishing large insertions and targeted gene disruptions. To achieve this, we engineered phage λ using homologous recombination paired with Cas13a-based counterselection for precise phage modifications. Initially, we established the utility of Cas13a in phage λ by conducting minimal recoding edits, deletions, and insertions. Subsequently, we scaled up the engineering to embed the comprehensive DNA-editing CRISPR-Cas transposase (DART) system within the phage genome, creating λ-DART phages. These modified λ-DART phages were then employed to infectEscherichia coli, generating CRISPR RNA-guided transposition events in the host genome. Applying our engineered λ-DART phages to monocultures and a mixed bacterial community comprising three genera led to efficient, precise, and specific gene knockouts and insertions in the targetedE. colicells, achieving editing efficiencies surpassing 50% of the population. This research enhances phage-mediated genome editing by enabling efficient in situ gene integrations in bacteria, offering an avenue for further application in microbial community contexts. This scalable method enables flexible microbial genome editing in situ to manipulate the function and composition of diverse ecosystems.

Science & Technology - Other Topics

Structure and sequence evolution in the pennycress ( Thlaspi arvense ) pangenome

Eukaryotic genomes harbor many forms of variation, including nucleotide diversity and structural polymorphisms, which experience natural selection and contribute to genome evolution and biodiversity. Harnessing this variation for agriculture hinges on our ability to detect, quantify, catalog, and deploy genetic diversity. Here, we explore seven complete genomes of the emerging biofuel crop pennycress ( Thlaspi arvense ) drawn from across the species' current genetic diversity to catalog variation in genome structure and content. Across this new pangenome resource, we find contrasting evolutionary modes in different genomic zones. Gene-poor, repeat-rich pericentromeric regions experience frequent rearrangements, including repeated centromere repositioning. By contrast, conserved gene-dense chromosome arms maintain large-scale synteny across accessions even in fast-evolving NOD-like receptor immune genes, where microsynteny breaks down across species, but gene cluster positioning macrosynteny is maintained. Our findings highlight that multiple elements of the genome experience dynamic evolution that conserves functional content on the chromosome scale but allows repositioning and presence–absence variation on a local scale. This diversity is invisible to classical reference-based strategies and highlights the strength and utility of pangenomic resources. These results provide a valuable case study of rapid genomic structural evolution within a species and powerful resources for crop development in an emerging biofuel crop.

Thlaspi arvense