Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic mutations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗

Chromosome-scale Genome Assembly of the Most Abundant Ectomycorrhizal Fungus Cenococcum Geophilum Reveals Massive TE Expansion and RIP Defense Mechanism

Transposable elements (TEs) play crucial roles in genome evolution and ecological adaptation in fungi, yet their dynamics in ectomycorrhizal species remain poorly understood. Cenococcum geophilum, the most widespread ectomycorrhizal fungus in boreal and temperate forests with its large, repeat-rich genome, represents an ideal system to investigate TE-mediated adaptation to the physical environment and symbiotic lifestyle. However, previous studies have been limited by fragmented genome assemblies that prevented the resolution of repeat-rich regions. We assembled a telomere-to-telomere reference genome of C. geophilum strain 1.58 using PacBio HiFi and Hi-C datasets, resulting in a 178.54 Mbp genome with seven contiguous chromosomes. We identified 14,145 genes and over 78% of the genome consists of transposable elements (TEs). Of these, 94% are affected by repeat-induced point mutations (RIP), a genome defense mechanism that acts during the sexual reproduction phase, indicating cryptic or ancient sexual reproduction in this putatively asexual fungus. Long terminal repeat retrotransposons, LINEs, and DNA transposons dominate, with three TE families (Ty3, Ty1, and Tad1) contributing over 60% of the genome size, indicating recent transposition bursts. Screening of 15 additional C. geophilum strains revealed recent and lineage-specific TE expansions, implying that several TEs escaped the RIP machinery and retained potential activity. Supporting TE activity in the context of symbiosis, we found 56 TEs differentially transcribed between ectomycorrhizal and free-living mycelium tissues. An even higher number (n = 66) of TEs were differentially expressed between stress resistance morphology (i.e. sclerotia) and free-living mycelium. This supports that TEs are differentially regulated as a response to symbiotic and stress-related conditions. Our results demonstrate that the C. geophilum genome expansion was driven by a few lineage-specific TE families in recent history, with high RIP activity attesting to sexual reproduction. We also provide insights how TEs could respond to lifestyle transitions and traits associated with desiccation resistance.

Cenococcum geophilum↗

Adapted laboratory evolution of Thermotoga sp. strain RQ7 under carbon starvation

Abstract Objective Adaptive laboratory evolution (ALE) is an effective approach to study the evolution behavior of bacterial cultures and to select for strains with desired metabolic features. In this study, we explored the possibility of evolving Thermotoga sp. strain RQ7 for cellulose-degrading abilities. Results Wild type RQ7 strain was subject to a series of transfers over six and half years with cellulose filter paper as the main and eventually the sole carbon source. Each transfer was accompanied with the addition of 50 μg of Caldicellulosiruptor saccharolyticus DSM 8903 genomic DNA. A total of 331 transfers were completed. No cellulose degradation was observed with the RQ7 cultures. Thirty three (33) isolates from six time points were sampled and sequenced. Nineteen (19) of the 33 isolates were unique, and the rest were duplicated clones. None of the isolates acquired C. saccharolyticus DNA, but all accumulated small-scale mutations throughout their genomes. Sequence analyses revealed 35 mutations that were preserved throughout the generations and another 15 mutations emerged near the end of the study. Many of the affected genes participate in phosphate metabolism, substrate transport, stress response, sensory transduction, and gene regulation.

54 ENVIRONMENTAL SCIENCES↗

Strategies to identify and edit improvements in synthetic genome segments episomally

Genome engineering projects often utilize bacterial artificial chromosomes (BACs) to carry multi-kilobase DNA segments at low copy number. However, all stages of whole-genome engineering have the potential to impose mutations on the synthetic genome that can reduce or eliminate the fitness of the final strain. Here, we describe improvements to a multiplex automated genome engineering (MAGE) protocol to improve recombineering frequency and multiplexability. This protocol was applied to recoding an Escherichia coli strain to replace seven codons with synonymous alternatives genome wide. Ten 44 402–47 179 bp de novo synthesized DNA segments contained in a BAC from the recoded strain were unable to complement deletion of the corresponding 33–61 wild-type genes using a single antibiotic resistance marker. Next-generation sequencing (NGS) was used to identify 1–7 non-recoding mutations in essential genes per segment, and MAGE in turn proved a useful strategy to repair these mutations on the recoded segment contained in the BAC when both the recoded and wild-type copies of the mutated genes had to exist by necessity during the repair process. Finally, two web-based tools were used to predict the impact of a subset of non-recoding missense mutations on strain fitness using protein structure and function calls.

59 BASIC BIOLOGICAL SCIENCES↗

Comprehensive characterization of protein–protein interactions perturbed by disease mutations

Technological and computational advances in genomics and interactomics have made it possible to identify how disease mutations perturb protein–protein interaction (PPI) networks within human cells. Here, we show that disease-associated germline variants are significantly enriched in sequences encoding PPI interfaces compared to variants identified in healthy participants from the projects 1000 Genomes and ExAC. Somatic missense mutations are also significantly enriched in PPI interfaces compared to noninterfaces in 10,861 tumor exomes. We computationally identified 470 putative oncoPPIs in a pan-cancer analysis and demonstrate that oncoPPIs are highly correlated with patient survival and drug resistance/sensitivity. Further, we experimentally validate the network effects of 13 oncoPPIs using a systematic binary interaction assay, and also demonstrate the functional consequences of two of these on tumor cell growth. In summary, this human interactome network framework provides a powerful tool for prioritization of alleles with PPI-perturbing mutations to inform pathobiological mechanism- and genotype-based therapeutic discovery.

59 BASIC BIOLOGICAL SCIENCES↗

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]↗

Stable hypermutators revealed by the genomic landscape of genes involved in genome stability among yeast species

Mutator phenotypes are short-lived due to the rapid accumulation of deleterious mutations. Yet, recent observations reveal that certain fungi can undergo prolonged accelerated evolution after losing genes involved in DNA repair. Here, we surveyed 1,154 yeast genomes representing nearly all known yeast species of the subphylum Saccharomycotina (phylum Ascomycota) to examine the relationship between reduced gene repertoires broadly associated with genome stability functions (e.g., DNA repair, cell cycle) and elevated evolutionary rates. We identified three distantly related lineages—encompassing 12% of species—that had both the most streamlined sets of genes involved in genome stability (specifically DNA repair) and the highest evolutionary rates in the entire subphylum. Two of these “faster-evolving lineages” (FELs)—a subclade within the order Pichiales and the Wickerhamiella/Starmerella (W/S) clade (order Dipodascales)—are described here for the first time, while the third corresponds to a previously documented Hanseniaspora FEL. Examination of genome stability gene repertoires revealed a set of genes predominantly absent in these three FELs, suggesting a potential role in the observed acceleration of evolutionary rates. In the W/S clade, genomic signatures are consistent with a substantial mutational burden, including pronounced A|T bias and endogenous DNA damage. Interestingly, we found that the W/S clade also contains DNA repair genes possibly acquired through horizontal gene transfer, including a photolyase of bacterial origin. These findings highlight how hypermutators can persist across macroevolutionary timescales, potentially linked to the loss of genes related with genome stability, with horizontal gene transfer as a possible avenue for partial functional compensation.

DNA repair↗

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic diversity loss in the Anthropocene

Anthropogenic habitat loss and climate change are reducing species’ geographic ranges, increasing extinction risk and losses of species’ genetic diversity. Although preserving genetic diversity is key to maintaining species’ adaptability, we lack predictive tools and global estimates of genetic diversity loss across ecosystems. We introduce a mathematical framework that bridges biodiversity theory and population genetics to understand the loss of naturally occurring DNA mutations with decreasing habitat. By analyzing genomic variation of 10,095 georeferenced individuals from 20 plant and animal species, we show that genome-wide diversity follows a mutations-area relationship power law with geographic area, which can predict genetic diversity loss from local population extinctions. We estimate that more than 10% of genetic diversity may already be lost for many threatened and nonthreatened species, surpassing the United Nations’ post-2020 targets for genetic preservation.

Science & Technology - Other Topics↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-guided functional suppression of AML-associated DNMT3A hotspot mutations

DNA methyltransferases DNMT3A- and DNMT3B-mediated DNA methylation critically regulate epigenomic and transcriptomic patterning during development. The hotspot DNMT3A mutations at the site of Arg822 (R882) promote polymerization, leading to aberrant DNA methylation that may contribute to the pathogenesis of acute myeloid leukemia (AML). However, the molecular basis underlying the mutation-induced functional misregulation of DNMT3A remains unclear. Here, we report the crystal structures of the DNMT3A methyltransferase domain, revealing a molecular basis for its oligomerization behavior distinct to DNMT3B, and the enhanced intermolecular contacts caused by the R882H or R882C mutation. Our biochemical, cellular, and genomic DNA methylation analyses demonstrate that introducing the DNMT3B-converting mutations inhibits the R882H-/R882C-triggered DNMT3A polymerization and enhances substrate access, thereby eliminating the dominant-negative effect of the DNMT3A R882 mutations in cells. Together, this study provides mechanistic insights into DNMT3A R882 mutations-triggered aberrant oligomerization and DNA hypomethylation in AML, with important implications in cancer therapy.

59 BASIC BIOLOGICAL SCIENCES↗

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley↗

Nuclear and chloroplast genome engineering of a productive non-model alga Desmodesmus armatus: Insights into unusual and selective acquisition mechanisms for foreign DNA

Despite the tremendous potential of algae to contribute to a future bioeconomy, there are practical and theoretical limitations to how well naturally sourced species and strains can perform in an outdoor setting. The application of biotechnology to modulate and engineer algae metabolism or to increase performance, resilience, or produce novel compounds, offers opportunities to overcome some of the major commercialization barriers. There are numerous approaches reported in the literature having variable success on genetic engineering of algae with non-model algae often presenting unique challenges to genetic engineering. We report here on successful nuclear and chloroplast genomic integration of selection marker resistance in the non-model alga Desmodesmus armatus. Nuclear transformation was accomplished using both electroporation and Agrobacterium-mediated approaches. However, in all surviving transformants, DNA integration was accompanied by excision and/or rearrangement of the gene of interest and fluorescence reporter coding sequences. Similarly, chloroplast transformation was successfully accomplished using a biolistic DNA delivery method. For these transformants, we also observed off-target mutations in the chloroplast genome, not previously observed in other, more routinely used, algae species. Finally, we present insights into potential mechanisms for these observed truncations, rearrangements, and mutations in D. armatus.

59 BASIC BIOLOGICAL SCIENCES↗

An orally bioavailable SARS-CoV-2 main protease inhibitor exhibits improved affinity and reduced sensitivity to mutations

Inhibitors of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) main protease (M pro ) such as nirmatrelvir (NTV) and ensitrelvir (ETV) have proven effective in reducing the severity of COVID-19, but the presence of resistance-conferring mutations in sequenced viral genomes raises concerns about future drug resistance. Second-generation oral drugs that retain function against these mutants are thus urgently needed. We hypothesized that the covalent hepatitis C virus protease inhibitor boceprevir (BPV) could serve as the basis for orally bioavailable drugs that inhibit SARS-CoV-2 M pro more efficiently than existing drugs. Performing structure-guided modifications of BPV, we developed a picomolar-affinity inhibitor, ML2006a4, with antiviral activity, oral pharmacokinetics, and therapeutic efficacy similar or superior to those of NTV. A crucial feature of ML2006a4 is a derivatization of the ketoamide reactive group that improves cell permeability and oral bioavailability. Last, ML2006a4 was found to be less sensitive to several mutations that cause resistance to NTV or ETV and occur in the natural SARS-CoV-2 population. Thus, anticipatory design can preemptively address potential resistance mechanisms to expand future treatment options against coronavirus variants.

Westberg, Michael↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

Molecular and biochemical characterization of rice developed through conventional integration of nDart1-0 transposon gene

Mutations, the genetic variations in genomic sequences, play an important role in molecular biology and biotechnology. During DNA replication or meiosis, one of the mutations is transposons or jumping genes. An indigenous transposon nDart1-0 was successfully introduced into local indica cultivar Basmati-370 from transposon-tagged line viz., GR-7895 (japonica genotype) through conventional breeding technique, successive backcrossing. Plants from segregating populationsshowed variegated phenotypes were tagged as BM-37 mutants. Blast analysis of the sequence data revealed that the GTP-binding protein, located on the BAC clone OJ1781_H11 of chromosome 5, contained an insertion of DNA transposon nDart1-0. The nDart1-0 has “A” at position 254 bp, whereas nDart1 homologs have “G”, which efficiently distinguishes nDart1-0 from its homologs. The histological analysis revealed that the chloroplast of mesophyll cells in BM-37 was disrupted with reduction in size of starch granules and higher number of osmophillic plastoglobuli, which resulted in decreased chlorophyll contents and carotenoids, gas exchange parameters (Pn, g, E, Ci), and reduced expression level of genes associated with chlorophyll biosynthesis, photosynthesis and chloroplast development. Along with the rise of GTP protein, the salicylic acid (SA) and gibberellic acid (GA) and antioxidant contents(SOD) and MDA levels significantly enhanced, while, the cytokinins (CK), ascorbate peroxidase (APX), catalase (CAT), total flavanoid contents (TFC) and total phenolic contents (TPC) significantly reduced in BM-37 mutant plants as compared with WT plants. These results support the notion that GTP-binding proteins influence the process underlying chloroplast formation. Therefore, it is anticipated that to combat biotic or abiotic stress conditions, the nDart1-0 tagged mutant (BM-37) of Basmati-370 would be beneficial.

60 APPLIED LIFE SCIENCES↗

Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor

Mutational signature analysis is commonly performed in cancer genomic studies. Here, we present SigProfilerExtractor, an automated tool for de novo extraction of mutational signatures, and benchmark it against another 13 bioinformatics tools by using 34 scenarios encompassing 2,500 simulated signatures found in 60,000 synthetic genomes and 20,000 synthetic exomes. For simulations with 5% noise, reflecting high-quality datasets, SigProfilerExtractor outperforms other approaches by elucidating between 20% and 50% more true-positive signatures while yielding 5-fold less false-positive signatures. Applying SigProfilerExtractor to 4,643 whole-genome- and 19,184 whole-exome-sequenced cancers reveals four novel signatures. Two of the signatures are confirmed in independent cohorts, and one of these signatures is associated with tobacco smoking. In summary, this report provides a reference tool for analysis of mutational signatures, a comprehensive benchmarking of bioinformatics tools for extracting signatures, and several novel mutational signatures, including one putatively attributed to direct tobacco smoking mutagenesis in bladder tissues.

59 BASIC BIOLOGICAL SCIENCES↗