Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic mutations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Transduction-like gene transfer in the methanogen Methanococcus voltae

Strain PS of Methanococcus voltae (a methanogenic, anaerobic archaebacterium) was shown to generate spontaneously 4.4-kbp chromosomal DNA fragments that are fully protected from DNase and that, upon contact with a cell, transform it genetically. This activity, here called VTA (voltae transfer agent), affects all markers tested: three different auxotrophies (histidine, purine, and cobalamin) and resistance to BES (2-bromoethanesulfonate, an inhibitor of methanogenesis). VTA was most effectively prepared by culture filtration. This process disrupted a fraction of the M. voltae cells (which have only an S-layer covering their cytoplasmic membrane). VTA was rapidly inactivated upon storage. VTA particles were present in cultures at concentrations of approximately two per cell. Gene transfer activity varied from a minimum of 2 x 10(-5) (BES resistance) to a maximum of 10(-3) (histidine independence) per donor cell. Very little VTA was found free in culture supernatants. The phenomenon is functionally similar to generalized transduction, but there is no evidence, for the time being, of intrinsically viral (i.e., containing a complete viral genome) particles. Consideration of VTA DNA size makes the existence of such viral particles unlikely. If they exist, they must be relatively few in number;perhaps they differ from VTA particles in size and other properties and thus escaped detection. Digestion of VTA DNA with the AluI restriction enzyme suggests that it is a random sample of the bacterial DNA, except for a 0.9-kbp sequence which is amplified relative to the rest of the bacterial chromosome. A VTA-sized DNA fraction was demonstrated in a few other isolates of M. voltae.

Methanococcus/genetics/growth & development/metabo↗

Genome-scale phylogeny and comparative genomics of the fungal order Sordariales

The order Sordariales is taxonomically diverse, and harbours many species with different lifestyles and large economic importance. Despite its importance, a robust genome-scale phylogeny, and associated comparative genomic analysis of the order is lacking. In this study, we examined whole-genome data from 99 Sordariales, including 52 newly sequenced genomes, and seven outgroup taxa. We inferred a comprehensive phylogeny that resolved several contentious relationships amongst families in the order, and cleared-up intrafamily relationships within the Podosporaceae. Extensive comparative genomics showed that genomes from the three largest families in the dataset (Chaetomiaceae, Podosporaceae and Sordariaceae) differ greatly in GC content, genome size, gene number, repeat percentage, evolutionary rate, and genome content affected by repeat-induced point mutations (RIP). All genomic traits showed phylogenetic signal, and ancestral state reconstruction revealed that the variation of the properties stems primarily from within-family evolution. Together, the results provide a thorough framework for understanding genome evolution in this important group of fungi.

59 BASIC BIOLOGICAL SCIENCES↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)↗

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗

Chromosome-scale Genome Assembly of the Most Abundant Ectomycorrhizal Fungus Cenococcum Geophilum Reveals Massive TE Expansion and RIP Defense Mechanism

Transposable elements (TEs) play crucial roles in genome evolution and ecological adaptation in fungi, yet their dynamics in ectomycorrhizal species remain poorly understood. Cenococcum geophilum, the most widespread ectomycorrhizal fungus in boreal and temperate forests with its large, repeat-rich genome, represents an ideal system to investigate TE-mediated adaptation to the physical environment and symbiotic lifestyle. However, previous studies have been limited by fragmented genome assemblies that prevented the resolution of repeat-rich regions. We assembled a telomere-to-telomere reference genome of C. geophilum strain 1.58 using PacBio HiFi and Hi-C datasets, resulting in a 178.54 Mbp genome with seven contiguous chromosomes. We identified 14,145 genes and over 78% of the genome consists of transposable elements (TEs). Of these, 94% are affected by repeat-induced point mutations (RIP), a genome defense mechanism that acts during the sexual reproduction phase, indicating cryptic or ancient sexual reproduction in this putatively asexual fungus. Long terminal repeat retrotransposons, LINEs, and DNA transposons dominate, with three TE families (Ty3, Ty1, and Tad1) contributing over 60% of the genome size, indicating recent transposition bursts. Screening of 15 additional C. geophilum strains revealed recent and lineage-specific TE expansions, implying that several TEs escaped the RIP machinery and retained potential activity. Supporting TE activity in the context of symbiosis, we found 56 TEs differentially transcribed between ectomycorrhizal and free-living mycelium tissues. An even higher number (n = 66) of TEs were differentially expressed between stress resistance morphology (i.e. sclerotia) and free-living mycelium. This supports that TEs are differentially regulated as a response to symbiotic and stress-related conditions. Our results demonstrate that the C. geophilum genome expansion was driven by a few lineage-specific TE families in recent history, with high RIP activity attesting to sexual reproduction. We also provide insights how TEs could respond to lifestyle transitions and traits associated with desiccation resistance.

Cenococcum geophilum↗

Adapted laboratory evolution of Thermotoga sp. strain RQ7 under carbon starvation

Abstract Objective Adaptive laboratory evolution (ALE) is an effective approach to study the evolution behavior of bacterial cultures and to select for strains with desired metabolic features. In this study, we explored the possibility of evolving Thermotoga sp. strain RQ7 for cellulose-degrading abilities. Results Wild type RQ7 strain was subject to a series of transfers over six and half years with cellulose filter paper as the main and eventually the sole carbon source. Each transfer was accompanied with the addition of 50 μg of Caldicellulosiruptor saccharolyticus DSM 8903 genomic DNA. A total of 331 transfers were completed. No cellulose degradation was observed with the RQ7 cultures. Thirty three (33) isolates from six time points were sampled and sequenced. Nineteen (19) of the 33 isolates were unique, and the rest were duplicated clones. None of the isolates acquired C. saccharolyticus DNA, but all accumulated small-scale mutations throughout their genomes. Sequence analyses revealed 35 mutations that were preserved throughout the generations and another 15 mutations emerged near the end of the study. Many of the affected genes participate in phosphate metabolism, substrate transport, stress response, sensory transduction, and gene regulation.

54 ENVIRONMENTAL SCIENCES↗

Ascorbate, added after irradiation, reduces the mutant yield and alters the spectrum of CD59- mutations in A(L) cells irradiated with high LET carbon ions

It has been reported that X-ray induced HPRT- mutation in cultured human cells is prevented by ascorbate added after irradiation. Mutation extinction is attributed to neutralization by ascorbate, of radiation-induced long-lived radicals (LLR) with half-lives of several hours. We here show that post-irradiation treatment with ascorbate (5 mM added 30 min after radiation) reduces, but does not eliminate, the induction of CD59- mutants in human-hamster hybrid A(L) cells exposed to high-LET carbon ions (LET of 100 KeV/microm). RibCys, [2(R,S)-D-ribo-1',2',3',4'-Tetrahydroxybutyl]-thiazolidene-4(R)-ca riboxylic acid] (4 mM) gave a similar but lesser effect. The lethality of the carbon ions was not altered by these chemicals. Preliminary data are presented that ascorbate also alters the spectrum of CD59- mutations induced by the carbon beam, mainly by reducing the incidence of small mutations and mutants displaying transmissible genomic instability (TGI), while large mutations are unaffected. Our results suggest that LLR are important in initiating TGI.

NASA Discipline Radiation Health↗

Strategies to identify and edit improvements in synthetic genome segments episomally

Genome engineering projects often utilize bacterial artificial chromosomes (BACs) to carry multi-kilobase DNA segments at low copy number. However, all stages of whole-genome engineering have the potential to impose mutations on the synthetic genome that can reduce or eliminate the fitness of the final strain. Here, we describe improvements to a multiplex automated genome engineering (MAGE) protocol to improve recombineering frequency and multiplexability. This protocol was applied to recoding an Escherichia coli strain to replace seven codons with synonymous alternatives genome wide. Ten 44 402–47 179 bp de novo synthesized DNA segments contained in a BAC from the recoded strain were unable to complement deletion of the corresponding 33–61 wild-type genes using a single antibiotic resistance marker. Next-generation sequencing (NGS) was used to identify 1–7 non-recoding mutations in essential genes per segment, and MAGE in turn proved a useful strategy to repair these mutations on the recoded segment contained in the BAC when both the recoded and wild-type copies of the mutated genes had to exist by necessity during the repair process. Finally, two web-based tools were used to predict the impact of a subset of non-recoding missense mutations on strain fitness using protein structure and function calls.

59 BASIC BIOLOGICAL SCIENCES↗

Deciphering Spaceflight Medical Risks Using High-Performance Computing and Next Generation Sequencing Data From Model Organisms

Somatic mutations are acquired point mutations and other forms of genetic alteration in the DNA of somatic cells in the body. Unlike germline mutations, which can be passed on from one individual to another, somatic mutations are not heritable. Somatic mutation (also called genetic sequence variation) has been recognized for decades as an important mechanism for initiating the development of cancer. Now, with the advent of next generation sequencing (NGS), the study of somatic mutation has become much more accessible, enabling genome scientists to characterize somatic mutations in a comprehensive fashion across the entire genome. From recent studies, it is now apparent that other disease processes besides cancer may be influenced by somatic mutation as well, including degenerative processes and inflammatory diseases. Cancer, degenerative diseases and inflammatory processes are all of concern in the setting of spaceflight. Study of somatic mutation analysis, therefore, is an important new tool for examining some of the earliest changes in the genome that lead to disease. This approach has tremendous potential for NASA, for analysis of both model organisms and humans.

Somatic mutation on ISS↗

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]↗

Stable hypermutators revealed by the genomic landscape of genes involved in genome stability among yeast species

Mutator phenotypes are short-lived due to the rapid accumulation of deleterious mutations. Yet, recent observations reveal that certain fungi can undergo prolonged accelerated evolution after losing genes involved in DNA repair. Here, we surveyed 1,154 yeast genomes representing nearly all known yeast species of the subphylum Saccharomycotina (phylum Ascomycota) to examine the relationship between reduced gene repertoires broadly associated with genome stability functions (e.g., DNA repair, cell cycle) and elevated evolutionary rates. We identified three distantly related lineages—encompassing 12% of species—that had both the most streamlined sets of genes involved in genome stability (specifically DNA repair) and the highest evolutionary rates in the entire subphylum. Two of these “faster-evolving lineages” (FELs)—a subclade within the order Pichiales and the Wickerhamiella/Starmerella (W/S) clade (order Dipodascales)—are described here for the first time, while the third corresponds to a previously documented Hanseniaspora FEL. Examination of genome stability gene repertoires revealed a set of genes predominantly absent in these three FELs, suggesting a potential role in the observed acceleration of evolutionary rates. In the W/S clade, genomic signatures are consistent with a substantial mutational burden, including pronounced A|T bias and endogenous DNA damage. Interestingly, we found that the W/S clade also contains DNA repair genes possibly acquired through horizontal gene transfer, including a photolyase of bacterial origin. These findings highlight how hypermutators can persist across macroevolutionary timescales, potentially linked to the loss of genes related with genome stability, with horizontal gene transfer as a possible avenue for partial functional compensation.

DNA repair↗

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic diversity loss in the Anthropocene

Anthropogenic habitat loss and climate change are reducing species’ geographic ranges, increasing extinction risk and losses of species’ genetic diversity. Although preserving genetic diversity is key to maintaining species’ adaptability, we lack predictive tools and global estimates of genetic diversity loss across ecosystems. We introduce a mathematical framework that bridges biodiversity theory and population genetics to understand the loss of naturally occurring DNA mutations with decreasing habitat. By analyzing genomic variation of 10,095 georeferenced individuals from 20 plant and animal species, we show that genome-wide diversity follows a mutations-area relationship power law with geographic area, which can predict genetic diversity loss from local population extinctions. We estimate that more than 10% of genetic diversity may already be lost for many threatened and nonthreatened species, surpassing the United Nations’ post-2020 targets for genetic preservation.

Science & Technology - Other Topics↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.

59 BASIC BIOLOGICAL SCIENCES↗

Developing a Genetic Variant Calling Pipeline for Quantifying the Complex Mutagenic Load Accumulated in BioNutrients-1 Production Pack Samples

Microorganisms hold great promise for on demand production of labile nutrients and pharmaceuticals as well recycling and in situ resource utilization. The utilization of microorganisms for such tasks on space missions is hindered by the limited data on how microbes respond to spaceflight. For example, the genetic stability of microorganisms, and the genomic engineered traits added to deliver desired functions, over long-term storage in the spacecraft environment is poorly understood. The BioNutrients-1 (BN-1) mission conducted a 5-year study of desiccated storage in Low Earth Orbit (LEO) to evaluate the suitability of eight synthetic biology chassis organisms for long-duration space missions. We are employing high-depth, whole genome sequencing (WGS) to determine the mutagenic load that accumulated during long-term storage. Mutation analysis pipelines are well established for homogenous culture grown from a single colony, but the mutational landscape of the BN-1 samples present a unique analysis challenge, as every cell in the BN-1 samples had a unique genetic journey of DNA damage and repair. Consequently, sequence variants are expected at low allele frequency within samples. To address this genetic complexity, we apply two distinct computational approaches to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO. For reference genome free mutation detection, we utilized DiscoSNP++, which is a de Bruijn graph approach. For reference genome-based mutation detection we utilize GATK for Microbes, which is a Bayesian probabilistic approach. We will benchmark these approaches against the mutations originally identified using CRISP, a method optimized for pooled samples. Ultimately, quantifying the mutation load imposed by storage or growth on the ISS will help identify chassis organisms with both high levels of genome stability and viability, which are desirable traits for implementation of bioproduction in long-duration missions.

SNP↗

Structure-guided functional suppression of AML-associated DNMT3A hotspot mutations

DNA methyltransferases DNMT3A- and DNMT3B-mediated DNA methylation critically regulate epigenomic and transcriptomic patterning during development. The hotspot DNMT3A mutations at the site of Arg822 (R882) promote polymerization, leading to aberrant DNA methylation that may contribute to the pathogenesis of acute myeloid leukemia (AML). However, the molecular basis underlying the mutation-induced functional misregulation of DNMT3A remains unclear. Here, we report the crystal structures of the DNMT3A methyltransferase domain, revealing a molecular basis for its oligomerization behavior distinct to DNMT3B, and the enhanced intermolecular contacts caused by the R882H or R882C mutation. Our biochemical, cellular, and genomic DNA methylation analyses demonstrate that introducing the DNMT3B-converting mutations inhibits the R882H-/R882C-triggered DNMT3A polymerization and enhances substrate access, thereby eliminating the dominant-negative effect of the DNMT3A R882 mutations in cells. Together, this study provides mechanistic insights into DNMT3A R882 mutations-triggered aberrant oligomerization and DNA hypomethylation in AML, with important implications in cancer therapy.

59 BASIC BIOLOGICAL SCIENCES↗