Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A toolkit for microbial community editing

This month’s Genome Watch highlights recently reported methods that enable genome editing of microorganisms within phylogenetically diverse communities.

59 BASIC BIOLOGICAL SCIENCES↗

Whole-Genome Resequencing to Evaluate Life History Variation in Anadromous Migration of Oncorhynchus mykiss

Anadromous fish experience physiological modifications necessary to migrate between vastly different freshwater and marine environments, but some species such as Oncorhynchus mykiss demonstrate variation in life history strategies with some individuals remaining exclusively resident in freshwater, whereas others undergo anadromous migration. Because there is limited understanding of genes involved in this life history variation across populations of this species, we evaluated the genomic difference between known anadromous ( n = 39) and resident ( n = 78) Oncorhynchus mykiss collected from the Klickitat River, WA, USA, with whole-genome resequencing methods. Sequencing of these collections yielded 5.64 million single-nucleotide polymorphisms that were tested for significant differences between resident and anadromous groups along with previously identified candidate gene regions. Although a few regions of the genome were marginally significant, there was one region on chromosome Omy12 that provided the most consistent signal of association with anadromy near two annotated genes in the reference assembly: COP9 signalosome complex subunit 6 (CSN6) and NACHT, LRR, and PYD domain–containing protein 3 (NLRP3). Previously identified candidate genes for anadromy within the inversion region of chromosome Omy05 in coastal steelhead and rainbow trout were not informative for this population as shown in previous studies. Results indicate that the significant region on chromosome Omy12 may represent a minor effect gene for male anadromy and suggests that this life history variation in Oncorhynchus mykiss is more strongly driven by other mechanisms related to environmental rearing such as epigenetic modification, gene expression, and phenotypic plasticity. Further studies into regulatory mechanisms of this trait are needed to understand drivers of anadromy in populations of this protected species.

Collins, Erin E.↗

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES↗

An efficient cre‐based workflow for genomic integration and expression of large biosynthetic pathways in Eubacterium limosum

Abstract Acetogenic Clostridia are obligate anaerobes that have emerged as promising microbes for the renewable production of biochemicals owing to their ability to efficiently metabolize sustainable single‐carbon feedstocks. Additionally, Clostridia are increasingly recognized for their biosynthetic potential, with recent discoveries of diverse secondary metabolites ranging from antibiotics to pigments to modulators of the human gut microbiota. Lack of efficient methods for genomic integration and expression of large heterologous DNA constructs remains a major challenge in studying biosynthesis in Clostridia and using them for metabolic engineering applications. To overcome this problem, we harnessed chassis‐independent recombinase‐assisted genome engineering (CRAGE) to develop a workflow for facile integration of large gene clusters (>10 kb) into the human gut acetogen Eubacterium limosum . We then integrated a non‐ribosomal peptide synthetase gene cluster from the gut anaerobe Clostridium leptum , which previously produced no detectable product in traditional heterologous hosts. Chromosomal expression in E. limosum without further optimization led to production of phevalin at 2.4 mg/L. These results further expand the molecular toolkit for a highly tractable member of the Clostridia, paving the way for sophisticated pathway engineering efforts, and highlighting the potential of E. limosum as a Clostridial chassis for exploration of anaerobic natural product biosynthesis.

Sanford, Patrick A.↗

Cas9-Based Metabolic Engineering of Issatchenkia orientalis for Enhanced Utilization of Cellulosic Hydrolysates

Issatchenkia orientalis, exhibiting high tolerance against harsh environmental conditions, is a promising metabolic engineering host for producing fuels and chemicals from cellulosic hydrolysates containing fermentation inhibitors under acidic conditions. Although genetic tools for I. orientalis exist, they require auxotrophic mutants so that the selection of a host strain is limited. We developed a drug resistance gene (cloNAT)-based genome-editing method for engineering any I. orientalis strains and engineered I. orientalis strains isolated from various sources for xylose fermentation. Specifically, xylose reductase, xylitol dehydrogenase, and xylulokinase from Scheffersomyces stipitis were integrated into an intended chromosomal locus in four I. orientalis strains (SD108, IO21, IO45, and IO46) through Cas9-based genome editing. Furthermore, the resulting strains (SD108X, IO21X, IO45X, and IO46X) efficiently produced ethanol from cellulosic and hemicellulosic hydrolysates even though the pH adjustment and nitrogen source were not provided. As they presented different fermenting capacities, selection of a host I. orientalis strain was crucial for producing fuels and chemicals using cellulosic hydrolysates.

09 BIOMASS FUELS↗

Genomic Signatures of a Major Adaptive Event in the Pathogenic Fungus Melampsora larici-populina

The recent availability of genome-wide sequencing techniques has allowed systematic screening for molecular signatures of adaptation, including in nonmodel organisms. Host–pathogen interactions constitute good models due to the strong selective pressures that they entail. We focused on an adaptive event which affected the poplar rust fungus Melampsora larici-populina when it overcame a resistance gene borne by its host, cultivated poplar. Based on 76 virulent and avirulent isolates framing narrowly the estimated date of the adaptive event, we examined the molecular signatures of selection. Using an array of genome scan methods based on different features of nucleotide diversity, we detected a single locus exhibiting a consistent pattern suggestive of a selective sweep in virulent individuals (excess of differentiation between virulent and avirulent samples, linkage disequilibrium, genotype–phenotype statistical association, and long-range haplotypes). Our study pinpoints a single gene and further a single amino acid replacement which may have allowed the adaptive event. Although our samples are nearly contemporary to the selective sweep, it does not seem to have affected genome diversity further than the immediate vicinity of the causal locus, which can be explained by a soft selective sweep (where selection acts on standing variation) and by the impact of recombination in mitigating the impact of selection. Therefore, it seems that properties of the life cycle of M. larici-populina, which entails both high genetic diversity and outbreeding, has facilitated its adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

The F-box protein gene exo-1 is a target for reverse engineering enzyme hypersecretion in filamentous fungi

Carbohydrate active enzymes (CAZymes) are vital for the lignocellulose-based biorefinery. The development of hypersecreting fungal protein production hosts is therefore a major aim for both academia and industry. However, despite advances in our understanding of their regulation, the number of promising candidate genes for targeted strain engineering remains limited. Here, we resequenced the genome of the classical hypersecreting Neurospora crassa mutant exo-1 and identified the causative point of mutation to reside in the F-box protein–encoding gene, NCU09899. The corresponding deletion strain displayed amylase and invertase activities exceeding those of the carbon catabolite derepressed strain ?cre-1, while glucose repression was still mostly functional in ?exo-1. Surprisingly, RNA sequencing revealed that while plant cell wall degradation genes are broadly misexpressed in ?exo-1, only a small fraction of CAZyme genes and sugar transporters are up-regulated, indicating that EXO-1 affects specific regulatory factors. Aiming to elucidate the underlying mechanism of enzyme hypersecretion, we found the high secretion of amylases and invertase in ?exo-1 to be completely dependent on the transcriptional regulator COL-26. Furthermore, misregulation of COL-26, CRE-1, and cellular carbon and nitrogen metabolism was confirmed by proteomics. Finally, we successfully transferred the hypersecretion trait of the exo-1 disruption by reverse engineering into the industrially deployed fungus Myceliophthora thermophila using CRISPR-Cas9. Our identification of an important F-box protein demonstrates the strength of classical mutants combined with next-generation sequencing to uncover unanticipated candidates for engineering. These data contribute to a more complete understanding of CAZyme regulation and will facilitate targeted engineering of hypersecretion in further organisms of interest.

Gabriel, Raphael↗

A Fluorescence‐Based Transient Expression Assay for the Analysis of Upstream Open Reading Frames in Plants

Upstream open reading frames (uORFs) are regulatory elements present in the 5′ leaders of mRNA that can significantly impact downstream gene expression in eukaryotes. In crop engineering, editing of uORFs can provide an avenue to upregulate expression of native genes without the need to add persistent transgenic copies. Even with genome-wide methods to identify translated uORFs such as ribosome profiling, their functional characterization depends on validation through reporter gene assays and mutagenesis studies. Current screening methods for plants use luciferases or protoplasts to measure differential gene expression between wild-type and mutated transcript leaders, which requires tissue processing and/or substrate addition. Here, we present a time- and cost-efficient alternative to investigate transcript leaders by co-expression of two fluorescent proteins in Nicotiana benthamiana leaf tissue and test our assay on genes involved in photoprotection, editing of which could provide a pathway to increase CO 2 assimilation during sun–shade transitions.

Nicotiana benthamiana↗

Transcriptional control of Clostridium autoethanogenum using CRISPRi

Gas fermentation by Clostridium autoethanogenum is a commercial process for the sustainable biomanufacturing of fuels and valuable chemicals using abundant, low-cost C1 feedstocks (CO and CO 2 ) from sources such as inedible biomass, unsorted and nonrecyclable municipal solid waste, and industrial emissions. Efforts toward pathway engineering and elucidation of gene function in this microbe have been limited by a lack of genetic tools to control gene expression and arduous genome engineering methods. To increase the pace of progress, here we developed an inducible CRISPR interference (CRISPRi) system for C. autoethanogenum and applied that system toward transcriptional repression of genes with ostensibly crucial functions in metabolism.

59 BASIC BIOLOGICAL SCIENCES↗

Population genomics provides insights into the genetic basis of adaptive evolution in the mushroom-forming fungus Lentinula edodes

Introduction: Mushroom-forming fungi comprise diverse species that develop complex multicellular structures. In cultivated species, both ecological adaptation and artificial selection have driven genome evolution. However, little is known about the connections among genotype, phenotype and adaptation in mushroom-forming fungi. Objectives: This study aimed to (1) uncover the population structure and demographic history of Lentinula edodes, (2) dissect the genetic basis of adaptive evolution in L. edodes, and (3) determine if genes related to fruiting body development are involved in adaptive evolution. Methods: We analyzed genomes and fruiting body-related traits (FBRTs) in 133 L. edodes strains and conducted RNA-seq analysis of fruiting body development in the YS69 strain. Combined methods of genomic scan for divergence, genome-wide association studies (GWAS), and RNA-seq were used to dissect the genetic basis of adaptive evolution. Results: We detected three distinct subgroups of L. edodes via single nucleotide polymorphisms, which showed robust phenotypic and temperature response differentiation and correlation with geographical distribution. Demographic history inference suggests that the subgroups diverged 36,871 generations ago. Moreover, L. edodes cultivars in China may have originated from the vicinity of Northeast China. A total of 942 genes were found to be related to genetic divergence by genomic scan, and 719 genes were identified to be candidates underlying FBRTs by GWAS. Integrating results of genomic scan and GWAS, 80 genes were detected to be related to phenotypic differentiation. A total of 364 genes related to fruiting body development were involved in genetic divergence and phenotypic differentiation. Conclusion: Adaptation to the local environment, especially temperature, triggered genetic divergence and phenotypic differentiation of L. edodes. A general model for genetic divergence and phenotypic differentiation during adaptive evolution in L. edodes, which involves in signal perception and transduction, transcriptional regulation, and fruiting body morphogenesis, was also integrated here.

59 BASIC BIOLOGICAL SCIENCES↗

Genome editing in Archaea

Methods of RNA-guided DNA endonuclease-mediated genome editing in Bacteria and Archaea are provided.

Metcalf, William W.↗

Comparative phylogenetics of repetitive elements in a diverse order of flowering plants (Brassicales)

Genome sizes of plants have long piqued the interest of researchers due to the vast differences among organisms. However, the mechanisms that drive size differences have yet to be fully understood. Two important contributing factors to genome size are expansions of repetitive elements, such as transposable elements (TEs), and whole-genome duplications (WGD). Although studies have found correlations between genome size and both TE abundance and polyploidy, these studies typically test for these patterns within a genus or species. The plant order Brassicales provides an excellent system to further test if genome size evolution patterns are consistent across larger time scales, as there are numerous WGDs. This order is also home to one of the smallest plant genomes, Arabidopsis thaliana—chosen as the model plant system for this reason—as well as to species with very large genomes. With new methods that allow for TE characterization from low-coverage genome shotgun data and 71 taxa across the Brassicales, we confirm the correlation between genome size and TE content, however, we are unable to reconstruct phylogenetic relationships and do not detect any shift in TE abundance associated with WGD.

59 BASIC BIOLOGICAL SCIENCES↗

Application of the metabolic modeling pipeline in KBase to categorize reactions, predict essential genes, and predict pathways in an isolate genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

DOE knowledgebase↗

Application of the Metabolic Modeling Pipeline in KBase to Categorize Reactions, Predict Essential Genes, and Predict Pathways in an Isolate Genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

Allen, Benjamin↗

A consensus-based ensemble approach to improve transcriptome assembly

Systems-level analyses, such as differential gene expression analysis, co-expression analysis, and metabolic pathway reconstruction, depend on the accuracy of the transcriptome. Multiple tools exist to perform transcriptome assembly from RNAseq data. However, assembling high quality transcriptomes is still not a trivial problem. This is especially the case for non-model organisms where adequate reference genomes are often not available. Different methods produce different transcriptome models and there is no easy way to determine which are more accurate. Furthermore, having alternative-splicing events exacerbates such difficult assembly problems. While benchmarking transcriptome assemblies is critical, this is also not trivial due to the general lack of true reference transcriptomes. In this study, we first provide a pipeline to generate a set of the simulated benchmark transcriptome and corresponding RNAseq data. Using the simulated benchmarking datasets, we compared the performance of various transcriptome assembly approaches including both de novo and genome-guided methods. The results showed that the assembly performance deteriorates significantly when alternative transcripts (isoforms) exist or for genome-guided methods when the reference is not available from the same genome. To improve the transcriptome assembly performance, leveraging the overlapping predictions between different assemblies, we present a new consensus-based ensemble transcriptome assembly approach, ConSemble. Without using a reference genome, ConSemble using four de novo assemblers achieved an accuracy up to twice as high as any de novo assemblers we compared. When a reference genome is available, ConSemble using four genome-guided assemblies removed many incorrectly assembled contigs with minimal impact on correctly assembled contigs, achieving higher precision and accuracy than individual genome-guided methods. Furthermore, ConSemble using de novo assemblers matched or exceeded the best performing genome-guided assemblers even when the transcriptomes included isoforms. We thus demonstrated that the ConSemble consensus strategy both for de novo and genome-guided assemblers can improve transcriptome assembly. The RNAseq simulation pipeline, the benchmark transcriptome datasets, and the script to perform the ConSemble assembly are all freely available from: http://bioinfolab.unl.edu/emlab/consemble/.

59 BASIC BIOLOGICAL SCIENCES↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Final Technical Report for DE-SC0022206

This project developed foundational genetic, genomic, and epigenetic tools for anaerobic fungi (Neocallimastigomycota), a group of microorganisms with exceptional natural abilities to deconstruct lignocellulosic biomass. Efficient biomass deconstruction remains a major barrier to economical production of renewable fuels, chemicals, and materials from agricultural and forestry residues. The project sought to enable mechanistic studies and future engineering of anaerobic fungi by improving genomic resources, establishing methods for gene expression, and investigating epigenetic regulation of biomass-degrading pathways. Major accomplishments included generation of the first chromosome-scale genome assemblies for multiple anaerobic fungal species, providing publicly available genomic resources that support both engineering and fundamental biological research. The project established the first reproducible system for heterologous gene expression in anaerobic fungi and identified genomic features and mobile genetic elements that may support future development of stable transformation technologies. In parallel, the project demonstrated direct conversion of untreated lignocellulosic biomass into fuels and specialty chemicals through a fungal-yeast bioprocess and identified anaerobic fungal enzymes with utility for metabolic engineering. The research also revealed that epigenetic regulation plays an important role in controlling fungal gene expression and enzyme production, identifying potential strategies for enhancing biomass degradation. Collectively, this work established anaerobic fungi as a tractable emerging platform for bioenergy and biomanufacturing research, generated valuable public resources, trained the next generation of researchers, and advanced DOE-BER goals related to predictive biology, sustainable bioprocessing, and the circular bioeconomy.

Solomon, Kevin [University of Delaware] (ORCID:000↗

Phylogenomics and genetic analysis of solvent-producing Clostridium species

Abstract The genus Clostridium is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 C. beijerinckii , 57 C. saccharobutylicum , 4 C. saccharoperbutylacetonicum , 5 C. butyricum , 7 C. acetobutylicum , and 3 C. tetanomorphum genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or Escherichia coli for enhanced chemical production. Recently, gene sequences from this collection were used to engineer Clostridium autoethanogenum , a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production, as well as butanol, butanoic acid, hexanol and hexanoic acid production.

59 BASIC BIOLOGICAL SCIENCES↗