Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Constructs and methods for genome editing and genetic engineering of fungi and protists

Provided herein are constructs for genome editing or genetic engineering its fungi or protists, methods of using the constructs and media for use in selecting cells. The construct include a polynucleotide encoding a thymidine kinase operably connected to a promoter, suitably a constitutive promoter; a polynucleotide encoding an endonuclease operably connected to an inducible promoter; and a recognition site for the endonuclease. The constructs may also include selectable markers for use in selecting recombinations.

Hittinger, Christopher Todd↗

Methods for rule-based genome design

Methods and systems for designing, testing, and validating genome designs based on rules or constraints or conditions or parameters or features and scoring are described herein. A computer-implemented method includes receiving data for a known genome and a list of alleles, identifying and removing occurrences of each allele in the known genome, determining a plurality of allele choices with which to replace occurrences in the known genome, generating a plurality of alternative gene sequences for a genome design based on the known genome, wherein each alternative gene sequence comprises a different allele choice, applying a plurality of rules or constraints or conditions or parameters or features to each alternative gene sequence by assigning a score for each rule or constraint or condition or parameter or feature in each alternative gene sequence, resulting in scores for the applied plurality of rules or constraints or conditions or parameters or features, scoring each alternative gene sequence based on a weighted combination of the scores for the plurality of rules or constraints or conditions or parameters or features, and selecting at least one alternative gene sequence as the genome design based on the scoring.

Kuznetsov, Gleb↗

Genome-Wide Association Study Reveals Complex Genetic Architecture of Cadmium and Mercury Accumulation and Tolerance Traits in Medicago truncatula

Heavy metals are an increasing problem due to contamination from human sources that and can enter the food chain by being taken up by plants. Understanding the genetic basis of accumulation and tolerance in plants is important for reducing the uptake of toxic metals in crops and crop relatives, as well as for removing heavy metals from soils by means of phytoremediation. Following exposure of Medicago truncatula seedlings to cadmium (Cd) and mercury (Hg), we conducted a genome-wide association study using relative root growth (RRG) and leaf accumulation measurements. Cd and Hg accumulation and RRG had heritability ranging 0.44 – 0.72 indicating high genetic diversity for these traits. The Cd and Hg trait associations were broadly distributed throughout the genome, indicated the traits are polygenic and involve several quantitative loci. For all traits, candidate genes included several membrane associated ATP-binding cassette transporters, P-type ATPase transporters, oxidative stress response genes, and stress related UDP-glycosyltransferases. The P-type ATPase transporters and ATP-binding cassette protein-families have roles in vacuole transport of heavy metals, and our findings support their wide use in physiological plant responses to heavy metals and abiotic stresses. We also found associations between Cd RRG with the genes CAX3 and PDR3, two linked adjacent genes, and leaf accumulation of Hg associated with the genes NRAMP6 and CAX9. When plant genotypes with the most extreme phenotypes were compared, we found significant divergence in genomic regions using population genomics methods that contained metal transport and stress response gene ontologies. Several of these genomic regions show high linkage disequilibrium (LD) among candidate genes suggesting they have evolved together. Minor allele frequency (MAF) and effect size of the most significant SNPs was negatively correlated with large effect alleles being most rare. This is consistent with purifying selection against alleles that increase toxicity and abiotic stress. Conversely, the alleles with large affect that had higher frequencies that were associated with the exclusion of Cd and Hg. Overall, macroevolutionary conservation of heavy metal and stress response genes is important for improvement of forage crops by harnessing wild genetic variants in gene banks such as the Medicago HapMap collection.

59 BASIC BIOLOGICAL SCIENCES↗

Multiplex RNA-guided genome engineering

Methods of multiplex genome engineering in cells using Cas9 is provided which includes a cycle of steps of introducing into the cell a first foreign nucleic acid encoding one or more RNAs complementary to the target DNA and which guide the enzyme to the target DNA, wherein the one or more RNAs and the enzyme are members of a co-localization complex for the target DNA, and introducing into the cell a second foreign nucleic acid encoding one or more donor nucleic acid sequences, and wherein the cycle is repeated a desired number of times to multiplex DNA engineering in cells.

Church, George M.↗

Unusual modifications of protein biomarkers expressed by plasmid, prophage, and bacterial host of pathogenic Escherichia coli identified using top‐down proteomic analysis

Rationale Pathogenic bacteria often carry prophage (bacterial viruses) and plasmids (small circular pieces of DNA) that may harbor toxin, antibacterial, and antibiotic resistance genes. Proteomic characterization of pathogenic bacteria should include the identification of host proteins and proteins produced by prophage and plasmid genomes. Methods Protein biomarkers of two strains of Shiga toxin–producingEscherichia coli(STEC) were identified using antibiotic induction, matrix‐assisted laser desorption/ionization tandem time‐of‐flight (MALDI‐TOF‐TOF) tandem mass spectrometry (MS/MS) with post‐source decay (PSD), top‐down proteomic (TDP) analysis, and plasmid sequencing. Alphafold2 was also used to compare predicted in silico structures of the identified proteins to prominent fragment ions generated using MS/MS‐PSD. Strain samples were also analyzed with and without chemical reduction treatment to detect the attachment of pendant groups bound by thioester or disulfide bonds. Results Shiga toxin was detected and/or identified in both STEC strains. For the first time, we also identified the osmotically inducible protein (OsmY) whose sequence unexpectedly had two forms: a full and a truncated sequence. The truncated OsmY terminates in the middle of an α‐helix as determined by Alphafold2. A plasmid‐encoded colicin immunity protein was also identified with and without attachment of an unidentified cysteine‐bound pendant group (~307 Da). Plasmid sequencing confirmed top‐down analysis and the identification of a promoter upstream of the immunity gene that is activated by antibiotic induction, that is, SOS box. Conclusions TDP analysis, coupled with other techniques (e.g., antibiotic induction, chemical reduction, plasmid sequencing, and in silico protein modeling), is a powerful tool to identify proteins (and their modifications), including prophage‐ and plasmid‐encoded proteins, produced by pathogenic microorganisms.

Biochemistry & Molecular Biology↗

Developing a pipeline to expand the genetic code of diverse bacteria for microbial engineering

Microbial biotechnologies are key to addressing grand challenges to promote human health, reverse carbon emissions, recycle mixed plastic waste, remediate contaminated soils, and achieve sustainable economies. Synthetic biology has enabled design of diverse microbes and their proteins for useful purposes, but the narrowness of the natural genetic code limits functional diversity (e.g., biosynthesis) of engineered microbes. The natural genetic code defines the fundamental rules of translating genetic information into proteins comprised of 22 ‘canonical’ amino acids. However, using a technique called genetic code expansion (GCE), the chemical properties and therefore functions of proteins can be transformed by incorporation of one or more of ~200 chemically diverse ‘non-canonical’ amino acids. The effective application of genetic code expansion in diverse microbes has the potential to revolutionize biotechnology. However, despite over 50 years of research and its transformative potential, the application of genetic code expansion has been limited to a handful of bacterial species. In this project, we will perform three tasks to both overcome the barriers that prevent wide spread adoption of GCE as molecular tool and demonstrate its potential for biotechnological applications. Specifically, we will (1) develop a genetic engineering methodology that will enable use of GCE in a broad range of bacterial hosts, (2) use high-throughput functional genomics methods to identify physiological responses to both genetic code expansion and exposure to non-canonical amino acids in three different bacteria, and (3) demonstrate an application of GCE by selectively incorporate non-canonical amino acids into surface displayed peptides such as those used for biomining.

59 BASIC BIOLOGICAL SCIENCES↗

Ecosystems and Networks Integrated with Genes and Molecular Assemblies (ENIGMA): Molecular and Computational Technologies for Environmental Microbiology (Final Scientific/Technical Report)

The ENIGMA science focus area (SFA) is a multi-disciplinary, multi-institutional research effort focused on addressing foundational knowledge gaps in environmental microbial communities by studying groundwater and sediment microbiomes in the shallow subsurface at the contaminated Oak Ridge Reservation (ORR). We seek to discover and characterize the reciprocal interactions between the microbial communities and the geochemical and geophysical parameters of the shallow subsurface within the contamination plume. The primary goal of this subcontract was to develop experimental and computational tools to advance our understanding of microbial adaptation and community assembly in contaminated environments, with specific efforts in high-throughput genomic methods, microbial ecology tools, and studies of heavy metal contamination impacts.

54 ENVIRONMENTAL SCIENCES↗

Hijacking a rapid and scalable metagenomic method reveals subgenome dynamics and evolution in polyploid plants

Premise: The genomes of polyploid plants archive the evolutionary events leading to their present forms. However, plant polyploid genomes present numerous hurdles to the genome comparison algorithms for classification of polyploid types and exploring genome dynamics. Methods: Here, the problem of intra- and inter-genome comparison for examining polyploid genomes is reframed as a metagenomic problem, enabling the use of the rapid and scalable MinHashing approach. To determine how types of polyploidy are described by this metagenomic approach, plant genomes were examined from across the polyploid spectrum for both k-mer composition and frequency with a range of k-mer sizes. In this approach, no subgenome-specific k-mers are identified; rather, whole-chromosome k-mer subspaces were utilized. Results: Given chromosome-scale genome assemblies with sufficient subgenome-specific repetitive element content, literature-verified subgenomic and genomic evolutionary relationships were revealed, including distinguishing auto- from allopolyploidy and putative progenitor genome assignment. The sequences responsible were the rapidly evolving landscape of transposable elements. An investigation into the MinHashing parameters revealed that the downsampled k-mer space (genomic signatures) produced excellent approximations of sequence similarity. Furthermore, the clustering approach used for comparison of the genomic signatures is scrutinized to ensure applicability of the metagenomics-based method. Discussion: The easily implementable and highly computationally efficient MinHashing-based sequence comparison strategy enables comparative subgenomics and genomics for large and complex polyploid plant genomes. Such comparisons provide evidence for polyploidy-type subgenomic assignments. In cases where subgenome-specific repeat signal may not be adequate given a chromosomes' global k-mer profile, alternative methods that are more specific but more computationally complex outperform this approach.

59 BASIC BIOLOGICAL SCIENCES↗

Data for A Fluorescence-Based Transient Expression Assay for the Analysis of Upstream Open Reading Frames in Plants

Scripts for the manuscript "A fluorescence-based transient expression assay for the analysis of upstream open reading frames in plant" by Haas et al. Upstream open reading frames (uORFs) are regulatory elements present in the 5′ leaders of mRNA that can significantly impact downstream gene expression in eukaryotes. In crop engineering, editing of uORFs can provide an avenue to upregulate expression of native genes without the need to add persistent transgenic copies. Even with genome- wide methods to identify translated uORFs such as ribosome profiling, their functional characterization depends on validation through reporter gene assays and mutagenesis studies. Current screening methods for plants use luciferases or protoplasts to measure differential gene expression between wild- type and mutated transcript leaders, which requires tissue processing and/or substrate addition. Here, we present a time- and cost- efficient alternative to investigate transcript leaders by co- expression of two fluorescent proteins in Nicotiana benthamiana leaf tissue and test our assay on genes involved in photoprotection, editing of which could provide a pathway to increase CO2 assimilation during sun–shade transitions.

Gene Editing↗

Detectability of Varied Hybridization Scenarios Using Genome-Scale Hybrid Detection Methods

Hybridization events complicate the accurate reconstruction of phylogenies, as they lead to patterns of genetic heritability that are unexpected under traditional, bifurcating models of species trees. This phenomenon has led to the development of methods to infer these varied hybridization events, both methods that reconstruct networks directly, as well as summary methods that predict individual hybridization events from a subset of taxa. However, a lack of empirical comparisons between methods – especially those pertaining to large networks with varied hybridization scenarios – hinders their practical use. Here, we provide a comprehensive review of popular summary methods: TICR, MSCquartets, HyDe, Patterson’s D-Statistic (ABBA-BABA), D3, and Dp. TICR and MSCquartets are based on quartet concordance factors gathered from gene tree topologies and HyDe, Patterson’s D-Statistic, D3, and Dp use site pattern frequencies to identify hybridization events between sets of three taxa. We then use simulated data to address questions of method accuracy and ideal use scenarios by testing methods against complex networks which depict gene flow events that differ in depth (timing), quantity (single vs. multiple, overlapping hybridizations), and rate of gene flow (γ). We find that deeper or multiple hybridization events may introduce noise and weaken the signal of hybridization, leading to higher relative false negative rates across all methods. Despite some forms of hybridization eluding quartet-based detection methods, MSCquartets displays high precision in most scenarios. While HyDe results in high false negative rates when tested on hybridizations involving extinct or unsampled ghost lineages, HyDe is the only method able to identify the direction of hybridization, distinguishing the source parental lineages from recipient hybrid lineages. Lastly, we test the methods on a dataset of ultraconserved elements from the bee subfamily Nomiinae, finding possible hybridization events between clades which correspond to regions of poor support in the species tree estimated in a previous study.

Bjorner, Marianne B.↗

Introgression and persistence of cultivar alleles in wild carrot ( Daucus carota ) populations in the United States

Abstract Premise Cultivated species and their wild relatives often hybridize in the wild, and the hybrids can survive and reproduce in some environments. However, it is unclear whether cultivar alleles are permanently incorporated into the wild genomes or whether they are purged by natural selection. This question is key to accurately assessing the risk of escape and spread of cultivar genes into wild populations. Methods We used genomic data and population genomic methods to study hybridization and introgression between cultivated and wild carrot (Daucus carota) in the United States. We used single nucleotide polymorphisms (SNPs) obtained via genotyping by sequencing for 450 wild individuals from 29 wild georeferenced populations in seven states and 144 cultivars from the United States, Europe, and Asia. Results Cultivated and wild carrot formed two genetically differentiated groups, and evidence of crop–wild admixture was detected in several but not all wild carrot populations in the United States. Two regions were identified where cultivar alleles were present in wild carrots: California and Nantucket Island (Massachusetts). Surprisingly, there was no evidence of introgression in some populations with a long‐known history of sympatry with the crop, suggesting that post‐hybridization barriers might prevent introgression in some areas. Conclusions Our results provide support for the introgression and long‐term persistence of cultivar alleles in wild carrots populations. We thus anticipate that the release of genetically engineered (GE) cultivars would lead to the introduction and spread of GE alleles in wild carrot populations.

Plant Sciences↗

Methods of making polypeptides with non-standard amino acids using genomically recoded organisms

A method of making a polypeptide including at least one covalent bond between a pair of reactive side chains of corresponding amino acids, wherein the covalent bond is insensitive to reduction is provided including genetically modifying a genomically recoded organism to express a corresponding synthetase, tRNA or synthetase/tRNA pair for translating mRNA encoding the corresponding amino acids having the reactive side chains into the polypeptide and to express the polypeptide including the at least one pair of the reactive side chains wherein the reactive side chains are oriented near one another when the expressed polypeptide is in a folded configuration, wherein the reactive side chains react to form the covalent bond that is insensitive to reduction.

Church, George M.↗

A method for achieving complete microbial genomes and improving bins from metagenomics data

Metagenomics facilitates the study of the genetic information from uncultured microbes and complex microbial communities. Assembling complete genomes from metagenomics data is difficult because most samples have high organismal complexity and strain diversity. Some studies have attempted to extract complete bacterial, archaeal, and viral genomes and often focus on species with circular genomes so they can help confirm completeness with circularity. However, less than 100 circularized bacterial and archaeal genomes have been assembled and published from metagenomics data despite the thousands of datasets that are available. Circularized genomes are important for (1) building a reference collection as scaffolds for future assemblies, (2) providing complete gene content of a genome, (3) confirming little or no contamination of a genome, (4) studying the genomic context and synteny of genes, and (5) linking protein coding genes to ribosomal RNA genes to aid metabolic inference in 16S rRNA gene sequencing studies. We developed a semi-automated method called Jorg to help circularize small bacterial, archaeal, and viral genomes using iterative assembly, binning, and read mapping. In addition, this method exposes potential misassemblies from k-mer based assemblies. We chose species of the Candidate Phyla Radiation (CPR) to focus our initial efforts because they have small genomes and are only known to have one ribosomal RNA operon. In addition to 34 circular CPR genomes, we present one circular Margulisbacteria genome, one circular Chloroflexi genome, and two circular megaphage genomes from 19 public and published datasets. We demonstrate findings that would likely be difficult without circularizing genomes, including that ribosomal genes are likely not operonic in the majority of CPR, and that some CPR harbor diverged forms of RNase P RNA. Code and a tutorial for this method is available at https://github.com/lmlui/Jorg and is available on the DOE Systems Biology KnowledgeBase as a beta app.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative mitogenomics of kingdom Fungi – evolutionary insights and metagenomic applications

Mitochondria are essential components of eukaryotic cells, responsible for ATP production through oxidative phosphorylation. Despite their biological importance, unique challenges have hindered the adoption of automated mitochondrial genome (mitogenome) annotation methods, obstructing mitochondrial comparative genomics in a broad evolutionary context. Using Fungi as a study system and a Joint Genome Institute (JGI) annotated high-quality reference set, we observed broad patterns of mitochondrial evolution across the kingdom. We found that the median fungal mitogenome size is 58 kb and identified exceptionally large examples over 1 Mb in Pezizomycetes. All 14 expected oxidative phosphorylation protein-coding genes, plus rps3, were generally conserved. We found evidence of major evolutionary transitions within the Ascomycota, including the transfer of mitochondrially encoded atp8 and atp9 to the nuclear genomes across the Pezizomycotina and shifts in mitogenome tRNA patterns across the kingdom. We found substantial concordance between mitochondrial and nuclear evolution, enabling us to document 3131 total fungal mitogenomes from JGI-derived metagenomic datasets. We also identified 6467 total undeclared mitogenomes embedded in Genbank fungal nuclear assemblies. We provide interactive tools for mitogenome analysis through the JGI MycoCosm platform. Collectively, this work generated nearly 10 000 new fungal mitogenome annotations, providing a foundation and resources for future exploration of comparative fungal mitogenomics.

Ahrendt, Steven R. [USDOE Joint Genome Institute (↗

Genome sequence, phylogenetic analysis, and structure-based annotation reveal metabolic potential of Chlorella sp. SLA-04

Algae are a broad class of photosynthetic eukaryotes that are phylogenetically and physiologically diverse. Most of the phylogenetic diversity has been inferred from 18S rDNA sequencing since there are only a few complete genomes available in public databases. Here we use ultra-long-read Nanopore sequencing to determine a gapless, telomere-to-telomere complete genome sequence of Chlorella sp. SLA-04, previously described as Chlorella sorokiniana SLA-04. Chlorella sp. SLA-04 is a green alga that grows to high cell density in a wide variety of environments - high and neutral pH, high and low alkalinity, and high and low salinity. SLA-04's ability to grow in high pH and high alkalinity media without external CO 2 supply is favorable for large-scale algal biomass production. Phylogenetic analysis performed using ribosomal DNA and conserved protein sequences consistently reveal that Chlorella sp. SLA-04 forms a distinct lineage from other strains of Chlorella sorokiniana. We complement traditional genome annotation methods with high throughput structural predictions and demonstrate that this approach expands functional prediction of the SLA-04 proteome. Genomic analysis of the SLA-04 genome identifies the genes capable of utilizing TCA cycle intermediates to replenish cytosolic acetyl-CoA pools for lipid production. We also identify a complete metabolic pathway for sphingolipid anabolism that may allow SLA-04 to readily adapt to changing environmental conditions and facilitate robust cultivation in mass production systems. Altogether, this work clarifies the phylogeny of Chlorella sp. SLA-04 within Trebouxiophyceae and demonstrates how structural predictions can be used to improve annotation beyond sequencebased methods.

59 BASIC BIOLOGICAL SCIENCES↗

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes

One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanotate) all use a protein-coding training model to predict protein-coding genes, with the latter three allowing for the training model to be created ab initio from the input genome. Different methods are available for creating the training model, and to increase the accuracy of such tools, we present here GOODORFS, a method for identifying protein-coding genes within a set of all possible open reading frames (ORFS). Our workflow begins with taking the amino acid frequencies of each ORF, calculating an entropy density profile (EDP), using KMeans to cluster the EDPs, and then selecting the cluster with the lowest variation as the coding ORFs. To test the efficacy of our method, we ran GOODORFS on 14,179 annotated phage genomes, and compared our results to the initial training-set creation step of four other similar methods (Glimmer, MED2, PHANOTATE, Prodigal). We found that GOODORFS was the most accurate (0.94) and had the best F1-score (0.85), while Glimmer had the highest precision (0.92) and PHANOTATE had the highest recall (0.96).

59 BASIC BIOLOGICAL SCIENCES↗

Phylogenomic Insights into the Evolution and Origin of Nematoda

Abstract The phylum Nematoda represents one of the most cosmopolitan and abundant metazoan groups on Earth. In this study, we reconstructed the phylogenomic tree for phylum Nematoda. A total of 60 genomes, belonging to 8 nematode orders, were newly sequenced, providing the first low-coverage genomes for the orders Dorylaimida, Mononchida, Monhysterida, Chromadorida, Triplonchida, and Enoplida. The resulting phylogeny is well-resolved across most clades, with topologies remaining consistent across various reconstruction parameters. The subclass Enoplia is placed as a sister group to the rest of Nematoda, agreeing with previously published phylogenies. While the order Triplonchida is monophyletic, it is not well-supported, and the order Enoplida is paraphyletic. Taxa possessing a stomatostylet form a monophyletic group; however, the superfamily Aphelenchoidea does not constitute a monophyletic clade. The genera Trichinella and Trichuris are inferred to have shared a common ancestor approximately 202 millions of years ago (Ma), a considerably later period than previously suggested. All stomatostylet-bearing nematodes are proposed to have originated ~305 Ma, corresponding to the transition from the Devonian to the Permian period. The genus Thornia is placed outside of Dorylaimina and Nygolaimina, disagreeing with its position in previous studies. In addition, we tested the whole genome amplification method and demonstrated that it is a promising strategy for obtaining sufficient DNA for phylogenomic studies of microscopic eukaryotes. This study significantly expanded the current nematode genome dataset, and the well-resolved phylogeny enhances our understanding of the evolution of Nematoda.

Qing, Xue (ORCID:0000000203559956)↗

Microbiome-enabled genomic selection improves prediction accuracy for nitrogen-related traits in maize

Root-associated microbiomes in the rhizosphere (rhizobiomes) are increasingly known to play an important role in nutrient acquisition, stress tolerance, and disease resistance of plants. However, it remains largely unclear to what extent these rhizobiomes contribute to trait variation for different genotypes and if their inclusion in the genomic selection protocol can enhance prediction accuracy. To address these questions, we developed a microbiome-enabled genomic selection method that incorporated host SNPs and amplicon sequence variants from plant rhizobiomes in a maize diversity panel under high and low nitrogen (N) field conditions. Our cross-validation results showed that the microbiome-enabled genomic selection model significantly outperformed the conventional genomic selection model for nearly all time-series traits related to plant growth and N responses, with an average relative improvement of 3.7%. The improvement was more pronounced under low N conditions (8.4–40.2% of relative improvement), consistent with the view that some beneficial microbes can enhance N nutrient uptake, particularly in low N fields. However, our study could not definitively rule out the possibility that the observed improvement is partially due to the amplicon sequence variants being influenced by microenvironments. Using a high-dimensional mediation analysis method, our study has also identified microbial mediators that establish a link between plant genotype and phenotype. Some of the detected mediator microbes were previously reported to promote plant growth. The enhanced prediction accuracy of the microbiome-enabled genomic selection models, demonstrated in a single environment, serves as a proof-of-concept for the potential application of microbiome-enabled plant breeding for sustainable agriculture.

60 APPLIED LIFE SCIENCES↗