Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A roadmap to understanding and anticipating microbial gene transfer in soil communities

Engineered microbes are being programmed using synthetic DNA for applications in soil to overcome global challenges related to climate change, energy, food security, and pollution. However, we cannot yet predict gene transfer processes in soil to assess the frequency of unintentional transfer of engineered DNA to environmental microbes when applying synthetic biology technologies at scale. This challenge exists because of the complex and heterogeneous characteristics of soils, which contribute to the fitness and transport of cells and the exchange of genetic material within communities. Here, we describe knowledge gaps about gene transfer across soil microbiomes. Here, we propose strategies to improve our understanding of gene transfer across soil communities, highlight the need to benchmark the performance of biocontainment measures in situ, and discuss responsibly engaging community stakeholders. We highlight opportunities to address knowledge gaps, such as creating a set of soil standards for studying gene transfer across diverse soil types and measuring gene transfer host range across microbiomes using emerging technologies. By comparing gene transfer rates, host range, and persistence of engineered microbes across different soils, we posit that community-scale, environment-specific models can be built that anticipate biotechnology risks. Such studies will enable the design of safer biotechnologies that allow us to realize the benefits of synthetic biology and mitigate risks associated with the release of such technologies.

bioccontainment↗

Addressing the pervasive scarcity of structural annotation in eukaryotic algae

Abstract Despite a continuous increase in algal genome sequencing, structural annotations of most algal genome assemblies remain unavailable. This pervasive scarcity of genome annotation has restricted rigorous investigation of these genomic resources and may have precipitated misleading biological interpretations. However, the annotation process for eukaryotic algal species is often challenging as genomic resources and transcriptomic evidence are not always available. To address this challenge, we benchmark the cutting-edge gene prediction methods that can be generalized for a broad range of non-model eukaryotes. Using the most accurate methods selected based on high-quality algal genomes, we predict structural annotations for 135 unannotated algal genomes. Using previously available genomic data pooled together with new data obtained in this study, we identified the core orthologous genes and the multi-gene phylogeny of eukaryotic algae, including of previously unexplored algal species. This study not only provides a benchmark for the use of structural annotation methods on a variety of non-model eukaryotes, but also compensates for missing data in the current spectrum of algal genomic resources. These results bring us one step closer to the full potential of eukaryotic algal genomics.

59 BASIC BIOLOGICAL SCIENCES↗

Improvement of eukaryotic protein predictions from soil metagenomes

During the last decades, metagenomics has highlighted the diversity of microorganisms from environmental or host-associated samples. Most metagenomics public repositories use annotation pipelines tailored for prokaryotes regardless of the taxonomic origin of contigs. Consequently, eukaryotic contigs with intrinsically different gene features, are not optimally annotated. Using a bioinformatics pipeline, we have filtered 7.9 billion contigs from 6,872 soil metagenomes in the JGI’s IMG/M database to identify eukaryotic contigs. We have re-annotated genes using eukaryote-tailored methods, yielding 8 million eukaryotic proteins and over 300,000 orphan proteins lacking homology in public databases. Comparing the gene predictions we made with initial JGI ones on the same contigs, we confirmed our pipeline improves eukaryotic proteins completeness and contiguity in soil metagenomes. The improved quality of eukaryotic proteins combined with a more comprehensive assignment method yielded more reliable taxonomic annotation. This dataset of eukaryotic soil proteins with improved completeness, quality and taxonomic annotation reliability is of interest for any scientist aiming at studying the composition, biological functions and gene flux in soil communities involving eukaryotes.

54 ENVIRONMENTAL SCIENCES↗

Modeling temporal and hormonal regulation of plant transcriptional response to wounding

Plants respond to wounding stress by changing gene expression patterns and inducing the production of hormones including jasmonic acid. This wounding transcriptional response activates specialized metabolism pathways such as the glucosinolate pathways in Arabidopsis thaliana. While the regulatory factors and sequences controlling a subset of wound-response genes are known, it remains unclear how wound response is regulated globally. Here, we how these responses are regulated by incorporating putative cis-regulatory elements, known transcription factor binding sites, in vitro DNA affinity purification sequencing, and DNase I hypersensitive sites to predict genes with different wound-response patterns using machine learning. We observed that regulatory sites and regions of open chromatin differed between genes upregulated at early and late wounding time-points as well as between genes induced by jasmonic acid and those not induced. Expanding on what we currently know, we identified cis-elements that improved model predictions of expression clusters over known binding sites. Using a combination of genome editing, in vitro DNA-binding assays, and transient expression assays using native and mutated cis-regulatory elements, we experimentally validated four of the predicted elements, three of which were not previously known to function in wound-response regulation. Our study provides a global model predictive of wound response and identifies new regulatory sequences important for wounding without requiring prior knowledge of the transcriptional regulators.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic adaptations of the green alga Dunaliella salina to life under high salinity

Life in high salinity environments poses challenges to cells in a variety of ways: maintenance of ion homeostasis and nutrient acquisition, often while concomitantly enduring saturating irradiances. Dunaliella salina has an exceptional ability to thrive even in saturated brine solutions. This ability has made it a model organism for studying responses to abiotic stress factors. Here we describe the occurrence of unique gene families, expansion of gene families, or gene losses that might be linked to osmoadaptive strategies. We discovered multiple unique genes coding for several of the homologous superfamily of the Ser-Thr-rich glycosyl-phosphatidyl-inositol-anchored membrane family and of the glycolipid 2-alpha-mannosyltransferase family, suggesting that such components on the cell surface are essential to life in high salt. Gene expansion was found in families that participate in sensing of abiotic stress and signal transduction in plants. One example is the patched family of the Sonic Hedgehog receptor proteins, supporting a previous hypothesis that plasma membrane sterols are important for sensing changes in salinities in D. salina. We also investigated genome-based capabilities regarding glycerol metabolism and present an extensive map for core carbon metabolism. We postulate that a second broader glycerol cycle exists that also connects to photorespiration, thus extending the previously described glycerol cycle. Further genome-based analysis of isoprenoid and carotenoid metabolism revealed duplications of genes for 1-deoxy-D-xylulose-5-phosphate synthase (DXS) and phytoene synthase (PSY), with the second gene copy of each enzyme being clustered together. Moreover, we identified two genes predicted to code for a prokaryotic-type phytoene desaturase (CRTI), indicating that D. salina may have eukaryotic and prokaryotic elements comprising its carotenoid biosynthesis pathways. In brief, our genomic data provide the basis for further gene discoveries regarding sensing abiotic stress, the metabolism of this halophilic alga, and its potential in biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide transcriptional analysis of flagellar regeneration in Chlamydomonas reinhardtii identifies orthologs of ciliary disease genes

The important role that cilia and flagella play in human disease creates an urgent need to identify genes involved in ciliary assembly and function. The strong and specific induction of flagellar-coding genes during flagellar regeneration in Chlamydomonas reinhardtii suggests that transcriptional profiling of such cells would reveal new flagella-related genes. We have conducted a genome-wide analysis of RNA transcript levels during flagellar regeneration in Chlamydomonas by using maskless photolithography method-produced DNA oligonucleotide microarrays with unique probe sequences for all exons of the 19,803 predicted genes. This analysis represents previously uncharacterized whole-genome transcriptional activity profiling study in this important model organism. Analysis of strongly induced genes reveals a large set of known flagellar components and also identifies a number of important disease-related proteins as being involved with cilia and flagella, including the zebrafish polycystic kidney genes Qilin, Reptin, and Pontin, as well as the testis-expressed tubby-like protein TULP2.

Polycystic Kidney Diseases/genetics↗

Transcriptome-wide association analysis identifies candidate susceptibility genes for prostate-specific antigen levels in men without prostate cancer

Deciphering the genetic basis of prostate-specific antigen (PSA) levels may improve their utility for prostate cancer (PCa) screening. Using genome-wide association study (GWAS) summary statistics from 95,768 PCa-free men, we conducted a transcriptome-wide association study (TWAS) to examine impacts of genetically predicted gene expression on PSA. Analyses identified 41 statistically significant (p < 0.05/12,192 = 4.10 × 10 –6 ) associations in whole blood and 39 statistically significant (p < 0.05/13,844 = 3.61 × 10 –6 ) associations in prostate tissue, with 18 genes associated in both tissues. Cross-tissue analyses identified 155 statistically significantly (p < 0.05/22,249 = 2.25 × 10 –6 ) genes. Out of 173 unique PSA-associated genes across analyses, we replicated 151 (87.3%) in a TWAS of 209,318 PCa-free individuals from the Million Veteran Program. Based on conditional analyses, we found 20 genes (11 single tissue, nine cross-tissue) that were associated with PSA levels in the discovery TWAS that were not attributable to a lead variant from a GWAS. Ten of these 20 genes replicated, and two of the replicated genes had colocalization probability of >0.5: CCNA2 and HIST1H2BN. Six of the 20 identified genes are not known to impact PCa risk. Fine-mapping based on whole blood and prostate tissue revealed five protein-coding genes with evidence of causal relationships with PSA levels. Of these five genes, four exhibited evidence of colocalization and one was conditionally independent of previous GWAS findings. These results yield hypotheses that should be further explored to improve understanding of genetic factors underlying PSA levels.

60 APPLIED LIFE SCIENCES↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

A Systems Biology Approach to Identify Essential Epigenetic Regulators for Specific Biological Processes in Plants

Upon sensing developmental or environmental cues, epigenetic regulators transform the chromatin landscape of a network of genes to modulate their expression and dictate adequate cellular and organismal responses. Knowledge of the specific biological processes and genomic loci controlled by each epigenetic regulator will greatly advance our understanding of epigenetic regulation in plants. To facilitate hypothesis generation and testing in this domain, we present EpiNet, an extensive gene regulatory network (GRN) featuring epigenetic regulators. EpiNet was enabled by (i) curated knowledge of epigenetic regulators involved in DNA methylation, histone modification, chromatin remodeling, and siRNA pathways; and (ii) a machine-learning network inference approach powered by a wealth of public transcriptome datasets. We applied GENIE3, a machine-learning network inference approach, to mine public Arabidopsis transcriptomes and construct tissue-specific GRNs with both epigenetic regulators and transcription factors as predictors. The resultant GRNs, named EpiNet, can now be intersected with individual transcriptomic studies on biological processes of interest to identify the most influential epigenetic regulators, as well as predicted gene targets of the epigenetic regulators. We demonstrate the validity of this approach using case studies of shoot and root apical meristem development.

root apical meristem↗

Scaffolded and annotated nuclear and organelle genomes of the North American brown alga Saccharina latissima

Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.

DeWeese, Kelly↗

Five key aspects of metaproteomics as a tool to understand functional interactions in host-associated microbiomes

Host-associated microbial communities (microbiomes) play critical roles in human, animal, and plant health and development. However, interactions between the host, members of the microbiome, and invading pathogens are in most cases still poorly understood. Such interactions are multidimensional and can alter the taxonomic composition and/or the functional metabolic activities of the microbiome in response to disease or treatment conditions. For example, after 2 days of antibiotic treatment, the mouse gut microbiome is altered and more susceptible to invasion by the pathogen Clostridioides difficile. Studies of these multidimensional interactions have been fueled by the ability to use high-throughput sequencing of phylogenetic marker genes to profile microbial community composition and shotgun metagenomics to profile functional potential. However, many protein-coding genes predicted from metagenomes are not necessarily expressed under a given condition, and thus, it is difficult to assess the activities and functional interactions in microbial communities based on DNA sequencing data alone. The physiological and pathological processes expressed in these communities under specific conditions are better reflected by the abundances of transcripts or proteins. In this Pearl, we provide a brief introduction to metaproteomics, which is a tool for the large-scale analysis of proteins in microbiomes that allows researchers to address a diversity of questions related to functions and interactions in microbiomes. The term “metaproteomics” was first used in 2004 for “the large-scale characterization of the entire protein complement of environmental microbiota at a given point in time”, and since then, a large array of metaproteomics approaches have been developed. Our objective in this Pearl is to highlight what we feel are 5 essential elements to be considered for a metaproteomics research campaign and to introduce nonexpert readers to the topic without going into too much technical detail.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

What Is in Umbilicaria pustulata? A Metagenomic Approach to Reconstruct the Holo-Genome of a Lichen

Lichens are valuable models in symbiosis research and promising sources of biosynthetic genes for biotechnological applications. Most lichenized fungi grow slowly, resist aposymbiotic cultivation, and are poor candidates for experimentation. Obtaining contiguous, high-quality genomes for such symbiotic communities is technically challenging. Here, we present the first assembly of a lichen holo-genome from metagenomic whole-genome shotgun data comprising both PacBio long reads and Illumina short reads. The nuclear genomes of the two primary components of the lichen symbiosis—the fungus Umbilicaria pustulata (33 Mb) and the green alga Trebouxia sp. (53 Mb)—were assembled at contiguities comparable to single-species assemblies. The analysis of the read coverage pattern revealed a relative abundance of fungal to algal nuclei of ~20:1. Gap-free, circular sequences for all organellar genomes were obtained. The bacterial community is dominated by Acidobacteriaceae and encompasses strains closely related to bacteria isolated from other lichens. Gene set analyses showed no evidence of horizontal gene transfer from algae or bacteria into the fungal genome. Our data suggest a lineage-specific loss of a putative gibberellin-20-oxidase in the fungus, a gene fusion in the fungal mitochondrion, and a relocation of an algal chloroplast gene to the algal nucleus. Major technical obstacles during reconstruction of the holo-genome were coverage differences among individual genomes surpassing three orders of magnitude. Moreover, we show that GC-rich inverted repeats paired with nonrandom sequencing error in PacBio data can result in missing gene predictions. This likely poses a general problem for genome assemblies based on long reads.

54 ENVIRONMENTAL SCIENCES↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

The first two chromosome‐scale genome assemblies of American hazelnut enable comparative genomic analysis of the genus Corylus

Summary The native, perennial shrub American hazelnut ( Corylus americana ) is cultivated in the Midwestern United States for its significant ecological benefits, as well as its high‐value nut crop. Implementation of modern breeding methods and quantitative genetic analyses of C. americana requires high‐quality reference genomes, a resource that is currently lacking. We therefore developed the first chromosome‐scale assemblies for this species using the accessions ‘Rush’ and ‘Winkler’. Genomes were assembled using HiFi PacBio reads and Arima Hi‐C data, and Oxford Nanopore reads and a high‐density genetic map were used to perform error correction. N50 scores are 31.9 Mb and 35.3 Mb, with 90.2% and 97.1% of the total genome assembled into the 11 pseudomolecules, for ‘Rush’ and ‘Winkler’, respectively. Gene prediction was performed using custom RNAseq libraries and protein homology data. ‘Rush’ has a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while ‘Winkler’ had corresponding scores of 96.9 and 96.5, indicating high‐quality assemblies. These two independent assemblies enable unbiased assessment of structural variation within C. americana , as well as patterns of syntenic relationships across the Corylus genus. Furthermore, we identified high‐density SNP marker sets from genotyping‐by‐sequencing data using 1343 C. americana , C. avellana and C. americana × C. avellana hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published Corylus genomes, were utilized to perform phylogenetic analysis of sporophytic self‐incompatibility (SSI) in hazelnut, providing evidence of unique molecular pathways governing self‐incompatibility in Corylus .

54 ENVIRONMENTAL SCIENCES↗

Complete Genome Sequence of Acidithiobacillus ferridurans JAGS, Isolated from Acidic Mine Drainage

We report a complete genome sequence of Acidithiobacillus ferridurans JAGS, determined using PacBio single-molecule real-time (SMRT) sequencing. The circular genome of JAGS (2,933,811 bp; GC content, 58.57%) contains 3,001 protein-coding sequences, 46 tRNAs, and 6 rRNAs. Predicted genes indicate the potential to fix CO 2 and N 2 and to utilize Fe 2+ , S0, and H 2 as energy sources.

59 BASIC BIOLOGICAL SCIENCES↗