Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Thousands of small, novel genes predicted in global phage genomes

Small genes (<150nucleotides) have been systematically overlooked in phage genomes. We employ a large scale comparative genomics approach to predict >40,000 small-gene families in 2.3 million phage genome contigs. We find that small genes in phage genomes are approximately 3-fold more prevalent than in host prokaryotic genomes. Our approach enriches for small genes that are translated in microbiomes, suggesting the small genes identified are coding. More than 9,000 families encode potentially secreted or transmembrane proteins, more than 5,000families encode predicted anti-CRISPR proteins, and more than500families encode predicted antimicrobial proteins. By combining homology and genomic-neighborhood analyses, we reveal substantial novelty and diversity within phage biology, including small phage genes found in multiple host phyla, small genes encoding proteins that play essential roles in host infection, and small genes that share genomic neighborhoods and whose encoded proteins may share related functions.

Fremin, Brayon↗

Gene Coexpression Connectivity Predicts Gene Targets Underlying High Ionic-Liquid Tolerance in Yarrowia lipolytica

Cellular robustness to cope with stressors is an important phenotype. Y. lipolytica is an industrial robust oleaginous yeast that has recently been discovered to tolerate record high concentrations of ILs, beneficial for novel biotransformation in organic solvents. However, genotypes that link to IL tolerance in Y. lipolytica are largely unknown.

59 BASIC BIOLOGICAL SCIENCES↗

Application of the metabolic modeling pipeline in KBase to categorize reactions, predict essential genes, and predict pathways in an isolate genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

DOE knowledgebase↗

Application of the Metabolic Modeling Pipeline in KBase to Categorize Reactions, Predict Essential Genes, and Predict Pathways in an Isolate Genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

Allen, Benjamin↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Predicting variable gene content in Escherichia coli using conserved genes

Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%–90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943–0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins” was accurately predicted (F1 = 0.902 [0.898–0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876–0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

Model of metabolism and gene expression predicts proteome allocation in Pseudomonas putida

Abstract The genome-scale model of metabolism and gene expression (ME-model) forPseudomonas putidaKT2440,iPpu1676-ME, provides a comprehensive representation of biosynthetic costs and proteome allocation. Compared to a metabolic-only model,iPpu1676-ME significantly expands on gene expression, macromolecular assembly, and cofactor utilization, enabling accurate growth predictions without additional constraints. Multi-omics analysis using RNA sequencing and ribosomal profiling data revealed translational prioritization inP. putida, with core pathways, such as nicotinamide biosynthesis and queuosine metabolism, exhibiting higher translational efficiency, while secondary pathways displayed lower priority. Notably, the ME-model significantly outperformed the M-model in alignment with multi-omics data, thereby validating its predictive capacity. Thus,iPpu1676-ME offers valuable insights intoP. putida’s proteome allocation and presents a powerful tool for understanding resource allocation in this industrially relevant microorganism.

Mathematical & Computational Biology↗

Plant diversity and functional identity drive grassland rhizobacterial community responses after 15 years of CO 2 and nitrogen enrichment

Abstract Improved understanding of bacterial community responses to multiple environmental filters over long time periods is a fundamental step to develop mechanistic explanations of plant–bacterial interactions as environmental change progresses. This is the first study to examine responses of grassland root‐associated bacterial communities to 15 years of experimental manipulations of plant species richness, functional group and factorial enrichment of atmospheric CO 2 (eCO 2 ) and soil nitrogen (+N). Across the experiment, plant species richness was the strongest predictor of rhizobacterial community composition, followed by +N, with no observed effect of eCO 2 . Monocultures of C 3 and C 4 grasses and legumes all exhibited dissimilar rhizobacterial communities within and among those groups. Functional responses were also dependent on plant functional group, where N 2 ‐fixation genes, NO 3− ‐reducing genes and P‐solubilizing predicted gene abundances increased under resource‐enriched conditions for grasses, but generally declined for legumes. In diverse plots with 16 plant species, the interaction of eCO 2 +N altered rhizobacterial composition, while +N increased the predicted abundance of nitrogenase‐encoding genes, and eCO 2 +N increased the predicted abundance of bacterial P‐solubilizing genes. Synthesis : Our findings suggest that rhizobacterial community structure and function will be affected by important global environmental change factors such as eCO 2 , but these responses are primarily contingent on plant species richness and the selective influence of different plant functional groups.

Revillini, Daniel↗

Tomato deploys defence and growth simultaneously to resist bacterial wilt disease

Abstract Plant disease limits crop production, and host genetic resistance is a major means of control. Plant pathogenic Ralstonia causes bacterial wilt disease and is best controlled with resistant varieties. Tomato wilt resistance is multigenic, yet the mechanisms of resistance remain largely unknown. We combined metaRNAseq analysis and functional experiments to identify core Ralstonia ‐responsive genes and the corresponding biological mechanisms in wilt‐resistant and wilt‐susceptible tomatoes. While trade‐offs between growth and defence are common in plants, wilt‐resistant plants activated both defence responses and growth processes. Measurements of innate immunity and growth, including reactive oxygen species production and root system growth, respectively, validated that resistant plants executed defence‐related processes at the same time they increased root growth. In contrast, in wilt‐susceptible plants roots senesced and root surface area declined following Ralstonia inoculation. Wilt‐resistant plants repressed genes predicted to negatively regulate water stress tolerance, while susceptible plants repressed genes predicted to promote water stress tolerance. Our results suggest that wilt‐resistant plants can simultaneously promote growth and defence by investing in resources that act in both processes. Infected susceptible plants activate defences, but fail to grow and so succumb to Ralstonia , likely because they cannot tolerate the water stress induced by vascular wilt.

Meline, Valerian↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

WVD2 and WDL1 modulate helical organ growth and anisotropic cell expansion in Arabidopsis

Wild-type Arabidopsis roots develop a wavy pattern of growth on tilted agar surfaces. For many Arabidopsis ecotypes, roots also grow askew on such surfaces, typically slanting to the right of the gravity vector. We identified a mutant, wvd2-1, that displays suppressed root waving and leftward root slanting under these conditions. These phenotypes arise from transcriptional activation of the novel WAVE-DAMPENED2 (WVD2) gene by the cauliflower mosaic virus 35S promoter in mutant plants. Seedlings overexpressing WVD2 exhibit constitutive right-handed helical growth in both roots and etiolated hypocotyls, whereas the petioles of WVD2-overexpressing rosette leaves exhibit left-handed twisting. Moreover, the anisotropic expansion of cells is impaired, resulting in the formation of shorter and stockier organs. In roots, the phenotype is accompanied by a change in the arrangement of cortical microtubules within peripheral cap cells and cells at the basal end of the elongation zone. WVD2 transcripts are detectable by reverse transcriptase-polymerase chain reaction in multiple organs of wild-type plants. Its predicted gene product contains a conserved region named "KLEEK," which is found only in plant proteins. The Arabidopsis genome possesses seven other genes predicted to encode KLEEK-containing products. Overexpression of one of these genes, WVD2-LIKE 1, which encodes a protein with regions of similarity to WVD2 extending beyond the KLEEK domain, results in phenotypes that are highly similar to wvd2-1. Silencing of WVD2 and its paralogs results in enhanced root skewing in the wild-type direction. Our observations suggest that at least two members of this gene family may modulate both rotational polarity and anisotropic cell expansion during organ growth.

NASA Discipline Plant Biology↗

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes

One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanotate) all use a protein-coding training model to predict protein-coding genes, with the latter three allowing for the training model to be created ab initio from the input genome. Different methods are available for creating the training model, and to increase the accuracy of such tools, we present here GOODORFS, a method for identifying protein-coding genes within a set of all possible open reading frames (ORFS). Our workflow begins with taking the amino acid frequencies of each ORF, calculating an entropy density profile (EDP), using KMeans to cluster the EDPs, and then selecting the cluster with the lowest variation as the coding ORFs. To test the efficacy of our method, we ran GOODORFS on 14,179 annotated phage genomes, and compared our results to the initial training-set creation step of four other similar methods (Glimmer, MED2, PHANOTATE, Prodigal). We found that GOODORFS was the most accurate (0.94) and had the best F1-score (0.85), while Glimmer had the highest precision (0.92) and PHANOTATE had the highest recall (0.96).

59 BASIC BIOLOGICAL SCIENCES↗

Mo than meets the eye: genomic insights into molybdoenzyme diversity of Seleniivibrio woodruffii strain S4T

Abstract Seleniivibrio woodruffii strain S4T is an obligate anaerobe belonging to the phylum Deferribacterota. It was isolated for its ability to respire selenate and was also found to respire arsenate. The high-quality draft genome of this bacterium is 2.9 Mbp, has a G+C content of 48%, 2762 predicted genes of which 2709 are protein-coding, and 53 RNA genes. An analysis of the genome focusing on the genes encoding for molybdenum-containing enzymes (molybdoenzymes) uncovered a remarkable number of genes encoding for members of the dimethylsulfoxide reductase family of proteins (DMSOR), including putative reductases for selenate and arsenate respiration, as well as genes for nitrogen fixation. Respiratory molybdoenzymes catalyze redox reactions that transfer electrons to a variety of substrates that can act as terminal electron acceptors for energy generation. Seleniivibrio woodruffii strain S4T also has essential genes for molybdate transporters and the biosynthesis of the molybdopterin guanine dinucleotide cofactors characteristic of the active centers of DMSORs. Phylogenetic analysis revealed candidate respiratory DMSORs spanning nine subfamilies encoded within the genome. Our analysis revealed the untapped potential of this interesting microorganism and expanded our knowledge of molybdoenzyme co-occurrence.

Louie, Tiffany S.↗

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics↗

Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production

Identifying gene targets for enhancing metabolite production in metabolic engineering is challenging due to the vast research literature and the approximation in genome-scale metabolic model (GEM) simulations. Here, to address this, we propose the Gene-Metabolite Association Prediction task, which automates gene discovery for given metabolite-gene pairs, accompanied by a benchmark dataset of 2474 metabolites and 1947 genes for Saccharomyces cerevisiae (SC) and Issatchenkia orientalis (IO). This task is complicated by incomplete metabolic graphs and metabolic heterogeneity. We introduce an Interactive Knowledge Transfer mechanism based on Metabolism Graphs (IKT4Meta) to enhance prediction accuracy by integrating cross-metabolism knowledge. Using Pretrained Language Models (PLMs) to generate inter-graph links mitigates heterogeneity issues, while intra-graph links are propagated via these anchors. Gene-metabolite predictions are then performed on the enriched graphs integrating multiple microorganisms’ knowledge. Experiments show that IKT4Meta outperforms baselines by up to 12.3% in link prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Galba: genome annotation with miniprot and AUGUSTUS

The Earth Biogenome Project has rapidly increased the number of available eukaryotic genomes, but most released genomes continue to lack annotation of protein-coding genes. In addition, no transcriptome data is available for some genomes. Various gene annotation tools have been developed but each has its limitations. Here, we introduce GALBA, a fully automated pipeline that utilizes miniprot, a rapid protein-to-genome aligner, in combination with AUGUSTUS to predict genes with high accuracy. Accuracy results indicate that GALBA is particularly strong in the annotation of large vertebrate genomes. We also present use cases in insects, vertebrates, and a land plant. GALBA is fully open source and available as a docker image for easy execution with Singularity in high-performance computing environments. Our pipeline addresses the critical need for accurate gene annotation in newly sequenced genomes, and we believe that GALBA will greatly facilitate genome annotation for diverse organisms.

59 BASIC BIOLOGICAL SCIENCES↗