Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Draft genome sequence of Gleimia europaea DSM 26657

Here, we report the draft genome sequence of Gleimia europaea DSM 26657, a pathogenic gram-positive bacillus, isolated from a patient with a subcutaneous fistula in 2007 in Germany. The genome is 2.0 Mb in size with 1,813 predicted genes, having only one putative antibiotic resistance gene.

antibiotic resistance↗

Draft genome of the switchgrass head smut pathogen Tilletia maclaganii

Tilletia maclaganii is a smut fungal pathogen that causes significant biomass reduction of switchgrass ( Panicum virgatum ) used for animal forage and biofuel production. Here we present the annotated genome of T. maclaganii , strain Tm001-NY21, estimated at 42.79 Mb in size, in 53 assembled contigs and encoding 10,235 predicted genes. This genome will be important for future comparative studies of Ustilaginales across its geographic and host range.

PacBio↗

RNAseq analysis of Cellvibrio japonicus during starch utilization differentiates between genes encoding carbohydrate active enzymes controlled by substrate detection or growth rate

ABSTRACT Bacterial utilization of starch is increasingly of interest as the importance and contributions of animal gut microbiomes become more defined. Consequently, identifying and characterizing the bacterial enzymes responsible for the degradation, transport, and metabolism of starch will enable developments in pharmaceutical, biotechnological, and culinary industries searching for novel prebiotics, carrier molecules, and low glycemic index sweeteners. The current challenge is that bacteria proficient at starch utilization often have hundreds of carbohydrate active enzymes, and it is unclear which are essential for starch utilization using only homology-based bioinformatics or computational methods. Complementary experimental data are also needed, especially to understand the regulation of bacterial starch utilization. We have completed an RNAseq analysis of the Gram-negative bacterium Cellvibrio japonicus and found that it has sophisticated regulation that includes substrate sensing and growth rate components for genes that encode starch-degrading enzymes. Among the 22 genes predicted to encode starch-active enzymes, C. japonicus has 10 alpha-amylases, 4 alpha-glucosidases, 2 pullulnases, and 2 cyclomaltodextrin glucanotransferases, 15 of which were up-regulated during exponential growth on starch and 8 up-regulated in stationary phase. Growth analyses with an enzyme secretion deficient mutant of C. japonicus suggested that secreted amylases are essential for this bacterium to degrade starch. Our approach of coupling a physiological growth assay with transcriptomic data provides a platform to identify targets for further genetic or biochemical analysis that can be broadly applied to other starch-utilizing bacteria. IMPORTANCE Understanding the bacterial metabolism of starch is important as this polysaccharide is a ubiquitous ingredient in foods, supplements, and medicines, all of which influence gut microbiome composition and health. Our RNAseq and growth data set provides a valuable resource to those who want to better understand the regulation of starch utilization in Gram-negative bacteria. These data are also useful as they provide an example of how to approach studying a starch-utilizing bacterium that has many putative amylases by coupling transcriptomic data with growth assays to overcome the potential challenges of functional redundancy. The RNAseq data can also be used as a part of larger meta-analyses to compare how C. japonicus regulates carbohydrate active enzymes, or how this bacterium compares to gut microbiome constituents in terms of starch utilization potential.

59 BASIC BIOLOGICAL SCIENCES↗

PayamDiba/Manuscript_tools

CoNSEPT is a tool to predict gene expression in various cis and trans contexts. Inputs to CoNSEPT are enhancer sequence, transcription factor levels in one or many trans conditions, TF motifs (PWMs), and any prior knowledge of TF-TF interactions.

Dibaeinia, Payam↗

The core autophagy machinery is not required for chloroplast singlet oxygen-mediated cell death in the Arabidopsis thaliana plastid ferrochelatase two mutant

Chloroplasts respond to stress and changes in the environment by producing reactive oxygen species (ROS) that have specific signaling abilities. The ROS singlet oxygen ( 1 O 2 ) is unique in that it can signal to initiate cellular degradation including the selective degradation of damaged chloroplasts. This chloroplast quality control pathway can be monitored in the Arabidopsis thaliana mutant plastid ferrochelatase two ( fc2 ) that conditionally accumulates chloroplast 1 O 2 under diurnal light cycling conditions leading to rapid chloroplast degradation and eventual cell death. The cellular machinery involved in such degradation, however, remains unknown. Recently, it was demonstrated that whole damaged chloroplasts can be transported to the central vacuole via a process requiring autophagosomes and core components of the autophagy machinery. The relationship between this process, referred to as chlorophagy, and the degradation of 1 O 2 -stressed chloroplasts and cells has remained unexplored. Results To further understand 1 O 2 -induced cellular degradation and determine what role autophagy may play, the expression of autophagy-related genes was monitored in 1 O 2 -stressed fc2 seedlings and found to be induced. Although autophagosomes were present in fc2 cells, they did not associate with chloroplasts during 1 O 2 stress. Mutations affecting the core autophagy machinery ( atg5 , atg7 , and atg10 ) were unable to suppress 1 O 2 -induced cell death or chloroplast protrusion into the central vacuole, suggesting autophagosome formation is dispensable for such 1 O 2 –mediated cellular degradation. However, both atg5 and atg7 led to specific defects in chloroplast ultrastructure and photosynthetic efficiencies, suggesting core autophagy machinery is involved in protecting chloroplasts from photo-oxidative damage. Finally, genes predicted to be involved in microautophagy were shown to be induced in stressed fc2 seedlings, indicating a possible role for an alternate form of autophagy in the dismantling of 1 O 2 -damaged chloroplasts. Conclusions Our results support the hypothesis that 1 O 2 -dependent cell death is independent from autophagosome formation, canonical autophagy, and chlorophagy. Furthermore, autophagosome-independent microautophagy may be involved in degrading 1 O 2 -damaged chloroplasts. At the same time, canonical autophagy may still play a role in protecting chloroplasts from 1 O 2 -induced photo-oxidative stress. Together, this suggests chloroplast function and degradation is a complex process utilizing multiple autophagy and degradation machineries, possibly depending on the type of stress or damage incurred.

59 BASIC BIOLOGICAL SCIENCES↗

DRAM example narrative

DRAM example narrative DRAM on KBase let's anyone run annotations using DRAM in the cloud. DRAM is an annotation tool that can annotate bacterial, archaeal and viral genomes and distills those annotatios into represetations of the functional genomic potential of those organisms. If you want to read more about DRAM you can check out the GitHub, wiki and journal article. DRAM annotate assemblies In KBase Assembly objects contain nucleotide sequences from genomes or metagenomes. DRAM can predict genes and annotate their function from KBase Assembly objects which may be microbial isolate genomes, metagenome assembled genomes or metagenomes. This is done with the Annotate and Distill Assemblies with DRAM app. This app can also anntoate AssemblySet objects which contain collection of Assembly objects. It also generates a Genome object and a GenomeSet object which can be used for further analysis with other KBase apps. The full annotations and other DRAM files are also available for download in the app.

59 BASIC BIOLOGICAL SCIENCES↗

VirION2: a short- and long-read sequencing and informatics workflow to study the genomic diversity of viruses in nature

Microbes play fundamental roles in shaping natural ecosystem properties and functions, but do so under constraints imposed by their viral predators. However, studying viruses in nature can be challenging due to low biomass and the lack of universal gene markers. Though metagenomic short-read sequencing has greatly improved our virus ecology toolkit—and revealed many critical ecosystem roles for viruses—microdiverse populations and fine-scale genomic traits are missed. Some of these microdiverse populations are abundant and the missed regions may be of interest for identifying selection pressures that underpin evolutionary constraints associated with hosts and environments. Though long-read sequencing promises complete virus genomes on single reads, it currently suffers from high DNA requirements and sequencing errors that limit accurate gene prediction. Here we introduce VirION2, an integrated short- and long-read metagenomic wet-lab and informatics pipeline that updates our previous method (VirION) to further enhance the utility of long-read viral metagenomics. Using a viral mock community, we first optimized laboratory protocols (polymerase choice, DNA shearing size, PCR cycling) to enable 76% longer reads (now median length of 6,965 bp) from 100-fold less input DNA (now 1 nanogram). Using a virome from a natural seawater sample, we compared viromes generated with VirION2 against other library preparation options (unamplified, original VirION, and short-read), and optimized downstream informatics for improved long-read error correction and assembly. VirION2 assemblies combined with short-read based data (‘enhanced’ viromes), provided significant improvements over VirION libraries in the recovery of longer and more complete viral genomes, and our optimized error-correction strategy using long- and short-read data achieved 99.97% accuracy. In the seawater virome, VirION2 assemblies captured 5,161 viral populations (including all of the virus populations observed in the other assemblies), 30% of which were uniquely assembled through inclusion of long-reads, and 22% of the top 10% most abundant virus populations derived from assembly of long-reads. Viral populations unique to VirION2 assemblies had significantly higher microdiversity means, which may explain why short-read virome approaches failed to capture them. These findings suggest the VirION2 sample prep and workflow can help researchers better investigate the virosphere, even from challenging low-biomass samples. Our new protocols are available to the research community on protocols.io as a ‘living document’ to facilitate dissemination of updates to keep pace with the rapid evolution of long-read sequencing technology.

Long-reads↗

Towards replacement of animal tests with in vitro assays: a gene expression biomarker predicts in vitro and in vivo estrogen receptor activity

High-throughput transcriptomics (HTTr) has the potential to support efforts to reduce or replace some animal tests. In past studies, we described a computational approach utilizing a gene expression biomarker consisting of 46 genes to predict estrogen receptor (ER) activity after chemical exposure in ER-positive human breast cancer cells including the MCF-7 cell line. We hypothesized that the biomarker model could identify ER activities of chemicals examined by Endocrine Disruptor Screening Program (EDSP) Tier 1 screening assays in which transcript profiles of the same chemicals were examined in MCF-7 cells. For the 62 chemicals examined including 5 chemicals examined in this study using RNA-Seq, the ER biomarker model accuracy was 1) 97% for in vitro reference chemicals, 2) 76–85% for guideline uterotrophic assays, and 3) 87–88% for guideline and nonguideline uterotrophic assays. For the same chemicals, these accuracies were similar or slightly better than those of the ToxCast ER model based on 18 in vitro assays. The performance of the ER biomarker model indicates that HTTr interpreted using the ER biomarker correctly identifies active and inactive ER reference chemicals. Finally, as part of the HTTr screening program the approach could rapidly identify chemicals with potential ER bioactivities for additional screening and testing.

60 APPLIED LIFE SCIENCES↗

Clinically applicable 53-Gene prognostic assay predicts chemotherapy benefit in gastric cancer: A multicenter study

BACKGROUND:We previously established a 53-gene prognostic signature for overall survival (OS) of gastric cancer patients. This retrospective multi-center study aimed to develop a clinically applicable gene expression detection assay and to investigate the prognostic value of this signature. METHODS:A TCGA gastric adenocarcinoma cohort (TCGA-STAD) was used for comparing 53-gene signature with other gene signatures. A high-throughput mRNA hybridization gene expression assay was developed to quantify the expression of 53-genes in formalin-fixed paraffin-embedded tissues of 540 patients enrolled from three hospitals. 180 patents were randomly selected from two hospitals to build a prognostic prediction model based on the 53-gene signature using leave-p-out (one-third out) cross-validation method together with Cox regression and Kaplan-Meier analysis, and the model was assessed on three validation cohorts. FINDINGS:In the evaluation phase, studies based on TCGA-STAD showed that the 53-gene signature was significantly superior to other three prognostic signatures and was independent of TCGA molecular subtypes and clinical factors. For clinical validation and utility, the prognostic scores were generated using the newly developed assay, which was reliable and sensitive, in 100 sampling training sets and were significantly associated with OS in 100 sampling validation sets. The scores were significantly associated with OS in three independent and combined validation cohorts, and in patients with stages II and III/IV. The multivariate Cox regression demonstrated that the prognostic power of the score was independent of clinical factors, consistent with those findings in the TCGA dataset. Finally, patients with good prognostic scores exhibited significantly a better 5-year OS rate from adjuvant FOLFOX chemotherapy after surgery than from other chemotherapies. INTERPRETATION:The 53-gene prognostic score system is clinically applicable for predicting the OS of patients independent of clinical factors in gastric cancers, which could also be a promising predictive biomarker for FOLFOX regimen. FUNDING:Chinese National Science and Technology, National Natural Science Foundation and Natural Science Foundation of Jiangsu Province.

60 APPLIED LIFE SCIENCES↗

Logic-based analysis of gene expression data predicts association between TNF, TGFB1 and EGF pathways in basal-like breast cancer

For breast cancer, clinically important subtypes are well characterized at the molecular level in terms of gene expression profiles. In addition, signaling pathways in breast cancer have been extensively studied as therapeutic targets due to their roles in tumor growth and metastasis. However, it is challenging to put signaling pathways and gene expression profiles together to characterize biological mechanisms of breast cancer subtypes since many signaling events result from post-translational modifications, rather than gene expression differences. We designed a logic-based computational framework to explain the differences in gene expression profiles among breast cancer subtypes using Pathway Logic and transcriptional network information. Pathway Logic is a rewriting-logic-based formal system for modeling biological pathways including post-translational modifications. Our method demonstrated its utility by constructing subtype-specific path from key receptors (TNFR, TGFBR1 and EGFR) to key transcription factor (TF) regulators (RELA, ATF2, SMAD3 and ELK1) and identifying potential association between pathways via TFs in basal-specific paths, which could provide a novel insight on aggressive breast cancer subtypes. Lastly, codes and results are available at http://epigenomics.snu.ac.kr/PL/.

59 BASIC BIOLOGICAL SCIENCES↗

NCAPH drives breast cancer progression and identifies a gene signature that predicts luminal a tumour recurrence

Luminal A tumours generally have a favourable prognosis but possess the highest 10-year recurrence risk among breast cancers. Additionally, a quarter of the recurrence cases occur within 5 years post-diagnosis. Identifying such patients is crucial as long-term relapsers could benefit from extended hormone therapy, while early relapsers might require more aggressive treatment. We conducted a study to explore non-structural chromosome maintenance condensin I complex subunit H’s (NCAPH) role in luminal A breast cancer pathogenesis, both in vitro and in vivo, aiming to identify an intratumoural gene expression signature, with a focus on elevated NCAPH levels, as a potential marker for unfavourable progression. Our analysis included transgenic mouse models overexpressing NCAPH and a genetically diverse mouse cohort generated by backcrossing. A least absolute shrinkage and selection operator (LASSO) multivariate regression analysis was performed on transcripts associated with elevated intratumoural NCAPH levels. We found that NCAPH contributes to adverse luminal A breast cancer progression. The intratumoural gene expression signature associated with elevated NCAPH levels emerged as a potential risk identifier. Transgenic mice overexpressing NCAPH developed breast tumours with extended latency, and in Mouse Mammary Tumor Virus (MMTV)-NCAPH ErbB2 double-transgenic mice, luminal tumours showed increased aggressiveness. High intratumoural Ncaph levels correlated with worse breast cancer outcome and subpar chemotherapy response. A 10-gene risk score, termed Gene Signature for Luminal A 10 (GSLA10), was derived from the LASSO analysis, correlating with adverse luminal A breast cancer progression. The GSLA10 signature outperformed the Oncotype DX signature in discerning tumours with unfavourable outcomes, previously categorised as luminal A by Prediction Analysis of Microarray 50 (PAM50) across three independent human cohorts. This new signature holds promise for identifying luminal A tumour patients with adverse prognosis, aiding in the development of personalised treatment strategies to significantly improve patient outcomes.

60 APPLIED LIFE SCIENCES↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Accurate flux predictions using tissue-specific gene expression in plant metabolic modeling

The accurate prediction of complex phenotypes such as metabolic fluxes in living systems is a grand challenge for systems biology and central to efficiently identifying biotechnological interventions that can address pressing industrial needs. The application of gene expression data to improve the accuracy of metabolic flux predictions using mechanistic modeling methods such as flux balance analysis (FBA) has not been previously demonstrated in multi-tissue systems, despite their biotechnological importance. We hypothesized that a method for generating metabolic flux predictions informed by relative expression levels between tissues would improve prediction accuracy. Relative gene expression levels derived from multiple transcriptomic and proteomic datasets were integrated into FBA predictions of a multi-tissue, diel model of Arabidopsis thaliana’s central metabolism. This integration dramatically improved the agreement of flux predictions with experimentally based flux maps from 13 C metabolic flux analysis compared with a standard parsimonious FBA approach. Disagreement between FBA predictions and MFA flux maps was measured using weighted averaged percent error values, and for parsimonious FBA this was 169%–180% for high light conditions and 94%–103% for low light conditions, depending on the gene expression dataset used. This fell to 10%-13% and 9%-11% upon incorporating expression data into the modeling process, which also substantially altered the predicted carbon and energy economy of the plant.

59 BASIC BIOLOGICAL SCIENCES↗

FUN-PROSE: A deep learning approach to predict condition-specific gene expression in fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

59 BASIC BIOLOGICAL SCIENCES↗

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES↗

Genome, transcriptome and secretome analyses of the antagonistic, yeast-like fungus Aureobasidium pullulans to identify potential biocontrol genes

Aureobasidium pullulans is an extremotolerant, cosmopolitan yeast-like fungus that successfully colonises vastly different ecological niches. The species is widely used in biotechnology and successfully applied as a commercial biocontrol agent against postharvest diseases and fireblight. However, the exact mechanisms that are responsible for its antagonistic activity against diverse plant pathogens are not known at the molecular level. Thus, it is difficult to optimise and improve the biocontrol applications of this species. As a foundation for elucidating biocontrol mechanisms, we have de novo assembled a high-quality reference genome of a strongly antagonistic A. pullulans strain, performed dual RNA-seq experiments, and analysed proteins secreted during the interaction with the plant pathogen Fusarium oxysporum. Based on the genome annotation, potential biocontrol genes were predicted to encode secreted hydrolases or to be part of secondary metabolite clusters (e.g., NRPS-like, NRPS, T1PKS, terpene, and β-lactone clusters). Transcriptome and secretome analyses defined a subset of 79 A. pullulans genes (among the 10,925 annotated genes) that were transcriptionally upregulated or exclusively detected at the protein level during the competition with F. oxysporum. These potential biocontrol genes comprised predicted secreted hydrolases such as glycosylases, esterases, and proteases, as well as genes encoding enzymes, which are predicted to be involved in the synthesis of secondary metabolites. This study highlights the value of a sequential approach starting with genome mining and consecutive transcriptome and secretome analyses in order to identify a limited number of potential target genes for detailed, functional analyses.

transcriptome↗