Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato

Beneath the surface: Unsolved questions in soil virus ecology

Soil virus ecology is an exciting but still nascent field of research in soil microbiology. While there has been a recent surge in soil virus research studies, many fundamental questions remain unanswered, and a range of technical and bioinformatic challenges need to be overcome. In this perspective article, we present a series of key questions that highlight fruitful research areas for ongoing and future efforts. These include describing the challenges involved in understanding soil viral abundance and activity, spatiotemporal dynamics, life strategy prevalence, virus-mediated biogeochemical impacts, viral protein function, host prediction, and soil RNA virus discovery. In the near term, combining approaches (e.g., cultivation-based, meta-omics, biogeochemical, experimental, and bioinformatic) will be key to assessing the ecological and biogeochemical impacts of soil viruses from the microscopic to the field and global scales. Still, we stress that results must be tempered by current methodological limitations and highlight knowledge gaps that are most pressing to fill via new methods or measurements, such as the prevalence of different viral replication strategies across soils, the fate of microbial necromass carbon after viral lysis, the frequency of virus-host encounters that do not lead to successful infections yet could be bioinformatically mistaken as infections, and the diversity and ecological impacts of RNA viruses in soil.

59 BASIC BIOLOGICAL SCIENCES

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES

Cholesterol modulates membrane elasticity via unified biophysical laws

Cholesterol and lipid unsaturation underlie a balance of opposing forces that features prominently in adaptive cell responses to diet and environmental cues. These competing factors have resulted in contradictory observations of membrane elasticity across different measurement scales, requiring chemical specificity to explain incompatible structural and elastic effects. Here, we demonstrate that – unlike macroscopic observations – lipid membranes exhibit a unified elastic behavior in the mesoscopic regime between molecular and macroscopic dimensions. Using nuclear spin techniques and computational analysis, we find that mesoscopic bending moduli follow a universal dependence on the lipid packing density regardless of cholesterol content, lipid unsaturation, or temperature. Our observations reveal that compositional complexity can be explained by simple biophysical laws that directly map membrane elasticity to molecular packing associated with biological function, curvature transformations, and protein interactions. The obtained scaling laws closely align with theoretical predictions based on conformational chain entropy and elastic stress fields. These findings provide unique insights into the membrane design rules optimized by nature and unlock predictive capabilities for guiding the functional performance of lipid-based materials in synthetic biology and real-world applications.

Kumarage, Teshani [Virginia Polytechnic Inst. and

A non-canonical fungal peroxisome PTS-1 signal, SYM, and its evolutionary aspects

Abstract Proteins localized to peroxisomes, particularly those expressed under specific conditions or in low abundance, are often undetected by routine proteomics methods due to detection sensitivity limits. In silico identification and experimental validation of peroxisomal targeting signals (PTSs) offer a reliable alternative. We demonstrate that SYM, a non-canonical plant PTS-1 signal, functions similarly inAspergillus nidulans, as GFP tagged with a SYM C-terminal tripeptide localizes to peroxisomes. One of two nativeA. nidulansproteins with C-terminal SYM tripeptide shows weak peroxisomal localization alongside cytoplasmic presence, indicating that only a subset of proteins with non-canonical signals access peroxisomes.In silicoanalysis of 1,010 fungal genomes identified diverse SYM-proteins with variable functions, suggesting that non-canonical PTS-1 signals may evolve spontaneously. Two-thirds of SYM-proteins are predicted to localize to specific intracellular compartments other than the peroxisome. We propose that despite their predicted localization, these proteins possessing SYM as a non-canonical peroxisomal signal might also have peroxisomal presence. Among SYM-proteins, pectinesterases, known plant pathogen virulence factors, were frequent. Notably, 25% of fungal pectinesterases harbor non-canonical PTS-1 signals, suggesting that partial peroxisomal localization of pectinesterases has evolved convergently. This suggests that partial peroxisomal localization may enhance protein functional flexibility, contributing to the organism’s adaptability.

Science & Technology - Other Topics

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES

Genetic Transfer in Action: Uncovering DNA Flow in an Extremophilic Microbial Community

ABSTRACT Horizontal genetic transfer (HGT) is a significant driver of genomic novelty in all domains of life. HGT has been investigated in many studies however, the focus has been on conspicuous protein‐coding DNA transfers that often prove to be adaptive in recipient organisms and are therefore fixed longer‐term in lineages. These results comprise a subclass of HGTs and do not represent exhaustive (coding and non‐coding) DNA transfer and its impact on ecology. Uncovering exhaustive HGT can provide key insights into the connectivity of genomes in communities and how these transfers may occur. In this study, we use the term frequency‐inverse document frequency (TF‐IDF) technique, that has been used successfully to mine DNA transfers within real and simulated high‐quality prokaryote genomes, to search for exhaustive HGTs within an extremophilic microbial community. We establish a pipeline for validating transfers identified using this approach. We find that most DNA transfers are within‐domain and involve non‐coding DNA. A relatively high proportion of the predicted protein‐coding HGTs appear to encode transposase activity, restriction‐modification system components, and biofilm formation functions. Our study demonstrates the utility of the TF‐IDF approach for HGT detection and provides insights into the mechanisms of recent DNA transfer.

Microbiology

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N