Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Computational Prediction of Coiled–Coil Protein Gelation Dynamics and Structure

Protein hydrogels represent an important and growing biomaterial for a multitude of applications, including diagnostics and drug delivery. We have previously explored the ability to engineer the thermoresponsive supramolecular assembly of coiled–coil proteins into hydrogels with varying gelation properties, where we have defined important parameters in the coiled–coil hydrogel design. Using Rosetta energy scores and Poisson–Boltzmann electrostatic energies, we iterate a computational design strategy to predict the gelation of coiled–coil proteins while simultaneously exploring five new coiled–coil protein hydrogel sequences. Provided this library, we explore the impact of in silico energies on structure and gelation kinetics, where we also reveal a range of blue autofluorescence that enables hydrogel disassembly and recovery. As a result of this library, we identify the new coiled–coil hydrogel sequence, Q5, capable of gelation within 24 h at 4 °C, a more than 2-fold increase over that of our previous iteration Q2. The fast gelation time of Q5 enables the assessment of structural transition in real time using small-angle X-ray scattering (SAXS) that is correlated to coarse-grained and atomistic molecular dynamics simulations revealing the supramolecular assembling behavior of coiled–coils toward nanofiber assembly and gelation. This work represents the first system of hydrogels with predictable self-assembly, autofluorescent capability, and a molecular model of coiled–coil fiber formation.

36 MATERIALS SCIENCE↗

RCSB Protein Data Bank: visualizing groups of experimentally determined PDB structures alongside computed structure models of proteins

Recent advances in Artificial Intelligence and Machine Learning (e.g., AlphaFold, RosettaFold, and ESMFold) enable prediction of three-dimensional (3D) protein structures from amino acid sequences alone at accuracies comparable to lower-resolution experimental methods. These tools have been employed to predict structures across entire proteomes and the results of large-scale metagenomic sequence studies, yielding an exponential increase in available biomolecular 3D structural information. Given the enormous volume of this newly computed biostructure data, there is an urgent need for robust tools to manage, search, cluster, and visualize large collections of structures. Equally important is the capability to efficiently summarize and visualize metadata, biological/biochemical annotations, and structural features, particularly when working with vast numbers of protein structures of both experimental origin from the Protein Data Bank (PDB) and computationally-predicted models. Moreover, researchers require advanced visualization techniques that support interactive exploration of multiple sequences and structural alignments. This paper introduces a suite of tools provided on the RCSB PDB research-focused web portal RCSB. org, tailor-made for efficient management, search, organization, and visualization of this burgeoning corpus of 3D macromolecular structure data.

3D visualization↗

Modeling SARS-CoV-2 proteins in the CASP-commons experiment

Critical Assessment of Structure Prediction (CASP) is an organization aimed at advancing the state of the art in computing protein structure from sequence. In the spring of 2020, CASP launched a community project to compute the structures of the most structurally challenging proteins coded for in the SARS-CoV-2 genome. Forty-seven research groups submitted over 3000 three-dimensional models and 700 sets of accuracy estimates on 10 proteins. The resulting models were released to the public. CASP community members also worked together to provide estimates of local and global accuracy and identify structure-based domain boundaries for some proteins. Subsequently, two of these structures (ORF3a and ORF8) have been solved experimentally, allowing assessment of both model quality and the accuracy estimates. Models from the AlphaFold2 group were found to have good agreement with the experimental structures, with main chain GDT_TS accuracy scores ranging from 63 (a correct topology) to 87 (competitive with experiment).

59 BASIC BIOLOGICAL SCIENCES↗

AI-Based Protein Interaction Screening and Identification (AISID)

In this study, we presented an AISID method extending AlphaFold-Multimer’s success in structure prediction towards identifying specific protein interactions with an optimized AISIDscore. The method was tested to identify the binding proteins in 18 human TNFSF (Tumor Necrosis Factor superfamily) members for each of 27 human TNFRSF (TNF receptor superfamily) members. For each TNFRSF member, we ranked the AISIDscore among the 18 TNFSF members. The correct pairing resulted in the highest AISIDscore for 13 out of 24 TNFRSF members which have known interactions with TNFSF members. Out of the 33 correct pairing between TNFSF and TNFRSF members, 28 pairs could be found in the top five (including 25 pairs in the top three) seats in the AISIDscore ranking. Surprisingly, the specific interactions between TNFSF10 (TNF-related apoptosis-inducing ligand, TRAIL) and its decoy receptors DcR1 and DcR2 gave the highest AISIDscore in the list, while the structures of DcR1 and DcR2 are unknown. The data strongly suggests that AlphaFold-Multimer might be a useful computational screening tool to find novel specific protein bindings. This AISID method may have broad applications in protein biochemistry, extending the application of AlphaFold far beyond structure predictions.

59 BASIC BIOLOGICAL SCIENCES↗

Dual Targeting Factors Are Required for LXG Toxin Export by the Bacterial Type VIIb Secretion System

Bacterial type VIIb secretion systems (T7SSb) are multisubunit integral membrane protein complexes found in Firmicutes that play a role in both bacterial competition and virulence by secreting toxic effector proteins. The majority of characterized T7SSb effectors adopt a polymorphic domain architecture consisting of a conserved N-terminal Leu-X-Gly (LXG) domain and a variable C-terminal toxin domain. Recent work has started to reveal the diversity of toxic activities exhibited by LXG effectors; however, little is known about how these proteins are recruited to the T7SSb apparatus. In this work, we sought to characterize genes encoding domains of unknown function (DUFs) 3130 and 3958, which frequently cooccur with LXG effector-encoding genes. Using coimmunoprecipitation-mass spectrometry analyses, in vitro copurification experiments, and T7SSb secretion assays, we found that representative members of these protein families form heteromeric complexes with their cognate LXG domain and in doing so, function as targeting factors that promote effector export. Additionally, an X-ray crystal structure of a representative DUF3958 protein, combined with predictive modeling of DUF3130 using AlphaFold2, revealed structural similarity between these protein families and the ubiquitous WXG100 family of T7SS effectors. Interestingly, we identified a conserved FxxxD motif within DUF3130 that is reminiscent of the YxxxD/E “export arm” found in mycobacterial T7SSa substrates and mutation of this motif abrogates LXG effector secretion. Overall, our data experimentally link previously uncharacterized bacterial DUFs to type VIIb secretion and reveal a molecular signature required for LXG effector export.

24 antibacterial toxins↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

Nitrogen limitation causes a seismic shift in redox state and phosphorylation of proteins implicated in carbon flux and lipidome remodeling in Rhodotorula toruloides

Background: Oleaginous yeast are prodigious producers of oleochemicals, offering alternative and secure sources for applications in foodstuff, skincare, biofuels, and bioplastics. Nitrogen starvation is the primary strategy used to induce oil accumulation in oleaginous yeast as part of a global stress response. While research has demonstrated that post-translational modifications (PTMs), including phosphorylation and protein cysteine thiol oxidation (redox PTMs), are involved in signaling pathways that regulate stress responses in metazoa and algae, their role in oleaginous yeast remain understudied and unexplored. Results: Towards linking the yeast oleaginous phenotype to protein function, we integrated lipidomics, redox proteomics, and phosphoproteomics to investigate Rhodotorula toruloides under nitrogen-rich and starved conditions over time. Our lipidomics results unearthed interactions involving sphingolipids and cardiolipins with ER stress and mitophagy. Our redox and phosphoproteomics data highlighted the roles of the AMPK, TOR, and calcium signaling pathways in regulation of lipogenesis, autophagy, and oxidative stress response. As a first, we also demonstrated that lipogenic enzymes including fatty acid synthase are modified as a consequence of shifts in cellular redox states due to nutrient availability. Conclusions: We conclude that lipid accumulation is largely a consequence of carbon rerouting and autophagy governed by changes to PTMs, and not increases in the abundance of enzymes involved in central carbon metabolism and fatty acid biosynthesis. Our systems-level approach sets the stage for acquiring multidimensional data sets for protein structural modeling and predicting the functional relevance of PTMs using Artificial Intelligence/Machine Learning (AI/ML). Coupled to those bioinformatics approaches, the putative PTM switches that we delineate will enable advanced metabolic engineering strategies to decouple lipid accumulation from nitrogen limitation.

Lipid Signalling↗

Bioinformatic and Mechanistic Analysis of the Palmerolide PKS-NRPS Biosynthetic Pathway From the Microbiome of an Antarctic Ascidian

Complex interactions exist between microbiomes and their hosts. Increasingly, defensive metabolites that have been attributed to host biosynthetic capability are now being recognized as products of host-associated microbes. These unique metabolites often have bioactivity targets in human disease and can be purposed as pharmaceuticals. Polyketides are a complex family of natural products that often serve as defensive metabolites for competitive or pro-survival purposes for the producing organism, while demonstrating bioactivity in human diseases as cholesterol lowering agents, anti-infectives, and anti-tumor agents. Marine invertebrates and microbes are a rich source of polyketides. Palmerolide A, a polyketide isolated from the Antarctic ascidian Synoicum adareanum, is a vacuolar-ATPase inhibitor with potent bioactivity against melanoma cell lines. The biosynthetic gene clusters (BGCs) responsible for production of secondary metabolites are encoded in the genomes of the producers as discrete genomic elements. A candidate palmerolide BGC was identified from a S. adareanum microbiome-metagenome based on a high degree of congruence with a chemical structure-based retrobiosynthetic prediction. Protein family homology analysis, conserved domain searches, active site and motif identification were used to identify and propose the function of the ~75 kbp trans-acyltransferase (AT) polyketide synthase-non-ribosomal synthase (PKS-NRPS) domains responsible for the stepwise synthesis of palmerolide A. Though PKS systems often act in a predictable co-linear sequence, this BGC includes multiple trans-acting enzymatic domains, a non-canonical condensation termination domain, a bacterial luciferase-like monooxygenase (LLM), and is found in multiple copies within the metagenome-assembled genome (MAG). Detailed inspection of the five highly similar pal BGC copies suggests the potential for biosynthesis of other members of the palmerolide chemical family. This is the first delineation of a biosynthetic gene cluster from an Antarctic microbial species, recently proposed as Candidatus Synoicihabitans palmerolidicus. These findings have relevance for fundamental knowledge of PKS combinatorial biosynthesis and could enhance drug development efforts of palmerolide A through heterologous gene expression.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning-driven insights into super protein complexes for outer membrane protein biogenesis in bacteria

To reach their final destinations, outer membrane proteins (OMPs) of gram-negative bacteria undertake an eventful journey beginning in the cytosol. Multiple molecular machines, chaperones, proteases, and other enzymes facilitate the translocation and assembly of OMPs. These helpers usually associate, often transiently, forming large protein assemblies. They are not well understood due to experimental challenges in capturing and characterizing protein-protein interactions (PPIs), especially transient ones. Using AF2Complex, we introduce a high-throughput, deep learning pipeline to identify PPIs within the Escherichia coli cell envelope and apply it to several proteins from an OMP biogenesis pathway. Among the top confident hits obtained from screening ~1500 envelope proteins, we find not only expected interactions but also unexpected ones with profound implications. Subsequently, we predict atomic structures for these protein complexes. These structures, typically of high confidence, explain experimental observations and lead to mechanistic hypotheses for how a chaperone assists a nascent, precursor OMP emerging from a translocon, how another chaperone prevents it from aggregating and docks to a β-barrel assembly port, and how a protease performs quality control. This work presents a general strategy for investigating biological pathways by using structural insights gained from deep learning-based predictions.

60 APPLIED LIFE SCIENCES↗

Generative design of de novo proteins based on secondary-structure constraints using an attention-based diffusion model

We report two generative deep-learning models that predict amino acid sequences and 3D protein structures on the basis of secondary-structure design objectives via either the overall content or the per-residue structure. Both models are robust regarding imperfect inputs and offer de novo design capacity because they can discover new protein sequences not yet discovered from natural mechanisms or systems. The residue-level secondary-structure design model generally yields higher accuracy and more diverse sequences. These findings suggest unexplored opportunities for protein designs and functional outcomes within the vast amino acid sequences beyond known proteins. Our models, based on an attention-based diffusion model and trained on a dataset extracted from experimentally known 3D protein structures, offer numerous downstream applications in the conditional generative design of various biological or engineering systems. Future work could include additional conditioning and an exploration of other functional properties of the generated proteins for various properties beyond structural objectives.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

BIPSPI+: Mining Type-Specific Datasets of Protein Complexes to Improve Protein Binding Site Prediction

Computational approaches for predicting protein-protein interfaces are extremely useful for understanding and modelling the quaternary structure of protein assemblies. In particular, partner-specific binding site prediction methods allow delineating the specific residues that compose the interface of protein complexes. In recent years, new machine learning and other algorithmic approaches have been proposed to solve this problem. However, little effort has been made in finding better training datasets to improve the performance of these methods. With the aim of vindicating the importance of the training set compilation procedure, in this work we present BIPSPI+, a new version of our original server trained on carefully curated datasets that outperforms our original predictor. We show how prediction performance can be improved by selecting specific datasets that better describe particular types of protein interactions and interfaces (e.g. homo/hetero). In addition, our upgraded web server offers a new set of functionalities such as the sequence-structure prediction mode, hetero- or homo-complex specialization and the guided docking tool that allows to compute 3D quaternary structure poses using the predicted interfaces. BIPSPI+ is freely available at https://bipspi.cnb.csic.es.

59 BASIC BIOLOGICAL SCIENCES↗

Omics-guided metabolic pathway discovery in plants: Resources, approaches, and opportunities

Plants produce a vast array of metabolites, the biosynthetic routes of which remain largely undetermined. Genome-scale enzyme and pathway annotations and omics technologies have revolutionized research to decrypt plant metabolism and produced a growing list of functionally characterized metabolic genes and pathways. However, what is known is still a tiny fraction of the metabolic capacity harbored by plants. Here, in this work, we review plant enzyme and pathway annotation resources and cutting-edge omics approaches to guide discovery and characterization of plant metabolic pathways. We also discuss strategies for improving enzyme function prediction by integrating protein 3D structure information and single cell omics. This review aims to serve as a primer for plant biologists to leverage omics datasets to facilitate understanding and engineering plant metabolism.

59 BASIC BIOLOGICAL SCIENCES↗