Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein complex structure prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Machine learning sheds light on microbial dark proteins

In this article, metagenomics projects have revealed more than 8 billion non-redundant microbial protein sequences from across the Earth’s biosphere. Of these, 1.17 billion proteins do not have recognizable homologues in any of the more than 100,000 reference genomes available1. Understanding the function of these microbial proteins is a daunting task. Fortunately, machine learning has recently achieved unprecedented accuracy in modelling complex biological data and making predictions. At the forefront of these advancements are machine learning-based approaches that can confidently predict atomic-level protein structures for many (but not all) amino acid sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting the structural basis of targeted protein degradation by integrating molecular dynamics simulations with structural mass spectrometry

Targeted protein degradation (TPD) is a promising approach in drug discovery for degrading proteins implicated in diseases. A key step in this process is the formation of a ternary complex where a heterobifunctional molecule induces proximity of an E3 ligase to a protein of interest (POI), thus facilitating ubiquitin transfer to the POI. In this work, we characterize 3 steps in the TPD process. (1) We simulate the ternary complex formation of SMARCA2 bromodomain and VHL E3 ligase by combining hydrogen-deuterium exchange mass spectrometry with weighted ensemble molecular dynamics (MD). (2) We characterize the conformational heterogeneity of the ternary complex using Hamiltonian replica exchange simulations and small-angle X-ray scattering. (3) We assess the ubiquitination of the POI in the context of the full Cullin-RING Ligase, confirming experimental ubiquitinomics results. Differences in degradation efficiency can be explained by the proximity of lysine residues on the POI relative to ubiquitin.

59 BASIC BIOLOGICAL SCIENCES↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Deploying synthetic coevolution and machine learning to engineer protein-protein interactions

Fine-tuning of protein-protein interactions occurs naturally through coevolution, but this process is difficult to recapitulate in the laboratory. We describe a platform for synthetic protein-protein coevolution that can isolate matched pairs of interacting muteins from complex libraries. This large dataset of coevolved complexes drove a systems-level analysis of molecular recognition between Z domain–affibody pairs spanning a wide range of structures, affinities, cross-reactivities, and orthogonalities, and captured a broad spectrum of coevolutionary networks. Furthermore, we harnessed pretrained protein language models to expand, in silico, the amino acid diversity of our coevolution screen, predicting remodeled interfaces beyond the reach of the experimental library. Further, the integration of these approaches provides a means of simulating protein coevolution and generating protein complexes with diverse molecular recognition properties for biotechnology and synthetic biology.

59 BASIC BIOLOGICAL SCIENCES↗

Cryo-EM Structure of the Mnx Protein Complex Reveals a Tunnel Framework for the Mechanism of Manganese Biomineralization

The global manganese cycle relies on microbes to oxidize soluble Mn(II) to insoluble Mn(IV) oxides. Some microbes require peroxide or superoxide as oxidants, but others can use O 2 directly, via multicopper oxidase (MCO) enzymes. One of these, MnxG from Bacillus sp. strain PL-12, was isolated in tight association with small accessory proteins, MnxE and MnxF. The protein complex, called Mnx, has eluded crystallization efforts, but we now report the 3D structure of a point mutant using cryo-EM single particle analysis, cross-linking mass spectrometry, and AlphaFold Multimer prediction. The ß-sheet–rich complex features MnxG enzyme, capped by a heterohexameric ring of alternating MnxE and MnxF subunits, and a tunnel that runs through MnxG and its MnxE 3 F 3 cap. The tunnel dimensions and charges can accommodate the mechanistically inferred binuclear manganese intermediates. Furthermore, comparison with the Fe(II)-oxidizing MCO, ceruloplasmin, identifies likely coordinating groups for the Mn(II) substrate, at the entrance to the tunnel. Thus, the 3D structure provides a rationale for the established manganese oxidase mechanism, and a platform for further experiments to elucidate mechanistic details of manganese biomineralization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Binding Free Energy Analysis of Colicin D, E3 and E8 to Their Respective Cognate Immunity Proteins Using Computational Simulations

Colicins are antimicrobial proteins produced by bacteria for the purpose of destroying neighboring bacteria. Colicin activity is neutralized by a specific cognate immunity protein in order to protect the host. This study investigates the structural and binding mechanisms underlying the interaction of colicin-D, -E3 and -E8 to their respective immunity proteins (ImD, Im3 and Im8) using structure prediction, molecular dynamics (MD) simulations and MM-PBSA approach of free energy calculations. High-confidence colicin-immunity (Col-Im) complex structures predicted using AlphaFold2 were subjected to MD simulations of 150 ns with GROMACS and were analyzed for the binding free energy calculation using gmx_MMPBSA. Results showed that the complex of Col_E3-Im3 exhibited the most favorable binding free energy, driven by strong van der Waals and electrostatic interactions. Col_D-ImD and Col_E8-Im8 also showed the favorable binding. Electrostatics and hydrogen bonding emerged as a key factor driving binding and stability, while polar solvation acted as a destabilizing factor across all systems. These outcomes provide an understanding of the molecular mechanisms of Col-Im systems, with potential applications for developing natural antimicrobials for food safety.

Biochemistry & Molecular Biology↗

Dissecting the structural heterogeneity of proteins by native mass spectrometry

Abstract A single gene yields many forms of proteins via combinations of posttranscriptional/posttranslational modifications. Proteins also fold into higher‐order structures and interact with other molecules. The combined molecular diversity leads to the heterogeneity of proteins that manifests as distinct phenotypes. Structural biology has generated vast amounts of data, effectively enabling accurate structural prediction by computational methods. However, structures are often obtained heterologously under homogeneous states in vitro. The lack of native heterogeneity under cellular context creates challenges in precisely connecting the structural data to phenotypes. Mass spectrometry (MS) based proteomics methods can profile proteome composition of complex biological samples. Most MS methods follow the “bottom‐up” approach, which denatures and digests proteins into short peptide fragments for ease of detection. Coupled with chemical biology approaches, higher‐order structures can be probed via incorporation of covalent labels on native proteins that are maintained at the peptide level. Alternatively, native MS follows the “top‐down” approach and directly analyzes intact proteins under nondenaturing conditions. Various tandem MS activation methods can dissect the intact proteins for in‐depth structural elucidation. Herein, we review recent native MS applications for characterizing heterogeneous samples, including proteins binding to mixtures of ligands, homo/hetero‐complexes with varying stoichiometry, intrinsically disordered proteins with dynamic conformations, glycoprotein complexes with mixed modification states, and active membrane protein complexes in near‐native membrane environments. We summarize the benefits, challenges, and ongoing developments in native MS, with the hope to demonstrate an emerging technology that complements other tools by filling the knowledge gaps in understanding the molecular heterogeneity of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Assembly of JAZ–JAZ and JAZ–NINJA complexes in jasmonate signaling

Jasmonates (JAs) are plant hormones with crucial roles in development and stress resilience. They activate MYC transcription factors by mediating the proteolysis of MYC inhibitors called JAZ proteins. In the absence of JA, JAZ proteins bind and inhibit MYC through the assembly of MYC–JAZ–Novel Interactor of JAZ (NINJA)–TPL repressor complexes. However, JAZ and NINJA are predicted to be largely intrinsically unstructured, which has precluded their experimental structure determination. Through a combination of biochemical, mutational, and biophysical analyses and AlphaFold-derived ColabFold modeling, we characterized JAZ–JAZ and JAZ–NINJA interactions and generated models with detailed, high-confidence domain interfaces. We demonstrate that JAZ, NINJA, and MYC interface domains are dynamic in isolation and become stabilized in a stepwise order upon complex assembly. By contrast, most JAZ and NINJA regions outside of the interfaces remain highly dynamic and cannot be modeled in a single conformation. Our data indicate that the small JAZ Zinc finger expressed in Inflorescence Meristem (ZIM) motif mediates JAZ–JAZ and JAZ–NINJA interactions through separate surfaces, and our data further suggest that NINJA modulates JAZ dimerization. This study advances our understanding of JA signaling by providing insights into the dynamics, interactions, and structure of the JAZ–NINJA core of the JA repressor complex.

59 BASIC BIOLOGICAL SCIENCES↗

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

Dual Targeting Factors Are Required for LXG Toxin Export by the Bacterial Type VIIb Secretion System

Bacterial type VIIb secretion systems (T7SSb) are multisubunit integral membrane protein complexes found in Firmicutes that play a role in both bacterial competition and virulence by secreting toxic effector proteins. The majority of characterized T7SSb effectors adopt a polymorphic domain architecture consisting of a conserved N-terminal Leu-X-Gly (LXG) domain and a variable C-terminal toxin domain. Recent work has started to reveal the diversity of toxic activities exhibited by LXG effectors; however, little is known about how these proteins are recruited to the T7SSb apparatus. In this work, we sought to characterize genes encoding domains of unknown function (DUFs) 3130 and 3958, which frequently cooccur with LXG effector-encoding genes. Using coimmunoprecipitation-mass spectrometry analyses, in vitro copurification experiments, and T7SSb secretion assays, we found that representative members of these protein families form heteromeric complexes with their cognate LXG domain and in doing so, function as targeting factors that promote effector export. Additionally, an X-ray crystal structure of a representative DUF3958 protein, combined with predictive modeling of DUF3130 using AlphaFold2, revealed structural similarity between these protein families and the ubiquitous WXG100 family of T7SS effectors. Interestingly, we identified a conserved FxxxD motif within DUF3130 that is reminiscent of the YxxxD/E “export arm” found in mycobacterial T7SSa substrates and mutation of this motif abrogates LXG effector secretion. Overall, our data experimentally link previously uncharacterized bacterial DUFs to type VIIb secretion and reveal a molecular signature required for LXG effector export.

24 antibacterial toxins↗

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology↗

Structure and mechanism of human vesicular polyamine transporter

Polyamines play essential roles in gene expression and modulate neuronal transmission in mammals. Vesicular polyamine transporters (VPAT) from the SLC18 family exploit the transmembrane H + gradient to translocate polyamines into secretory vesicles, enabling the quantal release of polyamine neuromodulators and underpinning learning and memory formation. Here, we report the cryo-electron microscopy structures of human VPAT in complex with spermine, spermidine, H + , or tetrabenazine, elucidating discrete lumen-facing states of the antiporter and pivotal interactions between VPAT and its substrate or inhibitor. Leveraging structure-inspired mutagenesis studies and protein structure prediction, we deduce an unforeseen mechanism whereby the polyamine and H + compete for multiple acidic protein residues both directly and indirectly, and rationalize how the antidopaminergic therapeutic tetrabenazine impedes vesicular transport of polyamines. This study unravels the mechanism of an H + -coupled polyamine antiporter, reveals mechanistic diversity between VPAT and other SLC18 antiporters, and raises new prospects for combating human disorders of polyamine homeostasis.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-CompBio/pf-gnn_pli

We have developed a deep learning based method to decode the protein-ligand interactions and predict their probability of binding. We use 3-Dimensional structure of proteins and ligand molecules with the ligand docked to the receptor and use the information at the interaction region to assess the binding of the protein-ligand complex.

Bontha, Mridula↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ultraviolet Superradiance from Mega-Networks of Tryptophan in Biological Architectures

Networks of tryptophan (Trp)–an aromatic amino acid with strong fluorescence response–are ubiquitous in biological systems, forming diverse architectures in transmembrane proteins, cytoskeletal filaments, subneuronal elements, photoreceptor complexes, virion capsids, and other cellular structures. We analyze the cooperative effects induced by ultraviolet (UV) excitation of several biologically relevant Trp mega-networks, thus giving insights into novel mechanisms for cellular signaling and control. Our theoretical analysis in the single-excitation manifold predicts the formation of strongly superradiant states due to collective interactions among organized arrangements of up to >10 5 Trp UV-excited transition dipoles in microtubule architectures, which leads to an enhancement of the fluorescence quantum yield (QY) that is confirmed by our experiments. We demonstrate the observed consequences of this superradiant behavior in the fluorescence QY for hierarchically organized tubulin structures, which increases in different geometric regimes at thermal equilibrium before saturation, highlighting the effect’s persistence in the presence of disorder. Our work thus showcases the many orders of magnitude across which the brightest (hundreds of femtoseconds) and darkest (tens of seconds) states can coexist in these Trp lattices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structural Heterogeneity and Hydrodynamics of an Intrinsically Disordered Protein Condensate

Biology demonstrates precise control over the free-energy landscape through the selective partitioning of biomacromolecules into membraneless organelles, enabling essential functions such as biochemical transformations, signaling cascades, and mechanical reinforcement. Although the function of these condensates depends on their underlying structure and hydrodynamics, molecular-scale information on these systems remains sparse. Here, in this study, neutron scattering is used to probe the organization and dynamics of the intrinsically disordered N-terminal domain of Galectin-3, an extracellular lectin responsible for facilitating liquid–liquid phase separation on the cellular surface, in both dilute and condensed phases. Dilute solutions contain isolated protein chains in equilibrium with mesoscopic clusters, whereas the condensed phase adopts a bicontinuous, microemulsion-like morphology. The dilute phase behavior is quantitatively described by coarse-grained polymer models from soft-matter physics, demonstrating their predictive power for complex biological proteins. At elevated concentrations, the proteins self-assemble akin to block copolymers, microphase separating through the aggregation of hydrophobic domains along the protein contour. The resulting condensate remains fluid-like despite a 25-fold increase in concentration; its internal hydrodynamics slow by only a factor of 3 relative to dilute protein chains. These results provide a molecular-level framework for how disordered proteins achieve both the structural complexity and dynamic fluidity of biomolecular condensates.

Carrick, Brian R. [Massachusetts Inst. of Technolo↗

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran↗