Improved Protein–Ligand Binding Affinity Prediction with Structure-Based Deep Fusion Inference
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The SARS-CoV2 pandemic has highlighted the importance of efficient and effective methods for identification of therapeutic drugs, and in particular has laid bare the need for methods that allow exploration of the full diversity of synthesizable small molecules. While classical high-throughput screening methods may consider up to millions of molecules, virtual screening methods hold the promise of enabling appraisal of billions of candidate molecules, thus expanding the search space while concurrently reducing costs and speeding discovery. Here, we describe a new screening pipeline, called drugsniffer, that is capable of rapidly exploring drug candidates from a library of billions of molecules, and is designed to support distributed computation on cluster and cloud resources. As an example of performance, our pipeline required ~40,000 total compute hours to screen for potential drugs targeting three SARS-CoV2 proteins among a library of ~3.7 billion candidate molecules.
Organophosphorus hydrolase (OPH) is a metalloenzyme that can hydrolyze organophosphorus agents resulting in products that are generally of reduced toxicity. The best OPH substrate found to date is diethyl p-nitrophenyl phosphate (paraoxon). Most structural and kinetic studies assume that the binding orientation of paraoxon is identical to that of diethyl 4-methylbenzylphosphonate, which is the only substrate analog co-crystallized with OPH. In the current work, we used a combined docking and molecular dynamics (MD) approach to predict the likely binding mode of paraoxon. Then, we used the predicted binding mode to run MD simulations on the wild type (WT) OPH complexed with paraoxon, and OPH mutants complexed with paraoxon. Additionally, we identified three hot-spot residues (D253, H254, and I255) involved in the stability of the OPH active site. We then experimentally assayed single and double mutants involving these residues for paraoxon binding affinity. The binding free energy calculations and the experimental kinetics of the reactions between each OPH mutant and paraoxon show that mutated forms D253E, D253E-H254R, and D253E-I255G exhibit enhanced substrate binding affinity over WT OPH. Interestingly, our experimental results show that the substrate binding affinity of the double mutant D253E-H254R increased by 19-fold compared to WT OPH.
Abstract New drug production, from target identification to marketing approval, takes over 12 years and can cost around $2.6 billion. Furthermore, the COVID-19 pandemic has unveiled the urgent need for more powerful computational methods for drug discovery. Here, we review the computational approaches to predicting protein–ligand interactions in the context of drug discovery, focusing on methods using artificial intelligence (AI). We begin with a brief introduction to proteins (targets), ligands (e.g. drugs) and their interactions for nonexperts. Next, we review databases that are commonly used in the domain of protein–ligand interactions. Finally, we survey and analyze the machine learning (ML) approaches implemented to predict protein–ligand binding sites, ligand-binding affinity and binding pose (conformation) including both classical ML algorithms and recent deep learning methods. After exploring the correlation between these three aspects of protein–ligand interaction, it has been proposed that they should be studied in unison. We anticipate that our review will aid exploration and development of more accurate ML-based prediction strategies for studying protein–ligand interactions.
The binding affinity of nicotinoids to the binding residues of the a4ß2 variant of the nicotinic acetylcholine receptor (nAChR) was identified as a strong predictor of the nicotinoid’s addictive character. Using ab-iniito calculations for model binding pockets of increasing size comprising of 3, 6, and 14 amino acids (3AA, 6AA, and 14AA) that are derived from the crystal structure, the differences in binding affinity of 6 nicotinoids, namely nicotine (NIC), nornicotine (NOR), anabasine (ANB), anatabine (ANT), myosmine (MYO), and cotinine (COT) were correlated to their previously reported doses required for increases in intracranial self-stimulation (ICSS) thresholds, a metric for their addictive function. By employing the many body decomposition, the differences in the binding affinities of the various nicotinoids could be attributed mainly to the proton exchange energy between the Pyridine and non-Pyridine rings of the nicotinoids and the interactions between them and a handful of proximal amino acids, namely Trp156, Trpß57, Tyr100, and Tyr204. Interactions between the guest nicotinoid and the amino acids of the binding pocket were found to be mainly classical in nature, except for those between the nicotinoid and Trp156. The larger pockets were found to model binding structures more accurately and predicted the addictive character of all nicotinoids while smaller models, which are more computationally feasible, would only predict the addictive character of nicotinoids that are similar to nicotine. Here, the present study identifies the binding affinity of the guest nicotinoid to the host binding pocket as a strong descriptor of the nicotinoid’s addiction potential and as such it can be employed as a fast screening technique for the potential addiction of nicotine analogs.
During HIV infection, specific RNA-protein interaction between the Rev response element (RRE) and viral Rev protein is required for nuclear export of intron-containing viral mRNA transcripts. Rev initially binds the high-affinity site in stem-loop II, which promotes oligomerization of additional Rev proteins on RRE. Here, we present the crystal structure of RRE stem-loop II in distinct closed and open conformations. The high-affinity Rev-binding site is located within the three-way junction rather than the predicted stem IIB. The closed and open conformers differ in their non-canonical interactions within the three-way junction, and only the open conformation has the widened major groove conducive to initial Rev interaction. Rev binding assays show that RRE stem-loop II has high- and low-affinity binding sites, each of which binds a Rev dimer. We propose a binding model, wherein Rev-binding sites on RRE are sequentially created through structural rearrangements induced by Rev-RRE interactions.
KRAS has long been referred to as an ‘undruggable’ target due to its high affinity for its cognate ligands (GDP and GTP) and its lack of readily exploited allosteric binding pockets. Recent progress in the development of covalent inhibitors of KRAS G12C has revealed that occupancy of an allosteric binding site located between the α3-helix and switch-II loop of KRAS G12C —sometimes referred to as the ‘switch-II pocket’—holds great potential in the design of direct inhibitors of KRAS G12C . In studying diverse switch-II pocket binders during the development of sotorasib (AMG 510), the first FDA-approved inhibitor of KRAS G12C , we found the dramatic conformational flexibility of the switch-II pocket posing significant challenges toward the structure-based design of inhibitors. Here, we present our computational approaches for dealing with receptor flexibility in the prediction of ligand binding pose and binding affinity. For binding pose prediction, we modified the covalent docking program CovDock to allow for protein conformational mobility. This new docking approach, termed as FlexCovDock, improves success rates from 55 to 89% for binding pose prediction on a dataset of 10 cross-docking cases and has been prospectively validated across diverse ligand chemotypes. For binding affinity prediction, we found standard free energy perturbation (FEP) methods could not adequately handle the significant conformational change of the switch-II loop. We developed a new computational strategy to accelerate conformational transitions through the use of targeted protein mutations. Using this methodology, the mean unsigned error (MUE) of binding affinity prediction were reduced from 1.44 to 0.89 kcal/mol on a set of 14 compounds. These approaches were of significant use in facilitating the structure-based design of KRAS G12C inhibitors and are anticipated to be of further use in the design of covalent (and noncovalent) inhibitors of other conformationally labile protein targets.
Antigen-specific immunotherapies (ASI) require successful loading and presentation of antigen peptides into the major histocompatibility complex (MHC) binding cleft. One route of ASI design is to mutate native antigens for either stronger or weaker binding interaction to MHC. Exploring all possible mutations is costly both experimentally and computationally. To reduce experimental and computational expense, here we investigate the minimal amount of prior data required to accurately predict the relative binding affinity of point mutations for peptide-MHC class II (pMHCII) binding. Using data from different residue subsets, we interpolate pMHCII mutant binding affinities by Gaussian process (GP) regression of residue volume and hydrophobicity. We apply GP regression to an experimental data set from the Immune Epitope Database, and theoretical data sets from NetMHCIIpan and Free Energy Perturbation calculations. We find that GP regression can predict binding affinities of nine neutral residues from a six-residue subset with an average R 2 coefficient of determination value of 0.62 ± 0.04 (±95% CI), average error of 0.09 ± 0.01 kcal/mol (±95% CI), and with an receiver operating characteristic (ROC) AUC value of 0.92 for binary classification of enhanced or diminished binding affinity. Similarly, metrics increase to an R2 value of 0.69 ± 0.04, average error of 0.07 ± 0.01 kcal/mol, and an ROC AUC value of 0.94 for predicting seven neutral residues from an eight-residue subset. Our work finds that prediction is most accurate for neutral residues at anchor residue sites without register shift. This work holds relevance to predicting pMHCII binding and accelerating ASI design.
EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.
Family 1 Coenzyme A transferases (CtfAB) from the extremely thermophilic bacterium, Thermosipho melanesiensis, has been used for in vivo acetone production up to 70°C. This enzyme has tentatively been identified as the rate-limiting step, due to its relatively low-binding affinity for acetate. However, existing kinetic and mechanistic studies on this enzyme are insufficient to evaluate this hypothesis. Here, kinetic analysis of purified recombinant T. melanesiensis CtfAB showed that it has a ping-pong bi-bi mechanism typical of Coenzyme A (CoA) transferases with Km values for acetate and acetoacetyl-CoA of 85 mM and 135 μM, respectively. Product inhibition by acetyl-CoA was competitive with respect to acetoacetyl-CoA and non-competitive with respect to acetate. Crystal structures of wild-type and mutant T. melanesiensis CtfAB were solved in the presence of acetate and in the presence or absence of acetyl-CoA. These structures led to a proposed structural basis for the competitive and non-competitive inhibition of acetyl-CoA: acetate binds independently of acetyl-CoA in an apparent low-affinity binding pocket in CtfA that is directly adjacent to a catalytic glutamate in CtfB. Similar to other CoA transferases, acetyl-CoA is bound in an apparent high-affinity binding site in CtfB with most interactions occurring between the phospho-ADP of CoA and CtfB residues far from the acetate binding pocket. This structural-based mechanism also explains the organic acid promiscuity of CtfAB. High-affinity interactions are predominantly between the conserved phospho-ADP of CoA, and the variable organic acid binding site is a low-affinity binding site with few specific interactions.
The COVID-19 pandemic highlights the need for computational tools to automate and accelerate drug design for novel protein targets. We leverage deep learning language models to generate and score drug candidates based on predicted protein binding affinity. We pre-trained a deep learning language model (BERT) on ∼9.6 billion molecules and achieved peak performance of 603 petaflops in mixed precision. Our work reduces pre-training time from days to hours, compared to previous efforts with this architecture, while also increasing the dataset size by nearly an order of magnitude. For scoring, we fine-tuned the language model using an assembled set of thousands of protein targets with binding affinity data and searched for inhibitors of specific protein targets, SARS-CoV-2 Mpro and PLpro. We utilized a genetic algorithm approach for finding optimal candidates using the generation and scoring capabilities of the language model. Our generalizable models accelerate the identification of inhibitors for emerging therapeutic targets.
Many proteins bind transition metal ions as cofactors to carry out their biological functions. Despite binding affinities for divalent transition metal ions being predominantly dictated by the Irving-Williams series for wild-type proteins, in vivo metal ion binding specificity is ensured by intracellular mechanisms that regulate free metal ion concentrations. However, a growing area of biotechnology research considers the use of metal-binding proteins in vitro to purify specific metal ions from wastewater, where specificity is dictated by the protein's metal binding affinities. A goal of metalloprotein engineering is to modulate these affinities to improve a protein's specificity towards a particular metal; however, the quantitative relationship between the affinities and the equilibrium metal-bound protein fractions depends on the underlying binding mechanisms. Here we demonstrate a high-throughput intrinsic tryptophan fluorescence quenching method to validate binding models in multi-metal solutions for CcNikZ-II, a nickel-binding protein from Clostridium carboxidivorans. Using our validated models, we quantify the relationship between binding affinity and specificity in different classes of metal-binding models for CcNikZ-II. In conclusion, we further illustrate the potential relevance of data-informed models to predicting engineering targets for improved specificity.
Stereospecific recognition of chiral molecules plays a crucial role in biological systems. The μ-opioid receptor (MOR) exhibits binding affinity towards (-)-morphine, a well-established gold standard in pain management, while it shows minimal binding affinity for the (+)-morphine enantiomer, resulting in a lack of analgesic activity. Understanding how MOR stereoselectively recognizes morphine enantiomers has remained a puzzle in neuroscience and pharmacology for over half-a-century due to the lack of direct observation techniques. To unravel this mystery, we constructed the binding and unbinding processes of morphine enantiomers with MOR via molecular dynamics simulations to investigate the thermodynamics and kinetics governing MOR's stereoselective recognition of morphine enantiomers. Our findings reveal that the binding of (-)-morphine stabilizes MOR in its activated state, exhibiting a deep energy well and a prolonged residence time. In contrast, (+)-morphine fails to sustain the activation state of MOR. Furthermore, the results suggest that specific residues, namely D114 2.50 and D147 3.32 , are deprotonated in the active state of MOR bound to (-)-morphine. This work highlights that the selectivity in molecular recognition goes beyond binding affinities, extending into the realm of residence time.