Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Virtual drug screening”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

30 records · Page 2

Pose Classification Using Three-Dimensional Atomic Structure-Based Neural Networks Applied to Ion Channel–Ligand Docking

The identification of promising lead compounds showing pharmacological activities toward a biological target is essential in early stage drug discovery. With the recent increase in available small-molecule databases, virtual high-throughput screening using physics-based molecular docking has emerged as an essential tool in assisting fast and cost-efficient lead discovery and optimization. However, the best scored docking poses are often suboptimal, resulting in incorrect screening and chemical property calculation. We address the pose classification problem by leveraging data-driven machine learning approaches to identify correct docking poses from AutoDock Vina and Glide screens. To enable effective classification of docking poses, we present two convolutional neural network approaches: a three-dimensional convolutional neural network (3D-CNN) and an attention-based point cloud network (PCN) trained on the PDBbind refined set. We demonstrate the effectiveness of our proposed classifiers on multiple evaluation data sets including the standard PDBbind CASF-2016 benchmark data set and various compound libraries with structurally different protein targets including an ion channel data set extracted from Protein Data Bank (PDB) and an in-house KCa3.1 inhibitor data set. Our experiments show that excluding false positive docking poses using the proposed classifiers improves virtual high-throughput screening to identify novel molecules against each target protein compared to the initial screen based on the docking scores.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

CoarsenConf: Equivariant Coarsening with Aggregated Attention for Molecular Conformer Generation

Molecular conformer generation (MCG) is an important task in cheminformatics and drug discovery. The ability to efficiently generate low-energy 3D structures can avoid expensive quantum mechanical simulations, leading to accelerated virtual screenings and enhanced structural exploration. Several generative models have been developed for MCG, but many struggle to consistently produce high-quality conformers for meaningful downstream applications. To address these issues, we introduce CoarsenConf, which coarse-grains molecular graphs based on torsional angles and integrates them into an SE(3)-equivariant hierarchical variational autoencoder. Through equivariant coarse-graining, we aggregate the fine-grained atomic coordinates of subgraphs connected via rotatable bonds, creating a variable-length coarse-grained latent representation. Our model uses a novel aggregated attention mechanism to restore fine-grained coordinates from the coarse-grained latent representation, enabling efficient generation of accurate conformers. Furthermore, we evaluate the chemical and biochemical quality of our generated conformers on multiple downstream applications, including property prediction and large-scale oracle-based protein docking. Overall, CoarsenConf generates more accurate conformer ensembles compared to prior generative models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

IMPECCABLE: Integrated Modeling Pipeline for COVID Cure by Assessing Better Leads

ABSTRACT The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100× to 1000× improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ~10 11 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.

97 MATHEMATICS AND COMPUTING↗

Machine learning and ligand binding predictions: A review of data, methods, and obstacles

We report that computational predictions of ligand binding is a difficult problem, with more accurate methods being extremely computationally expensive. The use of machine learning for drug binding predictions could possibly leverage the use of biomedical big data in exchange for time-intensive simulations. This paper reviews current trends in the use of machine learning for drug binding predictions, data sources to develop machine learning algorithms, and potential problems that may lead to overfitting and ungeneralizable models. A few popular datasets that can be used to develop virtual high-throughput screening models are characterized using spatial statistics to quantify potential biases. We can see from evaluating some common benchmarks that good performance correlates with models with high-predicted bias scores and models with low bias scores do not have much predictive power. A better understanding of the limits of available data sources and how to fix them will lead to more generalizable models that will lead to novel drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

AI-accelerated protein-ligand docking for SARS-CoV-2 is 100-fold faster with no significant change in detection

Protein-ligand docking is a computational method for identifying drug leads. The method is capable of narrowing a vast library of compounds down to a tractable size for downstream simulation or experimental testing and is widely used in drug discovery. While there has been progress in accelerating scoring of compounds with artificial intelligence, few works have bridged these successes back to the virtual screening community in terms of utility and forward-looking development. We demonstrate the power of high-speed ML models by scoring 1 billion molecules in under a day (50 k predictions per GPU seconds). We showcase a workflow for docking utilizing surrogate AI-based models as a pre-filter to a standard docking workflow. Our workflow is ten times faster at screening a library of compounds than the standard technique, with an error rate less than 0.01% of detecting the underlying best scoring 0.1% of compounds. Our analysis of the speedup explains that another order of magnitude speedup must come from model accuracy rather than computing speed. In order to drive another order of magnitude of acceleration, we share a benchmark dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million “in-stock” molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome. We believe this is strong evidence for the community to begin focusing on improving the accuracy of surrogate models to improve the ability to screen massive compound libraries 100 × or even 1000 × faster than current techniques and reduce missing top hits. The technique outlined aims to be a fast drop-in replacement for docking for screening billion-scale molecular libraries.

59 BASIC BIOLOGICAL SCIENCES↗

Inhibition of the C1s Protease and the Classical Complement Pathway by 6-(4-Phenylpiperazin-1-yl)Pyridine-3-Carboximidamide and Chemical Analogs

Abstract The classical pathway (CP) is a potent mechanism for initiating complement activity and is a driver of pathology in many complement-mediated diseases. The CP is initiated via activation of complement component C1, which consists of the pattern recognition molecule C1q bound to a tetrameric assembly of proteases C1r and C1s. Enzymatically active C1s provides the catalytic basis for cleavage of the downstream CP components, C4 and C2, and is therefore an attractive target for therapeutic intervention in CP-driven diseases. Although an anti-C1s mAb has been Food and Drug Administration approved, identifying small-molecule C1s inhibitors remains a priority. In this study, we describe 6-(4-phenylpiperazin-1-yl)pyridine-3-carboximidamide (A1) as a selective, competitive inhibitor of C1s. A1 was identified through a virtual screen for small molecules that interact with the C1s substrate recognition site. Subsequent functional studies revealed that A1 dose-dependently inhibits CP activation by heparin-induced immune complexes, CP-driven lysis of Ab-sensitized sheep erythrocytes, CP activation in a pathway-specific ELISA, and cleavage of C2 by C1s. Biochemical experiments demonstrated that A1 binds directly to C1s with a K d of ∼9.8 μM and competitively inhibits its activity with an inhibition constant (K i) of ∼5.8 μM. A 1.8-Å-resolution crystal structure revealed the physical basis for C1s inhibition by A1 and provided information on the structure–activity relationship of the A1 scaffold, which was supported by evaluating a panel of A1 analogs. Taken together, our work identifies A1 as a new class of small-molecule C1s inhibitor and lays the foundation for development of increasingly potent and selective A1 analogs for both research and therapeutic purposes.

Immunology↗

Computational studies reveal Fluorine based quinolines to be potent inhibitors for proteins involved in SARS-CoV-2 assembly

World is witnessing one of the worst pandemics of this century caused by SARS-CoV-2 virus which has affected millions of individuals. Despite rapid efforts to develop vaccines and drugs for COVID-19, the disease is still not under control. Chloroquine (CQ) and Hydroxychloroquine (HCQ) are two very promising inhibitors which have shown positive effect in combating the disease in preliminary experimental studies, but their use was reduced due to severe side-effects. Here, we performed a theoretical investigation of the same by studying the binding of the molecules with SARS-COV-2 Spike protein, the complex formed by Spike and ACE2 human receptor and a human serine protease TMPRSS2 which aids in cleavage of the Spike protein to initiate the viral activation in the body. Both the molecules had shown very good docking energies in the range of -6kcal/mol. Subsequently, we did a high throughput screening for other potential quinoline candidates which could be used as inhibitors. From the large pool of ligand candidates, we shortlisted the top three ligands (binding energy -8kcal/mol). We tested the stability of the docked complexes by running Molecular Dynamics (MD) simulations where we observed the stability of the quinoline analogues with the Spike-ACE2 and TMPRSS2 nevertheless the quinolines were not stable with the Spike protein alone. Thus, although the inhibitors bond very well with the protein molecules their intrinsic binding affinity depends on the protein dynamics. Moreover, the quinolines were stable when bound to electronegative pockets of Spike-ACE2 or TMPRSS2 but not with Viral Spike protein. We also observed that a Fluoride based compound: 3-[3-(Trifluoromethyl)phenyl]quinoline helps the inhibitor to bind with both Spike-ACE2 and TMPRSS2 with equal probability. The molecular details presented in this study would be very useful for developing quinoline based drugs for COVID-19 treatment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

3D-Scaffold: A Deep Learning Framework to Generate 3D Coordinates of Drug-like Molecules with Desired Scaffolds

The prerequisite of therapeutic drug design is to identify novel molecules with desired biophysical and biochemical properties. Deep generative models have demonstrated their ability to find such molecules by exploring a huge chemical space efficiently. An effective way to obtain molecules with desired target properties is the preservation of critical scaffolds in the generation process. To this end, we propose a domain aware generative framework called 3D-Scaffold that takes 3D coordinates of a desired scaffold as an input and generates 3D coordinates of novel therapeutic candidates as an output while always preserving the desired scaffolds in generated structures. We show that our framework generates predominantly valid, unique, novel, and experimentally synthesizable molecules that have drug-like properties similar to the molecules in the training set. Using domain specific datasets, we generate covalent and non-covalent antiviral inhibitors. Therefore, to measure the success of our framework in generating therapeutic candidates, generated structures were subjected to high throughput virtual screening via docking simulations, which shows favorable interaction against SARS-CoV-2 main protease and non-structural protein endoribonuclease (NSP15) targets. Most importantly, our model performs well with relatively small volumes of training data and generalizes to new scaffolds, making it applicable to other domain.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Drugging the entire human proteome: Are we there yet?

Each of the ~20 000 proteins in the human proteome is a potential target for compounds that bind to it and modify its function. The 3D structures of most of these proteins are now available. Here, we discuss the prospects for using these structures to perform proteome-wide virtual HTS (VHTS). Furthermore, we compare physics-based (docking) and AI VHTS approaches, some of which are now being applied with large databases of compounds to thousands of targets. Although preliminary proteome-wide screens are now within our grasp, further methodological developments are expected to improve the accuracy of the results.

60 APPLIED LIFE SCIENCES↗

Discovery of an autoinhibited conformation in mesotrypsin reveals a strategy for selective serine protease inhibition

Selective inhibition of the more than 100 S1 family serine proteases is a long-standing challenge due to their active site similarity. Mesotrypsin, implicated in cancer progression, exemplifies these difficulties; no current inhibitors achieve selectivity over other human trypsins. We found an unexpected autoinhibited conformation of mesotrypsin via x-ray crystallography, revealing a cryptic pocket adjacent to the active site. Using high-throughput virtual screening targeting this cryptic pocket, we identified a conformationally selective small-molecule inhibitor that stabilizes the inactive state of mesotrypsin. This inhibitor demonstrates selectivity for mesotrypsin over other trypsins. Our findings challenge the accepted view of digestive trypsins as constitutively active enzymes lacking potential for allosteric regulation. Furthermore, analyses of other structures suggest that dynamic sampling of closed states with analogous allosteric cryptic pockets appears widespread among S1 serine proteases. These observations point to a potentially generalizable strategy to achieve selective inhibition, offering broad implications for drug development targeting serine proteases in cancer and other diseases.

Coban, Matt↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗