Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “drug discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

FL-DISCO: Federated Generative Adversarial Network for Graph-based Molecule Drug Discovery: Special Session Paper

The outbreak of the global COVID-19 pandemic emphasizes the importance of collaborative drug discovery for high effectiveness; however, due to the stringent data regulation, data privacy becomes an imminent issue needing to be addressed to enable collaborative drug discovery. In addition to the data privacy issue, the efficiency of drug discovery is another key objective since infectious diseases spread exponentially and effectively conducting drug discovery could save lives. Advanced Artificial Intelligence (AI) techniques are promising to solve these problems: (1) Federated Learning (FL) is born to keep data privacy while learning data from distributed clients; (2) graph neural network (GNN) can extract structural properties of molecules whose underlying architecture is the connected atoms; and (3) generative adversarial network (GAN) can generate novel molecules while retaining the properties learned from the training data. In this work, we make the first attempt to build a holistic collaborative and privacy-preserving FL framework, namely FL- DISCO, which integrates GAN and GNN to generate molecular graphs. Experimental results demonstrate the effectiveness of FL- DISCO on: (1) IID data for ESOL and QM9, where FL-DISCO can generate highly novel compounds with high drug-likeliness, uniqueness and LogP scores compared to the baseline; (2) non- IID data for ESOL and QM9, where FL-DISCO generates 100% novel compounds with high validity and LogP scores compared to the baseline. We also demonstrate how different fractions of clients, generator and discriminator architectures affect our evaluation scores.

Manu, Daniel↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Deep generative molecular design reshapes drug discovery

Recent advances and accomplishments of artificial intelligence (AI) and deep generative models have established their usefulness in medicinal applications, especially in drug discovery and development. To correctly apply AI, the developer and user face questions such as which protocols to consider, which factors to scrutinize, and how the deep generative models can integrate the relevant disciplines. This review summarizes classical and newly developed AI approaches, providing an updated and accessible guide to the broad computational drug discovery and development community. We introduce deep generative models from different standpoints and describe the theoretical frameworks for representing chemical and biological structures and their applications. We discuss the data and technical challenges and highlight future directions of multimodal deep generative models for accelerating drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Pandemic drugs at pandemic speed: infrastructure for accelerating COVID-19 drug discovery with hybrid machine learning- and physics-based simulations on high-performance computers

The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottleneck screening the vast number of potential small molecules to shortlist lead compounds for antiviral drug development. New opportunities to accelerate drug discovery lie at the interface between machine learning methods, in this case, developed for linear accelerators, and physics-based methods. The two in silico methods, each have their own advantages and limitations which, interestingly, complement each other. Here, we present an innovative infrastructural development that combines both approaches to accelerate drug discovery. The scale of the potential resulting workflow is such that it is dependent on supercomputing to achieve extremely high throughput. We have demonstrated the viability of this workflow for the study of inhibitors for four COVID-19 target proteins and our ability to perform the required large-scale calculations to identify lead antiviral compounds through repurposing on a variety of supercomputers.

97 MATHEMATICS AND COMPUTING↗

An artificial intelligence accelerated virtual screening platform for drug discovery

Abstract Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we develop a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screen multi-billion compound libraries against two unrelated targets, a ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel Na V 1.7. For both targets, we discover hit compounds, including seven hits (14% hit rate) to KLHDC2 and four hits (44% hit rate) to Na V 1.7, all with single digit micromolar binding affinities. Screening in both cases is completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

Science & Technology - Other Topics↗

A New Drug Discovery Platform: Application to DNA Polymerase Eta and Apurinic/Apyrimidinic Endonuclease 1

The ability to quickly discover reliable hits from screening and rapidly convert them into lead compounds, which can be verified in functional assays, is central to drug discovery. The expedited validation of novel targets and the identification of modulators to advance to preclinical studies can significantly increase drug development success. Our SaXPyTM (“SAR by X-ray Poses Quickly”) platform, which is applicable to any X-ray crystallography-enabled drug target, couples the established methods of protein X-ray crystallography and fragment-based drug discovery (FBDD) with advanced computational and medicinal chemistry to deliver small molecule modulators or targeted protein degradation ligands in a short timeframe. Our approach, especially for elusive or “undruggable” targets, allows for (i) hit generation; (ii) the mapping of protein–ligand interactions; (iii) the assessment of target ligandability; (iv) the discovery of novel and potential allosteric binding sites; and (v) hit-to-lead execution. These advances inform chemical tractability and downstream biology and generate novel intellectual property. We describe here the application of SaXPy in the discovery and development of DNA damage response inhibitors against DNA polymerase eta (Pol η or POLH) and apurinic/apyrimidinic endonuclease 1 (APE1 or APEX1). Notably, our SaXPy platform allowed us to solve the first crystal structures of these proteins bound to small molecules and to discover novel binding sites for each target.

59 BASIC BIOLOGICAL SCIENCES↗

Annual Report for Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we plan to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We also plan to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task to identify pan-coronavirus protease inhibitors such as SARS-CoV-2. While the overall goals and milestones remain consistent with the original proposal, certain technical details have been modified, which we will describe in this report.

97 MATHEMATICS AND COMPUTING↗

Naegleria fowleri: Protein structures to facilitate drug discovery for the deadly, pathogenic free-living amoeba

Naegleria fowleri is a pathogenic, thermophilic, free-living amoeba which causes primary amebic meningoencephalitis (PAM). Penetrating the olfactory mucosa, the brain-eating amoeba travels along the olfactory nerves, burrowing through the cribriform plate to its destination: the brain’s frontal lobes. The amoeba thrives in warm, freshwater environments, with peak infection rates in the summer months and has a mortality rate of approximately 97%. A major contributor to the pathogen’s high mortality is the lack of sensitivity of N . fowleri to current drug therapies, even in the face of combination-drug therapy. To enable rational drug discovery and design efforts we have pursued protein production and crystallography-based structure determination efforts for likely drug targets from N . fowleri . The genes were selected if they had homology to drug targets listed in Drug Bank or were nominated by primary investigators engaged in N . fowleri research. In 2017, 178 N . fowleri protein targets were queued to the Seattle Structural Genomics Center of Infectious Disease (SSGCID) pipeline, and to date 89 soluble recombinant proteins and 19 unique target structures have been produced. Many of the new protein structures are potential drug targets and contain structural differences compared to their human homologs, which could allow for the development of pathogen-specific inhibitors. Five of the structures were analyzed in more detail, and four of five show promise that selective inhibitors of the active site could be found. The 19 solved crystal structures build a foundation for future work in combating this devastating disease by encouraging further investigation to stimulate drug discovery for this neglected pathogen.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗

High-throughput virtual laboratory for drug discovery using massive datasets

Time-to-solution for structure-based screening of massive chemical databases for COVID-19 drug discovery has been decreased by an order of magnitude, and a virtual laboratory has been deployed at scale on up to 27,612 GPUs on the Summit supercomputer, allowing an average molecular docking of 19,028 compounds per second. Over one billion compounds were docked to two SARS-CoV-2 protein structures with full optimization of ligand position and 20 poses per docking, each in under 24 hours. GPU acceleration and high-throughput optimizations of the docking program produced 350× mean speedup over the CPU version (50× speedup per node). GPU acceleration of both feature calculation for machine-learning based scoring and distributed database queries reduced processing of the 2.4 TB output by orders of magnitude. The resulting 50× speedup for the full pipeline reduces an initial 43 day runtime to 21 hours per protein for providing high-scoring compounds to experimental collaborators for validation assays.

97 MATHEMATICS AND COMPUTING↗

Early Drug Discovery and Development of Novel Cancer Therapeutics Targeting DNA Polymerase Eta (POLH)

Polymerase eta (or Pol η or POLH) is a specialized DNA polymerase that is able to bypass certain blocking lesions, such as those generated by ultraviolet radiation (UVR) or cisplatin, and is deployed to replication foci for translesion synthesis as part of the DNA damage response (DDR). Inherited defects in the gene encoding POLH (a.k.a., XPV) are associated with the rare, sun-sensitive, cancer-prone disorder, xeroderma pigmentosum, owing to the enzyme’s ability to accurately bypass UVR-induced thymine dimers. In standard-of-care cancer therapies involving platinum-based clinical agents, e.g., cisplatin or oxaliplatin, POLH can bypass platinum-DNA adducts, negating benefits of the treatment and enabling drug resistance. POLH inhibition can sensitize cells to platinum-based chemotherapies, and the polymerase has also been implicated in resistance to nucleoside analogs, such as gemcitabine. POLH overexpression has been linked to the development of chemoresistance in several cancers, including lung, ovarian, and bladder. Co-inhibition of POLH and the ATR serine/threonine kinase, another DDR protein, causes synthetic lethality in a range of cancers, reinforcing that POLH is an emerging target for the development of novel oncology therapeutics. Using a fragment-based drug discovery approach in combination with an optimized crystallization screen, we have solved the first X-ray crystal structures of small novel drug-like compounds, i.e., fragments, bound to POLH, as starting points for the design of POLH inhibitors. The intrinsic molecular resolution afforded by the method can be quickly exploited in fragment growth and elaboration as well as analog scoping and scaffold hopping using medicinal and computational chemistry to advance hits to lead. An initial small round of medicinal chemistry has resulted in inhibitors with a range of functional activity in an in vitro biochemical assay, leading to the rapid identification of an inhibitor to advance to subsequent rounds of chemistry to generate a lead compound. Importantly, our chemical matter is different from the traditional nucleoside analog-based approaches for targeting DNA polymerases.

60 APPLIED LIFE SCIENCES↗

Drug discovery efforts at George Mason University

With over 39,000 students, and research expenditures in excess of $200 million, George Mason University (GMU) is the largest R1 (Carnegie Classification of very high research activity) university in Virginia. Mason scientists have been involved in the discovery and development of novel diagnostics and therapeutics in areas as diverse as infectious diseases and cancer. Below are highlights of the efforts being led by Mason researchers in the drug discovery arena. To enable targeted cellular delivery, and non-biomedical applications, Veneziano and colleagues have developed a synthesis strategy that enables the design of self-assembling DNA nanoparticles (DNA origami) with prescribed shape and size in the 10 to 100 nm range. The nanoparticles can be loaded with molecules of interest such as drugs, proteins and peptides, and are a promising new addition to the drug delivery platforms currently in use. The investigators also recently used the DNA origami nanoparticles to fine tune the spatial presentation of immunogens to study the impact on B cell activation. These studies are an important step towards the rational design of vaccines for a variety of infectious agents. To elucidate the parameters for optimizing the delivery efficiency of lipid nanoparticles (LNPs), Buschmann, Paige and colleagues have devised methods for predicting and experimentally validating the pKa of LNPs based on the structure of the ionizable lipids used to formulate the LNPs. These studies may pave the way for the development of new LNP delivery vehicles that have reduced systemic distribution and improved endosomal release of their cargo post administration. To better understand protein-protein interactions and identify potential drug targets that disrupt such interactions, Luchini and colleagues have developed a methodology that identifies contact points between proteins using small molecule dyes. The dye molecules noncovalently bind to the accessible surfaces of a protein complex with very high affinity, but are excluded from contact regions. When the complex is denatured and digested with trypsin, the exposed regions covered by the dye do not get cleaved by the enzyme, whereas the contact points are digested. The resulting fragments can then be identified using mass spectrometry. The data generated can serve as the basis for designing small molecules and peptides that can disrupt the formation of protein complexes involved in disease processes. For example, using peptides based on the interleukin 1 receptor accessory protein (IL-1RAcP), Luchini, Liotta, Paige and colleagues disrupted the formation of IL-1/IL-R/IL-1RAcP complex and demonstrated that the inhibition of complex formation reduced the inflammatory response to IL-1B. Working on the discovery of novel antimicrobial agents, Bishop, van Hoek and colleagues have discovered a number of antimicrobial peptides from reptiles and other species. DRGN-1, is a synthetic peptide based on a histone H1-derived peptide that they had identified from Komodo Dragon plasma. DRGN-1 was shown to disrupt bacterial biofilms and promote wound healing in an animal model. The peptide, along with others, is being developed and tested in preclinical studies. Other research by van Hoek and colleagues focuses on in silico antimicrobial peptide discovery, screening of small molecules for antibacterial properties, as well as assessment of diffusible signal factors (DFS) as future therapeutics. The above examples provide insight into the cutting-edge studies undertaken by GMU scientists to develop novel methodologies and platform technologies important to drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

AI-based language models powering drug discovery and development

The discovery and development of new medicines is expensive, time-consuming, and often inefficient, with many failures along the way. Powered by artificial intelligence (AI), language models (LMs) have changed the landscape of natural language processing (NLP), offering possibilities to transform treatment development more effectively. Here, we summarize advances in AI-powered LMs and their potential to aid drug discovery and development. We highlight opportunities for AI-powered LMs in target identification, clinical design, regulatory decision-making, and pharmacovigilance. We specifically emphasize the potential role of AI-powered LMs for developing new treatments for Coronavirus 2019 (COVID-19) strategies, including drug repurposing, which can be extrapolated to other infectious diseases that have the potential to cause pandemics. Finally, we set out the remaining challenges and propose possible solutions for improvement.

60 APPLIED LIFE SCIENCES↗

American Heart Association Protein Atlas Portal for Accelerated Drug Discovery

The American Heart Association's Protein Atlas Data Portal is a cutting-edge, open-source atlas of the structural determinants of the binding of drugs to (all of) their target proteins as an unbiased means to accelerate drug discovery. The portal will provide researchers with access to comprehensive data that can help increase the selection of effective drugs and cut the time-to-market of new drugs by half.

drug targets↗

Green genes from blue greens: challenges and solutions to unlocking the potential of cyanobacteria in drug discovery

Cyanobacteria are prolific producers of biologically active compounds that are important in influencing ecology, behavior of interacting organisms, and as leads in drug discovery efforts. Here we discuss the challenges faced by all natural product researchers, especially those that focus on cyanobacteria, and then describe progress that has been made in these areas. We also propose some solutions, paths forward, and thoughts for consideration on these challenges.

Philmus, Benjamin↗

Rational approach to drug discovery for human schistosomiasis

Human schistosomiasis is a debilitating, life-threatening disease affecting more than 229 million people in as many as 78 countries. There is only one drug of choice effective against all three major species of Schistosoma, praziquantel (PZQ). However, as with many monotherapies, evidence for resistance is emerging in the field and can be selected for in the laboratory. Previously used therapies include oxamniquine (OXA), but shortcomings such as drug resistance and affordability resulted in discontinuation. Employing a genetic, biochemical and molecular approach, a sulfotransferase (SULT-OR) was identified as responsible for OXA drug resistance. By crystallizing SmSULT- OR with OXA, the mode of action of OXA was determined. This information allowed a rational approach to novel drug design. Our team approach with schistosome biologists, medicinal chemists, structural biologists and geneticists has enabled us to develop and test novel drug derivatives of OXA to treat this disease. Using an iterative process for drug development, we have successfully identified derivatives that are effective against all three species of the parasite. One derivative CIDD-0149830 kills 100% of all three human schistosome species within 5 days. The goal is to generate a second therapeutic with a different mode of action that can be used in conjunction with praziquantel to overcome the ever-growing threat of resistance and improve efficacy. The ability and need to design, screen, and develop future, affordable therapeutics to treat human schistosomiasis is critical for successful control program outcomes.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial Graph Attention and Curiosity-driven Policy for Antiviral Drug Discovery

We developed Distilled Graph Attention Policy Network (DGAPN), a reinforcement learning model to generate novel graph-structured chemical representations that optimize user-defined objectives by efficiently navigating a physically constrained domain. The framework is examined on the task of generating molecules that are designed to bind, noncovalently, to functional sites of SARS-CoV-2 proteins. We present a spatial Graph Attention (sGAT) mechanism that leverages self-attention over both node and edge attributes as well as encoding the spatial structure --- this capability is of considerable interest in synthetic biology and drug discovery. An attentional policy network is introduced to learn the decision rules for a dynamic, fragment-based chemical environment, and state-of-the-art policy gradient techniques are employed to train the network with stability. Exploration is driven by the stochasticity of the action space design and the innovation reward bonuses learned and proposed by random network distillation. In experiments, our framework achieved outstanding results compared to state-of-the-art algorithms, while reducing the complexity of paths to chemical synthesis.

Wu, Yulun↗

SPATIAL GRAPH ATTENTION AND CURIOSITY-DRIVEN POLICY FOR ANTIVIRAL DRUG DISCOVERY

We developed Distilled Graph Attention Policy Network (DGAPN), a reinforcement learning model to generate novel graph-structured chemical representations that optimize user-defined objectives by efficiently navigating a physically constrained domain. The framework is examined on the task of generating molecules that are designed to bind, noncovalently, to functional sites of SARS-CoV-2 proteins. We present a spatial Graph Attention (sGAT) mechanism that leverages self-attention over both node and edge attributes as well as encoding the spatial structure - this capability is of considerable interest in synthetic biology and drug discovery. An attentional policy network is introduced to learn the decision rules for a dynamic, fragment-based chemical environment, and state-of-the-art policy gradient techniques are employed to train the network with stability. Exploration is driven by the stochasticity of the action space design and the innovation reward bonuses learned and proposed by random network distillation. In experiments, our framework achieved outstanding results compared to state-of-the-art algorithms, while reducing the complexity of paths to chemical synthesis.

Wu, Y↗