Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “proteins binding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Proteome-wide analysis of protein stability in Escherichia coli under acid stress

Knowledge of protein acid sensitivity remains sparse and is largely derived from low-throughput, enzyme-specific assays. We used a scalable framework to map acid stability across the Escherichia coli proteome to assess the acid stability of 1,675 unique proteins, estimating pH 50 values for over 90% of them. The parameter pH50 was defined as the pH value at which only 50% of the initial protein remains in solution following acid treatment. Proteome-wide pH 50 values ranged from 2.28 to 6.33 (median 5.11). Approximately 9% of detected proteins remained stable across all tested pH conditions. Our results align with published data and the assay of citrate synthase (GltA) performed here. Protein acid stability differed significantly by subcellular localization: periplasmic proteins were relatively more abundant in the acid-stable group, cytoplasmic proteins were abundant at pH 50 values 4.5–5.5, and inner membrane proteins at higher pH 50 between 5.5 and 6.0. Outer membrane proteins were too few to draw strong conclusions regarding enrichment within specific pH 50 groups. Notably, the periplasmic binding protein of the molybdate ABC transporter (ModA), was enriched after incubation at low pH. Estimated pH 50 values showed no correlation with protein isoelectric point and molecular weight. Together, this work provides the first proteome-wide map of protein acid stability and establishes a general framework for studying different chemical stressors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Integrated multi-omic characterizations of the synapse reveal RNA processing factors and ubiquitin ligases associated with neurodevelopmental disorders

The molecular composition of the excitatory synapse is incompletely defined due to its dynamic nature across developmental stages and neuronal populations. To address this gap, we apply proteomic mass spectrometry to characterize the synapse in multiple biological models including the fetal human brain and hiPSC-derived neurons. To prioritize the identified proteins, we develop an orthogonal multi-omic screen of genomic, transcriptomic, interactomic, and structural data. This data-driven framework identifies proteins with key molecular features intrinsic to the synapse, including characteristic patterns of biophysical interactions and cross-tissue expression. The multi-omic analysis captures synaptic proteins across developmental stages and experimental systems, including 493 synaptic candidates supported by proteomics. We further investigate three such proteins that are associated with neurodevelopmental disorders – the CUL3 E3 ubiquitin ligase, the DDX3X and YBX1 nucleic-acid binding proteins – by mapping their networks of physically interacting synapse proteins or transcripts. Our study demonstrates the potential of an integrated multi-omic approach to systematically and more comprehensively resolve the synaptic architecture.

59 BASIC BIOLOGICAL SCIENCES↗

Pumping Iron: A Multi-omics Analysis of Two Extremophilic Algae Reveals Iron Economy Management

Marine algae are responsible for half of the world's primary productivity, but this critical carbonsink is often constrained by insufficient iron. One species of marine algae, Dunaliella tertiolecta, isremarkable for its ability to maintain photosynthesis and thrive in low-iron environments. A relatedspecies, Dunaliella salina Bardawil, shares this attribute but is an extremophile found in hypersaline environments. To elucidate how algae manage their iron requirements, we produced highquality genome assemblies and transcriptomes for both species to serve as a foundation for acomparative multi-omics analysis. We identified a host of iron-uptake proteins in both species,including a massive expansion of transferrins and a novel family of siderophore-iron uptakeproteins. Complementing these multiple iron-uptake routes, ferredoxin functions as a large ironreservoir that can be released by induction of flavodoxin. Proteomic analysis revealed reducedinvestment in the photosynthetic apparatus coupled with remodeling of antenna proteins bydramatic iron-deficiency induction of TIDI1, a light harvesting complex protein found also in otherchlorophytes. These combinatorial iron scavenging and sparing strategies make Dunaliellaunique among photosynthetic organisms

iron homeostasis, phytoplankton, Iron starvation i↗

Comprehensive analysis of the human ESCRT-III-MIT domain interactome reveals new cofactors for cytokinetic abscission

The 12 related human ESCRT-III proteins form filaments that constrict membranes and mediate fission, including during cytokinetic abscission. The C-terminal tails of polymerized ESCRT-III subunits also bind proteins that contain Microtubule-Interacting and Trafficking (MIT) domains. MIT domains can interact with ESCRT-III tails in many different ways to create a complex binding code that is used to recruit essential cofactors to sites of ESCRT activity. Here, we have comprehensively and quantitatively mapped the interactions between all known ESCRT-III tails and 19 recombinant human MIT domains. We measured 228 pairwise interactions, quantified 60 positive interactions, and discovered 18 previously unreported interactions. We also report the crystal structure of the SPASTIN MIT domain in complex with the IST1 C-terminal tail. Three MIT enzymes were studied in detail and shown to: (1) localize to cytokinetic midbody membrane bridges through interactions with their specific ESCRT-III binding partners (SPASTIN-IST1, KATNA1-CHMP3, and CAPN7-IST1), (2) function in abscission (SPASTIN, KATNA1, and CAPN7), and (3) function in the ‘NoCut’ abscission checkpoint (SPASTIN and CAPN7). Our studies define the human MIT-ESCRT-III interactome, identify new factors and activities required for cytokinetic abscission and its regulation, and provide a platform for analyzing ESCRT-III and MIT cofactor interactions in all ESCRT-mediated processes.

59 BASIC BIOLOGICAL SCIENCES↗

Cancer-specific loss of TERT activation sensitizes glioblastoma to DNA damage

Most glioblastomas (GBMs) achieve cellular immortality by acquiring a mutation in the telomerase reverse transcriptase ( TERT ) promoter. TERT promoter mutations create a binding site for a GA binding protein (GABP) transcription factor complex, whose assembly at the promoter is associated with TERT reactivation and telomere maintenance. Here, we demonstrate increased binding of a specific GABPB1L-isoform–containing complex to the mutant TERT promoter. Furthermore, we find that TERT promoter mutant GBM cells, unlike wild-type cells, exhibit a critical near-term dependence on GABPB1L for proliferation, notably also posttumor establishment in vivo. Up-regulation of the protein paralogue GABPB2, which is normally expressed at very low levels, can rescue this dependence. More importantly, when combined with frontline temozolomide (TMZ) chemotherapy, inducible GABPB1L knockdown and the associated TERT reduction led to an impaired DNA damage response that resulted in profoundly reduced growth of intracranial GBM tumors. Together, these findings provide insights into the mechanism of cancer-specific TERT regulation, uncover rapid effects of GABPB1L-mediated TERT suppression in GBM maintenance, and establish GABPB1L inhibition in combination with chemotherapy as a therapeutic strategy for TERT promoter mutant GBM.

59 BASIC BIOLOGICAL SCIENCES↗

ATM–dependent phosphorylation of CHD7 regulates morphogenesis-coupled DSB stress response in fetal radiation exposure

Following radiation exposure, unrepaired DNA double-strand breaks (DSBs) persist to some extent in a subset of cells as residual damage; they can exert adverse effects, including late-onset diseases. In search of the factor(s) that characterize(s) cells bearing such damage, we discovered ataxia-telangiectasia mutated (ATM)-dependent phosphorylation of the transcription factor chromodomain helicase DNA binding protein 7 (CHD7). CHD7 controls the morphogenesis of cell populations derived from neural crest cells during vertebrate early development. Indeed, malformations in various fetal bodies are attributable to CHD7 haploinsufficiency. Following radiation exposure, CHD7 becomes phosphorylated, ceases promoter/enhancer binding to target genes, and relocates to the DSB-repair protein complex, where it remains until the damage is repaired. Thus, ATM-dependent CHD7 phosphorylation appears to act as a functional switch. As such stress responses contribute to improved cell survival and canonical nonhomologous end joining, we conclude that CHD7 is involved in both morphogenetic and DSB-response functions. Thus, we propose that higher vertebrates have evolved intrinsic mechanisms underlying the morphogenesis-coupled DSB stress response. In fetal exposure, if the function of CHD7 becomes primarily shifted toward DNA repair, morphogenic activity is reduced, resulting in malformations.

59 BASIC BIOLOGICAL SCIENCES↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

Mechanism of karyopherin-β2 binding and nuclear import of ALS variants FUS(P525L) and FUS(R495X)

Mutations in the RNA-binding protein FUS cause familial amyotropic lateral sclerosis (ALS). Several mutations that affect the proline-tyrosine nuclear localization signal (PY-NLS) of FUS cause severe juvenile ALS. FUS also undergoes liquid–liquid phase separation (LLPS) to accumulate in stress granules when cells are stressed. In unstressed cells, wild type FUS resides predominantly in the nucleus as it is imported by the importin Karyopherin-β2 (Kapβ2), which binds with high affinity to the C-terminal PY-NLS of FUS. Here, we analyze the interactions between two ALS-related variants FUS(P525L) and FUS(R495X) with importins, especially Kapβ2, since they are still partially localized to the nucleus despite their defective/missing PY-NLSs. The crystal structure of the Kapβ2·FUS(P525L) PY-NLS complex shows the mutant peptide making fewer contacts at the mutation site, explaining decreased affinity for Kapβ2. Biochemical analysis revealed that the truncated FUS(R495X) protein, although missing the PY-NLS, can still bind Kapβ2 and suppresses LLPS. FUS(R495X) uses its C-terminal tandem arginine-glycine-glycine regions, RGG2 and RGG3, to bind the PY-NLS binding site of Kapβ2 for nuclear localization in cells when arginine methylation is inhibited. These findings suggest the importance of the C-terminal RGG regions in nuclear import and LLPS regulation of ALS variants of FUS that carry defective PY-NLSs.

59 BASIC BIOLOGICAL SCIENCES↗

Involvement of the Streptococcus mutans PgfE and GalE 4-epimerases in protein glycosylation, carbon metabolism, and cell division

Streptococcus mutans is a key pathogen associated with dental caries and is often implicated in infective endocarditis. This organism forms robust biofilms on tooth surfaces and can use collagen-binding proteins (CBPs) to efficiently colonize collagenous substrates, including dentin and heart valves. One of the best characterized CBPs of S. mutans is Cnm, which contributes to adhesion and invasion of oral epithelial and heart endothelial cells. These virulence properties were subsequently linked to post-translational modification (PTM) of the Cnm threonine-rich repeat region by the Pgf glycosylation machinery, which consists of 4 enzymes: PgfS, PgfM1, PgfE, and PgfM2. Inactivation of the S. mutans pgf genes leads to decreased collagen binding, reduced invasion of human coronary artery endothelial cells, and attenuated virulence in the Galleria mellonella invertebrate model. The present study aimed to better understand Cnm glycosylation and characterize the predicted 4-epimerase, PgfE. Using a truncated Cnm variant containing only 2 threonine-rich repeats, mass spectrometric analysis revealed extensive glycosylation with HexNAc2. Compositional analysis, complemented with lectin blotting, identified the HexNAc2 moieties as GlcNAc and GalNAc. Comparison of PgfE with the other S. mutans 4-epimerase GalE through structural modeling, nuclear magnetic resonance, and capillary electrophoresis demonstrated that GalE is a UDP-Glc-4-epimerase, while PgfE is a GlcNAc-4-epimerase. While PgfE exclusively participates in protein O-glycosylation, we found that GalE affects galactose metabolism and cell division. Furthermore, this study further emphasizes the importance of O-linked protein glycosylation and carbohydrate metabolism in S. mutans and identifies the PTM modifications of the key CBP, Cnm.

4-epimerases↗

Inhibition mechanisms of AcrF9, AcrF8, and AcrF6 against type I-F CRISPR–Cas complex revealed by cryo-EM

Prokaryotes and viruses have fought a long battle against each other. Prokaryotes use CRISPR–Cas-mediated adaptive immunity, while conversely, viruses evolve multiple anti-CRISPR (Acr) proteins to defeat these CRISPR–Cas systems. The type I-F CRISPR–Cas system in Pseudomonas aeruginosa requires the crRNA-guided surveillance complex (Csy complex) to recognize the invading DNA. Although some Acr proteins against the Csy complex have been reported, other relevant Acr proteins still need studies to understand their mechanisms. As such, here, we obtain three structures of previously unresolved Acr proteins (AcrF9, AcrF8, and AcrF6) bound to the Csy complex using electron cryo-microscopy (cryo-EM), with resolution at 2.57 Å, 3.42 Å, and 3.15 Å, respectively. The 2.57-Å structure reveals fine details for each molecular component within the Csy complex as well as the direct and water-mediated interactions between proteins and CRISPR RNA (crRNA). Our structures also show unambiguously how these Acr proteins bind differently to the Csy complex. AcrF9 binds to key DNA-binding sites on the Csy spiral backbone. AcrF6 binds at the junction between Cas7.6f and Cas8f, which is critical for DNA duplex splitting. AcrF8 binds to a distinct position on the Csy spiral backbone and forms interactions with crRNA, which has not been seen in other Acr proteins against the Csy complex. Our structure-guided mutagenesis and biochemistry experiments further support the anti-CRISPR mechanisms of these Acr proteins. Our findings support the convergent consequence of inhibiting degradation of invading DNA by these Acr proteins, albeit with different modes of interactions with the type I-F CRISPR–Cas system.

59 BASIC BIOLOGICAL SCIENCES↗

mTOR regulates aerobic glycolysis through NEAT1 and nuclear paraspeckle-mediated mechanism in hepatocellular carcinoma

Hepatocellular Carcinoma (HCC) is a major form of liver cancer and a leading cause of cancer-related death worldwide. New insights into HCC pathobiology and mechanism of drug actions are urgently needed to improve patient outcomes. HCC undergoes metabolic reprogramming of glucose metabolism from respiration to aerobic glycolysis, a phenomenon known as the ‘Warburg Effect’ that supports rapid cancer cell growth, survival, and invasion. mTOR is known to promote Warburg Effect, but the underlying mechanism(s) remains poorly defined. The aim of this study is to understand the mechanism(s) and significance of mTOR regulation of aerobic glycolysis in HCC. We profiled mTORC1-dependent long non-coding RNAs (lncRNAs) by RNA-seq of HCC cells treated with rapamycin. Chromatin immunoprecipitation (ChIP) and luciferase reporter assays were used to explore the transcriptional regulation of NEAT1 by mTORC1. [U- 13 C]-glucose labeling and metabolomic analysis, extracellular acidification Rate (ECAR) by Seahorse XF Analyzer, and glucose uptake assay were used to investigate the role of mTOR-NEAT1-NONO signaling in the regulation of aerobic glycolysis. RNA immunoprecipitation (RIP) and NONO-binding motif scanning were performed to identify the regulatory mechanism of pre-mRNA splicing by mTOR-NEAT1. Myristoylated AKT1 (mAKT1)/NRASV 12 -driven HCC model developed by hydrodynamic transfection (HDT) was employed to explore the significance of mTOR-NEAT1 signaling in HCC tumorigenesis and mTOR-targeted therapy. mTOR regulates lncRNA transcriptome in HCC and that NEAT1 is a major mTOR transcriptional target. Interestingly, although both NEAT1_1 and NEAT1_2 are down-regulated in HCC, only NEAT1_2 is significantly correlated with poor overall survival of HCC patients. NEAT1_2 is the organizer of nuclear paraspeckles that sequester the RNA-binding proteins NONO and SFPQ. We show that upon oncogenic activation, mTORC1 suppresses NEAT1_2 expression and paraspeckle biogenesis, liberating NONO/SFPQ, which in turn, binds to U5 within the spliceosome, stimulating mRNA splicing and expression of key glycolytic enzymes. This series of actions lead to enhanced glucose transport, aerobic glycolytic flux, lactate production, and HCC growth both in vitro and in vivo. Furthermore, the paraspeckle-mediated mechanism is important for the anticancer action of US FDA-approved drugs rapamycin/temsirolimus. These findings reveal a molecular mechanism by which mTOR promotes the ‘Warburg Effect’, which is important for the metabolism and development of HCC, and anticancer response of mTOR-targeted therapy.

59 BASIC BIOLOGICAL SCIENCES↗

EF-Hand Battle Royale: Hetero-ion Complexation in Lanmodulin

The lanmodulin (LanM) protein has emerged as an effective means for rare earth element (REE) extraction and separation from complex feedstocks without the use of organic solvents. Whereas the binding of LanM to individual REEs has been well characterized, little is known about the thermodynamics of mixed metal binding complexes (i.e., heterogeneous ion complexes), which limits the ability to accurately predict separation performance for a given metal ion mixture. In this paper, we employ the law of mass action to establish a theory of perfect cooperativity for LanM-REE complexation at the two highest-affinity binding sites. The theory is then used to derive an equation that explains the nonintuitive REE binding behavior of LanM, where separation factors for binary pairs of ions vary widely based on the ratio of ions in the aqueous phase, a phenomenon that is distinct from single-ion-binding chemical chelators. We then experimentally validate this theory and perform the first quantitative characterization of LanM complexation with heterogeneous ion pairs using resin-immobilized LanM. Importantly, the resulting homogeneous and heterogeneous constants enable accurate prediction of the equilibrium state of LanM in the presence of mixtures of up to 10 REEs, confirming that the perfect cooperativity model is an accurate mechanistic description of REE complexation by LanM. We further employ the model to simulate separation performance over a range of homogeneous and heterogeneous binding constants, revealing important insights into how mixed binding differentially impacts REE separations based on the relative positioning of the ion pairs within the lanthanide series. In addition to informing REE separation process optimization, these results provide mathematical and experimental insight into competition dynamics in other ubiquitous and medically relevant, cooperative binding proteins, such as calmodulin.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structural basis for cloverleaf RNA-initiated viral genome replication

The genomes of positive-strand RNA viruses serve as a template for both protein translation and genome replication. In enteroviruses, a cloverleaf RNA structure at the 5' end of the genome functions as a switch to transition from viral translation to replication by interacting with host poly(C)-binding protein 2 (PCBP2) and the viral 3CD pro protein. We determined the structures of cloverleaf RNA from coxsackievirus and poliovirus. Cloverleaf RNA folds into an H-type four-way junction and is stabilized by a unique adenosine-cytidine-uridine (A•C-U) base triple involving the conserved pyrimidine mismatch region. The two PCBP2 binding sites are spatially proximal and are located on the opposite end from the 3CD pro binding site on cloverleaf. We determined that the A•C-U base triple restricts the flexibility of the cloverleaf stem–loops resulting in partial occlusion of the PCBP2 binding site, and elimination of the A•C-U base triple increases the binding affinity of PCBP2 to the cloverleaf RNA. Based on the cloverleaf structures and biophysical assays, we propose a new mechanistic model by which enteroviruses use the cloverleaf structure as a molecular switch to transition from viral protein translation to genome replication.

59 BASIC BIOLOGICAL SCIENCES↗

Proteomics identifies complement protein signatures in patients with alcohol-associated hepatitis

Diagnostic challenges continue to impede development of effective therapies for successful management of alcohol-associated hepatitis (AH), creating an unmet need to identify noninvasive biomarkers for AH. In murine models, complement contributes to ethanol-induced liver injury. Therefore, we hypothesized that complement proteins could be rational diagnostic/prognostic biomarkers in AH. Here, we performed a comparative analysis of data derived from human hepatic and serum proteome to identify and characterize complement protein signatures in severe AH (sAH). The quantity of multiple complement proteins was perturbed in liver and serum proteome of patients with sAH. Multiple complement proteins differentiated patients with sAH from those with alcohol cirrhosis (AC) or alcohol use disorder (AUD) and healthy controls (HCs). Serum collectin 11 and C1q binding protein were strongly associated with sAH and exhibited good discriminatory performance among patients with sAH, AC, or AUD and HCs. Furthermore, complement component receptor 1-like protein was negatively associated with pro-inflammatory cytokines. Additionally, lower serum MBL associated serine protease 1 and coagulation factor II independently predicted 90-day mortality. In summary, meta-analysis of proteomic profiles from liver and circulation revealed complement protein signatures of sAH, highlighting a complex perturbation of complement and identifying potential diagnostic and prognostic biomarkers for patients with sAH.

60 APPLIED LIFE SCIENCES↗

Mono-mix strategy enables comparative proteomics of a cross-kingdom microbial symbiosis

Cross-kingdom microbial symbioses, such as those between algae and bacteria, are key players in biogeochemical cycles. The molecular changes during initiation and establishment of symbiosis are of great interest, but quantitatively monitoring such changes can be challenging, particularly when the microorganisms differ greatly in size or are intimately associated. Here, we analyze output from label-free, data-dependent acquisition (DDA) LC-MS/MS proteomics experiments investigating the well-studied interaction between the alga Chlamydomonas reinhardtii and the heterotrophic bacterium Mesorhizobium japonicum. We found that detection of bacterial proteins decreased in coculture by 50% proteome-wide due to the abundance of algal proteins. As a result, standard differential expression analysis led to numerous false-positive reports of significantly downregulated proteins, where it was not possible to distinguish meaningful biological responses to symbiosis from artifacts of the reduced protein detection in coculture relative to monoculture. We show that data normalization alone does not eliminate the impact of altered detection on differential expression analysis of the cross-kingdom symbiosis. We assessed two additional strategies to overcome this methodological artifact inherent to DDA proteomics. In the first, we combined algal and bacterial monocultures at a relative abundance that mimicked the coculture, creating a “mono-mix” control to which the coculture could be compared. This approach enabled comparable detection of bacterial proteins in the coculture and the monoculture control. In the second strategy, we enhanced detection of lowly abundant bacterial proteins by using sample fractionation upstream of LC-MS/MS analysis. When these simple approaches were combined, they allowed for meaningful comparisons of nearly 10,000 algal proteins and over 4,000 bacterial proteins in response to symbiosis by DDA. They successfully recovered expected changes in the bacterial proteome in response to algal coculture, including upregulation of sugar-binding proteins and transporters. They also revealed novel proteomic responses to coculture that guide hypotheses about algal-bacterial interactions.

Dupuis, Sunnyjoy [University of California, Berkel↗

BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data

Abstract BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB’s positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset and plans for future growth and development.

Liu, Tiqing↗