Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein Binding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Dynamical Signatures of Thermotoga maritima Maltose-Binding Proteins Affected by Ligand Binding

Functional segregation among protein isoforms depends on the interplay of their overall structures and the molecular dynamics of these structures. Thermotoga maritima maltose-binding protein (tmMBP) isoforms show size-dependent differential binding of maltose and malto-oligomers while maintaining remarkable fold conservation. This differential behavior needs detailed characterization in native-like aqueous conditions to understand the effects of protein dynamics on ligand binding and recognition. Small-angle neutron scattering (SANS), neutron spin echo (NSE) spectroscopy, and dynamic light scattering (DLS) were used in conjunction with previously published computational molecular dynamics (MD) simulations to understand the dynamic behavior of tmMBPs experimentally. SANS provided information on the overall structure of the molecules, while NSE was used to determine the dynamics in the nanosecond time scale. Both tmMBP2 and tmMBP3 have a bidomain architecture linked with a flexible hinge, with the binding pocket sitting in the cleft between the two domains. tmMBP2 and tmMBP3 showed different solution dynamics, with the translational and rotational components dominating the dynamics of both systems, resulting in a clear differentiation of their diffusion pattern. A faster dynamics component was also observed and was attributed to segmental dynamics. Differences observed between the ligand-free (apo) and ligand-bound (holo) states of the two proteins are attributed to conformational entropy. Our results highlight the intricacies of how structure and dynamics can together shape binding to a repertoire of substrates in structurally similar proteins.

Diffusion

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Functional role of myosin-binding protein H in thick filaments of developing vertebrate fast-twitch skeletal muscle

Myosin-binding protein H (MyBP-H) is a component of the vertebrate skeletal muscle sarcomere with sequence and domain homology to myosin-binding protein C (MyBP-C). Whereas skeletal muscle isoforms of MyBP-C (fMyBP-C, sMyBP-C) modulate muscle contractility via interactions with actin thin filaments and myosin motors within the muscle sarcomere “C-zone,” MyBP-H has no known function. This is in part due to MyBP-H having limited expression in adult fast-twitch muscle and no known involvement in muscle disease. Quantitative proteomics reported here reveal that MyBP-H is highly expressed in prenatal rat fast-twitch muscles and larval zebrafish, suggesting a conserved role in muscle development and prompting studies to define its function. We take advantage of the genetic control of the zebrafish model and a combination of structural, functional, and biophysical techniques to interrogate the role of MyBP-H. Transgenic, FLAG-tagged MyBP-H or fMyBP-C both localize to the C-zones in larval myofibers, whereas genetic depletion of endogenous MyBP-H or fMyBP-C leads to increased accumulation of the other, suggesting competition for C-zone binding sites. Does MyBP-H modulate contractility in the C-zone? Globular domains critical to MyBP-C’s modulatory functions are absent from MyBP-H, suggesting that MyBP-H may be functionally silent. However, our results suggest an active role. In vitro motility experiments indicate MyBP-H shares MyBP-C’s capacity as a molecular “brake.” These results provide new insights and raise questions about the role of the C-zone during muscle development.

59 BASIC BIOLOGICAL SCIENCES

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES

The RNA-binding protein Modulo promotes neural stem cell maintenance in Drosophila

A small population of stem cells in the developing Drosophila central nervous system generates the large number of different cell types that make up the adult brain. To achieve this, these neural stem cells (neuroblasts, NBs) divide asymmetrically to produce non-identical daughter cells. The balance between stem cell self-renewal and neural differentiation is regulated by various cellular machinery, including transcription factors, chromatin remodelers, and RNA-binding proteins. The list of these components remains incomplete, and the mechanisms regulating their function are not fully understood, however. Here, we identify a role for the RNA-binding protein Modulo (Mod; nucleolin in humans) in NB maintenance. We employ transcriptomic analyses to identify RNA targets of Mod and assess changes in global gene expression following its knockdown, results of which suggest a link with notable proneural genes and those essential for neurogenesis. Mod is expressed in larval brains and its loss leads to a significant decrease in the number of central brain NBs. Stem cells that remain lack expression of key NB identity factors and exhibit cell proliferation defects. Mechanistically, our analysis suggests these deficiencies arise at least in part from altered cell cycle progression, with a proportion of NBs arresting prior to mitosis. Overall, our data show that Mod function is essential for neural stem cell maintenance during neurogenesis.

Parra, Amalia S.

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES

SEC ‐ SAXS / MC Ensemble Structural Studies of the Microtubule Binding Protein Cdt1 Show Monomeric, Folded‐Over Conformations

ABSTRACT Cdt1 is a mixed folded protein critical for DNA replication licensing and it also has a “moonlighting” role at the kinetochore via direct binding to microtubules and the Ndc80 complex. However, it is unknown how the structure and conformations of Cdt1 could allow it to participate in these multiple, unique sets of protein complexes. While robust methods exist to study entirely folded or unfolded proteins, structure–function studies of combined, mixed folded/disordered proteins remain challenging. In this work, we employ orthogonal biophysical and computational techniques to provide structural characterization of mitosis‐competent human Cdt1. Thermal stability analyses shows that both folded winged helix domains1 are unstable. CD and NMR show that the N‐terminal and linker regions are intrinsically disordered. DLS shows that Cdt1 is monomeric and polydisperse, while SEC‐MALS confirms that it is monomeric at high concentrations, but without any apparent inter‐molecular self‐association. SEC‐SAXS enabled computational modeling of the protein structures. Using the program SASSIE, we performed rigid body Monte Carlo simulations to generate a conformational ensemble of structures. We observe that neither fully extended nor extremely compact Cdt1 conformations are consistent with SAXS. The best‐fit models have the N‐terminal and linker disordered regions extended into the solution and the two folded domains close to each other in apparent “folded over” conformations. We hypothesize the best‐fit Cdt1 conformations could be consistent with a function as a scaffold protein that may be sterically blocked without binding partners. Our study also provides a template for combining experimental and computational techniques to study mixed‐folded proteins.

Cell Biology

A minimal complex of KHNYN and zinc-finger antiviral protein binds and degrades single-stranded RNA

Detecting viral infection is a key role of the innate immune system. The genomes of some RNA viruses have a high CpG dinucleotide content relative to most vertebrate cell RNAs, making CpGs a molecular marker of infection. The human zinc-finger antiviral protein (ZAP) recognizes CpG, mediates clearance of the foreign CpG-rich RNA, and causes attenuation of CpG-rich RNA viruses. While ZAP binds RNA, it lacks enzymatic activity that might be responsible for RNA degradation and thus requires interacting cofactors for its function. One of these cofactors, KHNYN, has a predicted nuclease domain. Using biochemical approaches, we found that the KHNYN NYN domain is a single-stranded RNA ribonuclease that does not have sequence specificity and digests RNA with or without CpG dinucleotides equivalently in vitro. We show that unlike most KH domains, the KHNYN KH domain does not bind RNA. Indeed, a crystal structure of the KH region revealed a double-KH domain with a negatively charged surface that accounts for the lack of RNA binding. Rather, the KHNYN C-terminal domain (CTD) interacts with the ZAP RNA-binding domain (RBD) to provide target RNA specificity. We define a minimal complex composed of the ZAP RBD and the KHNYN NYN-CTD and use a fluorescence polarization assay to propose a model for how this complex interacts with a CpG dinucleotide-containing RNA. In the context of the cell, this module would represent the minimum ZAP and KHNYN domains required for CpG-recognition and ribonuclease activity essential for attenuation of viruses with clusters of CpG dinucleotides.

Yeoh, Zoe C. (ORCID:0000000226949068)

In silico λ-dynamics predicts protein binding specificities to modified RNAs

Abstract RNA modifications shape gene expression through a variety of chemical changes to canonical RNA bases. Although numbering in the hundreds, only a few RNA modifications are well characterized, in part due to the absence of methods to identify modification sites. Antibodies remain a common tool to identify modified RNA and infer modification sites through straightforward applications. However, specificity issues can result in off-target binding and confound conclusions. This work utilizes in silico λ-dynamics to efficiently estimate binding free energy differences of modification-targeting antibodies between a variety of naturally occurring RNA modifications. Crystal structures of inosine and N6-methyladenosine (m6A) targeting antibodies bound to their modified ribonucleosides were determined and served as structural starting points. λ-Dynamics was utilized to predict RNA modifications that permit or inhibit binding to these antibodies. In vitro RNA-antibody binding assays supported the accuracy of these in silico results. High agreement between experimental and computed binding propensities demonstrated that λ-dynamics can serve as a predictive screen for antibody specificity against libraries of RNA modifications. More importantly, this strategy is an innovative way to elucidate how hundreds of known RNA modifications interact with biological molecules without the limitations imposed by in vitro or in vivo methodologies.

Biochemistry & Molecular Biology

De Novo Design of Proteins That Bind Naphthalenediimides, Powerful Photooxidants with Tunable Photophysical Properties

De novo protein design provides a framework to test our understanding of protein function and build proteins with cofactors and functions not found in nature. Here, we report the design of proteins designed to bind powerful photooxidants and the evaluation of the use of these proteins to generate diffusible small-molecule reactive species. Because excited-state dynamics are influenced by the dynamics and hydration of a photooxidant’s environment, it was important to not only design a binding site but also to evaluate its dynamic properties. Thus, we used computational design in conjunction with molecular dynamics (MD) simulations to design a protein, designated NBP (NDI Binding Protein), that held a naphthalenediimide (NDI), a powerful photooxidant, in a programmable molecular environment. Solution NMR confirmed the structure of the complex. We evaluated two NDI cofactors in this de novo protein using ultrafast pump–probe spectroscopy to evaluate light-triggered intra- and intermolecular electron transfer function. Moreover, we demonstrated the utility of this platform to activate multiple molecular probes for protein labeling.

carbonyls

CAML: Commutative Algebra Machine Learning─A Case Study on Protein–Ligand Binding Affinity Prediction

Recently, Suwayyid and Wei introduced commutative algebra as an emerging paradigm for machine learning and data science. In this work, we propose commutative algebra machine learning (CAML) for the prediction of protein−ligand binding affinities. Specifically, we apply persistent Stanley−Reisner theory, a key concept in combinatorial commutative algebra, to the affinity predictions of protein−ligand binding and metalloprotein−ligand binding. We present three new algorithms, i.e., element-specific commutative algebra, category-specific commutative algebra, and commutative algebra on bipartite complexes, to tackle the complexity of data involved in (metallo) protein−ligand complexes. We show that the proposed CAML outperforms other state-of-theart methods in (metallo) protein−ligand binding affinity predictions, indicating the great potential of commutative algebra learning.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Expression, purification, and characterization of diacylated Lipo-YcjN from Escherichia coli

YcjN is a putative substrate binding protein expressed from a cluster of genes involved in carbohydrate import and metabolism in Escherichia coli. Here, we determine the crystal structure of YcjN to a resolution of 1.95 Å, revealing that its three-dimensional structure is similar to substrate binding proteins in subcluster D-I, which includes the well-characterized maltose binding protein. Furthermore, we found that recombinant overexpression of YcjN results in the formation of a lipidated form of YcjN that is posttranslationally diacylated at cysteine 21. Comparisons of size-exclusion chromatography profiles and dynamic light scattering measurements of lipidated and nonlipidated YcjN proteins suggest that lipidated YcjN aggregates in solution via its lipid moiety. Additionally, bioinformatic analysis indicates that YcjN-like proteins may exist in both Bacteria and Archaea, potentially in both lipidated and nonlipidated forms. Together, our results provide a better understanding of the aggregation properties of recombinantly expressed bacterial lipoproteins in solution and establish a foundation for future studies that aim to elucidate the role of these proteins in bacterial physiology.

Escherichia coli

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING

QM Investigation of Rare Earth Ion Interactions with First Hydration Shell Waters and Protein-Based Coordination Models

Here, conventional methods for extracting rare earth metals (REMs) from mined mineral ores are inefficient, expensive, and environmentally damaging. Recent discovery of lanmodulin (LanM), a protein that coordinates REMs with high-affinity and selectivity over competing ions, provides inspiration for new REM refinement methods. Here, we used quantum mechanical (QM) methods to investigate trivalent lanthanide cation (Ln 3+ ) interactions with coordination systems representing bulk solvent water and protein binding sites. Energy decomposition analysis (EDA) showed differences in the energetic components of Ln 3+ interaction with representatives of solvent (water, H 2 O) and protein binding sites (acetate, CH 3 COO – ), highlighting the importance of accurate description of electrostatics and polarization in computational modeling of REM interactions with biological and bioinspired molecules. Relative binding free energies were obtained for Ln 3+ with coordination complexes originating from binding sites in PDB structures of a lanthanum binding peptide (PDB entry 7CCO) and LanM, with explicit consideration of the first hydration shell waters, according to quasi-chemical theory (QCT). Beyond the first shell, the bulk solvent environment was represented with an implicit continuum model. Ln 3+ interactions with (H 2 O) 9 and both binding site models became more favorable, moving down the periodic series. This trend was more pronounced with the protein binding site models than with water, resulting in affinity increasing with periodic number, except for the last REM, Lu 3+ , which bound less favorably than the preceding element, Yb 3+ . Using the truncated 7CCO binding site model, the magnitude and trend of the experimental Ln 3+ relative binding free energies for the whole 7CCO peptide were reproduced. Conversely, the previously reported experimental data for LanM show a preference for the earlier lanthanides; this is likely due to longer-range interactions and cooperative effects, which are not represented by the reduced models. Using the truncated 7CCO binding site model, the magnitude and trend of the experimental Ln 3+ relative binding free energies for the whole 7CCO peptide were reproduced. In contrast to the previously reported experimental data for LanM, the peptide preferentially binds the earlier lanthanides. This difference likely arises due to longer-range interactions and cooperative effects not represented by the peptide. Further investigation of Ln 3+ interactions with whole proteins using polarizable molecular mechanics models with explicit solvent is warranted to understand the influence of longer-ranged interactions, cooperativity, and bulk solvent. Nevertheless, the present work provides new insights into Ln 3+ interactions with biomolecules and presents an effective computational platform for designing specific single-site REM binding peptides more efficiently.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Quick-and-Easy Validation of Protein–Ligand Binding Models Using Fragment-Based Semiempirical Quantum Chemistry

Electronic structure calculations in enzymes converge very slowly with respect to the size of the model region that is described using quantum mechanics (QM), requiring hundreds of atoms to obtain converged results and exhibiting substantial sensitivity (at least in smaller models) to which amino acids are included in the QM region. As such, there is considerable interest in developing automated procedures to construct a QM model region based on well-defined criteria. However, testing such procedures is burdensome due to the cost of large-scale electronic structure calculations. Here, we show that semiempirical methods can be used as alternatives to density functional theory (DFT) to assess convergence in sequences of models generated by various automated protocols. The cost of these convergence tests is reduced even further by means of a many-body expansion. We use this approach to examine convergence (with respect to model size) of protein–ligand binding energies. Fragment-based semiempirical calculations afford well-converged interaction energies in a tiny fraction of the cost required for DFT calculations. Two-body interactions between the ligand and single-residue amino acid fragments afford a low-cost way to construct a “QM-informed” enzyme model of reduced size, furnishing an automatable active-site model-building procedure. This provides a streamlined, user-friendly approach for constructing ligand binding-site models that needs neither a priori information nor manual adjustments. Extension to model-building for thermochemical calculations should be straightforward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Cysteine Rich Intestinal Protein 2 is a copper-responsive regulator of skeletal muscle differentiation and metal homeostasis

Copper (Cu) is essential for respiration, neurotransmitter synthesis, oxidative stress response, and transcription regulation, with imbalances leading to neurological, cognitive, and muscular disorders. Here we show the role of a novel Cu-binding protein (Cu-BP) in mammalian transcriptional regulation, specifically on skeletal muscle differentiation using murine primary myoblasts. Utilizing synchrotron X-ray fluorescence-mass spectrometry, we identified murine cysteine-rich intestinal protein 2 (mCrip2) as a key Cu-BP abundant in both nuclear and cytosolic fractions. mCrip2 binds two to four Cu + ions with high affinity and presents limited redox potential. CRISPR/Cas9-mediated deletion of mCrip2 impaired myogenesis, likely due to Cu accumulation in cells. CUT&RUN and transcriptome analyses revealed its association with gene promoters, including MyoD1 and metallothioneins, suggesting a novel Cu-responsive regulatory role for mCrip2. Our work describes the significance of mCrip2 in skeletal muscle differentiation and metal homeostasis, expanding understanding of the Cu-network in myoblasts. Copper (Cu) is essential for various cellular processes, including respiration and stress response, but imbalances can cause serious health issues. This study reveals a new Cu-binding protein (Cu-BP) involved in muscle development in primary myoblasts. Using unbiased metalloproteomic techniques and high throughput sequencing, we identified mCrip2 as a key Cu-BP found in cell nuclei and cytoplasm. mCrip2 binds up to four Cu + ions and has a limited redox potential. Deleting mCrip2 using CRISPR/Cas9 disrupted muscle formation due to Cu accumulation. Further analyses showed that mCrip2 regulates the expression of genes like MyoD1, essential for muscle differentiation, and metallothioneins in response to copper supplementation. This research highlights the importance of mCrip2 in muscle development and metal homeostasis, providing new insights into the Cu-network in cells.

59 BASIC BIOLOGICAL SCIENCES