Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein complex structure prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Structural Heterogeneity and Hydrodynamics of an Intrinsically Disordered Protein Condensate

Biology demonstrates precise control over the free-energy landscape through the selective partitioning of biomacromolecules into membraneless organelles, enabling essential functions such as biochemical transformations, signaling cascades, and mechanical reinforcement. Although the function of these condensates depends on their underlying structure and hydrodynamics, molecular-scale information on these systems remains sparse. Here, in this study, neutron scattering is used to probe the organization and dynamics of the intrinsically disordered N-terminal domain of Galectin-3, an extracellular lectin responsible for facilitating liquid–liquid phase separation on the cellular surface, in both dilute and condensed phases. Dilute solutions contain isolated protein chains in equilibrium with mesoscopic clusters, whereas the condensed phase adopts a bicontinuous, microemulsion-like morphology. The dilute phase behavior is quantitatively described by coarse-grained polymer models from soft-matter physics, demonstrating their predictive power for complex biological proteins. At elevated concentrations, the proteins self-assemble akin to block copolymers, microphase separating through the aggregation of hydrophobic domains along the protein contour. The resulting condensate remains fluid-like despite a 25-fold increase in concentration; its internal hydrodynamics slow by only a factor of 3 relative to dilute protein chains. These results provide a molecular-level framework for how disordered proteins achieve both the structural complexity and dynamic fluidity of biomolecular condensates.

Carrick, Brian R. [Massachusetts Inst. of Technolo↗

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran↗

Sas20 is a highly flexible starch-binding protein in the Ruminococcus bromii cell-surface amylosome

Ruminococcus bromii is a keystone species in the human gut that has the rare ability to degrade dietary resistant starch (RS). This bacterium secretes a suite of starch-active proteins that work together within larger complexes called amylosomes that allow R. bromii to bind and degrade RS. Starch adherence system protein 20 (Sas20) is one of the more abundant proteins assembled within amylosomes, but little could be predicted about its molecular features based on amino acid sequence. Here, we performed a structure–function analysis of Sas20 and determined that it features two discrete starch-binding domains separated by a flexible linker. We show that Sas20 domain 1 contains an N-terminal β-sandwich followed by a cluster of α-helices, and the nonreducing end of maltooligosaccharides can be captured between these structural features. Furthermore, the crystal structure of a close homolog of Sas20 domain 2 revealed a unique bilobed starch-binding groove that targets the helical α1,4-linked glycan chains found in amorphous regions of amylopectin and crystalline regions of amylose. Affinity PAGE and isothermal titration calorimetry demonstrated that both domains bind maltoheptaose and soluble starch with relatively high affinity (K d ≤ 20 μM) but exhibit limited or no binding to cyclodextrins. Finally, small-angle X-ray scattering analysis of the individual and combined domains support that these structures are highly flexible, which may allow the protein to adopt conformations that enhance its starch-targeting efficiency. Taken together, we conclude that Sas20 binds distinct features within the starch granule, facilitating the ability of R. bromii to hydrolyze dietary RS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bivalent recognition of fatty acyl-CoA by a human integral membrane palmitoyltransferase

S-acylation, also known as palmitoylation, is the most abundant form of protein lipidation in humans. This reversible posttranslational modification, which targets thousands of proteins, is catalyzed by 23 members of the DHHC family of integral membrane enzymes. DHHC enzymes use fatty acyl-CoA as the ubiquitous fatty acyl donor and become autoacylated at a catalytic cysteine; this intermediate subsequently transfers the fatty acyl group to a cysteine in the target protein. Protein S-acylation intersects with almost all areas of human physiology, and several DHHC enzymes are considered as possible therapeutic targets against diseases such as cancer. These efforts would greatly benefit from a detailed understanding of the molecular basis for this crucial enzymatic reaction. Here, in this study, we combine X-ray crystallography with all-atom molecular dynamics simulations to elucidate the structure of the precatalytic complex of human DHHC20 in complex with palmitoyl CoA. The resulting structure reveals that the fatty acyl chain inserts into a hydrophobic pocket within the transmembrane spanning region of the protein, whereas the CoA headgroup is recognized by the cytosolic domain through polar and ionic interactions. Biochemical experiments corroborate the predictions from our structural model. We show, using both computational and experimental analyses, that palmitoyl CoA acts as a bivalent ligand where the interaction of the DHHC enzyme with both the fatty acyl chain and the CoA headgroup is important for catalytic chemistry to proceed. This bivalency explains how, in the presence of high concentrations of free CoA under physiological conditions, DHHC enzymes can efficiently use palmitoyl CoA as a substrate for autoacylation.

59 BASIC BIOLOGICAL SCIENCES↗

Dual domain recognition determines SARS-CoV-2 PLpro selectivity for human ISG15 and K48-linked di-ubiquitin

The Papain-like protease (PLpro) is a domain of a multi-functional, non-structural protein 3 of coronaviruses. PLpro cleaves viral polyproteins and posttranslational conjugates with poly-ubiquitin and protective ISG15, composed of two ubiquitin-like (UBL) domains. Across coronaviruses, PLpro showed divergent selectivity for recognition and cleavage of posttranslational conjugates despite sequence conservation. We show that SARS-CoV-2 PLpro binds human ISG15 and K48-linked di-ubiquitin (K48-Ub 2 ) with nanomolar affinity and detect alternate weaker-binding modes. Crystal structures of untethered PLpro complexes with ISG15 and K48-Ub 2 combined with solution NMR and cross-linking mass spectrometry revealed how the two domains of ISG15 or K48-Ub 2 are differently utilized in interactions with PLpro. Analysis of protein interface energetics predicted differential binding stabilities of the two UBL/Ub domains that were validated experimentally. We emphasize how substrate recognition can be tuned to cleave specifically ISG15 or K48-Ub 2 modifications while retaining capacity to cleave mono-Ub conjugates. These results highlight alternative druggable surfaces that would inhibit PLpro function.

60 APPLIED LIFE SCIENCES↗

Accurate prediction of protein structures and interactions using a three-track neural network

DeepMind presented notably accurate predictions at the recent 14th Critical Assessment of Structure Prediction (CASP14) conference. We explored network architectures that incorporate related ideas and obtained the best performance with a three-track network in which information at the one-dimensional (1D) sequence level, the 2D distance map level, and the 3D coordinate level is successively transformed and integrated. The three-track network produces structure predictions with accuracies approaching those of DeepMind in CASP14, enables the rapid solution of challenging x-ray crystallography and cryo–electron microscopy structure modeling problems, and provides insights into the functions of proteins of currently unknown structure. The network also enables rapid generation of accurate protein-protein complex models from sequence information alone, short-circuiting traditional approaches that require modeling of individual subunits followed by docking. Here, we make the method available to the scientific community to speed biological research.

59 BASIC BIOLOGICAL SCIENCES↗

Transient local secondary structure in the intrinsically disordered C‐term of the Albino3 insertase

Albino3 (Alb3) is an integral membrane protein fundamental to the targeting and insertion of light‐harvesting complex (LHC) proteins into the thylakoid membrane. Alb3 contains a stroma‐exposed C‐terminus (Alb3‐Cterm) that is responsible for binding the LHC‐loaded transit complex before LHC membrane insertion. Alb3‐Cterm has been reported to be intrinsically disordered, but precise mechanistic details underlying how it recognizes and binds to the transit complex are lacking, and the functional roles of its four different motifs have been debated. Using a novel combination of experimental and computational techniques such as single‐molecule fluorescence resonance energy transfer, circular dichroism with deconvolution analysis, site‐directed mutagenesis, trypsin digestion assays, and all‐atom molecular dynamics simulations in conjunction with enhanced sampling techniques, we show that Alb3‐Cterm contains transient secondary structure in motifs I and II. The excellent agreement between the experimental and computational data provides a quantitatively consistent picture and allows us to identify a heterogeneous structural ensemble that highlights the local and transient nature of the secondary structure. This structural ensemble was used to predict both the inter‐residue distance distributions of single molecules and the apparent unfolding free energy of the transient secondary structure, which were both in excellent agreement with those determined experimentally. We hypothesize that this transient local secondary structure may play an important role in the recognition of Alb3‐Cterm for the LHC‐loaded transit complex, and these results should provide a framework to better understand protein targeting by the Alb3‐Oxa1‐YidC family of insertases.

Okoto, Patience S.↗

A Clustering-based biased Monte Carlo Approach to Protein Titration Curve Prediction

We develop and implement a novel approach to computing the ensemble averages in systems characterized by pair-wise interactions between the entities. Methods involving full enumeration of the configuration space result in exponential complexity. Sampling methods such as Markov Chain Monte Carlo (MCMC) algorithms have been proposed to tackle the exponential complexity of these problems. In certain scenarios where significant energetic coupling exists between the entities, the accuracy of the such algorithms can be diminished. We propose a strategy to improve the accuracy of the MCMC runs by taking advantage of the cluster structure in the interaction energy matrix. We propose two different schemes for performing the biased MCMC runs on the partitioned systems and show that they are valid MCMC schemes. We then apply these algorithms to the problem of computing the protonation fractions and hence the titration curves of titratable protein residues that constitute a given protein. We leverage both synthesized and real-world systems and show the improved performance of our biased MCMC methods when compared to the regular MCMC method.

Visweswara Sathanur, Arun↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying amyloid-related diseases by mapping mutations in low-complexity protein domains to pathologies

Proteins including FUS, hnRNPA2, and TDP-43 reversibly aggregate into amyloid-like fibrils through interactions of their low-complexity domains (LCDs). Mutations in LCDs can promote irreversible amyloid aggregation and disease. We introduce a computational approach to identify mutations in LCDs of disease-associated proteins predicted to increase propensity for amyloid aggregation. We identify several disease-related mutations in the intermediate filament protein keratin-8 (KRT8). Atomic structures of wild-type and mutant KRT8 segments confirm the transition to a pleated strand capable of amyloid formation. Biochemical analysis reveals KRT8 forms amyloid aggregates, and the identified mutations promote aggregation. Aggregated KRT8 is found in Mallory–Denk bodies, observed in hepatocytes of livers with alcoholic steatohepatitis (ASH). We demonstrate that ethanol promotes KRT8 aggregation, and KRT8 amyloids co-crystallize with alcohol. Lastly, KRT8 aggregation can be seeded by liver extract from people with ASH, consistent with the amyloid nature of KRT8 aggregates and the classification of ASH as an amyloid-related condition.

59 BASIC BIOLOGICAL SCIENCES↗

Cyanobacterial Phycobilisome Allostery as Revealed by Quantitative Mass Spectrometry

Phycobilisomes (PBSs) are the major photosynthetic light-harvesting complexes in cyanobacteria and red algae. PBS, a multisubunit protein complex, has two major interfaces that comprise intrinsically disordered regions (IDRs): rod–core and core-membrane. IDRs do not form regular, three-dimensional structures on their own. Their presence in the photosynthetic pigment–protein complexes portends their structural and functional importance. A recent model suggests that PB-loop, an IDR located on the PBS subunit ApcE and C-terminal extension (CTE) of the PBS subunit ApcG, forms a structural protrusion on the PBS core-membrane side, facing the thylakoid membrane. Here, the structural synergy between the rod–core region and the core-membrane region was investigated using quantitative mass spectrometry (MS). The AlphaFold-predicted CpcG-CTE structure was first modeled onto the PBS rod–core region, guided and justified by the isotopically encoded structural MS data. Quantitative cross-linking MS analysis revealed that the structural proximity of the PB-loop in ApcE and ApcG-CTE is significantly disturbed in the absence of six PBS rods, which are attached to PBS via CpcG-CTE, indicative of drastic conformational changes and decreased structural integrity. Furthermore, these results suggest that CpcG-rod attachment on the PBS rod–core side is essentially required for the PBS core-membrane structural assembly. The hypothesized long-range synergy between the rod–core interface (where the orange carotenoid protein also functions) and the terminal energy emitter of PBS must have important regulatory roles in PBS core assembly, light-harvesting, and excitation energy transmission. These data also lend strategies that genetic truncation of the light-harvesting antennas aimed for improved photosynthetic productivity must rely on an in-depth understanding of their global structural integrity.

59 BASIC BIOLOGICAL SCIENCES↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Structural basis for the calmodulin-mediated activation of eukaryotic elongation factor 2 kinase

Translation is a tightly regulated process that ensures optimal protein quality and enables adaptation to energy/nutrient availability. The α-kinase eukaryotic elongation factor 2 kinase (eEF-2K), a key regulator of translation, specifically phosphorylates the guanosine triphosphatase eEF-2, thereby reducing its affinity for the ribosome and suppressing the elongation phase of protein synthesis. eEF-2K activation requires calmodulin binding and autophosphorylation at the primary stimulatory site, T348. Biochemical studies predict a calmodulin-mediated activation mechanism for eEF-2K distinct from other calmodulin-dependent kinases. Here, we resolve the atomic details of this mechanism through a 2.3-Å crystal structure of the heterodimeric complex of calmodulin and the functional core of eEF-2K (eEF-2K TR ). This structure, which represents the activated T348-phosphorylated state of eEF-2K TR , highlights an intimate association of the kinase with the calmodulin C-lobe, creating an “activation spine” that connects its amino-terminal calmodulin-targeting motif to its active site through a conserved regulatory element.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the Rapid Calculation of Binding Affinities for Antigen and Antibody Design and Affinity Maturation Simulations

The accurate and efficient calculation of protein-protein binding affinities is an essential component in antibody and antigen design and optimization, and in computer modeling of antibody affinity maturation. Such calculations remain challenging despite advances in computer hardware and algorithms, primarily because proteins are flexible molecules, and thus, require explicit or implicit incorporation of multiple conformational states into the computational procedure. The astronomical size of the amino acid sequence space further compounds the challenge by requiring predictions to be computed within a short time so that many sequence variants can be tested. In this study, we compare three classes of methods for antibody/antigen (Ab/Ag) binding affinity calculations: (i) a method that relies on the physical separation of the Ab/Ag complex in equilibrium molecular dynamics (MD) simulations, (ii) a collection of 18 scoring functions that act on an ensemble of structures created using homology modeling software, and (iii) methods based on the molecular mechanics-generalized Born surface area (MM-GBSA) energy decomposition, in which the individual contributions of the energy terms are scaled to optimize agreement with the experiment. When applied to a set of 49 antibody mutations in two Ab/HIV gp120 complexes, all of the methods are found to have modest accuracy, with the highest Pearson correlations reaching about 0.6. In particular, the most computationally intensive method, i.e., MD simulation, did not outperform several scoring functions. The optimized energy decomposition methods provided marginally higher accuracy, but at the expense of requiring experimental data for parametrization. Within each method class, we examined the effect of the number of independent computational replicates, i.e., modeled structures or reinitialized MD simulations, on the prediction accuracy. We suggest using about ten modeled structures for scoring methods, and about five simulation replicates for MD simulations as a rule of thumb for obtaining reasonable convergence. We anticipate that our study will be a useful resource for practitioners working to incorporate binding affinity calculations within their protein design and optimization process.

59 BASIC BIOLOGICAL SCIENCES↗

Computational Study on Full-length Human Ku70 with Double Stranded DNA: Dynamics, Interactions and Functional Implications

The Ku70/80 heterodimer is the first repair protein in the initial binding of double-strand break (DSB) ends following DNA damage, and is a component of nonhomologous end joining repair, the primary pathway for DSB repair in mammalian cells. In this study we constructed a full-length human Ku70 structure based on its crystal structure, and performed 20 ns conventional molecular dynamic (CMD) simulations on this protein and several other complexes with short DNA duplexes of different sequences. The trajectories of these simulations indicated that, without the topological support of Ku80, the residues in the bridge and C-terminal arm of Ku70 are more flexible than other experimentally identified domains. We studied the two missing loops in the crystal structure and predicted that they are also very flexible. Simulations revealed that they make an important contribution to the Ku70 interaction with DNA. Dislocation of the previously studied SAP domain was observed in several systems, implying its role in DNA binding. Targeted molecular dynamic (TMD) simulation was also performed for one system with a far-away 14bp DNA duplex. The TMD trajectory and energetic analysis disclosed detailed interactions of the DNA-binding residues during the DNA dislocation, and revealed a possible conformational transition for a DSB end when encountering Ku70 in solution. Compared to experimentally based analysis, this study identified more detailed interactions between DNA and Ku70. Free energy analysis indicated Ku70 alone is able to bind DNA with relatively high affinity, with consistent contributions from various domains of Ku70 in different systems. The functional implications of these domains in the processes of Ku heterodimerization and DNA damage recognition and repair can be characterized in detail based upon this analysis.

Hu, Shaowen↗

Characterization of alternate encounter assemblies of SARS-CoV-2 main protease

The assembly of two monomeric constructs spanning segments 1-199 (MPro 1-199 ) and 10-306 (MPro 10-306 ) of SARS-CoV-2 main protease (MPro) was examined to assess the existence of a transient heterodimer intermediate in the N-terminal autoprocessing pathway of MPro model precursor. Together, they form a heterodimer population accompanied by a 13-fold increase in catalytic activity. Addition of inhibitor GC373 to the proteins increases the activity further by ~7-fold with a 1:1 complex and higher order assemblies approaching 1:2 and 2:2 molecules of MPro 1-199 and MPro 10-306 detectable by analytical ultracentrifugation and native mass estimation by light scattering. Assemblies larger than a heterodimer (1:1) are discussed in terms of alternate pathways of domain III association, either through switching the location of helix 201 to 214 onto a second helical domain of MPro 10-306 and vice versa or direct interdomain III contacts like that of the native dimer, based on known structures and AlphaFold 3 prediction, respectively. At a constant concentration of MPro 1-199 with molar excess of GC373, the rate of substrate hydrolysis displays first order dependency on the MPro 10-306 concentration and vice versa. An equimolar composition of the two proteins with excess GC373 exhibits half-maximal activity at ~6 μM MPro 1-199 . Catalytic activity arises primarily from MPro 1-199 and is dependent on the interface interactions involving the N-finger residues 1 to 9 of MPro 1-199 and E290 of MPro 10-306 . Importantly, our results confirm that a single N-finger region with its associated intersubunit contacts is sufficient to form a heterodimeric MPro intermediate with enhanced catalytic activity.

60 APPLIED LIFE SCIENCES↗

PigmentHunter: A point-and-click application for automated chlorophyll-protein simulations

Chlorophyll proteins (CPs) are the workhorses of biological photosynthesis, working together to absorb solar energy, transfer it to chemically active reaction centers, and control the charge-separation process that drives its storage as chemical energy. Yet predicting CP optical and electronic properties remains a serious challenge, driven by the computational difficulty of treating large, electronically coupled molecular pigments embedded in a dynamically structured protein environment. To address this challenge, we introduce here an analysis tool called PigmentHunter, which automates the process of preparing CP structures for molecular dynamics (MD), running short MD simulations on the nanoHUB.org science gateway, and then using electrostatic and steric analysis routines to predict optical absorption, fluorescence, and circular dichroism spectra within a Frenkel exciton model. Inter-pigment couplings are evaluated using point-dipole or transition-charge coupling models, while site energies can be estimated using both electrostatic and ring-deformation approaches. The package is built in a Jupyter Notebook environment, with a point-and-click interface that can be used either to manually prepare individual structures or to batch-process many structures at once. Here, we illustrate PigmentHunter’s capabilities with example simulations on spectral line shapes in the light harvesting 2 complex, site energies in the Fenna–Matthews–Olson protein, and ring deformation in photosystems I and II.

14 SOLAR ENERGY↗

Cyclic Peptides for Lanthanide Binding

Lanthanide ions are difficult to separate from one another due to their similar chemical properties. The discovery of lanthanide-binding peptides and proteins in nature has led to an increased interest in the possibility of utilizing the strong binding of peptides to lanthanide ions for their separations; as such, there has been an effort to identify or design peptides with improved lanthanide binding and selectivity toward particular lanthanide ions. Here, in this study, we designed and characterized lanthanide-binding cyclic peptides (LBCPs) with molecular dynamics simulations, electronic structure calculations, and emission spectroscopy. Luminescent decay measurements were done to determine the number of water molecules coordinated to the Eu 3+ ion in Eu-LBCP complexes and compare to the predicted number of water molecules by computation to assess the lanthanide-binding affinity of LBCPs. Measured stability constants show binding of the LBCPs to the Eu 3+ ion with stronger than micromolar affinity. We were able to identify multiple peptides that selectively bind to middle lanthanides. We describe the structural basis of the lanthanide-binding selectivity trend with strongest binding to the middle lanthanides, followed by the heavier lanthanides, and finally to the lighter ions.

ions↗