Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Comparative Analysis of TCR and TCR-pMHC Complex Structure Prediction Tools

The rapid development of computational approaches for predicting the structures of T cell receptors (TCRs) and TCR-peptide-major histocompatibility (TCR-pMHC) complexes, accelerated by AI breakthroughs such as AlphaFold, has made it feasible to calculate these structures with increasing accuracy. Although these tools show great potential, their relative accuracy and limitations remain unclear due to the lack of standardized benchmarks. Here, we systematically evaluate seven tools for predicting isolated TCR structures together with six tools for predicting TCR-pMHC complex structures. The methods include homology-based approaches, general prediction tools using AlphaFold, TCR-specific tools derived from AlphaFold2, and the newly developed tFold-TCR model. The evaluation uses a post-training data set comprising 40 αβ TCRs and 27 TCR-pMHC complexes (21 Class I and 6 Class II). Model accuracy is assessed at global, local, and interface levels using a variety of metrics. We find that each tool offers distinct advantages in various aspects of its predictions. AlphaFold2, AlphaFold3, and tFold-TCR excel in overall accuracy of TCR structure prediction, and TCRmodel2 and AlphaFold2 perform well in overall accuracy of TCR-pMHC structure prediction. However, TCR-specific tools derived from AlphaFold2 show lower accuracy in the framework region than both homology-based methods and general-purpose tools such as AlphaFold, and challenges remain for all in modeling CDR3 loops, docking orientations, TCR-peptide interfaces, and Class II MHC-peptide interfaces. Furthermore, these findings will guide researchers in selecting appropriate tools, emphasize the importance of using multiple evaluation metrics to assess model performance, and offer suggestions for improving TCR and TCR-pMHC structure prediction tools.

Chemical structure↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

An extended motif in the SARS-CoV-2 spike modulates binding and release of host coatomer in retrograde trafficking

β-Coronaviruses such as SARS-CoV-2 hijack coatomer protein-I (COPI) for spike protein retrograde trafficking to the progeny assembly site in endoplasmic reticulum-Golgi intermediate compartment (ERGIC). However, limited residue-level details are available into how the spike interacts with COPI. Here we identify an extended COPI binding motif in the spike that encompasses the canonical K-x-H dibasic sequence. This motif demonstrates selectivity for αCOPI subunit. Guided by an in silico analysis of dibasic motifs in the human proteome, we employ mutagenesis and binding assays to show that the spike motif terminal residues are critical modulators of complex dissociation, which is essential for spike release in ERGIC. αCOPI residues critical for spike motif binding are elucidated by mutagenesis and crystallography and found to be conserved in the zoonotic reservoirs, bats, pangolins, camels, and in humans. Collectively, our investigation on the spike motif identifies key COPI binding determinants with implications for retrograde trafficking.

59 BASIC BIOLOGICAL SCIENCES↗

OpenMDlr Dataset

Protein structure dataset created with OpenMDlr. Restraints applied to fold proteins into native-like conformations were obtained from a variety of methods; results from these different restraint sets are contained in respective subdirectories. Supplements OpenMDlr: Parallel, open-source tools for general protein structure modeling and refinement from pairwise distances (https://doi.org/10.1093/bioinformatics/btac307).

59 BASIC BIOLOGICAL SCIENCES↗

Structural models and functional annotations for the Sphagnum divinum proteome

This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models of the Rhodopseudomonas palustris Proteome

This dataset contains the structural models for the primary transcripts of the Rhodopseudomonas palustris proteome. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. palustris proteome to those available in the AlphaFold Protein Structure Database.

59 BASIC BIOLOGICAL SCIENCES↗

Toho-1 β-lactamase: backbone chemical shift assignments and changes in dynamics upon binding with avibactam

Backbone chemical shift assignments for the Toho-1 β-lactamase (263 amino acids, 28.9 kDa) are reported based on triple resonance solution-state NMR experiments performed on a uniformly 2 H, 13 C, 15 N-labeled sample. These assignments allow for subsequent site-specific characterization at the chemical, structural, and dynamical levels. At the chemical level, titration with the non-β-lactam β-lactamase inhibitor avibactam is found to give chemical shift perturbations indicative of tight covalent binding that allow for mapping of the inhibitor binding site. At the structural level, protein secondary structure is predicted based on the backbone chemical shifts and protein residue sequence using TALOS-N and found to agree well with structural characterization from X-ray crystallography. At the dynamical level, model-free analysis of 15 N relaxation data at a single field of 16.4 T reveals well-ordered structures for the ligand-free and avibactam-bound enzymes with generalized order parameters of ~0.85. Complementary relaxation dispersion experiments indicate that there is an escalation in motions on the millisecond timescale in the vicinity of the active site upon substrate binding. The combination of high rigidity on short timescales and active site flexibility on longer timescales is consistent with hypotheses for achieving both high catalytic efficiency and broad substrate specificity: the induced active site dynamics allows variously sized substrates to be accommodated and increases the probability that the optimal conformation for catalysis will be sampled.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Proteome-scale Structure Prediction Data - Pseudodesulfovibrio mercurii

The number of proteins predicted for Pseudodesulfovibrio mercurii is 3,446, each of which have five predicted structures from an AlphaFold run, as well as structural alignment results using the TMscore-based structural alignment method within the APoc program. Specifically, AlphaFold outputs the atoms and coordinates of the protein model in human-readable PDB files and quantitative prediction metrics in Python PICKLE files. The 5 models have been ranked based on the predicted TM-score (pTMS), a quantitative confidence metric output by AlphaFold that reports on protein model quality. The top ranked model has undergone an energy minimization calculation to relax and remove any potential clashes in the atomic coordinates. Structural alignment results are stored in two files for each protein; the top ranked model (as discussed above) is used for all alignment analyses. Both are compressed gzip files that, once unpacked, are human readable. The first file is the TMalign score results and contains the quantitative metrics for the top alignments between the predicted structure and experimental structures from the PDB70, a curated non-redundant database of about 80,000 experimental structures developed by the Soding lab. Each data point in this file is directly associated with one experimental structure; PDB ID and brief meta-data about the protein taken from the PDB70 file are reported alongside the quantitative metrics. The second results file contains the raw results associated with each alignment reported in the score results file. Specifically, the translation and rotation arrays for each alignment are provided so that the structural alignment can be recreated. Additionally, residue-level scores are reported to quantify the closeness of the aligned residues between the predicted and experimental models.

59 BASIC BIOLOGICAL SCIENCES↗

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Structural determination of a full-length plant cellulose synthase informed by experimental and in silico methods

Three-dimensional structure determination and prediction of proteins with intrinsically disordered regions, unstructured regions, conformational flexibility, and lacking homologous structures are challenging. We previously predicted and refined an in silico structure of a plant cellulose synthase from cotton (GhCESA1), and more recently, cryo-electron microscopy (cryo-EM) has resolved a majority of the lengths of two CESA structures from poplar (PttCESA8) and cotton (GhCESA7). However, 26–30% of these cryo-EM structures remain unresolved, including the N-terminal domain, half of the class-specific region, the gating loop region, and the C-terminal domain. Here, we describe the generation and evaluation of a full-length hybrid GhCESA1 model based on this cryo-EM PttCESA8 structure, with unresolved regions completed using this in silico refined GhCESA1 model. All-atom molecular dynamics simulations and subsequent energy minimizations were performed for the in silico and hybrid GhCESA1 models in a lipid bilayer-water-ion environment, and structural stability, dynamics, energetics, contacts, and quality were evaluated. The unresolved regions were found to be the most dynamic, in agreement with their poor electron density with cryo-EM. The hybrid model exhibited a higher total secondary structure content, more favorable intra-protein and protein-lipid interaction energies, and improved quality metrics. Moreover, hydrogen bonding was revealed to be a primary mechanism for intra-protein and protein-lipid contacts. These results demonstrate that in silico structure prediction and refinement may be useful to augment experimental structure determination, especially for disordered and unstructured regions. Furthermore, this hybrid model can serve as a steppingstone to derive full-length homology models of other CESAs found in more experimentally tractable organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Computational Prediction of Coiled–Coil Protein Gelation Dynamics and Structure

Protein hydrogels represent an important and growing biomaterial for a multitude of applications, including diagnostics and drug delivery. We have previously explored the ability to engineer the thermoresponsive supramolecular assembly of coiled–coil proteins into hydrogels with varying gelation properties, where we have defined important parameters in the coiled–coil hydrogel design. Using Rosetta energy scores and Poisson–Boltzmann electrostatic energies, we iterate a computational design strategy to predict the gelation of coiled–coil proteins while simultaneously exploring five new coiled–coil protein hydrogel sequences. Provided this library, we explore the impact of in silico energies on structure and gelation kinetics, where we also reveal a range of blue autofluorescence that enables hydrogel disassembly and recovery. As a result of this library, we identify the new coiled–coil hydrogel sequence, Q5, capable of gelation within 24 h at 4 °C, a more than 2-fold increase over that of our previous iteration Q2. The fast gelation time of Q5 enables the assessment of structural transition in real time using small-angle X-ray scattering (SAXS) that is correlated to coarse-grained and atomistic molecular dynamics simulations revealing the supramolecular assembling behavior of coiled–coils toward nanofiber assembly and gelation. This work represents the first system of hydrogels with predictable self-assembly, autofluorescent capability, and a molecular model of coiled–coil fiber formation.

36 MATERIALS SCIENCE↗