Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AlphaFold”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Phenix‐AlphaFold webservice: Enabling AlphaFold predictions for use in Phenix

Abstract Advances in machine learning have enabled sufficiently accurate predictions of protein structure to be used in macromolecular structure determination with crystallography and cryo‐electron microscopy data. The Phenix software suite has AlphaFold predictions integrated into an automated pipeline that can start with an amino acid sequence and data, and automatically perform model‐building and refinement to return a protein model fitted into the data. Due to the steep technical requirements of running AlphaFold efficiently, we have implemented a Phenix‐AlphaFold webservice that enables all Phenix users to run AlphaFold predictions remotely from the Phenix GUI starting with the official 1.21 release. This webservice will be improved based on how it is used by the research community and the future research directions for Phenix.

Poon, Billy K.↗

AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination

Abstract Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing alphafold-multimer-based protein complex structure prediction with MULTICOM in CASP15

To enhance the AlphaFold-Multimer-based protein complex structure prediction, we developed a quaternary structure prediction system (MULTICOM) to improve the input fed to AlphaFold-Multimer and evaluate and refine its outputs. MULTICOM samples diverse multiple sequence alignments (MSAs) and templates for AlphaFold-Multimer to generate structural predictions by using both traditional sequence alignments and Foldseek-based structure alignments, ranks structural predictions through multiple complementary metrics, and refines the structural predictions via a Foldseek structure alignment-based refinement method. The MULTICOM system with different implementations was blindly tested in the assembly structure prediction in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 as both server and human predictors. MULTICOM_qa ranked 3 rd among 26 CASP15 server predictors and MULTICOM_human ranked 7 th among 87 CASP15 server and human predictors. The average TM-score of the first predictions submitted by MULTICOM_qa for CASP15 assembly targets is ~0.76, 5.3% higher than ~0.72 of the standard AlphaFold-Multimer. The average TM-score of the best of top 5 predictions submitted by MULTICOM_qa is ~0.80, about 8% higher than ~0.74 of the standard AlphaFold-Multimer. Moreover, the Foldseek Structure Alignment-based Multimer structure Generation (FSAMG) method outperforms the widely used sequence alignment-based multimer structure generation.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating AlphaFold and deep learning for atomistic interpretation of cryo-EM maps

Abstract Interpretation of cryo-electron microscopy (cryo-EM) maps requires building and fitting 3D atomic models of biological molecules. AlphaFold-predicted models generate initial 3D coordinates; however, model inaccuracy and conformational heterogeneity often necessitate labor-intensive manual model building and fitting into cryo-EM maps. In this work, we designed a protein model-building workflow, which combines a deep-learning cryo-EM map feature enhancement tool, CryoFEM (Cryo-EM Feature Enhancement Model) and AlphaFold. A benchmark test using 36 cryo-EM maps shows that CryoFEM achieves state-of-the-art performance in optimizing the Fourier Shell Correlations between the maps and the ground truth models. Furthermore, in a subset of 17 datasets where the initial AlphaFold predictions are less accurate, the workflow significantly improves their model accuracy. Our work demonstrates that the integration of modern deep learning image enhancement and AlphaFold may lead to automated model building and fitting for the atomistic interpretation of cryo-EM maps.

59 BASIC BIOLOGICAL SCIENCES↗

Putting AlphaFold models to work with phenix.process_predicted_model and ISOLDE

AlphaFold has recently become an important tool in providing models for experimental structure determination by X-ray crystallography and cryo-EM. Large parts of the predicted models typically approach the accuracy of experimentally determined structures, although there are frequently local errors and errors in the relative orientations of domains. Importantly, residues in the model of a protein predicted by AlphaFold are tagged with a predicted local distance difference test score, informing users about which regions of the structure are predicted with less confidence. AlphaFold also produces a predicted aligned error matrix indicating its confidence in the relative positions of each pair of residues in the predicted model. The phenix.process_predicted_model tool downweights or removes low-confidence residues and can break a model into confidently predicted domains in preparation for molecular replacement or cryo-EM docking. These confidence metrics are further used in ISOLDE to weight torsion and atom–atom distance restraints, allowing the complete AlphaFold model to be interactively rearranged to match the docked fragments and reducing the need for the rebuilding of connecting regions.

59 BASIC BIOLOGICAL SCIENCES↗

Improved AlphaFold modeling with implicit experimental information

Machine-learning prediction algorithms such as AlphaFold and RoseTTAFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including new experimental information such as a density map, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt on the basis of experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We show that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for interpretation of crystallographic and electron cryo-microscopy maps.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating crystal structure determination with iterative AlphaFold prediction

Experimental structure determination can be accelerated with artificial intelligence (AI)-based structure-prediction methods such as AlphaFold . Here, an automatic procedure requiring only sequence information and crystallographic data is presented that uses AlphaFold predictions to produce an electron-density map and a structural model. Iterating through cycles of structure prediction is a key element of this procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. This procedure was applied to X-ray data for 215 structures released by the Protein Data Bank in a recent six-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2 Å. Predictions from the iterative template-guided prediction procedure were more accurate than those obtained without templates. It is concluded that AlphaFold predictions obtained based on sequence information alone are usually accurate enough to solve the crystallographic phase problem with molecular replacement, and a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization is suggested.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of cytoplasmic ring of nuclear pore complex by integrative cryo-EM and AlphaFold

The nuclear pore complex (NPC) is the conduit for bidirectional cargo traffic between the cytoplasm and the nucleus. We determined a near-complete structure of the cytoplasmic ring of the NPC from Xenopus oocytes using single-particle cryo–electron microscopy and AlphaFold prediction. Structures of nucleoporins were predicted with AlphaFold and fit into the medium-resolution map by using the prominent secondary structural density as a guide. Certain molecular interactions were further built or confirmed by complex prediction by using AlphaFold. Here, we identified the binding modes of five copies of Nup358, the largest NPC subunit with Phe-Gly repeats for cargo transport, and predicted it to contain a coiled-coil domain that may provide avidity to assist its role as a nucleation center for NPC formation under certain conditions

59 BASIC BIOLOGICAL SCIENCES↗

AlphaFold and Structural Mass Spectrometry Enable Interrogations on the Intrinsically Disordered Regions in Cyanobacterial Light-harvesting Complex Phycobilisome

Intrinsically disordered proteins/regions (IDPRs) are a very large and functionally important class of proteins that participate in weak multivalent interactions in protein complexes. They are recalcitrant for interrogations using X-ray crystallography and cryo-EM. The IDPRs observed at the interface of the photosynthetic pigment protein complexes (PPCs) remain much less clear, e.g., the major cyanobacterial light-harvesting complex (PBS) contains an unstructured PB-loop insertion in the phycocyanobilin domain (PB domain) of ApcE (the largest polypeptide in PBS). Here, a joint platform is built to probe such structural domains. This platform is characterized by two-round progressive justifications of in silico models by using the structural mass spectrometry data. First, the AlphaFold-generated 3D structure of the PB domain (containing PB-loop) was justified in the context of PBS. Second, docking the AlphaFold-generated ApcG (a ligand) into the first-step justified structure (a receptor). The final ligand-receptor complex was then subjected to a second-round justification, again, by using unequivocal isotopically-encoded cross-links identified in LC-MS/MS. This work reveals a full-length PB-loop structure modelled in the PBS basal cylinder, free from any spatial conflicts against the other subunits in PBS. The structure of PB domain highlights the close associations of the intrinsically disordered PB-loop with its binding partners in PBS, including ApcG, another IDPR. The PB-loop region involved in the binding of photosystem II (PSII) is also discussed in the context of excitation energy transfer regulation. Finally, this work calls attention to the highly disordered, yet interrogatable interface between the light-harvesting antenna complexes and the reaction centers.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaFold -assisted structure determination of a bacterial protein of unknown function using X-ray and electron crystallography

Macromolecular crystallography generally requires the recovery of missing phase information from diffraction data to reconstruct an electron-density map of the crystallized molecule. Most recent structures have been solved using molecular replacement as a phasing method, requiring an a priori structure that is closely related to the target protein to serve as a search model; when no such search model exists, molecular replacement is not possible. New advances in computational machine-learning methods, however, have resulted in major advances in protein structure predictions from sequence information. Methods that generate predicted structural models of sufficient accuracy provide a powerful approach to molecular replacement. Taking advantage of these advances, AlphaFold predictions were applied to enable structure determination of a bacterial protein of unknown function (UniProtKB Q63NT7, NCBI locus BPSS0212) based on diffraction data that had evaded phasing attempts using MIR and anomalous scattering methods. Using both X-ray and micro-electron (microED) diffraction data, it was possible to solve the structure of the main fragment of the protein using a predicted model of that domain as a starting point. The use of predicted structural models importantly expands the promise of electron diffraction, where structure determination relies critically on molecular replacement.

molecular replacement↗

Allosteric autoregulation of DNA binding via a DNA-mimicking protein domain: a biophysical study of ZNF410–DNA interaction using small angle X-ray scattering

ZNF410 is a highly-conserved transcription factor, remarkable in that it recognizes a 15-base pair DNA element but has just a single responsive target gene in mammalian erythroid cells. ZNF410 includes a tandem array of five zinc-fingers (ZFs), surrounded by uncharacterized N- and C-terminal regions. Unexpectedly, full-length ZNF410 has reduced DNA binding affinity, compared to that of the isolated DNA binding ZF array, both in vitro and in cells. AlphaFold predicts a partially-folded N-terminal subdomain that includes a 30-residue long helix, preceded by a hairpin loop rich in acidic (aspartate/glutamate) and serine/threonine residues. This hairpin loop is predicted by AlphaFold to lie against the DNA binding interface of the ZF array. In solution, ZNF410 is a monomer and binds to DNA with 1:1 stoichiometry. Surprisingly, the single best-fit model for the experimental small angle X-ray scattering profile, in the absence of DNA, is the original AlphaFold model with the N-terminal long-helix and the hairpin loop occupying the ZF DNA binding surface. For DNA binding, the hairpin loop presumably must be displaced. After combining biophysical, biochemical, bioinformatic and artificial intelligence-based AlphaFold analyses, we suggest that the hairpin loop mimics the structure and electrostatics of DNA, and provides an additional mechanism, supplementary to sequence specificity, of regulating ZNF410 DNA binding.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Proteome-scale Structure Prediction Data - Pseudodesulfovibrio mercurii

The number of proteins predicted for Pseudodesulfovibrio mercurii is 3,446, each of which have five predicted structures from an AlphaFold run, as well as structural alignment results using the TMscore-based structural alignment method within the APoc program. Specifically, AlphaFold outputs the atoms and coordinates of the protein model in human-readable PDB files and quantitative prediction metrics in Python PICKLE files. The 5 models have been ranked based on the predicted TM-score (pTMS), a quantitative confidence metric output by AlphaFold that reports on protein model quality. The top ranked model has undergone an energy minimization calculation to relax and remove any potential clashes in the atomic coordinates. Structural alignment results are stored in two files for each protein; the top ranked model (as discussed above) is used for all alignment analyses. Both are compressed gzip files that, once unpacked, are human readable. The first file is the TMalign score results and contains the quantitative metrics for the top alignments between the predicted structure and experimental structures from the PDB70, a curated non-redundant database of about 80,000 experimental structures developed by the Soding lab. Each data point in this file is directly associated with one experimental structure; PDB ID and brief meta-data about the protein taken from the PDB70 file are reported alongside the quantitative metrics. The second results file contains the raw results associated with each alignment reported in the score results file. Specifically, the translation and rotation arrays for each alignment are provided so that the structural alignment can be recreated. Additionally, residue-level scores are reported to quantify the closeness of the aligned residues between the predicted and experimental models.

59 BASIC BIOLOGICAL SCIENCES↗

AI-Based Protein Interaction Screening and Identification (AISID)

In this study, we presented an AISID method extending AlphaFold-Multimer’s success in structure prediction towards identifying specific protein interactions with an optimized AISIDscore. The method was tested to identify the binding proteins in 18 human TNFSF (Tumor Necrosis Factor superfamily) members for each of 27 human TNFRSF (TNF receptor superfamily) members. For each TNFRSF member, we ranked the AISIDscore among the 18 TNFSF members. The correct pairing resulted in the highest AISIDscore for 13 out of 24 TNFRSF members which have known interactions with TNFSF members. Out of the 33 correct pairing between TNFSF and TNFRSF members, 28 pairs could be found in the top five (including 25 pairs in the top three) seats in the AISIDscore ranking. Surprisingly, the specific interactions between TNFSF10 (TNF-related apoptosis-inducing ligand, TRAIL) and its decoy receptors DcR1 and DcR2 gave the highest AISIDscore in the list, while the structures of DcR1 and DcR2 are unknown. The data strongly suggests that AlphaFold-Multimer might be a useful computational screening tool to find novel specific protein bindings. This AISID method may have broad applications in protein biochemistry, extending the application of AlphaFold far beyond structure predictions.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Green Unannotated Protein Structures

The Deep Green list is based on the identification and curation of conserved unannotated proteins in three green lineage (Viridiplantae) model organisms; Arabidopsis thaliana, Chlamydomonas reinhardtii, and Setaria viridis. Preliminary characterization of Deep Green proteins and genes was done using various informatics tools and published data sets and is presented in Knoshaug, Sun, et al., 2023, submitted. The structures of these unannotated proteins were also predicted using AlphaFold (Jumper et al., 2021). The data deposited here are the AlphaFold structural predictions having the highest pLDDT score and thus identified as the best folded structure (ranked_0). These data enable others to do in-depth structural characterizations to aid in functional characterization leading to deeper understanding of plant biology. References: Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P. and Hassabis, D. (2021) Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. Knoshaug, E. P., Sun, P., Nag, A., Nguyen, H., Mattoon, E. M., Zhang, N., Liu, J., Chen, C., Cheng, J., Zhang, R., St. John, P., and Umen, J. (submitted) Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii, Arabidopsis thaliana, and Setaria viridis.

09 BIOMASS FUELS↗