Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

Artificial intelligence methods for protein structure and interaction prediction: Recent advances and challenges

Recent advances in artificial intelligence have introduced novel methods for high-accuracy prediction of protein tertiary structures, protein complex structures, and interactions between proteins and other biomolecules, such as small molecules and nucleic acids. Such advancements are accelerating biomedical research and the development of new protein design and bioengineering methods among many other important biotechnology applications. Here, in this review, we outline the recent advances in protein-centric biomolecular structure and interaction prediction, highlight some major challenges in the field, and discuss potential directions to address them.

Morehead, Alex [Lawrence Berkeley National Laborat↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Deep Green Unannotated Protein Structures

The Deep Green list is based on the identification and curation of conserved unannotated proteins in three green lineage (Viridiplantae) model organisms; Arabidopsis thaliana, Chlamydomonas reinhardtii, and Setaria viridis. Preliminary characterization of Deep Green proteins and genes was done using various informatics tools and published data sets and is presented in Knoshaug, Sun, et al., 2023, submitted. The structures of these unannotated proteins were also predicted using AlphaFold (Jumper et al., 2021). The data deposited here are the AlphaFold structural predictions having the highest pLDDT score and thus identified as the best folded structure (ranked_0). These data enable others to do in-depth structural characterizations to aid in functional characterization leading to deeper understanding of plant biology. References: Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P. and Hassabis, D. (2021) Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. Knoshaug, E. P., Sun, P., Nag, A., Nguyen, H., Mattoon, E. M., Zhang, N., Liu, J., Chen, C., Cheng, J., Zhang, R., St. John, P., and Umen, J. (submitted) Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii, Arabidopsis thaliana, and Setaria viridis.

09 BIOMASS FUELS↗

Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15

Abstract Estimating the accuracy of quaternary structural models of protein complexes and assemblies (EMA) is important for predicting quaternary structures and applying them to studying protein function and interaction. The pairwise similarity between structural models is proven useful for estimating the quality of protein tertiary structural models, but it has been rarely applied to predicting the quality of quaternary structural models. Moreover, the pairwise similarity approach often fails when many structural models are of low quality and similar to each other. To address the gap, we developed a hybrid method (MULTICOM_qa) combining a pairwise similarity score (PSS) and an interface contact probability score (ICPS) based on the deep learning inter‐chain contact prediction for estimating protein complex model accuracy. It blindly participated in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 and performed very well in estimating the global structure accuracy of assembly models. The average per‐target correlation coefficient between the model quality scores predicted by MULTICOM_qa and the true quality scores of the models of CASP15 assembly targets is 0.66. The average per‐target ranking loss in using the predicted quality scores to rank the models is 0.14. It was able to select good models for most targets. Moreover, several key factors (i.e., target difficulty, model sampling difficulty, skewness of model quality, and similarity between good/bad models) for EMA are identified and analyzed. The results demonstrate that combining the multi‐model method (PSS) with the complementary single‐model method (ICPS) is a promising approach to EMA.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaFold Protein Structure Database for Sequence-Independent Molecular Replacement

Crystallographic phasing recovers the phase information that is lost during a diffraction experiment. Molecular replacement is a commonly used phasing method for crystal structures in the protein data bank. In one form it uses a protein sequence to search a structure database to find suitable templates for phasing. However, sequence information is not always available, such as when proteins are crystallized with unknown binding partner proteins or when the crystal is of a contaminant. The recent development of AlphaFold published the predicted protein structures for every protein from twenty distinct species. In this work, we tested whether AlphaFold-predicted E. coli protein structures were accurate enough to enable sequence-independent phasing of diffraction data from two crystallization contaminants of unknown sequence. Using each of more than 4000 predicted structures as a search model, robust molecular replacement solutions were obtained, which allowed the identification and structure determination of YncE and YadF. Our results demonstrate the general utility of the AlphaFold-predicted structure database with respect to sequence-independent crystallographic phasing.

59 BASIC BIOLOGICAL SCIENCES↗

Protein model quality assessment using rotation–equivariant transformations on point clouds

Machine learning research concerning protein structure has seen a surge in popularity over the last years with promising advances for basic science and drug discovery. Working with macromolecular structure in a machine learning context requires an adequate numerical representation, and researchers have extensively studied representations such as graphs, discretized 3D grids, and distance maps. As part of CASP14, we explored a new and conceptually simple representation in a blind experiment: atoms as points in 3D, each with associated features. These features—initially just the basic element type of each atom—are updated through a series of neural network layers featuring rotation-equivariant convolutions. Starting from all atoms, we further aggregate information at the level of alpha carbons before making a prediction at the level of the entire protein structure. We find that this approach yields competitive results in protein model quality assessment despite its simplicity and despite the fact that it incorporates minimal prior information and is trained on relatively little data. As a result, its performance and generality are particularly noteworthy in an era where highly complex, customized machine learning methods such as AlphaFold 2 have come to dominate protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

The crystal structure of Grindelia robusta 7,13-copalyl diphosphate synthase reveals active site features controlling catalytic specificity

Diterpenoid natural products serve critical functions in plant development and ecological adaptation and many diterpenoids have economic value as bioproducts. The family of class II diterpene synthases catalyzes the committed reactions in diterpenoid biosynthesis, converting a common geranylgeranyl diphosphate precursor into different bicyclic prenyl diphosphate scaffolds. Enzymatic rearrangement and modification of these precursors generate the diversity of bioactive diterpenoids. We report the crystal structure of Grindelia robusta 7,13-copalyl diphosphate synthase, GrTPS2, at 2.1 Å of resolution. GrTPS2 catalyzes the committed reaction in the biosynthesis of grindelic acid, which represents the signature metabolite in species of gumweed (Grindelia spp., Asteraceae). Grindelic acid has been explored as a potential source for drug leads and biofuel production. The GrTPS2 crystal structure adopts the conserved three-domain fold of class II diterpene synthases featuring a functional active site in the γβ-domain and a vestigial α-domain. Substrate docking into the active site of the GrTPS2 apo protein structure predicted catalytic amino acids. Biochemical characterization of protein variants identified residues with impact on enzyme activity and catalytic specificity. Specifically, mutagenesis of Y457 provided mechanistic insight into the position-specific deprotonation of the intermediary carbocation to form the characteristic 7,13 double bond of 7,13-copalyl diphosphate.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust deep learning–based protein sequence design using ProteinMPNN

Although deep learning has revolutionized protein structure prediction, almost all experimentally characterized de novo protein designs have been generated using physically based approaches such as Rosetta. Here, we describe a deep learning–based protein sequence design method, ProteinMPNN, that has outstanding performance in both in silico and experimental tests. On native protein backbones, ProteinMPNN has a sequence recovery of 52.4% compared with 32.9% for Rosetta. The amino acid sequence at different positions can be coupled between single or multiple chains, enabling application to a wide range of current protein design challenges. We demonstrate the broad utility and high accuracy of ProteinMPNN using x-ray crystallography, cryo–electron microscopy, and functional studies by rescuing previously failed designs, which were made using Rosetta or AlphaFold, of protein monomers, cyclic homo-oligomers, tetrahedral nanoparticles, and target-binding proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Enzymes in 3D: Synthesis, remodelling, and hydrolysis of cell wall (1,3;1,4)-β-glucans

Abstract Recent breakthroughs in structural biology have provided valuable new insights into enzymes involved in plant cell wall metabolism. More specifically, the molecular mechanism of synthesis of (1,3;1,4)-β-glucans, which are widespread in cell walls of commercially important cereals and grasses, has been the topic of debate and intense research activity for decades. However, an inability to purify these integral membrane enzymes or apply transgenic approaches without interpretative problems associated with pleiotropic effects has presented barriers to attempts to define their synthetic mechanisms. Following the demonstration that some members of the CslF sub-family of GT2 family enzymes mediate (1,3;1,4)-β-glucan synthesis, the expression of the corresponding genes in a heterologous system that is free of background complications has now been achieved. Biochemical analyses of the (1,3;1,4)-β-glucan synthesized in vitro, combined with 3-dimensional (3D) cryogenic-electron microscopy and AlphaFold protein structure predictions, have demonstrated how a single CslF6 enzyme, without exogenous primers, can incorporate both (1,3)- and (1,4)-β-linkages into the nascent polysaccharide chain. Similarly, 3D structures of xyloglucan endo-transglycosylases and (1,3;1,4)-β-glucan endo- and exohydrolases have allowed the mechanisms of (1,3;1,4)-β-glucan modification and degradation to be defined. X-ray crystallography and multi-scale modeling of a broad specificity GH3 β-glucan exohydrolase recently revealed a previously unknown and remarkable molecular mechanism with reactant trajectories through which a polysaccharide exohydrolase can act with a processive action pattern. The availability of high-quality protein 3D structural predictions should prove invaluable for defining structures, dynamics, and functions of other enzymes involved in plant cell wall metabolism in the immediate future.

59 BASIC BIOLOGICAL SCIENCES↗

Distance‐based reconstruction of protein quaternary structures from inter‐chain contacts

Abstract Predicting the quaternary structure of protein complex is an important problem. Inter‐chain residue‐residue contact prediction can provide useful information to guide the ab initio reconstruction of quaternary structures. However, few methods have been developed to build quaternary structures from predicted inter‐chain contacts. Here, we develop the first method based on gradient descent optimization (GD) to build quaternary structures of protein dimers utilizing inter‐chain contacts as distance restraints. We evaluate GD on several datasets of homodimers and heterodimers using true/predicted contacts and monomer structures as input. GD consistently performs better than both simulated annealing and Markov Chain Monte Carlo simulation. Starting from an arbitrarily quaternary structure randomly initialized from the tertiary structures of protein chains and using true inter‐chain contacts as input, GD can reconstruct high‐quality structural models for homodimers and heterodimers with average TM‐score ranging from 0.92 to 0.99 and average interface root mean square distance from 0.72 Å to 1.64 Å. On a dataset of 115 homodimers, using predicted inter‐chain contacts as restraints, the average TM‐score of the structural models built by GD is 0.76. For 46% of the homodimers, high‐quality structural models with TM‐score ≥ 0.9 are reconstructed from predicted contacts. There is a strong correlation between the quality of the reconstructed models and the precision and recall of predicted contacts. Only a moderate precision or recall of inter‐chain contact prediction is needed to build good structural models for most homodimers. Moreover, GD improves the quality of quaternary structures predicted by AlphaFold2 on a Critical Assessment of Techniques for Protein Structure Prediction–Critical Assessments of Predictions of Interactions dataset.

59 BASIC BIOLOGICAL SCIENCES↗

OpenMDlr: parallel, open-source tools for general protein structure modeling and refinement from pairwise distances

Easy-to-use, open-source, general-purpose programs for modeling a protein structure from inter-atomic distances are needed for modeling from experimental data and refinement of predicted protein structures. OpenMDlr is an open-source Python package for modeling protein structures from pairwise distances between any atoms, and optionally, dihedral angles. Finally, we provide a user-friendly input format for harnessing modern biomolecular force fields in an easy-to-install package that can efficiently make use of multiple compute cores.

59 BASIC BIOLOGICAL SCIENCES↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes to Structure and Function Workshop Report 2022

The goal of the U.S. Department of Energy (DOE) Biological and Environmental Research (BER) Program is to achieve a predictive understanding of complex biological, earth, and environmental systems with the aim of advancing the nation’s energy and infrastructure security. (https://www.energy.gov/science/ ber/biological-and-environmental-research). To pursue this goal, collaborations among experts in diverse research areas that lead to multidisciplinary projects are indispensable. The roles of DOE’s User Facilities, which offer unique and powerful resources for such research projects, are evolving, and expectations for the facilities are increasing. To respond to Users’ needs, the Joint Genome Institute (JGI) and Environmental Molecular Sciences Laboratory (EMSL) initiated the Facilities Integrating Collaborations for User Science (FICUS) program in 2014. This collaboration has grown into a popular and successful program, advancing more than 100 multidisciplinary projects to date. Similarly, the new interFacility collaborations among the JGI, EMSL, and User resources for BER structural biology and imaging at the Basic Energy Science (BES) Program’s synchrotron and neutron facilities are becoming essential for cutting-edge transdisciplinary science. To further explore the need for the BER research community to combine genomic, functional, and structural approaches to advance their research, an organizing committee was formed to develop and jointly host a 3-part workshop. The committee’s members represented seven DOE National Laboratory User Facilities (Appendix 1 lists the members). The “Genomes to Structure and Function” virtual workshop (see Appendices 2–5) was composed of three sessions. The first session, titled “Molecular Structures” (October 27– 28, 2021), highlighted diverse integrative experimental and computational approaches correlating structural data with sequencing and functional information, as well as predicting protein structures to model complex biological systems. The second session, “Intracellular Organization, and Material Synthesis and Decomposition” (December 15–16, 2021), covered imaging methods for observing, quantifying, and manipulating biosystems. The third session, “Imaging the Rhizosphere and Cellular Organization” (January 26–27, 2022) emphasized advanced and non-invasive imaging techniques applied to plant root-microbe-soil interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Development of a Systematic and Extensible Force Field for Peptoids (STEPs)

Peptoids (N-substituted glycines) are a class of biomimetic polymers that have attracted significant attention due to their accessible synthesis and enzymatic and thermal stability relative to their naturally occurring counterparts (polypeptides). While these polymers provide the promise of more robust functional materials via hierarchical approaches, they present a new challenge for computational structure prediction for material design. The reliability of calculations hinges on the accuracy of interactions represented in the force field used to model peptoids. For proteins, structure prediction based on sequence and de novo design has made dramatic progress in recent years; however, these models are not readily transferable for peptoids. Current efforts to develop and implement peptoid-specific force fields are spread out, leading to replicated efforts and a fragmented collection of parameterized sidechains. Here, we developed a peptoid-specific force field containing 70 different side chains, using GAFF2 as starting point. The new model is validated based on the generation of Ramachandran-like plots from DFT optimization compared against force field reproduced potential energy and free energy surfaces as well as the reproduction of equilibrium cis/trans values for some residues experimentally known to form helical structures. In conclusion, equilibrium cis/trans distributions (Kct) are estimated for all parameterized residues to identify which residues have an intrinsic propensity for cis or trans states in the monomeric state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Geometry-complete perceptron networks for 3D molecular graphs

Abstract Motivation The field of geometric deep learning has recently had a profound impact on several scientific domains such as protein structure prediction and design, leading to methodological advancements within and outside of the realm of traditional machine learning. Within this spirit, in this work, we introduce GCPNet, a new chirality-aware SE(3)-equivariant graph neural network designed for representation learning of 3D biomolecular graphs. We show that GCPNet, unlike previous representation learning methods for 3D biomolecules, is widely applicable to a variety of invariant or equivariant node-level, edge-level, and graph-level tasks on biomolecular structures while being able to (1) learn important chiral properties of 3D molecules and (2) detect external force fields. Results Across four distinct molecular-geometric tasks, we demonstrate that GCPNet’s predictions (1) for protein–ligand binding affinity achieve a statistically significant correlation of 0.608, more than 5%, greater than current state-of-the-art methods; (2) for protein structure ranking achieve statistically significant target-local and dataset-global correlations of 0.616 and 0.871, respectively; (3) for Newtownian many-body systems modeling achieve a task-averaged mean squared error less than 0.01, more than 15% better than current methods; and (4) for molecular chirality recognition achieve a state-of-the-art prediction accuracy of 98.7%, better than any other machine learning method to date. Availability and implementation The source code, data, and instructions to train new models or reproduce our results are freely available at https://github.com/BioinfoMachineLearning/GCPNet.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and mechanism of human vesicular polyamine transporter

Polyamines play essential roles in gene expression and modulate neuronal transmission in mammals. Vesicular polyamine transporters (VPAT) from the SLC18 family exploit the transmembrane H + gradient to translocate polyamines into secretory vesicles, enabling the quantal release of polyamine neuromodulators and underpinning learning and memory formation. Here, we report the cryo-electron microscopy structures of human VPAT in complex with spermine, spermidine, H + , or tetrabenazine, elucidating discrete lumen-facing states of the antiporter and pivotal interactions between VPAT and its substrate or inhibitor. Leveraging structure-inspired mutagenesis studies and protein structure prediction, we deduce an unforeseen mechanism whereby the polyamine and H + compete for multiple acidic protein residues both directly and indirectly, and rationalize how the antidopaminergic therapeutic tetrabenazine impedes vesicular transport of polyamines. This study unravels the mechanism of an H + -coupled polyamine antiporter, reveals mechanistic diversity between VPAT and other SLC18 antiporters, and raises new prospects for combating human disorders of polyamine homeostasis.

59 BASIC BIOLOGICAL SCIENCES↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗