Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Robust deep learning–based protein sequence design using ProteinMPNN

Although deep learning has revolutionized protein structure prediction, almost all experimentally characterized de novo protein designs have been generated using physically based approaches such as Rosetta. Here, we describe a deep learning–based protein sequence design method, ProteinMPNN, that has outstanding performance in both in silico and experimental tests. On native protein backbones, ProteinMPNN has a sequence recovery of 52.4% compared with 32.9% for Rosetta. The amino acid sequence at different positions can be coupled between single or multiple chains, enabling application to a wide range of current protein design challenges. We demonstrate the broad utility and high accuracy of ProteinMPNN using x-ray crystallography, cryo–electron microscopy, and functional studies by rescuing previously failed designs, which were made using Rosetta or AlphaFold, of protein monomers, cyclic homo-oligomers, tetrahedral nanoparticles, and target-binding proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Enzymes in 3D: Synthesis, remodelling, and hydrolysis of cell wall (1,3;1,4)-β-glucans

Abstract Recent breakthroughs in structural biology have provided valuable new insights into enzymes involved in plant cell wall metabolism. More specifically, the molecular mechanism of synthesis of (1,3;1,4)-β-glucans, which are widespread in cell walls of commercially important cereals and grasses, has been the topic of debate and intense research activity for decades. However, an inability to purify these integral membrane enzymes or apply transgenic approaches without interpretative problems associated with pleiotropic effects has presented barriers to attempts to define their synthetic mechanisms. Following the demonstration that some members of the CslF sub-family of GT2 family enzymes mediate (1,3;1,4)-β-glucan synthesis, the expression of the corresponding genes in a heterologous system that is free of background complications has now been achieved. Biochemical analyses of the (1,3;1,4)-β-glucan synthesized in vitro, combined with 3-dimensional (3D) cryogenic-electron microscopy and AlphaFold protein structure predictions, have demonstrated how a single CslF6 enzyme, without exogenous primers, can incorporate both (1,3)- and (1,4)-β-linkages into the nascent polysaccharide chain. Similarly, 3D structures of xyloglucan endo-transglycosylases and (1,3;1,4)-β-glucan endo- and exohydrolases have allowed the mechanisms of (1,3;1,4)-β-glucan modification and degradation to be defined. X-ray crystallography and multi-scale modeling of a broad specificity GH3 β-glucan exohydrolase recently revealed a previously unknown and remarkable molecular mechanism with reactant trajectories through which a polysaccharide exohydrolase can act with a processive action pattern. The availability of high-quality protein 3D structural predictions should prove invaluable for defining structures, dynamics, and functions of other enzymes involved in plant cell wall metabolism in the immediate future.

59 BASIC BIOLOGICAL SCIENCES↗

Distance‐based reconstruction of protein quaternary structures from inter‐chain contacts

Abstract Predicting the quaternary structure of protein complex is an important problem. Inter‐chain residue‐residue contact prediction can provide useful information to guide the ab initio reconstruction of quaternary structures. However, few methods have been developed to build quaternary structures from predicted inter‐chain contacts. Here, we develop the first method based on gradient descent optimization (GD) to build quaternary structures of protein dimers utilizing inter‐chain contacts as distance restraints. We evaluate GD on several datasets of homodimers and heterodimers using true/predicted contacts and monomer structures as input. GD consistently performs better than both simulated annealing and Markov Chain Monte Carlo simulation. Starting from an arbitrarily quaternary structure randomly initialized from the tertiary structures of protein chains and using true inter‐chain contacts as input, GD can reconstruct high‐quality structural models for homodimers and heterodimers with average TM‐score ranging from 0.92 to 0.99 and average interface root mean square distance from 0.72 Å to 1.64 Å. On a dataset of 115 homodimers, using predicted inter‐chain contacts as restraints, the average TM‐score of the structural models built by GD is 0.76. For 46% of the homodimers, high‐quality structural models with TM‐score ≥ 0.9 are reconstructed from predicted contacts. There is a strong correlation between the quality of the reconstructed models and the precision and recall of predicted contacts. Only a moderate precision or recall of inter‐chain contact prediction is needed to build good structural models for most homodimers. Moreover, GD improves the quality of quaternary structures predicted by AlphaFold2 on a Critical Assessment of Techniques for Protein Structure Prediction–Critical Assessments of Predictions of Interactions dataset.

59 BASIC BIOLOGICAL SCIENCES↗

OpenMDlr: parallel, open-source tools for general protein structure modeling and refinement from pairwise distances

Easy-to-use, open-source, general-purpose programs for modeling a protein structure from inter-atomic distances are needed for modeling from experimental data and refinement of predicted protein structures. OpenMDlr is an open-source Python package for modeling protein structures from pairwise distances between any atoms, and optionally, dihedral angles. Finally, we provide a user-friendly input format for harnessing modern biomolecular force fields in an easy-to-install package that can efficiently make use of multiple compute cores.

59 BASIC BIOLOGICAL SCIENCES↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes to Structure and Function Workshop Report 2022

The goal of the U.S. Department of Energy (DOE) Biological and Environmental Research (BER) Program is to achieve a predictive understanding of complex biological, earth, and environmental systems with the aim of advancing the nation’s energy and infrastructure security. (https://www.energy.gov/science/ ber/biological-and-environmental-research). To pursue this goal, collaborations among experts in diverse research areas that lead to multidisciplinary projects are indispensable. The roles of DOE’s User Facilities, which offer unique and powerful resources for such research projects, are evolving, and expectations for the facilities are increasing. To respond to Users’ needs, the Joint Genome Institute (JGI) and Environmental Molecular Sciences Laboratory (EMSL) initiated the Facilities Integrating Collaborations for User Science (FICUS) program in 2014. This collaboration has grown into a popular and successful program, advancing more than 100 multidisciplinary projects to date. Similarly, the new interFacility collaborations among the JGI, EMSL, and User resources for BER structural biology and imaging at the Basic Energy Science (BES) Program’s synchrotron and neutron facilities are becoming essential for cutting-edge transdisciplinary science. To further explore the need for the BER research community to combine genomic, functional, and structural approaches to advance their research, an organizing committee was formed to develop and jointly host a 3-part workshop. The committee’s members represented seven DOE National Laboratory User Facilities (Appendix 1 lists the members). The “Genomes to Structure and Function” virtual workshop (see Appendices 2–5) was composed of three sessions. The first session, titled “Molecular Structures” (October 27– 28, 2021), highlighted diverse integrative experimental and computational approaches correlating structural data with sequencing and functional information, as well as predicting protein structures to model complex biological systems. The second session, “Intracellular Organization, and Material Synthesis and Decomposition” (December 15–16, 2021), covered imaging methods for observing, quantifying, and manipulating biosystems. The third session, “Imaging the Rhizosphere and Cellular Organization” (January 26–27, 2022) emphasized advanced and non-invasive imaging techniques applied to plant root-microbe-soil interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Development of a Systematic and Extensible Force Field for Peptoids (STEPs)

Peptoids (N-substituted glycines) are a class of biomimetic polymers that have attracted significant attention due to their accessible synthesis and enzymatic and thermal stability relative to their naturally occurring counterparts (polypeptides). While these polymers provide the promise of more robust functional materials via hierarchical approaches, they present a new challenge for computational structure prediction for material design. The reliability of calculations hinges on the accuracy of interactions represented in the force field used to model peptoids. For proteins, structure prediction based on sequence and de novo design has made dramatic progress in recent years; however, these models are not readily transferable for peptoids. Current efforts to develop and implement peptoid-specific force fields are spread out, leading to replicated efforts and a fragmented collection of parameterized sidechains. Here, we developed a peptoid-specific force field containing 70 different side chains, using GAFF2 as starting point. The new model is validated based on the generation of Ramachandran-like plots from DFT optimization compared against force field reproduced potential energy and free energy surfaces as well as the reproduction of equilibrium cis/trans values for some residues experimentally known to form helical structures. In conclusion, equilibrium cis/trans distributions (Kct) are estimated for all parameterized residues to identify which residues have an intrinsic propensity for cis or trans states in the monomeric state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Geometry-complete perceptron networks for 3D molecular graphs

Abstract Motivation The field of geometric deep learning has recently had a profound impact on several scientific domains such as protein structure prediction and design, leading to methodological advancements within and outside of the realm of traditional machine learning. Within this spirit, in this work, we introduce GCPNet, a new chirality-aware SE(3)-equivariant graph neural network designed for representation learning of 3D biomolecular graphs. We show that GCPNet, unlike previous representation learning methods for 3D biomolecules, is widely applicable to a variety of invariant or equivariant node-level, edge-level, and graph-level tasks on biomolecular structures while being able to (1) learn important chiral properties of 3D molecules and (2) detect external force fields. Results Across four distinct molecular-geometric tasks, we demonstrate that GCPNet’s predictions (1) for protein–ligand binding affinity achieve a statistically significant correlation of 0.608, more than 5%, greater than current state-of-the-art methods; (2) for protein structure ranking achieve statistically significant target-local and dataset-global correlations of 0.616 and 0.871, respectively; (3) for Newtownian many-body systems modeling achieve a task-averaged mean squared error less than 0.01, more than 15% better than current methods; and (4) for molecular chirality recognition achieve a state-of-the-art prediction accuracy of 98.7%, better than any other machine learning method to date. Availability and implementation The source code, data, and instructions to train new models or reproduce our results are freely available at https://github.com/BioinfoMachineLearning/GCPNet.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and mechanism of human vesicular polyamine transporter

Polyamines play essential roles in gene expression and modulate neuronal transmission in mammals. Vesicular polyamine transporters (VPAT) from the SLC18 family exploit the transmembrane H + gradient to translocate polyamines into secretory vesicles, enabling the quantal release of polyamine neuromodulators and underpinning learning and memory formation. Here, we report the cryo-electron microscopy structures of human VPAT in complex with spermine, spermidine, H + , or tetrabenazine, elucidating discrete lumen-facing states of the antiporter and pivotal interactions between VPAT and its substrate or inhibitor. Leveraging structure-inspired mutagenesis studies and protein structure prediction, we deduce an unforeseen mechanism whereby the polyamine and H + compete for multiple acidic protein residues both directly and indirectly, and rationalize how the antidopaminergic therapeutic tetrabenazine impedes vesicular transport of polyamines. This study unravels the mechanism of an H + -coupled polyamine antiporter, reveals mechanistic diversity between VPAT and other SLC18 antiporters, and raises new prospects for combating human disorders of polyamine homeostasis.

59 BASIC BIOLOGICAL SCIENCES↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

Data-Efficient Generation of Protein Conformational Ensembles with Backbone-to-Side-Chain Transformers

Excitement at the prospect of using data-driven generative models to sample configurational ensembles of biomolecular systems stems from the extraordinary success of these models on a diverse set of high-dimensional sampling tasks. Unlike image generation or even the closely related problem of protein structure prediction, there are currently no data sources with sufficient breadth to parametrize generative models for conformational ensembles. To enable discovery, a fundamentally different approach to building generative models is required: models should be able to propose rare, albeit physical, conformations that may not arise in even the largest data sets. Here, in this work, we introduce a modular strategy to generate conformations based on “backmapping” from a fixed protein backbone that (1) maintains conformational diversity of the side chains and (2) couples the side-chain fluctuations using global information about the protein conformation. Our model combines simple statistical models of side-chain conformations based on rotamer libraries with the now ubiquitous transformer architecture to sample with atomistic accuracy. Together, these ingredients provide a strategy for rapid data acquisition and hence a crucial ingredient for scalable physical simulation with generative neural networks.

36 MATERIALS SCIENCE↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Unraveling the Molecular Origin of Prey-Wrapping Spider Silk's Unique Mechanical Properties and Assembly Process Using NMR

Prey wrapping spider silk's unique mechanical properties are investigated confirming the silk's high degree of extensibility and superior toughness compared to other types of spider silk. For the first time, the pre-spinning dope phase is studied in isotope-enriched intact aciniform (AC) silk glands using solution NMR that reveals a combination of α-helical domains linked by disordered random coil chains consistent with previously proposed “beads-on-a-string” models. The model is further refined through the AlphaFold2 protein structure prediction tool. Finally, extensive magic angle spinning (MAS) solid-state (SS) NMR data for isotopically-enriched fibers is used to refine the structural model for AC silk from two species, A. aurantia and A. argentata. The SSNMR data shows that the AC silk fibers are highly α-helical, coiled-coil in structure but, also exhibit significant β-sheet components that can be traced back to the Gly-rich disordered linker regions in the pre-spinning dope phase that are converted to β-sheet structures during fiber formation. This combination of mechanical and structural characterization enhances the understanding of AC silk's liquid-to-solid transition and structure-mechanics relationship. In conclusion, these prey wrap silk results and models will provide the basis for the design of biomimetic materials inspired by the AC spider silk system.

36 MATERIALS SCIENCE↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Universally Accessible Structural Data on Macromolecular Conformation, Assembly, and Dynamics by Small Angle X-Ray Scattering for DNA Repair Insights.

Structures provide a critical breakthrough step for biological analyses, and small angle X-ray scattering (SAXS) is a powerful structural technique to study dynamic DNA repair proteins. As toxic and mutagenic repair intermediates need to be prevented from inadvertently harming the cell, DNA repair proteins often chaperone these intermediates through dynamic conformations, coordinated assemblies, and allosteric regulation. By measuring structural conformations in solution for both proteins, DNA, RNA, and their complexes, SAXS provides insight into initial DNA damage recognition, mechanisms for validation of their substrate, and pathway regulation. Here, we describe exemplary SAXS analyses of a DNA damage response protein spanning from what can be derived directly from the data to obtaining super resolution through the use of SAXS selection of atomic models. We outline strategies and tactics for practical SAXS data collection and analysis. Making these structural experiments in reach of any basic and clinical researchers who have protein, SAXS data can readily be collected at government-funded synchrotrons, typically at no cost for academic researchers. In addition to discussing how SAXS complements and enhances cryo-electron microscopy, X-ray crystallography, NMR, and computational modeling, we furthermore discuss taking advantage of recent advances in protein structure prediction in combination with SAXS analysis.

Chinnam, Naga Babu↗

Characterization of aromatic acid/proton symporters in Pseudomonas putida KT2440 toward efficient microbial conversion of lignin-related aromatics

Pseudomonas putida KT2440 (hereafter KT2440) is a well-studied platform bacterium for the production of industrially valuable chemicals from heterogeneous mixtures of aromatic compounds obtained from lignin depolymerization. KT2440 can grow on lignin-related monomers, such as ferulate (FA), 4-coumarate (4CA), vanillate (VA), 4-hydroxybenzoate (4HBA), and protocatechuate (PCA). Genes associated with their catabolism are known, but knowledge about the uptake systems remains limited. In this work, we studied the KT2440 transporters of lignin-related monomers and their substrate selectivity. Based on the inhibition by protonophores, we focused on five genes encoding aromatic acid/H+ symporter family transporters categorized into major facilitator superfamily that uses the proton motive force. Furthermore, the mutants of PP_1376 (pcaK) and PP_3349 (hcnK) exhibited significantly reduced growth on PCA/4HBA and FA/4CA, respectively, while no change was observed on VA for any of the five gene mutants. At pH 9.0, the conversion of these compounds by hcnK mutant (FA/4CA) and vanK mutant (VA) was dramatically reduced, revealing that these transporters are crucial for the uptake of the anionic substrates at high pH. Uptake assays using 14 C-labeled substrates in Escherichia coli and biosensor-based assays confirmed that PcaK, HcnK, and VanK have ability to take up PCA, FA/4CA, and VA/PCA, respectively. Additionally, analyses of the predicted protein structures suggest that the size and hydropathic properties of the substrate-binding sites of these transporters determine their substrate preferences. Overall, this study reveals that at physiological pH, PcaK and HcnK have a major role in the uptake of PCA/4HBA and FA/4CA, respectively, and VanK is a VA/PCA transporter. This information can contribute to the engineering of strains for the efficient conversion of lignin-related monomers to value-added chemicals.

59 BASIC BIOLOGICAL SCIENCES↗