Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

ppdx : Automated modeling of protein–protein interaction descriptors for use with machine learning

This paper describes ppdx, a python workflow tool that combines protein sequence alignment, homology modeling, and structural refinement, to compute a broad array of descriptors for characterizing protein–protein interactions. The descriptors can be used to predict various properties of interest, such as protein–protein binding affinities, or inhibitory concentrations (IC 50 ), using approaches that range from simple regression to more complex machine learning models. The software is highly modular. It supports different protocols for generating structures, and 95 descriptors can be currently computed. More protocols and descriptors can be easily added. The implementation is highly parallel and can fully exploit the available cores in a single workstation, or multiple nodes on a supercomputer, allowing many systems to be analyzed simultaneously. As an illustrative application, ppdx is used to parametrize a model that predicts the IC 50 of a set of antigens and a class of antibodies directed to the influenza hemagglutinin stalk.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genome sequence and characterization of a novel Pseudomonas putida phage, MiCath

Abstract Pseudomonads are ubiquitous bacteria with importance in medicine, soil, agriculture, and biomanufacturing. We report a novel Pseudomonas putida phage, MiCath, which is the first known phage infecting P. putida S12, a strain increasingly used as a synthetic biology chassis. MiCath was isolated from garden soil under a tomato plant using P. putida S12 as a host and was also found to infect four other P. putida strains. MiCath has a ~ 61 kbp double-stranded DNA genome which encodes 97 predicted open reading frames (ORFs); functions could only be predicted for 48 ORFs using comparative genomics. Functions include structural phage proteins, other common phage proteins (e.g., terminase), a queuosine gene cassette, a cas4 exonuclease, and an endosialidase. Restriction digestion analysis suggests the queuosine gene cassette encodes a pathway capable of modification of guanine residues. When compared to other phage genomes, MiCath shares at most 74% nucleotide identity over 2% of the genome with any sequenced phage. Overall, MiCath is a novel phage with no close relatives, encoding many unique gene products.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Assessment of Pose Prediction Accuracy in RNA–Ligand Docking

Structure-based virtual high-throughput screening is used in early-stage drug discovery. Over the years, docking protocols and scoring functions for protein–ligand complexes have evolved to improve the accuracy in the computation of binding strengths and poses. In the past decade, RNA has also emerged as a target class for new small-molecule drugs. However, most ligand docking programs have been validated and tested for proteins and not RNA. Here, we test the docking power (pose prediction accuracy) of three state-of-the-art docking protocols on 173 RNA–small molecule crystal structures. The programs are AutoDock4 (AD4) and AutoDock Vina (Vina), which were designed for protein targets, and rDock, which was designed for both protein and nucleic acid targets. AD4 performed relatively poorly. For RNA targets for which a crystal structure of a bound ligand used to limit the docking search space is available and for which the goal is to identify new molecules for the same pocket, rDock performs slightly better than Vina, with success rates of 48% and 63%, respectively. However, in the more common type of early-stage drug discovery setting, in which no structure of a ligand–target complex is known and for which a larger search space is defined, rDock performed similarly to Vina, with a low success rate of ~27%. Further, Vina was found to have bias for ligands with certain physicochemical properties, whereas rDock performs similarly for all ligand properties. Thus, for projects where no ligand–protein structure already exists, Vina and rDock are both applicable. However, the relatively poor performance of all methods relative to protein–target docking illustrates a need for further methods refinement.

59 BASIC BIOLOGICAL SCIENCES↗

An evolutionarily conserved tryptophan cage promotes folding of the extended RNA recognition motif in the hnRNPR ‐like protein family

Abstract The heterogeneous nuclear ribonucleoprotein (hnRNP) R‐like family is a class of RNA binding proteins in the hnRNP superfamily with diverse functions in RNA processing. Here, we present the 1.90 Å X‐ray crystal structure and solution NMR studies of the first RNA recognition motif (RRM) of human hnRNPR. We find that this domain adopts an extended RRM (eRRM1) featuring a canonical RRM with a structured N‐terminal extension (N ext ) motif that docks against the RRM and extends the β‐sheet surface. The adjoining loop is structured and forms a tryptophan cage motif to position the N ext motif for docking to the RRM. Combining mutagenesis, solution NMR spectroscopy, and thermal denaturation studies, we evaluate the importance of residues in the N ext –RRM interface and adjoining loop on eRRM folding and conformational dynamics. We find that these sites are essential for protein solubility, conformational ordering, and thermal stability. Consistent with their importance, mutations in the N ext –RRM interface and loop are associated with several cancers in a survey of somatic mutations in cancer studies. Sequence and structure comparison of the human hnRNPR eRRM1 to experimentally verified and predicted hnRNPR‐like proteins reveals conserved features in the eRRM.

Biochemistry & Molecular Biology↗

Transforming our understanding of chloroplast-associated genes through comprehensive characterization of protein localizations and protein-protein interactions

Bioenergy crops are a renewable source of fuels and are a critical base for building a carbon-neutral economy. Rational engineering of bioenergy crops has the potential to enhance the yields. However, our ability to engineer plants is limited because the functions of most genes remain unknown. Systematic characterization of gene function in plants thus has the potential to greatly accelerate bioenergy research. Here, we focus on the chloroplast, an underexplored energy-producing organelle that is a hallmark of plants. The chloroplast is one of the promising targets of biofuel crop engineering efforts because of its central role in photosynthesis, metabolism, and intracellular signaling. However, the protein composition of the chloroplast and the functions of most of its proteins remain poorly characterized. At the core of this project, we sought to comprehensively determine the localization of chloroplast-associated proteins and generate a spatially defined protein-protein interaction network for chloroplast. For this purpose, we used the leading model alga Chlamydomonas reinhardtii, which greatly increased experimental speed and throughput. We illustrated the value of our findings to land plants by determining the localization of Arabidopsis thaliana land plant homologs of the Chlamydomonas proteins. Altogether, we were successful in determining the localization of 1,034 chloroplast-associated proteins in Chlamydomonas. The localizations provide numerous insights into the spatial organization of chloroplasts and how they function to support photosynthesis. The localization patterns of distinct proteins revealed new chloroplast structures and revealed new spatial organization inside the chloroplast. We also identified new components of known chloroplast structures, such as the chloroplast envelope, nucleoid, plastoglobuli, and pyrenoid. We identified these new components by investigating the interacting partners of known proteins. Many proteins localized in both the chloroplast and other cellular structures, thereby hinting at new functions and communication between cellular structures. We also applied machine learning on the atlas to generate predictions for the location of all of the proteins in Chlamydomonas. This enabled us to assign putative functions to many uncharacterized proteins based on their cellular location. Altogether, this research establishes a rich resource that opens new avenues of investigation and guides future work in deciphering and manipulating chloroplast function. Next, we developed an extensive protein-protein interaction network for the chloroplast by performing affinity purification-mass spectrometry on ~1,150 tagged chloroplast-associated proteins, the first such large-scale study in any photosynthetic organism. This dataset reveals 4,694 high-confidence protein-protein interactions, offering insights into the functions of thousands of conserved poorly-characterized chloroplast proteins. This systematic identification of protein-protein interactions in the chloroplast also provides multiple exciting new research directions and a detailed blueprint of the chloroplast's operation. This research lays the groundwork to decipher the inner workings of the chloroplast, the cell structure at the heart of photosynthesis. The spatial atlas and protein-protein interactions reveal chloroplast organizational features that would not have been accessible with traditional approaches. The localization mapping, insights into the function, and research materials generated further provide a rich resource for the research community to advance the understanding of how the chloroplast is organized to enable engineering of enhanced photosynthetic organisms.

59 BASIC BIOLOGICAL SCIENCES↗

PYK-SubstitutionOME: an integrated database containing allosteric coupling, ligand affinity and mutational, structural, pathological, bioinformatic and computational information about pyruvate kinase isozymes

Interpreting changes in patient genomes, understanding how viruses evolve and engineering novel protein function all depend on accurately predicting the functional outcomes that arise from amino acid substitutions. To that end, the development of first-generation prediction algorithms was guided by historic experimental datasets. However, these datasets were heavily biased toward substitutions at positions that have not changed much throughout evolution (i.e. conserved). Although newer datasets include substitutions at positions that span a range of evolutionary conservation scores, these data are largely derived from assays that agglomerate multiple aspects of function. To facilitate predictions from the foundational chemical properties of proteins, large substitution databases with biochemical characterizations of function are needed. We report here a database derived from mutational, biochemical, bioinformatic, structural, pathological and computational studies of a highly studied protein family—pyruvate kinase (PYK). A centerpiece of this database is the biochemical characterization—including quantitative evaluation of allosteric regulation—of the changes that accompany substitutions at positions that sample the full conservation range observed in the PYK family. We have used these data to facilitate critical advances in the foundational studies of allosteric regulation and protein evolution and as rigorous benchmarks for testing protein predictions. We trust that the collected dataset will be useful for the broader scientific community in the further development of prediction algorithms.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Accelerating crystal structure determination with iterative AlphaFold prediction

Experimental structure determination can be accelerated with artificial intelligence (AI)-based structure-prediction methods such as AlphaFold . Here, an automatic procedure requiring only sequence information and crystallographic data is presented that uses AlphaFold predictions to produce an electron-density map and a structural model. Iterating through cycles of structure prediction is a key element of this procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. This procedure was applied to X-ray data for 215 structures released by the Protein Data Bank in a recent six-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2 Å. Predictions from the iterative template-guided prediction procedure were more accurate than those obtained without templates. It is concluded that AlphaFold predictions obtained based on sequence information alone are usually accurate enough to solve the crystallographic phase problem with molecular replacement, and a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization is suggested.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling receptor flexibility in the structure-based design of KRAS G12C inhibitors

KRAS has long been referred to as an ‘undruggable’ target due to its high affinity for its cognate ligands (GDP and GTP) and its lack of readily exploited allosteric binding pockets. Recent progress in the development of covalent inhibitors of KRAS G12C has revealed that occupancy of an allosteric binding site located between the α3-helix and switch-II loop of KRAS G12C —sometimes referred to as the ‘switch-II pocket’—holds great potential in the design of direct inhibitors of KRAS G12C . In studying diverse switch-II pocket binders during the development of sotorasib (AMG 510), the first FDA-approved inhibitor of KRAS G12C , we found the dramatic conformational flexibility of the switch-II pocket posing significant challenges toward the structure-based design of inhibitors. Here, we present our computational approaches for dealing with receptor flexibility in the prediction of ligand binding pose and binding affinity. For binding pose prediction, we modified the covalent docking program CovDock to allow for protein conformational mobility. This new docking approach, termed as FlexCovDock, improves success rates from 55 to 89% for binding pose prediction on a dataset of 10 cross-docking cases and has been prospectively validated across diverse ligand chemotypes. For binding affinity prediction, we found standard free energy perturbation (FEP) methods could not adequately handle the significant conformational change of the switch-II loop. We developed a new computational strategy to accelerate conformational transitions through the use of targeted protein mutations. Using this methodology, the mean unsigned error (MUE) of binding affinity prediction were reduced from 1.44 to 0.89 kcal/mol on a set of 14 compounds. These approaches were of significant use in facilitating the structure-based design of KRAS G12C inhibitors and are anticipated to be of further use in the design of covalent (and noncovalent) inhibitors of other conformationally labile protein targets.

59 BASIC BIOLOGICAL SCIENCES↗

PigmentHunter: A point-and-click application for automated chlorophyll-protein simulations

Chlorophyll proteins (CPs) are the workhorses of biological photosynthesis, working together to absorb solar energy, transfer it to chemically active reaction centers, and control the charge-separation process that drives its storage as chemical energy. Yet predicting CP optical and electronic properties remains a serious challenge, driven by the computational difficulty of treating large, electronically coupled molecular pigments embedded in a dynamically structured protein environment. To address this challenge, we introduce here an analysis tool called PigmentHunter, which automates the process of preparing CP structures for molecular dynamics (MD), running short MD simulations on the nanoHUB.org science gateway, and then using electrostatic and steric analysis routines to predict optical absorption, fluorescence, and circular dichroism spectra within a Frenkel exciton model. Inter-pigment couplings are evaluated using point-dipole or transition-charge coupling models, while site energies can be estimated using both electrostatic and ring-deformation approaches. The package is built in a Jupyter Notebook environment, with a point-and-click interface that can be used either to manually prepare individual structures or to batch-process many structures at once. Here, we illustrate PigmentHunter’s capabilities with example simulations on spectral line shapes in the light harvesting 2 complex, site energies in the Fenna–Matthews–Olson protein, and ring deformation in photosystems I and II.

14 SOLAR ENERGY↗

UniKP: a unified framework for the prediction of enzyme kinetic parameters

Prediction of enzyme kinetic parameters is essential for designing and optimizing enzymes for various biotechnological and industrial applications, but the limited performance of current prediction tools on diverse tasks hinders their practical applications. Here, we introduce UniKP, a unified framework based on pretrained language models for the prediction of enzyme kinetic parameters, including enzyme turnover number (k cat ), Michaelis constant (K m ), and catalytic efficiency (k cat / K m ), from protein sequences and substrate structures. A two-layer framework derived from UniKP (EF-UniKP) has also been proposed to allow robust k cat prediction in considering environmental factors, including pH and temperature. In addition, four representative re-weighting methods are systematically explored to successfully reduce the prediction error in high-value prediction tasks. We have demonstrated the application of UniKP and EF-UniKP in several enzyme discovery and directed evolution tasks, leading to the identification of new enzymes and enzyme mutants with higher activity. UniKP is a valuable tool for deciphering the mechanisms of enzyme kinetics and enables novel insights into enzyme engineering and their industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Biomineralization Potential of Metal-Sequestering Bacteria for Rare Earth Elements (REE) Material Synthesis

Extremophilic methylotrophs, such as Methylotuvimicrobium alcaliphilum 20ZR are known for their natural ability to carry out REE capture and uptake. It has been predicted that proteins associated with surface layers facilitate transport and homeostasis of essential metals. Indeed, electron microscopic analysis coupled with antibody staining revealed that cell envelop of metal-grown M. alcaliphilum 20ZR contains metal-binding proteins in the base of the cup-shaped structures of S-layers. While less explored, the proteins associated with S-layers are also predicted to contribute to scavenging of other minerals essential for core metabolism. Here, we employed bottom-up proteomic study to compare the protein content of S-layer fractions of M. alcaliphilum 20ZR cultures grown in the presence and absence trace metals to identify putative components REE-scavenging machinery.

36 MATERIALS SCIENCE↗

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

Dissecting the structural heterogeneity of proteins by native mass spectrometry

Abstract A single gene yields many forms of proteins via combinations of posttranscriptional/posttranslational modifications. Proteins also fold into higher‐order structures and interact with other molecules. The combined molecular diversity leads to the heterogeneity of proteins that manifests as distinct phenotypes. Structural biology has generated vast amounts of data, effectively enabling accurate structural prediction by computational methods. However, structures are often obtained heterologously under homogeneous states in vitro. The lack of native heterogeneity under cellular context creates challenges in precisely connecting the structural data to phenotypes. Mass spectrometry (MS) based proteomics methods can profile proteome composition of complex biological samples. Most MS methods follow the “bottom‐up” approach, which denatures and digests proteins into short peptide fragments for ease of detection. Coupled with chemical biology approaches, higher‐order structures can be probed via incorporation of covalent labels on native proteins that are maintained at the peptide level. Alternatively, native MS follows the “top‐down” approach and directly analyzes intact proteins under nondenaturing conditions. Various tandem MS activation methods can dissect the intact proteins for in‐depth structural elucidation. Herein, we review recent native MS applications for characterizing heterogeneous samples, including proteins binding to mixtures of ligands, homo/hetero‐complexes with varying stoichiometry, intrinsically disordered proteins with dynamic conformations, glycoprotein complexes with mixed modification states, and active membrane protein complexes in near‐native membrane environments. We summarize the benefits, challenges, and ongoing developments in native MS, with the hope to demonstrate an emerging technology that complements other tools by filling the knowledge gaps in understanding the molecular heterogeneity of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Mercury methylation by metabolically versatile and cosmopolitan marine bacteria

Microbes transform aqueous mercury (Hg) into methylmercury (MeHg), a potent neurotoxin that accumulates in terrestrial and marine food webs, with potential impacts on human health. This process requires the gene pair hgcAB, which encodes for proteins that actuate Hg methylation, and has been well described for anoxic environments. However, recent studies report potential MeHg formation in suboxic seawater, although the microorganisms involved remain poorly understood. In this study, we conducted large-scale multi-omic analyses to search for putative microbial Hg methylators along defined redox gradients in Saanich Inlet, British Columbia, a model natural ecosystem with previously measured Hg and MeHg concentration profiles. Analysis of gene expression profiles along the redoxcline identified several putative Hg methylating microbial groups, including Calditrichaeota, SAR324 and Marinimicrobia, with the last the most active based on hgc transcription levels. Marinimicrobia hgc genes were identified from multiple publicly available marine metagenomes, consistent with a potential key role in marine Hg methylation. Computational homology modelling predicts that Marinimicrobia HgcAB proteins contain the highly conserved amino acid sites and folding structures required for functional Hg methylation. Furthermore, a number of terminal oxidases from aerobic respiratory chains were associated with several putative novel Hg methylators. Our findings thus reveal potential novel marine Hg-methylating microorganisms with a greater oxygen tolerance and broader habitat range than previously recognized.

59 BASIC BIOLOGICAL SCIENCES↗