Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Multi-Omics integration can be used to rescue metabolic information for some of the dark region of the Pseudomonas putida proteome

In every omics experiment, genes or their products are identified for which even state of the art tools are unable to assign a function. In the biotechnology chassis organism Pseudomonas putida, these proteins of unknown function make up 14% of the proteome. This missing information can bias analyses since these proteins can carry out functions which impact the engineering of organisms. As a consequence of predicting protein function across all organisms, function prediction tools generally fail to use all of the types of data available for any specific organism, including protein and transcript expression information. Additionally, the release of Alphafold predictions for all Uniprot proteins provides a novel opportunity for leveraging structural information. We constructed a bespoke machine learning model to predict the function of recalcitrant proteins of unknown function in Pseudomonas putida based on these sources of data, which annotated 1079 terms to 213 proteins. Among the predicted functions supplied by the model, we found evidence for a significant overrepresentation of nitrogen metabolism and macromolecule processing proteins. These findings were corroborated by manual analyses of selected proteins which identified, among others, a functionally unannotated operon that likely encodes a branch of the shikimate pathway.

60 APPLIED LIFE SCIENCES↗

Using diverse potentials and scoring functions for the development of improved machine-learned models for protein–ligand affinity and docking pose prediction

The advent of computational drug discovery holds the promise of significantly reducing the effort of experimentalists, along with monetary cost. More generally, predicting the binding of small organic molecules to biological macromolecules has far-reaching implications for a range of problems, including metabolomics. However, problems such as predicting the bound structure of a protein–ligand complex along with its affinity have proven to be an enormous challenge. In recent years, machine learning-based methods have proven to be more accurate than older methods, many based on simple linear regression. Nonetheless, there remains room for improvement, as these methods are often trained on a small set of features, with a single functional form for any given physical effect, and often with little mention of the rationale behind choosing one functional form over another. Moreover, it is not entirely clear why one machine learning method is favored over another. Here, we endeavor to undertake a comprehensive effort towards developing high-accuracy, machine-learned scoring functions, systematically investigating the effects of machine learning method and choice of features, and, when possible, providing insights into the relevant physics using methods that assess feature importance. Here, we show synergism among disparate features, yielding adjusted R 2 with experimental binding affinities of up to 0.871 on an independent test set and enrichment for native bound structures of up to 0.913. When purely physical terms that model enthalpic and entropic effects are used in the training, we use feature importance assessments to probe the relevant physics and hopefully guide future investigators working on this and other computational chemistry problems.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-based group A streptococcal vaccine design: Helical wheel homology predicts antibody cross-reactivity among streptococcal M protein–derived peptides

Group A streptococcus (Strep A) surface M protein, an α-helical coiled-coil dimer, is a vaccine target and a major determinant of streptococcal virulence. The sequence-variable N-terminal region of the M protein defines the M type and also contains epitopes that promote opsonophagocytic killing of streptococci. Recent reports have reported considerable cross-reactivity among different M types, suggesting the prospect of identifying cross-protective epitopes that would constitute a broadly protective multivalent vaccine against Strep A isolates. In this work, we have used a combination of immunological assays, structural biology, and cheminformatics to construct a recombinant M protein–based vaccine that included six Strep A M peptides that were predicted to elicit antisera that would cross-react with an additional 15 nonvaccine M types of Strep A. Rabbit antisera against this recombinant vaccine cross-reacted with 10 of the 15 nonvaccine M peptides. Two of the five nonvaccine M peptides that did not cross-react shared high sequence identity (≥50%) with the vaccine peptides, implying that high sequence identity alone was insufficient for cross-reactivity among the M peptides. Additional structural analyses revealed that the sequence identity at corresponding polar helical-wheel heptad sites between vaccine and nonvaccine peptides accurately distinguishes cross-reactive from non–cross-reactive peptides. On the basis of these observations, we developed a scoring algorithm based on the sequence identity at polar heptad sites. When applied to all epidemiologically important M types, this algorithm should enable the selection of a minimal number of M peptide–based vaccine candidates that elicit broadly protective immunity against Strep A.

60 APPLIED LIFE SCIENCES↗

Subdomain cryo-EM structure of nodaviral replication protein A crown complex provides mechanistic insights into RNA genome replication

For positive-strand RNA [(+)RNA] viruses, the major target for antiviral therapies is genomic RNA replication, which occurs at poorly understood membrane-bound viral RNA replication complexes. Recent cryoelectron microscopy (cryo-EM) of nodavirus RNA replication complexes revealed that the viral double-stranded RNA replication template is coiled inside a 30- to 90-nm invagination of the outer mitochondrial membrane, whose necked aperture to the cytoplasm is gated by a 12-fold symmetric, 35-nm diameter “crown” complex that contains multifunctional viral RNA replication protein A. Here we report optimizing cryo-EM tomography and image processing to improve crown resolution from 33 to 8.5 Å. This resolves the crown into 12 distinct vertical segments, each with 3 major subdomains: A membrane-connected basal lobe and an apical lobe that together comprise the ~19-nm-diameter central turret, and a leg emerging from the basal lobe that connects to the membrane at ~35-nm diameter. Despite widely varying replication vesicle diameters, the resulting two rings of membrane interaction sites constrain the vesicle neck to a highly uniform shape. Labeling protein A with a His-tag that binds 5-nm Ni-nanogold allowed cryo-EM tomography mapping of the C terminus of protein A to the apical lobe, which correlates well with the predicted structure of the C-proximal polymerase domain of protein A. These and other results indicate that the crown contains 12 copies of protein A arranged basally to apically in an N-to-C orientation. Moreover, the apical polymerase localization has significant mechanistic implications for template RNA recruitment and (-) and (+)RNA synthesis.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and Antigenicity of the Porcine Astrovirus 4 Capsid Spike

Porcine astrovirus 4 (PoAstV4) has been recently associated with respiratory disease in pigs. In order to understand the scope of PoAstV4 infections and to support the development of a vaccine to combat PoAstV4 disease in pigs, we designed and produced a recombinant PoAstV4 capsid spike protein for use as an antigen in serological assays and for potential future use as a vaccine antigen. Structural prediction of the full-length PoAstV4 capsid protein guided the design of the recombinant PoAstV4 capsid spike domain expression plasmid. The recombinant PoAstV4 capsid spike was expressed in Escherichia coli, purified by affinity and size-exclusion chromatography, and its crystal structure was determined at 1.85 Å resolution, enabling structural comparisons to other animal and human astrovirus capsid spike structures. The recombinant PoAstV4 capsid spike protein was also used as an antigen for the successful development of a serological assay to detect PoAstV4 antibodies, demonstrating that the recombinant PoAstV4 capsid spike retains antigenic epitopes found on the native PoAstV4 capsid. These studies lay a foundation for seroprevalence studies and the development of a PoAstV4 vaccine for swine.

Virology↗

A Surface Exposed, Two-Domain Lipoprotein Cargo of a Type XI Secretion System Promotes Colonization of Host Intestinal Epithelia Expressing Glycans

The only known required component of the newly described Type XI secretion system (TXISS) is an outer membrane protein (OMP) of the DUF560 family. TXISS OMPs are broadly distributed across proteobacteria, but properties of the cargo proteins they secrete are largely unexplored. We report biophysical, histochemical, and phenotypic evidence that Xenorhabdus nematophila NilC is surface exposed. Biophysical data and structure predictions indicate that NilC is a two-domain protein with a C-terminal, 8-stranded β-barrel. This structure has been noted as a common feature of TXISS effectors and may be important for interactions with the TXISS OMP . The NilC N-terminal domain is more enigmatic, but our results indicate it is ordered and forms a β-sheet structure, and bioinformatics suggest structural similarities to carbohydrate-binding proteins. X. nematophila NilC and its presumptive TXISS OMP partner NilB are required for colonizing the anterior intestine of Steinernema carpocapsae nematodes: the receptacle of free-living, infective juveniles and the anterior intestinal cecum (AIC) in juveniles and adults. We show that, in adult nematodes, the AIC expresses a Wheat Germ Agglutinin (WGA)-reactive material, indicating the presence of N-acetylglucosamine or N-acetylneuraminic acid sugars on the AIC surface. A role for this material in colonization is supported by the fact that exogenous addition of WGA can inhibit AIC colonization by X. nematophila. Conversely, the addition of exogenous purified NilC increases the frequency with which X. nematophila is observed at the AIC, demonstrating that abundant extracellular NilC can enhance colonization. NilC may facilitate X. nematophila adherence to the nematode intestinal surface by binding to host glycans, it might support X. nematophila nutrition by cleaving sugars from the host surface, or it might help protect X. nematophila from nematode host immunity. Proteomic and metabolomic analyses of wild type X. nematophila compared to those lacking nilB and nilC revealed differences in cell wall and secreted polysaccharide metabolic pathways. Additionally, purified NilC is capable of binding peptidoglycan, suggesting that periplasmic NilC may interact with the bacterial cell wall. Overall, these findings support a model that NilB-regulated surface exposure of NilC mediates interactions between X. nematophila and host surface glycans during colonization. This is a previously unknown function for a TXISS.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning-assisted elucidation of CD81–CD44 interactions in promoting cancer stemness and extracellular vesicle integrity

Tumor-initiating cells with reprogramming plasticity or stem-progenitor cell properties (stemness) are thought to be essential for cancer development and metastatic regeneration in many cancers; however, elucidation of the underlying molecular network and pathways remains demanding. Combining machine learning and experimental investigation, here we report CD81, a tetraspanin transmembrane protein known to be enriched in extracellular vesicles (EVs), as a newly identified driver of breast cancer stemness and metastasis. Using protein structure modeling and interface prediction-guided mutagenesis, we demonstrate that membrane CD81 interacts with CD44 through their extracellular regions in promoting tumor cell cluster formation and lung metastasis of triple negative breast cancer (TNBC) in human and mouse models. In-depth global and phosphoproteomic analyses of tumor cells deficient with CD81 or CD44 unveils endocytosis-related pathway alterations, leading to further identification of a quality-keeping role of CD44 and CD81 in EV secretion as well as in EV-associated stemness-promoting function. CD81 is coexpressed along with CD44 in human circulating tumor cells (CTCs) and enriched in clustered CTCs that promote cancer stemness and metastasis, supporting the clinical significance of CD81 in association with patient outcomes. Our study highlights machine learning as a powerful tool in facilitating the molecular understanding of new molecular targets in regulating stemness and metastasis of TNBC.

59 BASIC BIOLOGICAL SCIENCES↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Auto_PDI

Protein-DNA Interaction Workflow (PDI Workflow), a pipeline that focuses on generating high-quality docking and molecular dynamics simulations for Protein-DNA complexes. This allows us to take DNA sequences with unknown tertiary structures, accurately predict their structure, dock them with the desired target protein, and then simulate their interactions using molecular dynamics simulations

Kumar, Neeraj↗

BMC Caller: a webtool to identify and analyze bacterial microcompartment types in sequence data

Bacterial microcompartments (BMCs) are protein-based organelles found across the bacterial tree of life. They consist of a shell, made of proteins that oligomerize into hexagonally and pentagonally shaped building blocks, that surrounds enzymes constituting a segment of a metabolic pathway. The proteins of the shell are unique to BMCs. They also provide selective permeability; this selectivity is dictated by the requirements of their cargo enzymes. We have recently surveyed the wealth of different BMC types and their occurrence in all available genome sequence data by analyzing and categorizing their components found in chromosomal loci using HMM (Hidden Markov Model) protein profiles. To make this a “do-it yourself” analysis for the public we have devised a webserver, BMC Caller (https://bmc-caller.prl.msu.edu), that compares user input sequences to our HMM profiles, creates a BMC locus visualization, and defines the functional type of BMC, if known. Shell proteins in the input sequence data are also classified according to our function-agnostic naming system and there are links to similar proteins in our database as well as an external link to a structure prediction website to easily generate structural models of the shell proteins, which facilitates understanding permeability properties of the shell. Additionally, the BMC Caller website contains a wealth of information on previously analyzed BMC loci with links to detailed data for each BMC protein and phylogenetic information on the BMC shell proteins. Our tools greatly facilitate BMC type identification to provide the user information about the associated organism’s metabolism and enable discovery of new BMC types by providing a reference database of all currently known examples.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

Unusual modifications of protein biomarkers expressed by plasmid, prophage, and bacterial host of pathogenic Escherichia coli identified using top‐down proteomic analysis

Rationale Pathogenic bacteria often carry prophage (bacterial viruses) and plasmids (small circular pieces of DNA) that may harbor toxin, antibacterial, and antibiotic resistance genes. Proteomic characterization of pathogenic bacteria should include the identification of host proteins and proteins produced by prophage and plasmid genomes. Methods Protein biomarkers of two strains of Shiga toxin–producingEscherichia coli(STEC) were identified using antibiotic induction, matrix‐assisted laser desorption/ionization tandem time‐of‐flight (MALDI‐TOF‐TOF) tandem mass spectrometry (MS/MS) with post‐source decay (PSD), top‐down proteomic (TDP) analysis, and plasmid sequencing. Alphafold2 was also used to compare predicted in silico structures of the identified proteins to prominent fragment ions generated using MS/MS‐PSD. Strain samples were also analyzed with and without chemical reduction treatment to detect the attachment of pendant groups bound by thioester or disulfide bonds. Results Shiga toxin was detected and/or identified in both STEC strains. For the first time, we also identified the osmotically inducible protein (OsmY) whose sequence unexpectedly had two forms: a full and a truncated sequence. The truncated OsmY terminates in the middle of an α‐helix as determined by Alphafold2. A plasmid‐encoded colicin immunity protein was also identified with and without attachment of an unidentified cysteine‐bound pendant group (~307 Da). Plasmid sequencing confirmed top‐down analysis and the identification of a promoter upstream of the immunity gene that is activated by antibiotic induction, that is, SOS box. Conclusions TDP analysis, coupled with other techniques (e.g., antibiotic induction, chemical reduction, plasmid sequencing, and in silico protein modeling), is a powerful tool to identify proteins (and their modifications), including prophage‐ and plasmid‐encoded proteins, produced by pathogenic microorganisms.

Biochemistry & Molecular Biology↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-CompBio/pf-gnn_pli

We have developed a deep learning based method to decode the protein-ligand interactions and predict their probability of binding. We use 3-Dimensional structure of proteins and ligand molecules with the ligand docked to the receptor and use the information at the interaction region to assess the binding of the protein-ligand complex.

Bontha, Mridula↗

The leucine-rich repeats in allelic barley MLA immune receptors define specificity towards sequence-unrelated powdery mildew avirulence effectors with a predicted common RNase-like fold

Nucleotide-binding domain leucine-rich repeat-containing receptors (NLRs) in plants can detect avirulence (AVR) effectors of pathogenic microbes. The Mildew locus a (Mla) NLR gene has been shown to confer resistance against diverse fungal pathogens in cereal crops. In barley, Mla has undergone allelic diversification in the host population and confers isolate-specific immunity against the powdery mildew-causing fungal pathogen Blumeria graminis forma specialis hordei (Bgh). We previously isolated the Bgh effectors AVR A1 , AVR A7 , AVR A9 , AVR A13 , and allelic AVR A10 /AVR A22 , which are recognized by matching MLA1, MLA7, MLA9, MLA13, MLA10 and MLA22, respectively. Here, we extend our knowledge of the Bgh effector repertoire by isolating the AVR A6 effector, which belongs to the family of catalytically inactive RNase-Like Proteins expressed in Haustoria (RALPHs). Using structural prediction, we also identified RNase-like folds in AVR A1 , AVR A7 , AVR A10 /AVR A22 , and AVR A13 , suggesting that allelic MLA recognition specificities could detect structurally related avirulence effectors. To better understand the mechanism underlying the recognition of effectors by MLAs, we deployed chimeric MLA1 and MLA6, as well as chimeric MLA10 and MLA22 receptors in plant co-expression assays, which showed that the recognition specificity for AVR A1 and AVR A6 as well as allelic AVR A10 and AVR A22 is largely determined by the receptors’ C-terminal leucine-rich repeats (LRRs). The design of avirulence effector hybrids allowed us to identify four specific AVR A10 and five specific AVR A22 aa residues that are necessary to confer MLA10- and MLA22-specific recognition, respectively. This suggests that the MLA LRR mediates isolate-specific recognition of structurally related AVR A effectors. Thus, functional diversification of multi-allelic MLA receptors may be driven by a common structural effector scaffold, which could be facilitated by proliferation of the RALPH effector family in the pathogen genome.

59 BASIC BIOLOGICAL SCIENCES↗

Binding Free Energy Analysis of Colicin D, E3 and E8 to Their Respective Cognate Immunity Proteins Using Computational Simulations

Colicins are antimicrobial proteins produced by bacteria for the purpose of destroying neighboring bacteria. Colicin activity is neutralized by a specific cognate immunity protein in order to protect the host. This study investigates the structural and binding mechanisms underlying the interaction of colicin-D, -E3 and -E8 to their respective immunity proteins (ImD, Im3 and Im8) using structure prediction, molecular dynamics (MD) simulations and MM-PBSA approach of free energy calculations. High-confidence colicin-immunity (Col-Im) complex structures predicted using AlphaFold2 were subjected to MD simulations of 150 ns with GROMACS and were analyzed for the binding free energy calculation using gmx_MMPBSA. Results showed that the complex of Col_E3-Im3 exhibited the most favorable binding free energy, driven by strong van der Waals and electrostatic interactions. Col_D-ImD and Col_E8-Im8 also showed the favorable binding. Electrostatics and hydrogen bonding emerged as a key factor driving binding and stability, while polar solvation acted as a destabilizing factor across all systems. These outcomes provide an understanding of the molecular mechanisms of Col-Im systems, with potential applications for developing natural antimicrobials for food safety.

Biochemistry & Molecular Biology↗