Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Machine learning sheds light on microbial dark proteins

In this article, metagenomics projects have revealed more than 8 billion non-redundant microbial protein sequences from across the Earth’s biosphere. Of these, 1.17 billion proteins do not have recognizable homologues in any of the more than 100,000 reference genomes available1. Understanding the function of these microbial proteins is a daunting task. Fortunately, machine learning has recently achieved unprecedented accuracy in modelling complex biological data and making predictions. At the forefront of these advancements are machine learning-based approaches that can confidently predict atomic-level protein structures for many (but not all) amino acid sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Atypical Divergence of SARS-CoV-2 Orf8 from Orf7a within the Coronavirus Lineage Suggests Potential Stealthy Viral Strategies in Immune Evasion

Orf8, one of the most puzzling genes in the SARS lineage of coronaviruses, marks a unique and striking difference in genome organization between SARS-CoV-2 and SARS-CoV-1. Here, using sequence comparisons, we unequivocally reveal the distant sequence similarities between SARS-CoV-2 Orf8 with its SARS-CoV-1 counterparts and the X4-like genes of coronaviruses, including its highly divergent “paralog” gene Orf7a, whose product is a potential immune antagonist of known structure. Supervised sequence space walks unravel identity levels that drop below 10% and yet exhibit subtle conservation patterns in this novel superfamily, characterized by an immunoglobulin-like beta sandwich topology. We document the high accuracy of the sequence space walk process in detail and characterize the subgroups of the superfamily in sequence space by systematic annotation of gene and taxon groups. While SARS-CoV-1 Orf7a and Orf8 genes are most similar to bat virus sequences, their SARS-CoV-2 counterparts are closer to pangolin virus homologs, reflecting the fine structure of conservation patterns within the SARS-CoV-2 genomes. The divergence between Orf7a and Orf8 is exceptionally idiosyncratic, since Orf7a is more constrained, whereas Orf8 is subject to rampant change, a peculiar feature that may be related to hitherto-unknown viral infection strategies. Despite their common origin, the Orf7a and Orf8 protein families exhibit different modes of evolutionary trajectories within the coronavirus lineage, which might be partly attributable to their complex interactions with the mammalian host cell, reflected by a multitude of functional associations of Orf8 in SARS-CoV-2 compared to a very small number of interactions discovered for Orf7a.

60 APPLIED LIFE SCIENCES↗

Foldy v1

We created a cloud-based application for running AlphaFold2 and related structural programs called Foldy. Foldy is built on Helm / Kubernetes, which enables a facile production deployment and a scalable backend. Once set up by an institution, Foldy requires no software expertise from its users. It can be used to predict the structure of large proteins (up to 3000 amino acids), to visualize Pfam and antiSMASH annotations, and to perform ligand docking with Autodock Vina.

Roberts, Jacob↗

Multi-Omics integration can be used to rescue metabolic information for some of the dark region of the Pseudomonas putida proteome

In every omics experiment, genes or their products are identified for which even state of the art tools are unable to assign a function. In the biotechnology chassis organism Pseudomonas putida, these proteins of unknown function make up 14% of the proteome. This missing information can bias analyses since these proteins can carry out functions which impact the engineering of organisms. As a consequence of predicting protein function across all organisms, function prediction tools generally fail to use all of the types of data available for any specific organism, including protein and transcript expression information. Additionally, the release of Alphafold predictions for all Uniprot proteins provides a novel opportunity for leveraging structural information. We constructed a bespoke machine learning model to predict the function of recalcitrant proteins of unknown function in Pseudomonas putida based on these sources of data, which annotated 1079 terms to 213 proteins. Among the predicted functions supplied by the model, we found evidence for a significant overrepresentation of nitrogen metabolism and macromolecule processing proteins. These findings were corroborated by manual analyses of selected proteins which identified, among others, a functionally unannotated operon that likely encodes a branch of the shikimate pathway.

60 APPLIED LIFE SCIENCES↗

Using diverse potentials and scoring functions for the development of improved machine-learned models for protein–ligand affinity and docking pose prediction

The advent of computational drug discovery holds the promise of significantly reducing the effort of experimentalists, along with monetary cost. More generally, predicting the binding of small organic molecules to biological macromolecules has far-reaching implications for a range of problems, including metabolomics. However, problems such as predicting the bound structure of a protein–ligand complex along with its affinity have proven to be an enormous challenge. In recent years, machine learning-based methods have proven to be more accurate than older methods, many based on simple linear regression. Nonetheless, there remains room for improvement, as these methods are often trained on a small set of features, with a single functional form for any given physical effect, and often with little mention of the rationale behind choosing one functional form over another. Moreover, it is not entirely clear why one machine learning method is favored over another. Here, we endeavor to undertake a comprehensive effort towards developing high-accuracy, machine-learned scoring functions, systematically investigating the effects of machine learning method and choice of features, and, when possible, providing insights into the relevant physics using methods that assess feature importance. Here, we show synergism among disparate features, yielding adjusted R 2 with experimental binding affinities of up to 0.871 on an independent test set and enrichment for native bound structures of up to 0.913. When purely physical terms that model enthalpic and entropic effects are used in the training, we use feature importance assessments to probe the relevant physics and hopefully guide future investigators working on this and other computational chemistry problems.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-based group A streptococcal vaccine design: Helical wheel homology predicts antibody cross-reactivity among streptococcal M protein–derived peptides

Group A streptococcus (Strep A) surface M protein, an α-helical coiled-coil dimer, is a vaccine target and a major determinant of streptococcal virulence. The sequence-variable N-terminal region of the M protein defines the M type and also contains epitopes that promote opsonophagocytic killing of streptococci. Recent reports have reported considerable cross-reactivity among different M types, suggesting the prospect of identifying cross-protective epitopes that would constitute a broadly protective multivalent vaccine against Strep A isolates. In this work, we have used a combination of immunological assays, structural biology, and cheminformatics to construct a recombinant M protein–based vaccine that included six Strep A M peptides that were predicted to elicit antisera that would cross-react with an additional 15 nonvaccine M types of Strep A. Rabbit antisera against this recombinant vaccine cross-reacted with 10 of the 15 nonvaccine M peptides. Two of the five nonvaccine M peptides that did not cross-react shared high sequence identity (≥50%) with the vaccine peptides, implying that high sequence identity alone was insufficient for cross-reactivity among the M peptides. Additional structural analyses revealed that the sequence identity at corresponding polar helical-wheel heptad sites between vaccine and nonvaccine peptides accurately distinguishes cross-reactive from non–cross-reactive peptides. On the basis of these observations, we developed a scoring algorithm based on the sequence identity at polar heptad sites. When applied to all epidemiologically important M types, this algorithm should enable the selection of a minimal number of M peptide–based vaccine candidates that elicit broadly protective immunity against Strep A.

60 APPLIED LIFE SCIENCES↗

Subdomain cryo-EM structure of nodaviral replication protein A crown complex provides mechanistic insights into RNA genome replication

For positive-strand RNA [(+)RNA] viruses, the major target for antiviral therapies is genomic RNA replication, which occurs at poorly understood membrane-bound viral RNA replication complexes. Recent cryoelectron microscopy (cryo-EM) of nodavirus RNA replication complexes revealed that the viral double-stranded RNA replication template is coiled inside a 30- to 90-nm invagination of the outer mitochondrial membrane, whose necked aperture to the cytoplasm is gated by a 12-fold symmetric, 35-nm diameter “crown” complex that contains multifunctional viral RNA replication protein A. Here we report optimizing cryo-EM tomography and image processing to improve crown resolution from 33 to 8.5 Å. This resolves the crown into 12 distinct vertical segments, each with 3 major subdomains: A membrane-connected basal lobe and an apical lobe that together comprise the ~19-nm-diameter central turret, and a leg emerging from the basal lobe that connects to the membrane at ~35-nm diameter. Despite widely varying replication vesicle diameters, the resulting two rings of membrane interaction sites constrain the vesicle neck to a highly uniform shape. Labeling protein A with a His-tag that binds 5-nm Ni-nanogold allowed cryo-EM tomography mapping of the C terminus of protein A to the apical lobe, which correlates well with the predicted structure of the C-proximal polymerase domain of protein A. These and other results indicate that the crown contains 12 copies of protein A arranged basally to apically in an N-to-C orientation. Moreover, the apical polymerase localization has significant mechanistic implications for template RNA recruitment and (-) and (+)RNA synthesis.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and Antigenicity of the Porcine Astrovirus 4 Capsid Spike

Porcine astrovirus 4 (PoAstV4) has been recently associated with respiratory disease in pigs. In order to understand the scope of PoAstV4 infections and to support the development of a vaccine to combat PoAstV4 disease in pigs, we designed and produced a recombinant PoAstV4 capsid spike protein for use as an antigen in serological assays and for potential future use as a vaccine antigen. Structural prediction of the full-length PoAstV4 capsid protein guided the design of the recombinant PoAstV4 capsid spike domain expression plasmid. The recombinant PoAstV4 capsid spike was expressed in Escherichia coli, purified by affinity and size-exclusion chromatography, and its crystal structure was determined at 1.85 Å resolution, enabling structural comparisons to other animal and human astrovirus capsid spike structures. The recombinant PoAstV4 capsid spike protein was also used as an antigen for the successful development of a serological assay to detect PoAstV4 antibodies, demonstrating that the recombinant PoAstV4 capsid spike retains antigenic epitopes found on the native PoAstV4 capsid. These studies lay a foundation for seroprevalence studies and the development of a PoAstV4 vaccine for swine.

Virology↗

A Surface Exposed, Two-Domain Lipoprotein Cargo of a Type XI Secretion System Promotes Colonization of Host Intestinal Epithelia Expressing Glycans

The only known required component of the newly described Type XI secretion system (TXISS) is an outer membrane protein (OMP) of the DUF560 family. TXISS OMPs are broadly distributed across proteobacteria, but properties of the cargo proteins they secrete are largely unexplored. We report biophysical, histochemical, and phenotypic evidence that Xenorhabdus nematophila NilC is surface exposed. Biophysical data and structure predictions indicate that NilC is a two-domain protein with a C-terminal, 8-stranded β-barrel. This structure has been noted as a common feature of TXISS effectors and may be important for interactions with the TXISS OMP . The NilC N-terminal domain is more enigmatic, but our results indicate it is ordered and forms a β-sheet structure, and bioinformatics suggest structural similarities to carbohydrate-binding proteins. X. nematophila NilC and its presumptive TXISS OMP partner NilB are required for colonizing the anterior intestine of Steinernema carpocapsae nematodes: the receptacle of free-living, infective juveniles and the anterior intestinal cecum (AIC) in juveniles and adults. We show that, in adult nematodes, the AIC expresses a Wheat Germ Agglutinin (WGA)-reactive material, indicating the presence of N-acetylglucosamine or N-acetylneuraminic acid sugars on the AIC surface. A role for this material in colonization is supported by the fact that exogenous addition of WGA can inhibit AIC colonization by X. nematophila. Conversely, the addition of exogenous purified NilC increases the frequency with which X. nematophila is observed at the AIC, demonstrating that abundant extracellular NilC can enhance colonization. NilC may facilitate X. nematophila adherence to the nematode intestinal surface by binding to host glycans, it might support X. nematophila nutrition by cleaving sugars from the host surface, or it might help protect X. nematophila from nematode host immunity. Proteomic and metabolomic analyses of wild type X. nematophila compared to those lacking nilB and nilC revealed differences in cell wall and secreted polysaccharide metabolic pathways. Additionally, purified NilC is capable of binding peptidoglycan, suggesting that periplasmic NilC may interact with the bacterial cell wall. Overall, these findings support a model that NilB-regulated surface exposure of NilC mediates interactions between X. nematophila and host surface glycans during colonization. This is a previously unknown function for a TXISS.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning-assisted elucidation of CD81–CD44 interactions in promoting cancer stemness and extracellular vesicle integrity

Tumor-initiating cells with reprogramming plasticity or stem-progenitor cell properties (stemness) are thought to be essential for cancer development and metastatic regeneration in many cancers; however, elucidation of the underlying molecular network and pathways remains demanding. Combining machine learning and experimental investigation, here we report CD81, a tetraspanin transmembrane protein known to be enriched in extracellular vesicles (EVs), as a newly identified driver of breast cancer stemness and metastasis. Using protein structure modeling and interface prediction-guided mutagenesis, we demonstrate that membrane CD81 interacts with CD44 through their extracellular regions in promoting tumor cell cluster formation and lung metastasis of triple negative breast cancer (TNBC) in human and mouse models. In-depth global and phosphoproteomic analyses of tumor cells deficient with CD81 or CD44 unveils endocytosis-related pathway alterations, leading to further identification of a quality-keeping role of CD44 and CD81 in EV secretion as well as in EV-associated stemness-promoting function. CD81 is coexpressed along with CD44 in human circulating tumor cells (CTCs) and enriched in clustered CTCs that promote cancer stemness and metastasis, supporting the clinical significance of CD81 in association with patient outcomes. Our study highlights machine learning as a powerful tool in facilitating the molecular understanding of new molecular targets in regulating stemness and metastasis of TNBC.

59 BASIC BIOLOGICAL SCIENCES↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Auto_PDI

Protein-DNA Interaction Workflow (PDI Workflow), a pipeline that focuses on generating high-quality docking and molecular dynamics simulations for Protein-DNA complexes. This allows us to take DNA sequences with unknown tertiary structures, accurately predict their structure, dock them with the desired target protein, and then simulate their interactions using molecular dynamics simulations

Kumar, Neeraj↗

BMC Caller: a webtool to identify and analyze bacterial microcompartment types in sequence data

Bacterial microcompartments (BMCs) are protein-based organelles found across the bacterial tree of life. They consist of a shell, made of proteins that oligomerize into hexagonally and pentagonally shaped building blocks, that surrounds enzymes constituting a segment of a metabolic pathway. The proteins of the shell are unique to BMCs. They also provide selective permeability; this selectivity is dictated by the requirements of their cargo enzymes. We have recently surveyed the wealth of different BMC types and their occurrence in all available genome sequence data by analyzing and categorizing their components found in chromosomal loci using HMM (Hidden Markov Model) protein profiles. To make this a “do-it yourself” analysis for the public we have devised a webserver, BMC Caller (https://bmc-caller.prl.msu.edu), that compares user input sequences to our HMM profiles, creates a BMC locus visualization, and defines the functional type of BMC, if known. Shell proteins in the input sequence data are also classified according to our function-agnostic naming system and there are links to similar proteins in our database as well as an external link to a structure prediction website to easily generate structural models of the shell proteins, which facilitates understanding permeability properties of the shell. Additionally, the BMC Caller website contains a wealth of information on previously analyzed BMC loci with links to detailed data for each BMC protein and phylogenetic information on the BMC shell proteins. Our tools greatly facilitate BMC type identification to provide the user information about the associated organism’s metabolism and enable discovery of new BMC types by providing a reference database of all currently known examples.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

Chimeric Plant Calcium/Calmodulin-Dependent Protein Kinase Gene with a Neural Visinin-Like Calcium-Binding Domain

Calcium, a universal second messenger, regulates diverse cellular processes in eukaryotes. Ca-2(+) and Ca-2(+)/calmodulin-regulated protein phosphorylation play a pivotal role in amplifying and diversifying the action of Ca-2(+)- mediated signals. A chimeric Ca-2(+)/calmodulin-dependent protein kinase (CCaMK) gene with a visinin-like Ca-2(+)- binding domain was cloned and characterized from lily. The cDNA clone contains an open reading frame coding for a protein of 520 amino acids. The predicted structure of CCaMK contains a catalytic domain followed by two regulatory domains, a calmodulin-binding domain and a visinin-like Ca-2(+)-binding domain. The amino-terminal region of CCaMK contains all 11 conserved subdomains characteristic of serine/threonine protein kinases. The calmodulin-binding region of CCaMK has high homology (79%) to alpha subunit of mammalian Ca-2(+)/calmodulin-dependent protein kinase. The calmodulin-binding region is fused to a neural visinin-like domain that contains three Ca-2(+)-binding EF-hand motifs and a biotin-binding site. The Escherichia coli-expressed protein (approx. 56 kDa) binds calmodulin in a Ca-2(+)-dependent manner. Furthermore, Ca-45-binding assays revealed that CCaMK directly binds Ca-2(+). The CCaMK gene is preferentially expressed in developing anthers. Southern blot analysis revealed that CCaMK is encoded by a single gene. The structural features of the gene suggest that it has multiple regulatory controls and could play a unique role in Ca-2(+) signaling in plants.

Patil, Shameekumar↗

Unusual modifications of protein biomarkers expressed by plasmid, prophage, and bacterial host of pathogenic Escherichia coli identified using top‐down proteomic analysis

Rationale Pathogenic bacteria often carry prophage (bacterial viruses) and plasmids (small circular pieces of DNA) that may harbor toxin, antibacterial, and antibiotic resistance genes. Proteomic characterization of pathogenic bacteria should include the identification of host proteins and proteins produced by prophage and plasmid genomes. Methods Protein biomarkers of two strains of Shiga toxin–producingEscherichia coli(STEC) were identified using antibiotic induction, matrix‐assisted laser desorption/ionization tandem time‐of‐flight (MALDI‐TOF‐TOF) tandem mass spectrometry (MS/MS) with post‐source decay (PSD), top‐down proteomic (TDP) analysis, and plasmid sequencing. Alphafold2 was also used to compare predicted in silico structures of the identified proteins to prominent fragment ions generated using MS/MS‐PSD. Strain samples were also analyzed with and without chemical reduction treatment to detect the attachment of pendant groups bound by thioester or disulfide bonds. Results Shiga toxin was detected and/or identified in both STEC strains. For the first time, we also identified the osmotically inducible protein (OsmY) whose sequence unexpectedly had two forms: a full and a truncated sequence. The truncated OsmY terminates in the middle of an α‐helix as determined by Alphafold2. A plasmid‐encoded colicin immunity protein was also identified with and without attachment of an unidentified cysteine‐bound pendant group (~307 Da). Plasmid sequencing confirmed top‐down analysis and the identification of a promoter upstream of the immunity gene that is activated by antibiotic induction, that is, SOS box. Conclusions TDP analysis, coupled with other techniques (e.g., antibiotic induction, chemical reduction, plasmid sequencing, and in silico protein modeling), is a powerful tool to identify proteins (and their modifications), including prophage‐ and plasmid‐encoded proteins, produced by pathogenic microorganisms.

Biochemistry & Molecular Biology↗