Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Transforming our understanding of chloroplast-associated genes through comprehensive characterization of protein localizations and protein-protein interactions

Bioenergy crops are a renewable source of fuels and are a critical base for building a carbon-neutral economy. Rational engineering of bioenergy crops has the potential to enhance the yields. However, our ability to engineer plants is limited because the functions of most genes remain unknown. Systematic characterization of gene function in plants thus has the potential to greatly accelerate bioenergy research. Here, we focus on the chloroplast, an underexplored energy-producing organelle that is a hallmark of plants. The chloroplast is one of the promising targets of biofuel crop engineering efforts because of its central role in photosynthesis, metabolism, and intracellular signaling. However, the protein composition of the chloroplast and the functions of most of its proteins remain poorly characterized. At the core of this project, we sought to comprehensively determine the localization of chloroplast-associated proteins and generate a spatially defined protein-protein interaction network for chloroplast. For this purpose, we used the leading model alga Chlamydomonas reinhardtii, which greatly increased experimental speed and throughput. We illustrated the value of our findings to land plants by determining the localization of Arabidopsis thaliana land plant homologs of the Chlamydomonas proteins. Altogether, we were successful in determining the localization of 1,034 chloroplast-associated proteins in Chlamydomonas. The localizations provide numerous insights into the spatial organization of chloroplasts and how they function to support photosynthesis. The localization patterns of distinct proteins revealed new chloroplast structures and revealed new spatial organization inside the chloroplast. We also identified new components of known chloroplast structures, such as the chloroplast envelope, nucleoid, plastoglobuli, and pyrenoid. We identified these new components by investigating the interacting partners of known proteins. Many proteins localized in both the chloroplast and other cellular structures, thereby hinting at new functions and communication between cellular structures. We also applied machine learning on the atlas to generate predictions for the location of all of the proteins in Chlamydomonas. This enabled us to assign putative functions to many uncharacterized proteins based on their cellular location. Altogether, this research establishes a rich resource that opens new avenues of investigation and guides future work in deciphering and manipulating chloroplast function. Next, we developed an extensive protein-protein interaction network for the chloroplast by performing affinity purification-mass spectrometry on ~1,150 tagged chloroplast-associated proteins, the first such large-scale study in any photosynthetic organism. This dataset reveals 4,694 high-confidence protein-protein interactions, offering insights into the functions of thousands of conserved poorly-characterized chloroplast proteins. This systematic identification of protein-protein interactions in the chloroplast also provides multiple exciting new research directions and a detailed blueprint of the chloroplast's operation. This research lays the groundwork to decipher the inner workings of the chloroplast, the cell structure at the heart of photosynthesis. The spatial atlas and protein-protein interactions reveal chloroplast organizational features that would not have been accessible with traditional approaches. The localization mapping, insights into the function, and research materials generated further provide a rich resource for the research community to advance the understanding of how the chloroplast is organized to enable engineering of enhanced photosynthetic organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

Sequence and structural implications of a bovine corneal keratan sulfate proteoglycan core protein. Protein 37B represents bovine lumican and proteins 37A and 25 are unique

Amino acid sequence from tryptic peptides of three different bovine corneal keratan sulfate proteoglycan (KSPG) core proteins (designated 37A, 37B, and 25) showed similarities to the sequence of a chicken KSPG core protein lumican. Bovine lumican cDNA was isolated from a bovine corneal expression library by screening with chicken lumican cDNA. The bovine cDNA codes for a 342-amino acid protein, M(r) 38,712, containing amino acid sequences identified in the 37B KSPG core protein. The bovine lumican is 68% identical to chicken lumican, with an 83% identity excluding the N-terminal 40 amino acids. Location of 6 cysteine and 4 consensus N-glycosylation sites in the bovine sequence were identical to those in chicken lumican. Bovine lumican had about 50% identity to bovine fibromodulin and 20% identity to bovine decorin and biglycan. About two-thirds of the lumican protein consists of a series of 10 amino acid leucine-rich repeats that occur in regions of calculated high beta-hydrophobic moment, suggesting that the leucine-rich repeats contribute to beta-sheet formation in these proteins. Sequences obtained from 37A and 25 core proteins were absent in bovine lumican, thus predicting a unique primary structure and separate mRNA for each of the three bovine KSPG core proteins.

NASA Discipline Cell Biology↗

Inhibitor-3 inhibits Protein Phosphatase 1 via a metal binding dynamic protein–protein interaction

To achieve substrate specificity, protein phosphate 1 (PP1) forms holoenzymes with hundreds of regulatory and inhibitory proteins. Inhibitor-3 (I3) is an ancient inhibitor of PP1 with putative roles in PP1 maturation and the regulation of PP1 activity. Here, we show that I3 residues 27–68 are necessary and sufficient for PP1 binding and inhibition. In addition to a canonical RVxF motif, which is shared by nearly all PP1 regulators and inhibitors, and a non-canonical SILK motif, I3 also binds PP1 via multiple basic residues that bind directly in the PP1 acidic substrate binding groove, an interaction that provides a blueprint for how substrates bind this groove for dephosphorylation. Unexpectedly, this interaction positions a CCC (cys-cys-cys) motif to bind directly across the PP1 active site. Using biophysical and inhibition assays, we show that the I3 CCC motif binds and inhibits PP1 in an unexpected dynamic, fuzzy manner, via transient engagement of the PP1 active site metals. Together, these data not only provide fundamental insights into the mechanisms by which IDP protein regulators of PP1 achieve inhibition, but also shows that fuzzy interactions between IDPs and their folded binding partners, in addition to enhancing binding affinity, can also directly regulate enzyme activity.

59 BASIC BIOLOGICAL SCIENCES↗

A multiplexed bacterial two-hybrid for rapid characterization of protein–protein interactions and iterative protein design

Protein-protein interactions (PPIs) are crucial for biological functions and have applications ranging from drug design to synthetic cell circuits. Coiled-coils have been used as a model to study the sequence determinants of specificity. However, building well-behaved sets of orthogonal pairs of coiled-coils remains challenging due to inaccurate predictions of orthogonality and difficulties in testing at scale. To address this, we develop the next-generation bacterial two-hybrid (NGB2H) method, which allows for the rapid exploration of interactions of programmed protein libraries in a quantitative and scalable way using next-generation sequencing readout. We design, build, and test large sets of orthogonal synthetic coiled-coils, assayed over 8,000 PPIs, and used the dataset to train a more accurate coiled-coil scoring algorithm (iCipa). After characterizing nearly 18,000 new PPIs, we identify to the best of our knowledge the largest set of orthogonal coiled-coils to date, with fifteen on-target interactions. Our approach provides a powerful tool for the design of orthogonal PPIs.

59 BASIC BIOLOGICAL SCIENCES↗

Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function

Abstract Motivation Millions of protein sequences have been generated by numerous genome and transcriptome sequencing projects. However, experimentally determining the function of the proteins is still a time consuming, low-throughput, and expensive process, leading to a large protein sequence-function gap. Therefore, it is important to develop computational methods to accurately predict protein function to fill the gap. Even though many methods have been developed to use protein sequences as input to predict function, much fewer methods leverage protein structures in protein function prediction because there was lack of accurate protein structures for most proteins until recently. Results We developed TransFun—a method using a transformer-based protein language model and 3D-equivariant graph neural networks to distill information from both protein sequences and structures to predict protein function. It extracts feature embeddings from protein sequences using a pre-trained protein language model (ESM) via transfer learning and combines them with 3D structures of proteins predicted by AlphaFold2 through equivariant graph neural networks. Benchmarked on the CAFA3 test dataset and a new test dataset, TransFun outperforms several state-of-the-art methods, indicating that the language model and 3D-equivariant graph neural networks are effective methods to leverage protein sequences and structures to improve protein function prediction. Combining TransFun predictions and sequence similarity-based predictions can further increase prediction accuracy. Availability and implementation The source code of TransFun is available at https://github.com/jianlin-cheng/TransFun.

59 BASIC BIOLOGICAL SCIENCES↗

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

Engineering a new tripartite split-ccGFP system from Corynactis californica for detecting protein–protein interactions

Protein-protein interactions (PPIs) are critical to a range of biological processes and, consequently, aberrant interactions are implicated in many disorders. The study of the complex networks of PPIs promises to elucidate undiscovered roles in cellular processes and the mechanisms of disease. To accomplish this, tools to effectively sense PPIs are necessary. Effective PPI sensors must rapidly detect interactions in real-time with high sensitivity without perturbing the proteins of interest (POIs) under study. Split fluorescent proteins have previously been used to successfully monitor PPIs, in part due to the small size of the tags. Here, we developed an optimized tripartite split GFP system based on Corynactis californica GFP (ccGFP) to detect PPIs in vitro. In this sensor system, ccGFP fragments ccGFP10 and ccGFP11 are tagged to two POIs. PPIs can then be detected via fluorescence by complementation to the third fragment, ccGFP1-9, which reconstitutes functional ccGFP. The optimized ccGFP system shows improved detection kinetics and pH and temperature stability compared to a previous system. We then validated the sensor by monitoring PPIs in two model systems: attractive/repulsive coiled-coils and rapamycin-inducible FRB/FKBP heterodimerization. Finally, we developed an anti-tripartite ccGFP single-chain variable fragment (scFv), which could enable versatile detection of identified protein-protein complexes.

59 BASIC BIOLOGICAL SCIENCES↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

An Arabidopsis Ran-binding protein, AtRanBP1c, is a co-activator of Ran GTPase-activating protein and requires the C-terminus for its cytoplasmic localization

Ran-binding proteins (RanBPs) are a group of proteins that bind to Ran (Ras-related nuclear small GTP-binding protein), and thus either control the GTP/GDP-bound states of Ran or help couple the Ran GTPase cycle to a cellular process. AtRanBP1c is a Ran-binding protein from Arabidopsis thaliana (L.) Heynh. that was recently shown to be critically involved in the regulation of auxin-induced mitotic progression [S.-H. Kim et al. (2001) Plant Cell 13:2619-2630]. Here we report that AtRanBP1c inhibits the EDTA-induced release of GTP from Ran and serves as a co-activator of Ran-GTPase-activating protein (RanGAP) in vitro. Transient expression of AtRanBP1c fused to a beta-glucuronidase (GUS) reporter reveals that the protein localizes primarily to the cytosol. Neither the N- nor C-terminus of AtRanBP1c, which flank the Ran-binding domain (RanBD), is necessary for the binding of PsRan1-GTP to the protein, but both are needed for the cytosolic localization of GUS-fused AtRanBP1c. These findings, together with a previous report that AtRanBP1c is critically involved in root growth and development, imply that the promotion of GTP hydrolysis by the Ran/RanGAP/AtRanBP1c complex in the cytoplasm, and the resulting concentration gradient of Ran-GDP to Ran-GTP across the nuclear membrane could be important in the regulation of auxin-induced mitotic progression in root tips of A. thaliana.

NASA Discipline Plant Biology↗

Suppression of muscle protein turnover and amino acid degradation by dietary protein deficiency

To define the adaptations that conserve amino acids and muscle protein when dietary protein intake is inadequate, rats (60-70 g final wt) were fed a normal or protein-deficient (PD) diet (18 or 1% lactalbumin), and their muscles were studied in vitro. After 7 days on the PD diet, both protein degradation and synthesis fell 30-40% in skeletal muscles and atria. This fall in proteolysis did not result from reduced amino acid supply to the muscle and preceded any clear decrease in plasma amino acids. Oxidation of branched-chain amino acids, glutamine and alanine synthesis, and uptake of alpha-aminoisobutyrate also fell by 30-50% in muscles and adipose tissue of PD rats. After 1 day on the PD diet, muscle protein synthesis and amino acid uptake decreased by 25-40%, and after 3 days proteolysis and leucine oxidation fell 30-45%. Upon refeeding with the normal diet, protein synthesis also rose more rapidly (+30% by 1 day) than proteolysis, which increased significantly after 3 days (+60%). These different time courses suggest distinct endocrine signals for these responses. The high rate of protein synthesis and low rate of proteolysis during the first 3 days of refeeding a normal diet to PD rats contributes to the rapid weight gain ("catch-up growth") of such animals.

Non-NASA Center↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Soft matter approach for creating novel protein hydrogels using fractal whey protein assemblies as building blocks

In this study, mesoscopic sized fractal assembly (FA) particles were prepared using whey protein isolate; the coldset gelation properties of FA particles were investigated in-depth. Two types of FA particles (FA -62 and FA -90) with different mean sizes were synthesized through controlled thermal treatment of whey protein solutions at two concentrations (62 g/L and 90 g/L). Particle characteristics e.g. hydrodynamic radius, ζ-potential and surface hydrophobicity were dependent on pH and structure of FA. Transmission electron microscopy (TEM) observation confirmed the fractal morphology and small-angle X-ray scattering (SAXS) analysis suggested an internal fractal structure for the obtained FA particles. Eight cold-set FA protein gels (2 % w/v) were manufactured by controlling two gelling factors at two levels: pH (5.8 and 7.0) and Ca 2+ content (5 mM and 10 mM). Rheological characteristics in the large amplitude oscillatory shear regime revealed that pH 7.0 gels were softer and elastic while pH 5.8 gels were harder and brittle. Rheology Pipkin diagrams demonstrated that the strain softening/stiffening and the shear-thinning/thickening behaviors may be fine-tuned by manipulating the key gelation factors: e.g. FA structure, pH, and Ca 2+ . The entrainment speed-dependent friction coefficient curves showed that at an intermediate velocity regime (6-250 mm s -1 ), FA -90 particles induced hydrogels had superior lubrication effect compared to FA -62 gels. Further, this work demonstrated a food structure design approach regarding tuning texture and lubrication properties of protein gels without changing protein content and protein composition. The optimized protein hydrogels may be used as texturizer for "cleaner label" food formulas and/or as delivery system for carrying micronutrients.

60 APPLIED LIFE SCIENCES↗

Defense against phytopathogens relies on efficient antimicrobial protein secretion mediated by the microtubule-binding protein TGNap1

Plant immunity depends on the secretion of antimicrobial proteins, which occurs through yet-largely unknown mechanisms. The trans-Golgi network (TGN), a hub for intracellular and extracellular trafficking pathways, and the cytoskeleton, which is required for antimicrobial protein secretion, are emerging as pathogen targets to dampen plant immunity. In this work, we demonstrate that tgnap1-2, a loss-of-function mutant of Arabidopsis TGNap1, a TGN-associated and microtubule (MT)-binding protein, is susceptible to Pseudomonas syringae (Pst DC3000). Pst DC3000 infected tgnap1-2 is capable of mobilizing defense pathways, accumulating salicylic acid (SA), and expressing antimicrobial proteins. The susceptibility of tgnap1-2 is due to a failure to efficiently transport antimicrobial proteins to the apoplast in a partially MT-dependent pathway but independent from SA and is additive to the pathogen-antagonizing MIN7, a TGN-associated ARF-GEF protein. Therefore, our data demonstrate that plant immunity relies on TGNap1 for secretion of antimicrobial proteins, and that TGNap1 is a key immunity element that functionally links secretion and cytoskeleton in SA-independent pathogen responses.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting compatibility between ferredoxins and the Fe protein of nitrogenase using in silico protein modeling

Biological nitrogen fixation is the process by which certain bacteria and archaea use the enzyme nitrogenase to reduce atmospheric nitrogen into bioavailable ammonium. Engineering non‐nitrogen‐fixing organisms, like plants, to use nitrogenase could reduce dependency on synthetic fertilizer and mitigate the environmental impacts of industrial fertilizer production. However, nitrogenase activity requires delivery of reducing power by small electron carrying proteins known as ferredoxins and flavodoxins, and successfully engineering nitrogenase into new systems will require a mechanistic understanding of electron delivery by these proteins. Most organisms often have multiple ferredoxins, raising the question of which ferredoxin can support nitrogenase activity. The purpose of this study is to gain insight into how we can predict which ferredoxin is compatible with the Fe protein, the component of nitrogenase that interacts with ferredoxin or flavodoxin. Our in silico protein–protein docking simulations reveal that most ferredoxins and flavodoxins involved in nitrogen fixation have the shortest distance (≤10 Å) between their redox cofactor and the [4Fe‐4S] cluster of the Fe protein. We found shorter cofactor distance contributes to faster intermolecular electron tunneling rates. Bacterial ferredoxins that play a role in nitrogen fixation also exhibit more complementary interactions with the Fe protein than bacterial and plant ferredoxins not involved in this process. Heterologous expression of a set of ferredoxins from both nitrogen‐fixing and non‐nitrogen‐fixing bacteria in the diazotroph Rhodopseudomonas palustris supports our model‐derived prediction that shorter distances between the electron‐carrying cofactors favor nitrogenase compatibility. These findings offer a framework to predict and potentially enhance ferredoxin–nitrogenase compatibility, which will help to improve our ability to engineer nitrogen fixation into non‐nitrogen‐fixing organisms like plants.

59 BASIC BIOLOGICAL SCIENCES↗

Protein folds vs. protein folding: Differing questions, different challenges

We report protein fold prediction using deep-learning artificial intelligence (AI) has transformed the field of protein structure prediction. By combining physical and geometric constraints—and especially patterns extracted from the Protein Data Bank —these machine learning algorithms can predict protein structures at or near atomic resolution and do so in seconds. Today, these computational methods have now solved more than 200 million protein structures, which are accessible from the AlphaFold Protein Structure Database. This accomplishment seems all the more remarkable because few thought it possible or saw it coming. Deservedly, deep-learning AI was named Science magazine’s 2021 “breakthrough of the year”. Clearly, deep-learning AI represents a major advance in protein fold prediction.

54 ENVIRONMENTAL SCIENCES↗

Deploying synthetic coevolution and machine learning to engineer protein-protein interactions

Fine-tuning of protein-protein interactions occurs naturally through coevolution, but this process is difficult to recapitulate in the laboratory. We describe a platform for synthetic protein-protein coevolution that can isolate matched pairs of interacting muteins from complex libraries. This large dataset of coevolved complexes drove a systems-level analysis of molecular recognition between Z domain–affibody pairs spanning a wide range of structures, affinities, cross-reactivities, and orthogonalities, and captured a broad spectrum of coevolutionary networks. Furthermore, we harnessed pretrained protein language models to expand, in silico, the amino acid diversity of our coevolution screen, predicting remodeled interfaces beyond the reach of the experimental library. Further, the integration of these approaches provides a means of simulating protein coevolution and generating protein complexes with diverse molecular recognition properties for biotechnology and synthetic biology.

59 BASIC BIOLOGICAL SCIENCES↗