Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Structural- and Functional-Informed Machine Learning for Protein Function Prediction

In this project we aimed to extend methods for protein function prediction to include structural prediction data, and benchmark methods against existing tools. We proposed to apply the method to large metagenome datasets, and develop approaches to examine activity-based protein profiling results for protein function-structure patterns. Nitrogen cycle protein families were previously identified and are used here to provide a proof-of-principle for use of structure prediction in protein function classification.

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of buttressed loops for sculpting protein functions

In natural proteins, structured loops have central roles in molecular recognition, signal transduction and enzyme catalysis. However, because of the intrinsic flexibility and irregularity of loop regions, organizing multiple structured loops at protein functional sites has been very difficult to achieve by de novo protein design. Here we describe a solution to this problem that designs tandem repeat proteins with structured loops (9–14 residues) buttressed by extensive hydrogen bonding interactions. Experimental characterization shows that the designs are monodisperse, highly soluble, folded and thermally stable. Crystal structures are in close agreement with the design models, with the loops structured and buttressed as designed. We demonstrate the functionality afforded by loop buttressing by designing and characterizing binders for extended peptides in which the loops form one side of an extended binding pocket. The ability to design multiple structured loops should contribute generally to efforts to design new protein functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function

Abstract Motivation Millions of protein sequences have been generated by numerous genome and transcriptome sequencing projects. However, experimentally determining the function of the proteins is still a time consuming, low-throughput, and expensive process, leading to a large protein sequence-function gap. Therefore, it is important to develop computational methods to accurately predict protein function to fill the gap. Even though many methods have been developed to use protein sequences as input to predict function, much fewer methods leverage protein structures in protein function prediction because there was lack of accurate protein structures for most proteins until recently. Results We developed TransFun—a method using a transformer-based protein language model and 3D-equivariant graph neural networks to distill information from both protein sequences and structures to predict protein function. It extracts feature embeddings from protein sequences using a pre-trained protein language model (ESM) via transfer learning and combines them with 3D structures of proteins predicted by AlphaFold2 through equivariant graph neural networks. Benchmarked on the CAFA3 test dataset and a new test dataset, TransFun outperforms several state-of-the-art methods, indicating that the language model and 3D-equivariant graph neural networks are effective methods to leverage protein sequences and structures to improve protein function prediction. Combining TransFun predictions and sequence similarity-based predictions can further increase prediction accuracy. Availability and implementation The source code of TransFun is available at https://github.com/jianlin-cheng/TransFun.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism↗

Production of functional proteins: balance of shear stress and gravity

A method for the production of functional proteins including hormones by renal cells in a three dimensional culturing process responsive to shear stress uses a rotating wall vessel. Natural mixture of renal cells expresses the enzyme 1-.alpha.-hydroxylase which can be used to generate the active form of vitamin D: 1,25-diOH vitamin D.sub.3. The fibroblast cultures and co-culture of renal cortical cells express the gene for erythropoietin and secrete erythropoietin into the culture supernatant. Other shear stress response genes are also modulated by shear stress, such as toxin receptors megalin and cubulin (gp280). Also provided is a method of treating an in-need individual with the functional proteins produced in a three dimensional co-culture process responsive to shear stress using a rotating wall vessel.

Goodwin, Thomas John↗

Origins of Protein Functions in Cells

In modern organisms proteins perform a majority of cellular functions, such as chemical catalysis, energy transduction and transport of material across cell walls. Although great strides have been made towards understanding protein evolution, a meaningful extrapolation from contemporary proteins to their earliest ancestors is virtually impossible. In an alternative approach, the origin of water-soluble proteins was probed through the synthesis and in vitro evolution of very large libraries of random amino acid sequences. In combination with computer modeling and simulations, these experiments allow us to address a number of fundamental questions about the origins of proteins. Can functionality emerge from random sequences of proteins? How did the initial repertoire of functional proteins diversify to facilitate new functions? Did this diversification proceed primarily through drawing novel functionalities from random sequences or through evolution of already existing proto-enzymes? Did protein evolution start from a pool of proteins defined by a frozen accident and other collections of proteins could start a different evolutionary pathway? Although we do not have definitive answers to these questions yet, important clues have been uncovered. In one example (Keefe and Szostak, 2001), novel ATP binding proteins were identified that appear to be unrelated in both sequence and structure to any known ATP binding proteins. One of these proteins was subsequently redesigned computationally to bind GTP through introducing several mutations that introduce targeted structural changes to the protein, improve its binding to guanine and prevent water from accessing the active center. This study facilitates further investigations of individual evolutionary steps that lead to a change of function in primordial proteins. In a second study (Seelig and Szostak, 2007), novel enzymes were generated that can join two pieces of RNA in a reaction for which no natural enzymes are known. Recently it was found that, as in the previous case, the proteins have a structure unknown among modern enzymes. In this case, in vitro evolution started from a small, non-enzymatic protein. A similar selection process initiated from a library of random polypeptides is in progress. These results not only allow for estimating the occurrence of function in random protein assemblies but also provide evidence for the possibility of alternative protein worlds. Extant proteins might simply represent a frozen accident in the world of possible proteins. Alternative collections of proteins, even with similar functions, could originate alternative evolutionary paths.

Seelig, Burchard↗

Production of functional proteins: balance of shear stress and gravity

The present invention provides a method for production of functional proteins including hormones by renal cells in a three dimensional co-culture process responsive to shear stress using a rotating wall vessel. Natural mixture of renal cells expresses the enzyme 1-a-hydroxylase which can be used to generate the active form of vitamin D: 1,25-diOH vitamin D3. The fibroblast cultures and co-culture of renal cortical cells express the gene for erythropoietin and secrete erythropoietin into the culture supernatant. Other shear stress response genes are also modulated by shear stress, such as toxin receptors megalin and cubulin (gp280). Also provided is a method of treating in-need individual with the functional proteins produced in a three dimensional co-culture process responsive to shear stress using a rotating wall vessel.

Goodwin, Thomas John↗

Production of functional proteins: balance of shear stress and gravity

The present invention provides a method for production of functional proteins including hormones by renal cells in a three dimensional co-culture process responsive to shear stress using a rotating wall vessel. Natural mixture of renal cells expresses the enzyme 1-a-hydroxylase which can be used to generate the active form of vitamin D: 1,25-diOH vitamin D3. The fibroblast cultures and co-culture of renal cortical cells express the gene for erythropoietin and secrete erythropoietin into the culture supernatant. Other shear stress response genes are also modulated by shear stress, such as toxin receptors megalin and cubulin (gp280). Also provided is a method of treating in-need individual with the functional proteins produced in a three dimensional co-culture process responsive to shear stress using a rotating wall vessel.

Goodwin, Thomas John↗

Emergence of Complexity in Protein Functions and Metabolic Networks

In modern organisms proteins perform a majority of cellular functions, such as chemical catalysis, energy transduction and transport of material across cell walls. Although great strides have been made towards understanding protein evolution, a meaningful extrapolation from contemporary proteins to their earliest ancestors is virtually impossible. In an alternative approach, the origin of water-soluble proteins was probed through the synthesis of very large libraries of random amino acid sequences and subsequently subjecting them to in vitro evolution. In combination with computer modeling and simulations, these experiments allow us to address a number of fundamental questions about the origins of proteins. Can functionality emerge from random sequences of proteins? How did the initial repertoire of functional proteins diversify to facilitate new functions? Did this diversification proceed primarily through drawing novel functionalities from random sequences or through evolution of already existing proto-enzymes? Did protein evolution start from a pool of proteins defined by a frozen accident and other collections of proteins could start a different evolutionary pathway? Although we do not have definitive answers to these questions, important clues have been uncovered. Considerable progress has been also achieved in understanding the origins of membrane proteins. We will address this issue in the example of ion channels - proteins that mediate transport of ions across cell walls. Remarkably, despite overall complexity of these proteins in contemporary cells, their structural motifs are quite simple, with -helices being most common. By combining results of experimental and computer simulation studies on synthetic models and simple, natural channels, I will show that, even though architectures of membrane proteins are not nearly as diverse as those of water-soluble proteins, they are sufficiently flexible to adapt readily to the functional demands arising during evolution.

Pohorille, Andzej↗

Modelling protein functional domains in signal transduction using Maude

Modelling of protein-protein interactions in signal transduction is receiving increased attention in computational biology. This paper describes recent research in the application of Maude, a symbolic language founded on rewriting logic, to the modelling of functional domains within signalling proteins. Protein functional domains (PFDs) are a critical focus of modern signal transduction research. In general, Maude models can simulate biological signalling networks and produce specific testable hypotheses at various levels of abstraction. Developing symbolic models of signalling proteins containing functional domains is important because of the potential to generate analyses of complex signalling networks based on structure-function relationships.

Signal Transduction↗

Large language models generate functional protein sequences across diverse families

Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here, in this paper, we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

NAI Supported Research, a Case Study: the Origin of Protein Functions

In the 20 years of its existence, the NAI supported a host of studies that fundamentally advanced the field of astrobiology in terms of both understanding the underlying concepts and developing novel research techniques. In this presentation, I will recount one such case in the field of the origin of life, whereby application of new, powerful experimental techniques combined with state-of-the art computer simulations led to a surprising insight into the origin of protein functions.

Pohorille, Andrew↗

Watching a signaling protein function: What has been learned over four decades of time-resolved studies of photoactive yellow protein

Photoactive yellow protein (PYP) is a signaling protein whose internal p-coumaric acid chromophore undergoes reversible, light-induced trans-to-cis isomerization, which triggers a sequence of structural changes that ultimately lead to a signaling state. Since its discovery nearly 40 years ago, PYP has attracted much interest and has become one of the most extensively studied proteins found in nature. The method of time-resolved crystallography, pioneered by Keith Moffat, has successfully characterized intermediates in the PYP photocycle at near atomic resolution over 12 decades of time down to the sub-picosecond time scale, allowing one to stitch together a movie and literally watch a protein as it functions. But how close to reality is this movie? To address this question, results from numerous complementary time-resolved techniques including x-ray crystallography, x-ray scattering, and spectroscopy are discussed. Emerging from spectroscopic studies is a general consensus that three time constants are required to model the excited state relaxation, with a highly strained ground-state cis intermediate formed in less than 2.4 ps. Persistent strain drives the sequence of structural transitions that ultimately produce the signaling state. Crystal packing forces produce a restoring force that slows somewhat the rates of interconversion between the intermediates. Moreover, the solvent composition surrounding PYP can influence the number and structures of intermediates as well as the rates at which they interconvert. When chloride is present, the PYP photocycle in a crystal closely tracks that in solution, which suggests the epic movie of the PYP photocycle is indeed based in reality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Monomer-scale design of functional protein polymers using consensus repeat sequences

Protein-based polymers possess chemically defined sequences that can encode diverse properties and functions into a new class of biopolymeric materials. However, sequence variation that emerges from evolution can obscure the sequence–function relationships of naturally derived polymers. One strategy to clarify these relationships is to identify common sequences between proteins with similar functions. These conserved sequences often emerge from repeat proteins, and “consensus repeat sequences” provide a convenient platform for systematic investigations of biopolymer sequence–property relationships. In this review, we highlight recent approaches to engineer tunable polymeric materials using monomer-scale design of consensus repeat proteins. Here, we explore established and emerging protein-based materials with mechanical resilience, thermodynamic phase behavior, chemical responsiveness, biomolecular transport, and hierarchical structure. Overall, recent advances in the monomer-scale design of repetitive protein polymers present exciting fundamental and translational opportunities for polymer scientists and engineers.

36 MATERIALS SCIENCE↗

Controlling Mineralization with Protein–Functionalized Peptoid Nanotubes

Sequence-defined foldamers that self-assemble into well-defined architectures are promising scaffolds to template inorganic mineralization. However, it has been challenging to achieve robust control of nucleation and growth without sequence redesign or extensive experimentation. Here, peptoid nanotubes functionalized with a panel of solid-binding proteins are used to mineralize homogeneously distributed and monodisperse anatase nanocrystals from the water-soluble TiBALDH precursor. Crystallite size is systematically tuned between 1.4 and 4.4 nm by changing protein coverage and the identity and valency of the genetically engineered solid-binding segments. The approach is extended to the synthesis of gold nanoparticles and, using a protein encoding both material-binding specificities, to the fabrication of titania/gold nanocomposites capable of photocatalysis under visible-light illumination. Here, beyond uncovering critical roles for hierarchical organization and denticity on solid-binding protein mineralization outcomes, the strategy described herein should prove valuable for the fabrication of hierarchical hybrid materials incorporating a broad range of inorganic components.

36 MATERIALS SCIENCE↗

ER-associated VAP27-1 and VAP27-3 proteins functionally link the lipid-binding ORP2A at the ER-chloroplast contact sites

Abstract The plant endoplasmic reticulum (ER) contacts heterotypic membranes at membrane contact sites (MCSs) through largely undefined mechanisms. For instance, despite the well-established and essential role of the plant ER-chloroplast interactions for lipid biosynthesis, and the reported existence of physical contacts between these organelles, almost nothing is known about the ER-chloroplast MCS identity. Here we show that the Arabidopsis ER membrane-associated VAP27 proteins and the lipid-binding protein ORP2A define a functional complex at the ER-chloroplast MCSs. Specifically, through in vivo and in vitro association assays, we found that VAP27 proteins interact with the outer envelope membrane (OEM) of chloroplasts, where they bind to ORP2A. Through lipidomic analyses, we established that VAP27 proteins and ORP2A directly interact with the chloroplast OEM monogalactosyldiacylglycerol (MGDG), and we demonstrated that the loss of the VAP27-ORP2A complex is accompanied by subtle changes in the acyl composition of MGDG and PG. We also found that ORP2A interacts with phytosterols and established that the loss of the VAP27-ORP2A complex alters sterol levels in chloroplasts. We propose that, by interacting directly with OEM lipids, the VAP27-ORP2A complex defines plant-unique MCSs that bridge ER and chloroplasts and are involved in chloroplast lipid homeostasis.

59 BASIC BIOLOGICAL SCIENCES↗