Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism

Predicting functional divergence in protein evolution by site-specific rate shifts

Most modern tools that analyze protein evolution allow individual sites to mutate at constant rates over the history of the protein family. However, Walter Fitch observed in the 1970s that, if a protein changes its function, the mutability of individual sites might also change. This observation is captured in the "non-homogeneous gamma model", which extracts functional information from gene families by examining the different rates at which individual sites evolve. This model has recently been coupled with structural and molecular biology to identify sites that are likely to be involved in changing function within the gene family. Applying this to multiple gene families highlights the widespread divergence of functional behavior among proteins to generate paralogs and orthologs.

Review

Single residue substitutions that change the gating properties of a mechanosensitive channel in Escherichia coli

MscL is a channel that opens a large pore in the Escherichia coli cytoplasmic membrane in response to mechanical stress. Previously, we highly enriched the MscL protein by using patch clamp as a functional assay and cloned the corresponding gene. The predicted protein contains a largely hydrophobic core spanning two-thirds of the molecule and a more hydrophilic carboxyl terminal tail. Because MscL had no homology to characterized proteins, it was impossible to predict functional regions of the protein by simple inspection. Here, by mutagenesis, we have searched for functionally important regions of this molecule. We show that a short deletion from the amino terminus (3 amino acids), and a larger deletion of 27 amino acids from the carboxyl terminus of this protein, had little if any effect in channel properties. We have thus narrowed the search of the core mechanosensitive mechanism to 106 residues of this 136-amino acid protein. In contrast, single residue substitutions of a lysine in the putative first transmembrane domain or a glutamine in the periplasmic loop caused pronounced shifts in the mechano-sensitivity curves and/or large changes in the kinetics of channel gating, suggesting that the conformational structure in these regions is critical for normal mechanosensitive channel gating.

Non-NASA Center

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran

Simple Math is Enough: Two Examples of Inferring Functional Associations from Genomic Data

Non-random features in the genomic data are usually biologically meaningful. The key is to choose the feature well. Having a p-value based score prioritizes the findings. If two proteins share a unusually large number of common interaction partners, they tend to be involved in the same biological process. We used this finding to predict the functions of 81 un-annotated proteins in yeast.

Liang, Shoudan

Maize Rough Endosperm6 (rgh6) Encodes A Predicted Dead-Box RNA Helicase and Affects Mirna Processing in Endosperm Development

Maize rough endosperm (rgh) mutants have defective kernels with a rough, etched, or pitted endosperm surface. Molecular genetic analysis of this mutant class has identified multiple RNA processing proteins critical to endosperm development. Here, we report on the developmental and molecular function of the rgh6 locus. The rgh6 mutant was isolated from the UniformMu transposon tagging population. Mutant kernels have reduced endosperm size and defective embryos that develop in a more apical position than typical for defective embryos. TB translocation crosses revealed that rgh6 mutant endosperm inhibits normal embryo development. Positional cloning of the rgh6 locus found that it encodes a predicted DEAD-box RNA helicase. Consistent with a predicted function for RNA processing, transient expression of a RGH6-GFP fusion protein is localized to nucleolus and nuclear speckles in Nicotiana benthamiana leaves. Rgh6 transcripts are highly expressed in endosperm epidermal cell types such as the aleurone, basal endosperm cell layer, embryo surrounding region, and endosperm adjacent to scutellum. Markers of these cell types show increased levels in rgh6 mutant kernels. Mutant endosperm tissues have increased precursor microRNA (pre-miRNA) and decreased mature miRNA relative to normal sibling endosperm, indicating that rgh6 is required for miRNA processing. The transcript levels for most miRNA target genes accumulate to a higher level in rgh6 mutant tissue. These results suggest that miRNA processing and regulation of miRNA target genes are required for normal endosperm development.

Plant Sciences

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost

Comparison of Two Bioinformatics Tools Used to Characterize the Microbial Diversity and Predictive Functional Attributes of Microbial Mats from Lake Obersee, Antarctica

In this study, using NextGen sequencing of the collective 16S rRNA genes obtained from two sets of samples collected from Lake Obersee, Antarctica, we compared and contrasted two bioinformatics tools, PICRUSt and Tax4Fun. We then developed an R script to assess the taxonomic and predictive functional profiles of the microbial communities within the samples. Taxa such as Pseudoxanthomonas, Planctomycetaceae, Cyanobacteria Subsection III, Nitrosomonadaceae, Leptothrix, and Rhodobacter were exclusively identified by Tax4Fun that uses SILVA database; whereas PICRUSt that uses Greengenes database uniquely identified Pirellulaceae, Gemmatimonadetes A1-B1, Pseudanabaena, Salinibacterium and Sinibacteraceae. Predictive functional profiling of the microbial communities using Tax4Fun and PICRUSt separately revealed common metabolic capabilities, while also showing specific functional IDs not shared between the two approaches. Combining these functional predictions using a customized R script revealed a more inclusive metabolic profile, such as hydrolases, oxidoreductases, transferases; enzymes involved in carbohydrate and amino acid metabolisms; and membrane transport proteins known for nutrient uptake from the surrounding environment. Our results present the first molecular-phylogenetic characterization and predictive functional profiles of the microbial mat communities in Lake Obersee, while demonstrating the efficacy of combining both the taxonomic assignment information and functional IDs using the R script created in this study for a more streamlined evaluation of predictive functional profiles of microbial communities.

Hyunmin Koo

Co- and/or post-translational modifications are critical for TCH4 XET activity

TCH4 encodes a xyloglucan endotransglycosylase (XET) of Arabidopsis thaliana. XETs endolytically cleave and religate xyloglucan polymers; xyloglucan is one of the primary structural components of the plant cell wall. Therefore, XET function may affect cell shape and plant morphogenesis. To gain insight into the biochemical function of TCH4, we defined structural requirements for optimal XET activity. Recombinant baculoviruses were designed to produce distinct forms of TCH4. TCH4 protein engineered to be synthesized in the cytosol and thus lack normal co- and post-translational modifications is virtually inactive. TCH4 proteins, with and without a polyhistidine tag, that harbor an intact N-terminus are directed to the secretory pathway. Thus, as predicted, the N-terminal region of TCH4 functions as a signal peptide. TCH4 is shown to have at least one disulfide bond as monitored by a mobility shift in SDS-PAGE in the presence of dithiothreitol (DTT). This disulfide bond(s) is essential for full XET activity. TCH4 is glycosylated in vivo; glycosidases that remove N-linked glycosylation eliminated 98% of the XET activity. Thus, co- and/or post-translational modifications are critical for optimal TCH4 XET activity. Furthermore, using site-specific mutagenesis, we demonstrated that the first glutamate residue of the conserved DEIDFEFL motif (E97) is essential for activity. A change to glutamine at this position resulted in an inactive protein; a change to aspartic acid caused protein mislocalization. These data support the hypothesis that, in analogy to Bacillus beta-glucanases, this region may be the active site of XET enzymes.

NASA Discipline Cell Biology

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric

Evolutionary, structural and biochemical evidence for a new interaction site of the leptin obesity protein

The Leptin protein is central to the regulation of energy metabolism in mammals. By integrating evolutionary, structural, and biochemical information, a surface segment, outside of its known receptor contacts, is predicted as a second interaction site that may help to further define its roles in energy balance and its functional differences between humans and other mammals.

Evolution, Molecular

Purification of the small mechanosensitive channel of Escherichia coli (MscS): the subunit structure, conduction, and gating characteristics in liposomes

The small mechanosensitive channel, MscS, is a part of the turgor-driven solute efflux system that protects bacteria from lysis in the event of osmotic downshift. It has been identified in Escherichia coli as a product of the orphan yggB gene, now called mscS (Levina et al., 1999, EMBO J. 18:1730). Here I show that that the isolated 31-kDa MscS protein is sufficient to form a functional mechanosensitive channel gated directly by tension in the lipid bilayer. MscS-6His complexes purified in the presence of octylglucoside and lipids migrate in a high-resolution gel-filtration column as particles of approximately 200 kDa. Consistent with that, the protein cross-linking patterns predict a hexamer. The channel reconstituted in soybean asolectin liposomes was activated by pressures of 20-60 mm Hg and displayed the same asymmetric I-V curve and slight anionic preference as in situ. At the same time, the single-channel conductance is proportional to the buffer conductivity in a wide range of salt concentrations. The rate of channel activation in response to increasing pressure gradient across the patch was slower than the rate of closure in response to decreasing steps of pressure gradient. Therefore, the open probability curves were recorded with descending series of pressures. Determination of the curvature of patches by video imaging permitted measurements of the channel activity as a function of membrane tension (gamma). Po(gamma) curves had the midpoint at 5.5 +/- 0.1 dyne/cm and gave estimates for the energy of opening DeltaG = 11.4 +/- 0.5 kT, and the transition-related area change DeltaA = 8.4 +/- 0.4 nm(2) when fitted with a two-state Boltzmann model. The correspondence between channel properties in the native and reconstituted systems is discussed.

NASA Discipline Cell Biology

Protein crystal growth - Growth kinetics for tetragonal lysozyme crystals

Results are reported from theoretical and experimental studies of the growth rate of lysozyme as a function of diffusion in earth-gravity conditions. The investigations were carried out to form a comparison database for future studies of protein crystal growth in the microgravity environment of space. A diffusion-convection model is presented for predicting crystal growth rates in the presence of solutal concentration gradients. Techniques used to grow and monitor the growth of hen egg white lysozyme are detailed. The model calculations and experiment data are employed to discuss the effects of transport and interfacial kinetics in the growth of the crystals, which gradually diminished the free energy in the growth solution. Density gradient-driven convection, caused by presence of the gravity field, was a limiting factor in the growth rate.

Pusey, M. L.

Finding the global minimum: a fuzzy end elimination implementation

The 'fuzzy end elimination theorem' (FEE) is a mathematically proven theorem that identifies rotameric states in proteins which are incompatible with the global minimum energy conformation. While implementing the FEE we noticed two different aspects that directly affected the final results at convergence. First, the identification of a single dead-ending rotameric state can trigger a 'domino effect' that initiates the identification of additional rotameric states which become dead-ending. A recursive check for dead-ending rotameric states is therefore necessary every time a dead-ending rotameric state is identified. It is shown that, if the recursive check is omitted, it is possible to miss the identification of some dead-ending rotameric states causing a premature termination of the elimination process. Second, we examined the effects of removing dead-ending rotameric states from further considerations at different moments of time. Two different methods of rotameric state removal were examined for an order dependence. In one case, each rotamer found to be incompatible with the global minimum energy conformation was removed immediately following its identification. In the other, dead-ending rotamers were marked for deletion but retained during the search, so that they influenced the evaluation of other rotameric states. When the search was completed, all marked rotamers were removed simultaneously. In addition, to expand further the usefulness of the FEE, a novel method is presented that allows for further reduction in the remaining set of conformations at the FEE convergence. In this method, called a tree-based search, each dead-ending pair of rotamers which does not lead to the direct removal of either rotameric state is used to reduce significantly the number of remaining conformations. In the future this method can also be expanded to triplet and quadruplet sets of rotameric states. We tested our implementation of the FEE by exhaustively searching ten protein segments and found that the FEE identified the global minimum every time. For each segment, the global minimum was exhaustively searched in two different environments: (i) the segments were extracted from the protein and exhaustively searched in the absence of the surrounding residues; (ii) the segments were exhaustively searched in the presence of the remaining residues fixed at crystal structure conformations. We also evaluated the performance of the method for accurately predicting side chain conformations. We examined the influence of factors such as type and accuracy of backbone template used, and the restrictions imposed by the choice of potential function, parameterization and rotamer database. Conclusions are drawn on these results and future prospects are given.

NASA Program Exobiology

RNA catalysis and the origins of life

The role of RNA catalysis in the origins of life is considered in connection with the discovery of riboszymes, which are RNA molecules that catalyze sequence-specific hydrolysis and transesterification reactions of RNA substrates. Due to this discovery, theories positing protein-free replication as preceding the appearance of the genetic code are more plausible. The scope of RNA catalysis in biology and chemistry is discussed, and it is noted that the development of methods to select (or predict) RNA sequences with preassigned catalytic functions would be a major contribution to the study of life's origins.

Orgel, Leslie E.