Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

Human XPG nuclease structure, assembly, and activities with insights for neurodegeneration and cancer from pathogenic mutations

Xeroderma pigmentosum group G (XPG) protein is both a functional partner in multiple DNA damage responses (DDR) and a pathway coordinator and structure-specific endonuclease in nucleotide excision repair (NER). Different mutations in the XPG gene ERCC5 lead to either of two distinct human diseases: Cancer-prone xeroderma pigmentosum (XP-G) or the fatal neurodevelopmental disorder Cockayne syndrome (XP-G/CS). To address the enigmatic structural mechanism for these differing disease phenotypes and for XPG’s role in multiple DDRs, here we determined the crystal structure of human XPG catalytic domain (XPGcat), revealing XPG-specific features for its activities and regulation. Furthermore, XPG DNA binding elements conserved with FEN1 superfamily members enable insights on DNA interactions. Notably, all but one of the known pathogenic point mutations map to XPGcat, and both XP-G and XP-G/CS mutations destabilize XPG and reduce its cellular protein levels. Mapping the distinct mutation classes provides structure-based predictions for disease phenotypes: Residues mutated in XP-G are positioned to reduce local stability and NER activity, whereas residues mutated in XP-G/CS have implied long-range structural defects that would likely disrupt stability of the whole protein, and thus interfere with its functional interactions. Combined data from crystallography, biochemistry, small angle X-ray scattering, and electron microscopy unveil an XPG homodimer that binds, unstacks, and sculpts duplex DNA at internal unpaired regions (bubbles) into strongly bent structures, and suggest how XPG complexes may bind both NER bubble junctions and replication forks. Collective results support XPG scaffolding and DNA sculpting functions in multiple DDR processes to maintain genome stability.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of α $IIb$ $β$ 3 integrin structures along its minimum free energy activation pathway

The adhesion protein integrin is a transmembrane heterodimer that plays a pivotal role in cellular processes such as cell signaling and cell migration. To execute its function, integrin undergoes extensive conformational changes from a bent-closed to an extended-open state. Resolving the structures across these changes remains a challenge with both experimental and computational methods, but it is crucial for understanding the activation mechanism of integrin. We address this challenge for the platelet integrin α IIb β 3 by employing finite temperature string method with structures of the images along the initial guess path generated by a multiscale data-driven framework. The full-length all-atom structures along the resulting minimum free energy path between the inactive bent-closed and active extended-open states of α IIb β 3 integrin are consistent with a variety of experimentally resolved structures. Changes in these predicted structures along the path show that the extension and separation of the α and β subunits from the bent-closed to the extended-open state require correlated movements between the subdomain pairs in α IIb β 3 . Furthermore, these results provide new insights into integrin activation mechanism, and the predicted structures have potential applications in guiding the design of integrin-targeting therapeutics.

Dasetty, Siva [University of Chicago, IL (United S↗

RNA target highlights in CASP15 : Evaluation of predicted models by structure providers

Abstract The first RNA category of the Critical Assessment of Techniques for Structure Prediction competition was only made possible because of the scientists who provided experimental structures to challenge the predictors. In this article, these scientists offer a unique and valuable analysis of both the successes and areas for improvement in the predicted models. All 10 RNA‐only targets yielded predictions topologically similar to experimentally determined structures. For one target, experimentalists were able to phase their x‐ray diffraction data by molecular replacement, showing a potential application of structure predictions for RNA structural biologists. Recommended areas for improvement include: enhancing the accuracy in local interaction predictions and increased consideration of the experimental conditions such as multimerization, structure determination method, and time along folding pathways. The prediction of RNA–protein complexes remains the most significant challenge. Finally, given the intrinsic flexibility of many RNAs, we propose the consideration of ensemble models.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Heterologous expression of a fully active Azotobacter vinelandii nitrogenase Fe protein in Escherichia coli

ABSTRACT The functional versatility of the Fe protein, the reductase component of nitrogenase, makes it an appealing target for heterologous expression, which could facilitate future biotechnological adaptations of nitrogenase-based production of valuable chemical commodities. Yet, the heterologous synthesis of a fully active Fe protein of Azotobacter vinelandii ( Av NifH) in Escherichia coli has proven to be a challenging task. Here, we report the successful synthesis of a fully active Av NifH protein upon co-expression of this protein with Av IscS/U and Av NifM in E. coli . Our metal, activity, electron paramagnetic resonance, and X-ray absorption spectroscopy/extended X-ray absorption fine structure (EXAFS) data demonstrate that the heterologously expressed Av NifH protein has a high [Fe 4 S 4 ] cluster content and is fully functional in nitrogenase catalysis and assembly. Moreover, our phylogenetic analyses and structural predictions suggest that Av NifM could serve as a chaperone and assist the maturation of a cluster-replete Av NifH protein. Given the crucial importance of the Fe protein for the functionality of nitrogenase, this work establishes an effective framework for developing a heterologous expression system of the complete, two-component nitrogenase system; additionally, it provides a useful tool for further exploring the intricate biosynthetic mechanism of this structurally unique and functionally important metalloenzyme. IMPORTANCE The heterologous expression of a fully active Azotobacter vinelandii Fe protein (AvNifH) has never been accomplished. Given the functional importance of this protein in nitrogenase catalysis and assembly, the successful expression of AvNifH in Escherichia coli as reported herein supplies a key element for the further development of heterologous expression systems that explore the catalytic versatility of the Fe protein, either on its own or as a key component of nitrogenase, for nitrogenase-based biotechnological applications in the future. Moreover, the “clean” genetic background of the heterologous expression host allows for an unambiguous assessment of the effect of certain nif-encoded protein factors, such as AvNifM described in this work, in the maturation of AvNifH, highlighting the utility of this heterologous expression system in further advancing our understanding of the complex biosynthetic mechanism of nitrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis of the amidase ClbL central to the biosynthesis of the genotoxin colibactin

Colibactin is a genotoxic natural product produced by select commensal bacteria in the human gut microbiota. The compound is a bis-electrophile that is predicted to form interstrand DNA cross-links in target cells, leading to double-strand DNA breaks. The biosynthesis of colibactin is carried out by a mixed NRPS–PKS assembly line with several noncanonical features. An amidase, ClbL, plays a key role in the pathway, catalyzing the final step in the formation of the pseudodimeric scaffold. ClbL couples α-aminoketone and β-ketothioester intermediates attached to separate carrier domains on the NRPS–PKS assembly. Here, the 1.9 Å resolution structure of ClbL is reported, providing a structural basis for this key step in the colibactin biosynthetic pathway. The structure reveals an open hydrophobic active site surrounded by flexible loops, and comparison with homologous amidases supports its unusual function and predicts macromolecular interactions with pathway carrier-protein substrates. Modeling protein–protein interactions supports a predicted molecular basis for enzyme–carrier domain interactions. Overall, the work provides structural insight into this unique enzyme that is central to the biosynthesis of colibactin.

59 BASIC BIOLOGICAL SCIENCES↗

ppdx : Automated modeling of protein–protein interaction descriptors for use with machine learning

This paper describes ppdx, a python workflow tool that combines protein sequence alignment, homology modeling, and structural refinement, to compute a broad array of descriptors for characterizing protein–protein interactions. The descriptors can be used to predict various properties of interest, such as protein–protein binding affinities, or inhibitory concentrations (IC 50 ), using approaches that range from simple regression to more complex machine learning models. The software is highly modular. It supports different protocols for generating structures, and 95 descriptors can be currently computed. More protocols and descriptors can be easily added. The implementation is highly parallel and can fully exploit the available cores in a single workstation, or multiple nodes on a supercomputer, allowing many systems to be analyzed simultaneously. As an illustrative application, ppdx is used to parametrize a model that predicts the IC 50 of a set of antigens and a class of antibodies directed to the influenza hemagglutinin stalk.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genome sequence and characterization of a novel Pseudomonas putida phage, MiCath

Abstract Pseudomonads are ubiquitous bacteria with importance in medicine, soil, agriculture, and biomanufacturing. We report a novel Pseudomonas putida phage, MiCath, which is the first known phage infecting P. putida S12, a strain increasingly used as a synthetic biology chassis. MiCath was isolated from garden soil under a tomato plant using P. putida S12 as a host and was also found to infect four other P. putida strains. MiCath has a ~ 61 kbp double-stranded DNA genome which encodes 97 predicted open reading frames (ORFs); functions could only be predicted for 48 ORFs using comparative genomics. Functions include structural phage proteins, other common phage proteins (e.g., terminase), a queuosine gene cassette, a cas4 exonuclease, and an endosialidase. Restriction digestion analysis suggests the queuosine gene cassette encodes a pathway capable of modification of guanine residues. When compared to other phage genomes, MiCath shares at most 74% nucleotide identity over 2% of the genome with any sequenced phage. Overall, MiCath is a novel phage with no close relatives, encoding many unique gene products.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Assessment of Pose Prediction Accuracy in RNA–Ligand Docking

Structure-based virtual high-throughput screening is used in early-stage drug discovery. Over the years, docking protocols and scoring functions for protein–ligand complexes have evolved to improve the accuracy in the computation of binding strengths and poses. In the past decade, RNA has also emerged as a target class for new small-molecule drugs. However, most ligand docking programs have been validated and tested for proteins and not RNA. Here, we test the docking power (pose prediction accuracy) of three state-of-the-art docking protocols on 173 RNA–small molecule crystal structures. The programs are AutoDock4 (AD4) and AutoDock Vina (Vina), which were designed for protein targets, and rDock, which was designed for both protein and nucleic acid targets. AD4 performed relatively poorly. For RNA targets for which a crystal structure of a bound ligand used to limit the docking search space is available and for which the goal is to identify new molecules for the same pocket, rDock performs slightly better than Vina, with success rates of 48% and 63%, respectively. However, in the more common type of early-stage drug discovery setting, in which no structure of a ligand–target complex is known and for which a larger search space is defined, rDock performed similarly to Vina, with a low success rate of ~27%. Further, Vina was found to have bias for ligands with certain physicochemical properties, whereas rDock performs similarly for all ligand properties. Thus, for projects where no ligand–protein structure already exists, Vina and rDock are both applicable. However, the relatively poor performance of all methods relative to protein–target docking illustrates a need for further methods refinement.

59 BASIC BIOLOGICAL SCIENCES↗

An evolutionarily conserved tryptophan cage promotes folding of the extended RNA recognition motif in the hnRNPR ‐like protein family

Abstract The heterogeneous nuclear ribonucleoprotein (hnRNP) R‐like family is a class of RNA binding proteins in the hnRNP superfamily with diverse functions in RNA processing. Here, we present the 1.90 Å X‐ray crystal structure and solution NMR studies of the first RNA recognition motif (RRM) of human hnRNPR. We find that this domain adopts an extended RRM (eRRM1) featuring a canonical RRM with a structured N‐terminal extension (N ext ) motif that docks against the RRM and extends the β‐sheet surface. The adjoining loop is structured and forms a tryptophan cage motif to position the N ext motif for docking to the RRM. Combining mutagenesis, solution NMR spectroscopy, and thermal denaturation studies, we evaluate the importance of residues in the N ext –RRM interface and adjoining loop on eRRM folding and conformational dynamics. We find that these sites are essential for protein solubility, conformational ordering, and thermal stability. Consistent with their importance, mutations in the N ext –RRM interface and loop are associated with several cancers in a survey of somatic mutations in cancer studies. Sequence and structure comparison of the human hnRNPR eRRM1 to experimentally verified and predicted hnRNPR‐like proteins reveals conserved features in the eRRM.

Biochemistry & Molecular Biology↗

Transforming our understanding of chloroplast-associated genes through comprehensive characterization of protein localizations and protein-protein interactions

Bioenergy crops are a renewable source of fuels and are a critical base for building a carbon-neutral economy. Rational engineering of bioenergy crops has the potential to enhance the yields. However, our ability to engineer plants is limited because the functions of most genes remain unknown. Systematic characterization of gene function in plants thus has the potential to greatly accelerate bioenergy research. Here, we focus on the chloroplast, an underexplored energy-producing organelle that is a hallmark of plants. The chloroplast is one of the promising targets of biofuel crop engineering efforts because of its central role in photosynthesis, metabolism, and intracellular signaling. However, the protein composition of the chloroplast and the functions of most of its proteins remain poorly characterized. At the core of this project, we sought to comprehensively determine the localization of chloroplast-associated proteins and generate a spatially defined protein-protein interaction network for chloroplast. For this purpose, we used the leading model alga Chlamydomonas reinhardtii, which greatly increased experimental speed and throughput. We illustrated the value of our findings to land plants by determining the localization of Arabidopsis thaliana land plant homologs of the Chlamydomonas proteins. Altogether, we were successful in determining the localization of 1,034 chloroplast-associated proteins in Chlamydomonas. The localizations provide numerous insights into the spatial organization of chloroplasts and how they function to support photosynthesis. The localization patterns of distinct proteins revealed new chloroplast structures and revealed new spatial organization inside the chloroplast. We also identified new components of known chloroplast structures, such as the chloroplast envelope, nucleoid, plastoglobuli, and pyrenoid. We identified these new components by investigating the interacting partners of known proteins. Many proteins localized in both the chloroplast and other cellular structures, thereby hinting at new functions and communication between cellular structures. We also applied machine learning on the atlas to generate predictions for the location of all of the proteins in Chlamydomonas. This enabled us to assign putative functions to many uncharacterized proteins based on their cellular location. Altogether, this research establishes a rich resource that opens new avenues of investigation and guides future work in deciphering and manipulating chloroplast function. Next, we developed an extensive protein-protein interaction network for the chloroplast by performing affinity purification-mass spectrometry on ~1,150 tagged chloroplast-associated proteins, the first such large-scale study in any photosynthetic organism. This dataset reveals 4,694 high-confidence protein-protein interactions, offering insights into the functions of thousands of conserved poorly-characterized chloroplast proteins. This systematic identification of protein-protein interactions in the chloroplast also provides multiple exciting new research directions and a detailed blueprint of the chloroplast's operation. This research lays the groundwork to decipher the inner workings of the chloroplast, the cell structure at the heart of photosynthesis. The spatial atlas and protein-protein interactions reveal chloroplast organizational features that would not have been accessible with traditional approaches. The localization mapping, insights into the function, and research materials generated further provide a rich resource for the research community to advance the understanding of how the chloroplast is organized to enable engineering of enhanced photosynthetic organisms.

59 BASIC BIOLOGICAL SCIENCES↗

PYK-SubstitutionOME: an integrated database containing allosteric coupling, ligand affinity and mutational, structural, pathological, bioinformatic and computational information about pyruvate kinase isozymes

Interpreting changes in patient genomes, understanding how viruses evolve and engineering novel protein function all depend on accurately predicting the functional outcomes that arise from amino acid substitutions. To that end, the development of first-generation prediction algorithms was guided by historic experimental datasets. However, these datasets were heavily biased toward substitutions at positions that have not changed much throughout evolution (i.e. conserved). Although newer datasets include substitutions at positions that span a range of evolutionary conservation scores, these data are largely derived from assays that agglomerate multiple aspects of function. To facilitate predictions from the foundational chemical properties of proteins, large substitution databases with biochemical characterizations of function are needed. We report here a database derived from mutational, biochemical, bioinformatic, structural, pathological and computational studies of a highly studied protein family—pyruvate kinase (PYK). A centerpiece of this database is the biochemical characterization—including quantitative evaluation of allosteric regulation—of the changes that accompany substitutions at positions that sample the full conservation range observed in the PYK family. We have used these data to facilitate critical advances in the foundational studies of allosteric regulation and protein evolution and as rigorous benchmarks for testing protein predictions. We trust that the collected dataset will be useful for the broader scientific community in the further development of prediction algorithms.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Accelerating crystal structure determination with iterative AlphaFold prediction

Experimental structure determination can be accelerated with artificial intelligence (AI)-based structure-prediction methods such as AlphaFold . Here, an automatic procedure requiring only sequence information and crystallographic data is presented that uses AlphaFold predictions to produce an electron-density map and a structural model. Iterating through cycles of structure prediction is a key element of this procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. This procedure was applied to X-ray data for 215 structures released by the Protein Data Bank in a recent six-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2 Å. Predictions from the iterative template-guided prediction procedure were more accurate than those obtained without templates. It is concluded that AlphaFold predictions obtained based on sequence information alone are usually accurate enough to solve the crystallographic phase problem with molecular replacement, and a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization is suggested.

59 BASIC BIOLOGICAL SCIENCES↗

PigmentHunter: A point-and-click application for automated chlorophyll-protein simulations

Chlorophyll proteins (CPs) are the workhorses of biological photosynthesis, working together to absorb solar energy, transfer it to chemically active reaction centers, and control the charge-separation process that drives its storage as chemical energy. Yet predicting CP optical and electronic properties remains a serious challenge, driven by the computational difficulty of treating large, electronically coupled molecular pigments embedded in a dynamically structured protein environment. To address this challenge, we introduce here an analysis tool called PigmentHunter, which automates the process of preparing CP structures for molecular dynamics (MD), running short MD simulations on the nanoHUB.org science gateway, and then using electrostatic and steric analysis routines to predict optical absorption, fluorescence, and circular dichroism spectra within a Frenkel exciton model. Inter-pigment couplings are evaluated using point-dipole or transition-charge coupling models, while site energies can be estimated using both electrostatic and ring-deformation approaches. The package is built in a Jupyter Notebook environment, with a point-and-click interface that can be used either to manually prepare individual structures or to batch-process many structures at once. Here, we illustrate PigmentHunter’s capabilities with example simulations on spectral line shapes in the light harvesting 2 complex, site energies in the Fenna–Matthews–Olson protein, and ring deformation in photosystems I and II.

14 SOLAR ENERGY↗

Light-modulated abundance of an mRNA encoding a calmodulin-regulated, chromatin-associated NTPase in pea

A CDNA encoding a 47 kDa nucleoside triphosphatase (NTPase) that is associated with the chromatin of pea nuclei has been cloned and sequenced. The translated sequence of the cDNA includes several domains predicted by known biochemical properties of the enzyme, including five motifs characteristic of the ATP-binding domain of many proteins, several potential casein kinase II phosphorylation sites, a helix-turn-helix region characteristic of DNA-binding proteins, and a potential calmodulin-binding domain. The deduced primary structure also includes an N-terminal sequence that is a predicted signal peptide and an internal sequence that could serve as a bipartite-type nuclear localization signal. Both in situ immunocytochemistry of pea plumules and immunoblots of purified cell fractions indicate that most of the immunodetectable NTPase is within the nucleus, a compartment proteins typically reach through nuclear pores rather than through the endoplasmic reticulum pathway. The translated sequence has some similarity to that of human lamin C, but not high enough to account for the earlier observation that IgG against human lamin C binds to the NTPase in immunoblots. Northern blot analysis shows that the NTPase MRNA is strongly expressed in etiolated plumules, but only poorly or not at all in the leaf and stem tissues of light-grown plants. Accumulation of NTPase mRNA in etiolated seedlings is stimulated by brief treatments with both red and far-red light, as is characteristic of very low-fluence phytochrome responses. Southern blotting with pea genomic DNA indicates the NTPase is likely to be encoded by a single gene.

NASA Discipline Number 40-50↗