Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Hemiptera phylogenomic resources: Tree‐based orthology prediction and conserved exon identification

Abstract High‐throughput sequencing of transcriptomes and targeted genomic regions are advancing our knowledge of The Tree of Life. Building phylogenies with regions of the genome requires 1‐to‐1 orthologue resources of genes and noncoding loci. One organismal group that has received little attention in this area is the Hemiptera, the fifth largest insect order represented by ~103,590 named species. Here, we present a set of 3,872 Hemiptera 1‐to‐1 orthogroups based on tree‐based orthology inference of eight Hemiptera species with publicly available genome sequences. We also estimate a set of 406 orthologous exons with similar mRNA splice sites that can be used for Sanger sequencing and develop enrichment probes for targeted genome sequencing for phylogenomic inference. We show this novel set of orthologues is informative at the protein, coding sequence and exon molecular levels and provides robust branch support in both gene tree–species tree methods and concatenated sequence phylogenies. In addition, we demonstrate the utility of these loci to resolve relationships in whiteflies, Bemisia tabaci , a large species complex with few phylogenomic resources. Last, we compare our Hemiptera phylogeny with previously published phylogenies and other orthologue databases, while providing suggestions on further improvement to this phylogenomic resource.

Owen, Christopher L.↗

Biochemical and functional analysis of CTR1, a protein kinase that negatively regulates ethylene signaling in Arabidopsis

CTR1 encodes a negative regulator of the ethylene response pathway in Arabidopsis thaliana. The C-terminal domain of CTR1 is similar to the Raf family of protein kinases, but its first two-thirds encodes a novel protein domain. We used a variety of approaches to investigate the function of these two CTR1 domains. Recombinant CTR1 protein was purified from a baculoviral expression system, and shown to possess intrinsic Ser/Thr protein kinase activity with enzymatic properties similar to Raf-1. Deletion of the N-terminal domain did not elevate the kinase activity of CTR1, indicating that, at least in vitro, this domain does not autoinhibit kinase function. Molecular analysis of loss-of-function ctr1 alleles indicated that several mutations disrupt the kinase catalytic domain, and in vitro studies confirmed that at least one of these eliminates kinase activity, which indicates that kinase activity is required for CTR1 function. One missense mutation, ctr1-8, was found to result from an amino acid substitution within a new conserved motif within the N-terminal domain. Ctr1-8 has no detectable effect on the kinase activity of CTR1 in vitro, but rather disrupts the interaction with the ethylene receptor ETR1. This mutation also disrupts the dominant negative effect that results from overexpression of the CTR1 amino-terminal domain in transgenic Arabidopsis. These results suggest that CTR1 interacts with ETR1 in vivo, and that this association is required to turn off the ethylene-signaling pathway.

Non-NASA Center↗

Top-down mass spectrometry and assigning internal fragments for determining disulfide bond positions in proteins

Disulfide bonds in proteins have a substantial impact on protein structure, stability, and biological activity. Localizing disulfide bonds is critical for understanding protein folding and higher-order structure. Conventional top-down mass spectrometry (TD-MS), where only terminal fragments are assigned for disulfide-intact proteins, can access disulfide information, but suffers from low fragmentation efficiency, thereby limiting sequence coverage. Here, we show that assigning internal fragments generated from TD-MS enhances the sequence coverage of disulfide-intact proteins by 20–60% by returning information from the interior of the protein sequence, which cannot be obtained by terminal fragments alone. Further, the inclusion of internal fragments can extend the sequence information of disulfide-intact proteins to near complete sequence coverage. Importantly, the enhanced sequence information that arise from the assignment of internal fragments can be used to determine the relative position of disulfide bonds and the exact disulfide connectivity between cysteines. The data presented here demonstrates the benefits of incorporating internal fragment analysis into the TD-MS workflow for analyzing disulfide-intact proteins, which would be valuable for characterizing biotherapeutic proteins such as monoclonal antibodies and antibody–drug conjugates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hydrogen peroxide homeostasis: activation of plant catalase by calcium/calmodulin

Environmental stimuli such as UV, pathogen attack, and gravity can induce rapid changes in hydrogen peroxide (H(2)O(2)) levels, leading to a variety of physiological responses in plants. Catalase, which is involved in the degradation of H(2)O(2) into water and oxygen, is the major H(2)O(2)-scavenging enzyme in all aerobic organisms. A close interaction exists between intracellular H(2)O(2) and cytosolic calcium in response to biotic and abiotic stresses. Studies indicate that an increase in cytosolic calcium boosts the generation of H(2)O(2). Here we report that calmodulin (CaM), a ubiquitous calcium-binding protein, binds to and activates some plant catalases in the presence of calcium, but calcium/CaM does not have any effect on bacterial, fungal, bovine, or human catalase. These results document that calcium/CaM can down-regulate H(2)O(2) levels in plants by stimulating the catalytic activity of plant catalase. Furthermore, these results provide evidence indicating that calcium has dual functions in regulating H(2)O(2) homeostasis, which in turn influences redox signaling in response to environmental signals in plants.

NASA Discipline Plant Biology↗

The use of additive and subtractive approaches to examine the nuclear localization sequence of the polyomavirus major capsid protein VP1

A nuclear localization signal (NLS) has been identified in the N-terminal (Ala1-Pro-Lys-Arg-Lys-Ser-Gly-Val-Ser-Lys-Cys11) amino acid sequence of the polyomavirus major capsid protein VP1. The importance of this amino acid sequence for nuclear transport of VP1 protein was demonstrated by a genetic "subtractive" study using the constructs pSG5VP1 (full-length VP1) and pSG5 delta 5'VP1 (truncated VP1, lacking amino acids Ala1-Cys11). These constructs were used to transfect COS-7 cells, and expression and intracellular localization of the VP1 protein was visualized by indirect immunofluorescence. These studies revealed that the full-length VP1 was expressed and localized in the nucleus, while the truncated VP1 protein was localized in the cytoplasm and not transported to the nucleus. These findings were substantiated by an "additive" approach using FITC-labeled conjugates of synthetic peptides homologous to the NLS of VP1 cross-linked to bovine serum albumin or immunoglobulin G. Both conjugates localized in the nucleus after microinjection into the cytoplasm of 3T6 cells. The importance of individual amino acids found in the basic sequence (Lys3-Arg-Lys5) of the NLS was also investigated. This was accomplished by synthesizing three additional peptides in which lysine-3 was substituted with threonine, arginine-4 was substituted with threonine, or lysine-5 was substituted with threonine. It was found that lysine-3 was crucial for nuclear transport, since substitution of this amino acid with threonine prevented nuclear localization of the microinjected, FITC-labeled conjugate.

Non-NASA Center↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

Isolation, characterization, and amino acid sequences of auracyanins, blue copper proteins from the green photosynthetic bacterium Chloroflexus aurantiacus

Three small blue copper proteins designated auracyanin A, auracyanin B-1, and auracyanin B-2 have been isolated from the thermophilic green gliding photosynthetic bacterium Chloroflexus aurantiacus. All three auracyanins are peripheral membrane proteins. Auracyanin A was described previously (Trost, J. T., McManus, J. D., Freeman, J. C., Ramakrishna, B. L., and Blankenship, R. E. (1988) Biochemistry 27, 7858-7863) and is not glycosylated. The two B forms are glycoproteins and have almost identical properties to each other, but are distinct from the A form. The sodium dodecyl sulfate-polyacrylamide gel electrophoresis apparent monomer molecular masses are 14 (A), 18 (B-2), and 22 (B-1) kDa. The amino acid sequences of the B forms are presented. All three proteins have similar absorbance, circular dichroism, and resonance Raman spectra, but the electron spin resonance signals are quite different. Laser flash photolysis kinetic analysis of the reactions of the three forms of auracyanin with lumiflavin and flavin mononucleotide semiquinones indicates that the site of electron transfer is negatively charged and has an accessibility similar to that found in other blue copper proteins. Copper analysis indicates that all three proteins contain 1 mol of copper per mol of protein. All three auracyanins exhibit a midpoint redox potential of +240 mV. Light-induced absorbance changes and electron spin resonance signals suggest that auracyanin A may play a role in photosynthetic electron transfer. Kinetic data indicate that all three proteins can donate electrons to cytochrome c-554, the electron donor to the photosynthetic reaction center.

NASA Discipline Exobiology↗

Complete Genome Sequence of Sulfurospirillum sp. Strain ACS DCE , an Anaerobic Bacterium That Respires Tetrachloroethene under Acidic pH Conditions

Sulfurospirillum sp. strain ACS DCE couples growth with reductive dechlorination of tetrachloroethene to cis-1,2-dichloroethene at pH values as low as 5.5. The genome sequence of strain ACS DCE consists of a circular 2,737,849-bp chromosome and a 39,868-bp plasmid and carries 2,737 protein-coding sequences, including two reductive dehalogenase genes.

59 BASIC BIOLOGICAL SCIENCES↗

Towards predictive control of reversible nanoparticle assembly with solid-binding proteins

Although a broad range of ligand-functionalized nanoparticles and physico-chemical triggers have been exploited to create stimuli-responsive colloidal systems, little attention has been paid to the reversible assembly of unmodified nanoparticles with non-covalently bound proteins. Previously, we reported that a derivative of green fluorescent protein engineered with oppositely located silica-binding peptides mediates the repeated assembly and disassembly of 10-nm silica nanoparticles when pH is toggled between 7.5 and 8.5. We captured the subtle interplay between interparticle electrostatic repulsion and their protein-mediated short-range attraction with a multiscale model energetically benchmarked to collective system behavior captured by scattering experiments. Here, in this work, we show that both solution conditions (pH and ionic strength) and protein engineering (sequence and position of engineered silica-binding peptides) provide pathways for reversible control over growth and fragmentation, leading to clusters ranging in size from 25 nm protein-coated particles to micrometer-size aggregate. We further find that the higher electrolyte environment associated with successive cycles of base addition eventually eliminates reversibility. Our model accurately predicts these multiple length scales phenomena. The underpinning concepts provide design principles for the dynamic control of other protein- and particle-based nanocomposites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Suppressor Mutations in Type II Secretion Mutants of Vibrio cholerae : Inactivation of the VesC Protease

The type II secretion system (T2SS) is a conserved transport pathway responsible for the secretion of a range of virulence factors by many pathogens, including Vibrio cholerae. Disruption of the T2SS genes in V. cholerae results in loss of secretion, changes in cell envelope function, and growth defects. While T2SS mutants are viable, high-throughput genomic analyses have listed these genes among essential genes. To investigate whether secondary mutations arise as a consequence of T2SS inactivation, we sequenced the genomes of six V. cholerae T2SS mutants with deletions or insertions in either the epsG, epsL, or epsM genes and identified secondary mutations in all mutants. Two of the six T2SS mutants contain distinct mutations in the gene encoding the T2SS-secreted protease VesC. Other mutations were found in genes coding for V. cholerae cell envelope proteins. Subsequent sequence analysis of the vesC gene in 92 additional T2SS mutant isolates identified another 19 unique mutations including insertions or deletions, sequence duplications, and single-nucleotide changes resulting in amino acid substitutions in the VesC protein. Analysis of VesC variants and the X-ray crystallographic structure of wild-type VesC suggested that all mutations lead to loss of VesC production and/or function. One possible mechanism by which V. cholerae T2SS mutagenesis can be tolerated is through selection of vesC-inactivating mutations, which may, in part, suppress cell envelope damage, establishing permissive conditions for the disruption of the T2SS. Other mutations may have been acquired in genes encoding essential cell envelope proteins to prevent proteolysis by VesC.

59 BASIC BIOLOGICAL SCIENCES↗

The evolution of transcriptional regulation in eukaryotes

Gene expression is central to the genotype-phenotype relationship in all organisms, and it is an important component of the genetic basis for evolutionary change in diverse aspects of phenotype. However, the evolution of transcriptional regulation remains understudied and poorly understood. Here we review the evolutionary dynamics of promoter, or cis-regulatory, sequences and the evolutionary mechanisms that shape them. Existing evidence indicates that populations harbor extensive genetic variation in promoter sequences, that a substantial fraction of this variation has consequences for both biochemical and organismal phenotype, and that some of this functional variation is sorted by selection. As with protein-coding sequences, rates and patterns of promoter sequence evolution differ considerably among loci and among clades for reasons that are not well understood. Studying the evolution of transcriptional regulation poses empirical and conceptual challenges beyond those typically encountered in analyses of coding sequence evolution: promoter organization is much less regular than that of coding sequences, and sequences required for the transcription of each locus reside at multiple other loci in the genome. Because of the strong context-dependence of transcriptional regulation, sequence inspection alone provides limited information about promoter function. Understanding the functional consequences of sequence differences among promoters generally requires biochemical and in vivo functional assays. Despite these challenges, important insights have already been gained into the evolution of transcriptional regulation, and the pace of discovery is accelerating.

Review, Academic↗

Sequence and structural implications of a bovine corneal keratan sulfate proteoglycan core protein. Protein 37B represents bovine lumican and proteins 37A and 25 are unique

Amino acid sequence from tryptic peptides of three different bovine corneal keratan sulfate proteoglycan (KSPG) core proteins (designated 37A, 37B, and 25) showed similarities to the sequence of a chicken KSPG core protein lumican. Bovine lumican cDNA was isolated from a bovine corneal expression library by screening with chicken lumican cDNA. The bovine cDNA codes for a 342-amino acid protein, M(r) 38,712, containing amino acid sequences identified in the 37B KSPG core protein. The bovine lumican is 68% identical to chicken lumican, with an 83% identity excluding the N-terminal 40 amino acids. Location of 6 cysteine and 4 consensus N-glycosylation sites in the bovine sequence were identical to those in chicken lumican. Bovine lumican had about 50% identity to bovine fibromodulin and 20% identity to bovine decorin and biglycan. About two-thirds of the lumican protein consists of a series of 10 amino acid leucine-rich repeats that occur in regions of calculated high beta-hydrophobic moment, suggesting that the leucine-rich repeats contribute to beta-sheet formation in these proteins. Sequences obtained from 37A and 25 core proteins were absent in bovine lumican, thus predicting a unique primary structure and separate mRNA for each of the three bovine KSPG core proteins.

NASA Discipline Cell Biology↗

Finding the global minimum: a fuzzy end elimination implementation

The 'fuzzy end elimination theorem' (FEE) is a mathematically proven theorem that identifies rotameric states in proteins which are incompatible with the global minimum energy conformation. While implementing the FEE we noticed two different aspects that directly affected the final results at convergence. First, the identification of a single dead-ending rotameric state can trigger a 'domino effect' that initiates the identification of additional rotameric states which become dead-ending. A recursive check for dead-ending rotameric states is therefore necessary every time a dead-ending rotameric state is identified. It is shown that, if the recursive check is omitted, it is possible to miss the identification of some dead-ending rotameric states causing a premature termination of the elimination process. Second, we examined the effects of removing dead-ending rotameric states from further considerations at different moments of time. Two different methods of rotameric state removal were examined for an order dependence. In one case, each rotamer found to be incompatible with the global minimum energy conformation was removed immediately following its identification. In the other, dead-ending rotamers were marked for deletion but retained during the search, so that they influenced the evaluation of other rotameric states. When the search was completed, all marked rotamers were removed simultaneously. In addition, to expand further the usefulness of the FEE, a novel method is presented that allows for further reduction in the remaining set of conformations at the FEE convergence. In this method, called a tree-based search, each dead-ending pair of rotamers which does not lead to the direct removal of either rotameric state is used to reduce significantly the number of remaining conformations. In the future this method can also be expanded to triplet and quadruplet sets of rotameric states. We tested our implementation of the FEE by exhaustively searching ten protein segments and found that the FEE identified the global minimum every time. For each segment, the global minimum was exhaustively searched in two different environments: (i) the segments were extracted from the protein and exhaustively searched in the absence of the surrounding residues; (ii) the segments were exhaustively searched in the presence of the remaining residues fixed at crystal structure conformations. We also evaluated the performance of the method for accurately predicting side chain conformations. We examined the influence of factors such as type and accuracy of backbone template used, and the restrictions imposed by the choice of potential function, parameterization and rotamer database. Conclusions are drawn on these results and future prospects are given.

NASA Program Exobiology↗

Expression of blue pigment synthetase a from Streptomyces lavenduale reveals insights on the effects of refactoring biosynthetic megasynthases for heterologous expression in Escherichia coli .

High GC bacteria from the genus Streptomyces harbor expansive secondary metabolism. The expression of biosynthetic proteins and the characterization and identification of biological "parts" for synthetic biology purposes from such pathways are of interest. However, the high GC content of proteins from actinomycetes in addition to the large size and multi-domain architecture of many biosynthetic proteins (such as non-ribosomal peptide synthetases; NRPSs, and polyketide synthases; PKSs often called "megasynthases") often presents issues with full-length translation and folding. Here we evaluate a non-ribosomal peptide synthetase (NRPS) from Streptomyces lavenduale, a multidomain "megasynthase" gene that comes from a high GC (72.5%) genome. While a preliminary step in revealing differences, to our knowledge this presents the first head-to-head comparison of codon-optimized sequences versus a native sequence of proteins of streptomycete origin heterologously expressed in E. coli. We found that any disruption in co-translational folding from codon mismatch that reduces the titer of indigoidine is explainable via the formation of more inclusion bodies as opposed to compromising folding or posttranslational modification in the soluble fraction. In conclusion, this result supports that one could apply any refactoring strategies that improve soluble expression in E. coli without concern that the protein that reaches the soluble fraction is differentially folded.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence specificity analysis of the SETD2 protein lysine methyltransferase and discovery of a SETD2 super-substrate

SETD2 catalyzes methylation at lysine 36 of histone H3 and it has many disease connections. We investigated the substrate sequence specificity of SETD2 and identified nine additional peptide and one protein (FBN1) substrates. Our data showed that SETD2 strongly prefers amino acids different from those in the H3K36 sequence at several positions of its specificity profile. Based on this, we designed an optimized super-substrate containing four amino acid exchanges and show by quantitative methylation assays with SETD2 that the super-substrate peptide is methylated about 290-fold more efficiently than the H3K36 peptide. Protein methylation studies confirmed very strong SETD2 methylation of the super-substrate in vitro and in cells. We solved the structure of SETD2 with bound super-substrate peptide containing a target lysine to methionine mutation, which revealed better interactions involving three of the substituted residues. Our data illustrate that substrate sequence design can strongly increase the activity of protein lysine methyltransferases.

36 MATERIALS SCIENCE↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, overproduction and purification of Vibrio proteolyticus ribosomal protein L18 for in vitro and in vivo studies

A strategy suggested by comparative genomic studies was used to amplify the entire Vibrio proteolyticus (Vp) gene for ribosomal protein L18. Vp L18 and its flanking regions were sequenced and compared with the deduced amino acid (aa) sequences of other known L18 proteins. A 26-aa residue segment at the carboxy terminus contains many strongly conserved residues and may be critical for the L18 interaction with 5S rRNA. This approach should allow rapid characterization of L18 from large numbers of bacteria. Both Vp L18 and Escherichia coli (Ec) L18 were overproduced and purified using a T7 expression vector which fuses an N-terminal peptide segment (His-tag) containing 6 histidine residues to the recombinant protein. The purified fusion proteins, Vp His::L18 and Ec His::L18, were both found to bind to either the Vp 5S or Ec 5S rRNAs in vitro. Vp His::L18 protein was also shown to incorporate into Ec ribosomes in vivo. This His-tag strategy likely will have general applicability for the study of ribosomal proteins in vitro and in vivo.

Non-NASA Center↗