Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Water, Solute, and Ion Transport in De Novo-Designed Membrane Protein Channels

Biological organisms engineer peptide sequences to fold into membrane pore proteins capable of performing a wide variety of transport functions. Synthetic de novo-designed membrane pores can mimic this approach to achieve a potentially even larger set of functions. Here, in this work, we explore water, solute, and ion transport in three de novo designed β-barrel membrane channels in the 5–10 Å pore size range. We show that these proteins form passive membrane pores with high water transport efficiencies and size rejection characteristics consistent with the pore size encoded in the protein structure. Ion conductance and ion selectivity measurements also show trends consistent with the pore size, with the two larger pores showing weak cation selectivity. MD simulations of water and ion transport and solute size exclusion are consistent with the experimental trends and provide further insights into structure–function correlations in these membrane pores.

59 BASIC BIOLOGICAL SCIENCES↗

Cloning of the cDNA for U1 small nuclear ribonucleoprotein particle 70K protein from Arabidopsis thaliana

We cloned and sequenced a plant cDNA that encodes U1 small nuclear ribonucleoprotein (snRNP) 70K protein. The plant U1 snRNP 70K protein cDNA is not full length and lacks the coding region for 68 amino acids in the amino-terminal region as compared to human U1 snRNP 70K protein. Comparison of the deduced amino acid sequence of the plant U1 snRNP 70K protein with the amino acid sequence of animal and yeast U1 snRNP 70K protein showed a high degree of homology. The plant U1 snRNP 70K protein is more closely related to the human counter part than to the yeast 70K protein. The carboxy-terminal half is less well conserved but, like the vertebrate 70K proteins, is rich in charged amino acids. Northern analysis with the RNA isolated from different parts of the plant indicates that the snRNP 70K gene is expressed in all of the parts tested. Southern blotting of genomic DNA using the cDNA indicates that the U1 snRNP 70K protein is coded by a single gene.

NASA Discipline Plant Biology↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of a small tetraheme cytochrome c and a flavocytochrome c as two of the principal soluble cytochromes c in Shewanella oneidensis strain MR1

Two abundant, low-redox-potential cytochromes c were purified from the facultative anaerobe Shewanella oneidensis strain MR1 grown anaerobically with fumarate. The small cytochrome was completely sequenced, and the genes coding for both proteins were cloned and sequenced. The small cytochrome c contains 91 residues and four heme binding sites. It is most similar to the cytochromes c from Shewanella frigidimarina (formerly Shewanella putrefaciens) NCIMB400 and the unclassified bacterial strain H1R (64 and 55% identity, respectively). The amount of the small tetraheme cytochrome is regulated by anaerobiosis, but not by fumarate. The larger of the two low-potential cytochromes contains tetraheme and flavin domains and is regulated by anaerobiosis and by fumarate and thus most nearly corresponds to the flavocytochrome c-fumarate reductase previously characterized from S. frigidimarina to which it is 59% identical. However, the genetic context of the cytochrome genes is not the same for the two Shewanella species, and they are not located in multicistronic operons. The small cytochrome c and the cytochrome domain of the flavocytochrome c are also homologous, showing 34% identity. Structural comparison shows that the Shewanella tetraheme cytochromes are not related to the Desulfovibrio cytochromes c(3) but define a new folding motif for small multiheme cytochromes c.

Cytochrome c Group/chemistry/genetics/metabolism↗

Permeation of membranes by the neutral form of amino acids and peptides: relevance to the origin of peptide translocation

The flux of amino acids and other nutrient solutes such as phosphate across lipid bilayers (liposomes) is 10(5) slower than facilitated inward transport across biological membranes. This suggest that primitive cells lacking highly evolved transport systems would have difficulty transporting sufficient nutrients for cell growth to occur. There are two possible ways by which early life may have overcome this difficulty: (1) The membranes of the earliest cellular life-forms may have been intrinsically more permeable to solutes; or (2) some transport mechanism may have been available to facilitate transbilayer movement of solutes essential for cell survival and growth prior to the evolution of membrane transport proteins. Translocation of neutral species represents one such mechanism. The neutral forms of amino acids modified by methylation (creating protonated weak bases) permeate membranes up to 10(10) times faster than charged forms. This increased permeability when coupled to a transmembrane pH gradient can result in significantly increased rates of net unidirectional transport. Such pH gradients can be generated in vesicles used to model protocells that preceded and were presumably ancestral to early forms of life. This transport mechanism may still play a role in some protein translocation processes (e.g. for certain signal sequences, toxins and thylakoid proteins) in vivo.

NASA Discipline Exobiology↗

Properties of protein unfolded states suggest broad selection for expanded conformational ensembles

Much attention is being paid to conformational biases in the ensembles of intrinsically disordered proteins. However, it is currently unknown whether or how conformational biases within the disordered ensembles of foldable proteins affect function in vivo. Recently, we demonstrated that water can be a good solvent for unfolded polypeptide chains, even those with a hydrophobic and charged sequence composition typical of folded proteins. These results run counter to the generally accepted model that protein folding begins with hydrophobicity-driven chain collapse. Here we investigate what other features, beyond amino acid composition, govern chain collapse. We found that local clustering of hydrophobic and/or charged residues leads to significant collapse of the unfolded ensemble of pertactin, a secreted autotransporter virulence protein from Bordetella pertussis , as measured by small angle X-ray scattering (SAXS). Sequence patterns that lead to collapse also correlate with increased intermolecular polypeptide chain association and aggregation. Crucially, sequence patterns that support an expanded conformational ensemble enhance pertactin secretion to the bacterial cell surface. Similar sequence pattern features are enriched across the large and diverse family of autotransporter virulence proteins, suggesting sequence patterns that favor an expanded conformational ensemble are under selection for efficient autotransporter protein secretion, a necessary prerequisite for virulence. More broadly, we found that sequence patterns that lead to more expanded conformational ensembles are enriched across water-soluble proteins in general, suggesting protein sequences are under selection to regulate collapse and minimize protein aggregation, in addition to their roles in stabilizing folded protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, molecular properties, and chromosomal mapping of mouse lumican

PURPOSE. Lumican is a major proteoglycan of vertebrate cornea. This study characterizes mouse lumican, its molecular form, cDNA sequence, and chromosomal localization. METHODS. Lumican sequence was determined from cDNA clones selected from a mouse corneal cDNA expression library using a bovine lumican cDNA probe. Tissue expression and size of lumican mRNA were determined using Northern hybridization. Glycosidase digestion followed by Western blot analysis provided characterization of molecular properties of purified mouse corneal lumican. Chromosomal mapping of the lumican gene (Lcn) used Southern hybridization of a panel of genomic DNAs from an interspecific murine backcross. RESULTS. Mouse lumican is a 338-amino acid protein with high-sequence identity to bovine and chicken lumican proteins. The N-terminus of the lumican protein contains consensus sequences for tyrosine sulfation. A 1.9-kb lumican mRNA is present in cornea and several other tissues. Antibody against bovine lumican reacted with recombinant mouse lumican expressed in Escherichia coli and also detected high molecular weight proteoglycans in extracts of mouse cornea. Keratanase digestion of corneal proteoglycans released lumican protein, demonstrating the presence of sulfated keratan sulfate chains on mouse corneal lumican in vivo. The lumican gene (Lcn) was mapped to the distal region of mouse chromosome 10. The Lcn map site is in the region of a previously identified developmental mutant, eye blebs, affecting corneal morphology. CONCLUSIONS. This study demonstrates sulfated keratan sulfate proteoglycan in mouse cornea and describes the tools (antibodies and cDNA) necessary to investigate the functional role of this important corneal molecule using naturally occurring and induced mutants of the murine lumican gene.

NASA Discipline Cell Biology↗

PTM‐Psi : A python package to facilitate the computational investigation of p ost‐ t ranslational m odification on p rotein s tructures and their i mpacts on dynamics and functions

Abstract Post‐translational modification (PTM) of a protein occurs after it has been synthesized from its genetic template, and involves chemical modifications of the protein's specific amino acid residues. Despite of the central role played by PTM in regulating molecular interactions, particularly those driven by reversible redox reactions, it remains challenging to interpret PTMs in terms of protein dynamics and function because there are numerous combinatorially enormous means for modifying amino acids in response to changes in the protein environment. In this study, we provide a workflow that allows users to interpret how perturbations caused by PTMs affect a protein's properties, dynamics, and interactions with its binding partners based on inferred or experimentally determined protein structure. This Python‐based workflow, called PTM‐Psi , integrates several established open‐source software packages, thereby enabling the user to infer protein structure from sequence, develop force fields for non‐standard amino acids using quantum mechanics, calculate free energy perturbations through molecular dynamics simulations, and score the bound complexes via docking algorithms. Using the S ‐nitrosylation of several cysteines on the GAP2 protein as an example, we demonstrated the utility of PTM‐Psi for interpreting sequence–structure–function relationships derived from thiol redox proteomics data. We demonstrate that the S ‐nitrosylated cysteine that is exposed to the solvent indirectly affects the catalytic reaction of another buried cysteine over a distance in GAP2 protein through the movement of the two ligands. Our workflow tracks the PTMs on residues that are responsive to changes in the redox environment and lays the foundation for the automation of molecular and systems biology modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling SARS-CoV-2 proteins in the CASP-commons experiment

Critical Assessment of Structure Prediction (CASP) is an organization aimed at advancing the state of the art in computing protein structure from sequence. In the spring of 2020, CASP launched a community project to compute the structures of the most structurally challenging proteins coded for in the SARS-CoV-2 genome. Forty-seven research groups submitted over 3000 three-dimensional models and 700 sets of accuracy estimates on 10 proteins. The resulting models were released to the public. CASP community members also worked together to provide estimates of local and global accuracy and identify structure-based domain boundaries for some proteins. Subsequently, two of these structures (ORF3a and ORF8) have been solved experimentally, allowing assessment of both model quality and the accuracy estimates. Models from the AlphaFold2 group were found to have good agreement with the experimental structures, with main chain GDT_TS accuracy scores ranging from 63 (a correct topology) to 87 (competitive with experiment).

59 BASIC BIOLOGICAL SCIENCES↗

Nop9 recognizes structured and single-stranded RNA elements of preribosomal RNA

Nop9 is an essential factor in the processing of preribosomal RNA. Its absence in yeast is lethal, and defects in the human ortholog are associated with breast cancer, autoimmunity, and learning/language impairment. PUF family RNA-binding proteins are best known for sequence-specific RNA recognition, and most contain eight α-helical repeats that bind to the RNA bases of single-stranded RNA. Nop9 is an unusual member of this family in that it contains eleven repeats and recognizes both RNA structure and sequence. Here we report a crystal structure of Saccharomyces cerevisiae Nop9 in complex with its target RNA within the 20S preribosomal RNA. This structure reveals that Nop9 brings together a carboxy-terminal module recognizing the 5' single-stranded region of the RNA and a bifunctional amino-terminal module recognizing the central double-stranded stem region. We further show that the 3' single-stranded region of the 20S target RNA adds sequence-independent binding energy to the RNA–Nop9 interaction. Both the amino- and carboxy-terminal modules retain the characteristic sequence-specific recognition of PUF proteins, but the amino-terminal module has also evolved a distinct interface, which allows Nop9 to recognize either single-stranded RNA sequences or RNAs with a combination of single-stranded and structured elements.

59 BASIC BIOLOGICAL SCIENCES↗

Novel Approach to Quantification of Telomere Length with Direct Nanopore Sequencing and PCR Amplification

The ends of human chromosomes contain telomeres, or tandem arrays of repeating DNA sequences capped by multiple associated proteins that protect chromosomal ends from degradation. Telomeres function to preserve genomic stability by preventing natural chromosomal ends from being recognized as broken DNA double-strand breaks and triggering inappropriate DNA damage responses. Mounting evidence shows telomere length is an inherited trait that decreases with cellular division and normal aging. In addition, telomere length also appears to be influenced by other factors such as cellular oxidative stress, radiation and mechanical unloading of tissues as in microgravity. To measure these potential effects of the space environment on telomere lengths and cellular aging and regenerative potential we developed a novel telomere measurement approach based on nanopore sequencing of PCR amplified bar-coded chromosome termini. Specifically, telomeres can be directly enriched using barcode sequences ligated to the end of a free end- repaired telomere using the WetLab-2 facility SmartCycler on ISS. Prior to the ligation and amplification protocol a proteinase K digestion of capping proteins followed by a single 95-degree C heat denaturation of the protease is included. After digestion and bar-code ligation, PCR amplification will initiate with the ligated barcoded sequence, suppressing amplification of intra-genomic fragments and resulting in long read barcoded telomere amplicons including the nanopore motor protein sequences. Purified PCR amplicons are then used for nanopore sequencing library generation by simple addition of motor proteins and sequencing library is loaded into the MinION nanopore DNA-sequencer. Amplicon sequence reads from the nanopore device can be base-called quickly on ISS due to barcoding ligation and subsequent PCR amplification enhancing the telomere sequence resolution. If successfully implemented on ISS this technique will provide a novel means of measuring regenerative ability of somatic stem cells in astronauts, and of determining whether spaceflight in microgravity alters their telomere lengths and causes premature cellular aging.

Ma, Kristin R.↗

Identification and characterization of the WYL BrxR protein and its gene as separable regulatory elements of a BREX phage restriction system

Bacteriophage exclusion (‘BREX’) phage restriction systems are found in a wide range of bacteria. Various BREX systems encode unique combinations of proteins that usually include a site-specific methyltransferase; none appear to contain a nuclease. Here we describe the identification and characterization of a Type I BREX system from Acinetobacter and the effect of deleting each BREX ORF on growth, methylation, and restriction. We identified a previously uncharacterized gene in the BREX operon that is dispensable for methylation but involved in restriction. Biochemical and crystallographic analyses of this factor, which we term BrxR (‘BREX Regulator’), demonstrate that it forms a homodimer and specifically binds a DNA target site upstream of its transcription start site. Deletion of the BrxR gene causes cell toxicity, reduces restriction, and significantly increases the expression of BrxC. In contrast, the introduction of a premature stop codon into the BrxR gene, or a point mutation blocking its DNA binding ability, has little effect on restriction, implying that the BrxR coding sequence and BrxR protein play independent functional roles. We speculate that elements within the BrxR coding sequence are involved in cis regulation of anti-phage activity, while the BrxR protein itself plays an additional regulatory role, perhaps during horizontal transfer.

59 BASIC BIOLOGICAL SCIENCES↗

Bioinformatics Investigations of Universal Stress Proteins from Mercury-Methylating Desulfovibrionaceae

The presence of methylmercury in aquatic environments and marine food sources is of global concern. The chemical reaction for the addition of a methyl group to inorganic mercury occurs in diverse bacterial taxonomic groups including the Gram-negative, sulfate-reducing Desulfovibrionaceae family that inhabit extreme aquatic environments. The availability of whole-genome sequence datasets for members of the Desulfovibrionaceae presents opportunities to understand the microbial mechanisms that contribute to methylmercury production in extreme aquatic environments. We have applied bioinformatics resources and developed visual analytics resources to categorize a collection of 719 putative universal stress protein (USP) sequences predicted from 93 genomes of Desulfovibrionaceae. We have focused our bioinformatics investigations on protein sequence analytics by developing interactive visualizations to categorize Desulfovibrionaceae universal stress proteins by protein domain composition and functionally important amino acids. We identified 651 Desulfovibrionaceae universal stress protein sequences, of which 488 sequences had only one USP domain and 163 had two USP domains. The 488 single USP domain sequences were further categorized into 340 sequences with ATP-binding motif and 148 sequences without ATP-binding motif. The 163 double USP domain sequences were categorized into (1) both USP domains with ATP-binding motif (3 sequences); (2) both USP domains without ATP-binding motif (138 sequences); and (3) one USP domain with ATP-binding motif (21 sequences). We developed visual analytics resources to facilitate the investigation of these categories of datasets in the presence or absence of the mercury-methylating gene pair (hgcAB). Future research could utilize these functional categories to investigate the participation of universal stress proteins in the bacterial cellular uptake of inorganic mercury and methylmercury production, especially in anaerobic aquatic environments.

Isokpehi, Raphael D. (ORCID:0000000268770840)↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

The His-tag as a decoy modulating preferred orientation in cryoEM

The His-tag is a widely used affinity tag that facilitates purification by means of affinity chromatography of recombinant proteins for functional and structural studies. We show here that His-tag presence affects how coproheme decarboxylase interacts with the air-water interface during grid preparation for cryoEM. Depending on His-tag presence or absence, we observe significant changes in patterns of preferred orientation. Our analysis of particle orientations suggests that His-tag presence can mask the hydrophobic and hydrophilic patches on a protein’s surface that mediate the interactions with the air-water interface, while the hydrophobic linker between a His-tag and the coding sequence of the protein may enhance other interactions with the air-water interface. Our observations suggest that tagging, including rational design of the linkers between an affinity tag and a protein of interest, offer a promising approach to modulating interactions with the air-water interface.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular information theory meets protein folding

We propose an application of molecular information theory to analyze the folding of single domain proteins. We analyze results from various areas of protein science, such as sequence-based potentials, reduced amino acid alphabets, backbone configurational entropy, secondary structure content, residue burial layers, and mutational studies of protein stability changes. We found that the average information contained in the sequences of evolved proteins is very close to the average information needed to specify a fold ~2.2 ± 0.3 bits/(site operation). The effective alphabet size in evolved proteins equals the effective number of conformations of a residue in the compact unfolded state at around 5. We calculated an energy-to-information conversion efficiency upon folding of around 50%, lower than the theoretical limit of 70%, but much higher than human built macroscopic machines. We propose a simple mapping between molecular information theory and energy landscape theory and explore the connections between sequence evolution, configurational entropy and the energetics of protein folding.

Ignacio E. Sánchez↗

Functional characteristics of the calcium modulated proteins seen from an evolutionary perspective

We have constructed dendrograms relating 173 EF-hand proteins of known amino acid sequence. We aligned all of these proteins by their EF-hand domains, omitting interdomain regions. Initial dendrograms were computed by minimum mutation distance methods. Using these as starting points, we determined the best dendrogram by the method of maximum parsimony, scored by minimum mutation distance. We identified 14 distinct subfamilies as well as 6 unique proteins that are perhaps the sole representatives of other subfamilies. This information is given in tabular form. Within subfamilies one can easily align interdomain regions. The resulting dendrograms are very similar to those computed using domains only. Dendrograms constructed using pairs of domains show general congruence. However, there are enough exceptions to caution against an overly simple scheme in which one pair of gene duplications leads from one domain precurser to a four domain prototype from which all other forms evolved. The ability to bind calcium was lost and acquired several times during evolution. The distribution of introns does not conform to the dendrogram based on amino acid sequences. The rates of evolution appear to be much slower within subfamilies, especially within calmodulin, than those prior to the definition of subfamily.

Kretsinger, R. H.↗