SEARCH · Engineering Papers
Results for “Sequence Homology”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Evolution of EF-hand calcium-modulated proteins. III. Exon sequences confirm most dendrograms based on protein sequences: calmodulin dendrograms show significant lack of parallelism
In the first report in this series we presented dendrograms based on 152 individual proteins of the EF-hand family. In the second we used sequences from 228 proteins, containing 835 domains, and showed that eight of the 29 subfamilies are congruent and that the EF-hand domains of the remaining 21 subfamilies have diverse evolutionary histories. In this study we have computed dendrograms within and among the EF-hand subfamilies using the encoding DNA sequences. In most instances the dendrograms based on protein and on DNA sequences are very similar. Significant differences between protein and DNA trees for calmodulin remain unexplained. In our fourth report we evaluate the sequences and the distribution of introns within the EF-hand family and conclude that exon shuffling did not play a significant role in its evolution.
Evolution of EF-hand calcium-modulated proteins. II. Domains of several subfamilies have diverse evolutionary histories
In the first report in this series we described the relationships and evolution of 152 individual proteins of the EF-hand subfamilies. Here we add 66 additional proteins and define eight (CDC, TPNV, CLNB, LPS, DGK, 1F8, VIS, TCBP) new subfamilies and seven (CAL, SQUD, CDPK, EFH5, TPP, LAV, CRGP) new unique proteins, which we assume represent new subfamilies. The main focus of this study is the classification of individual EF-hand domains. Five subfamilies--calmodulin, troponin C, essential light chain, regulatory light chain, CDC31/caltractin--and three uniques--call, squidulin, and calcium-dependent protein kinase--are congruent in that all evolved from a common four-domain precursor. In contrast calpain and sarcoplasmic calcium-binding protein (SARC) each evolved from its own one-domain precursor. The remaining 19 subfamilies and uniques appear to have evolved by translocation and splicing of genes encoding the EF-hand domains that were precursors to the congruent eight and to calpain and to SARC. The rates of evolution of the EF-hand domains are slower following formation of the subfamilies and establishment of their functions. Subfamilies are not readily classified by patterns of calcium coordination, interdomain linker stability, and glycine and proline distribution. There are many homoplasies indicating that similar variants of the EF-hand evolved by independent pathways.
Comparative sequence analyses on the 16S rRNA (rDNA) of Bacillus acidocaldarius, Bacillus acidoterrestris, and Bacillus cycloheptanicus and proposal for creation of a new genus, Alicyclobacillus gen. nov
Comparative 16S rRNA (rDNA) sequence analyses performed on the thermophilic Bacillus species Bacillus acidocaldarius, Bacillus acidoterrestris, and Bacillus cycloheptanicus revealed that these organisms are sufficiently different from the traditional Bacillus species to warrant reclassification in a new genus, Alicyclobacillus gen. nov. An analysis of 16S rRNA sequences established that these three thermoacidophiles cluster in a group that differs markedly from both the obligately thermophilic organisms Bacillus stearothermophilus and the facultatively thermophilic organism Bacillus coagulans, as well as many other common mesophilic and thermophilic Bacillus species. The thermoacidophilic Bacillus species B. acidocaldarius, B. acidoterrestris, and B. cycloheptanicus also are unique in that they possess omega-alicylic fatty acid as the major natural membranous lipid component, which is a rare phenotype that has not been found in any other Bacillus species characterized to date. This phenotype, along with the 16S rRNA sequence data, suggests that these thermoacidophiles are biochemically and genetically unique and supports the proposal that they should be reclassified in the new genus Alicyclobacillus.
Purification and sequence analysis of two rat tissue inhibitors of metalloproteinases
Two protein inhibitors of metalloproteinases (TIMP) were isolated from medium conditioned by the clonal rat osteosarcoma line UMR 106-01. Initial purification of both a 30-kDa inhibitor and a 20-kDa inhibitor was accomplished using heparin-Sepharose chromatography with dextran sulfate elution followed by DEAE-Sepharose and CM-Sepharose chromatography. Purification of the 20-kDa inhibitor to homogeneity was completed with reverse-phase high-performance liquid chromatography. The 20-kDa inhibitor was identified as rat TIMP-2. The 30-kDa inhibitor, although not purified to homogeneity, was identified as rat TIMP-1. Amino terminal amino acid sequence analysis of the 30-kDa inhibitor demonstrated 86% identity to human TIMP-1 for the first 22 amino acids while the sequence of the 20-kDa inhibitor was identical to that of human TIMP-2 for the first 22 residues. Treatment with peptide:N-glycosidase F indicated that the 30-kDa rat inhibitor is glycosylated while the 20-kDa inhibitor is apparently unglycosylated. Inhibition of both rat and human interstitial collagenase by rat TIMP-2 was stoichiometric, with a 1:1 molar ratio required for complete inhibition. Exposure of UMR 106-01 cells to 10(-7) M parathyroid hormone resulted in approximately a 40% increase in total inhibitor production over basal levels.
A definition of the domains Archaea, Bacteria and Eucarya in terms of small subunit ribosomal RNA characteristics
The number of small subunit rRNA sequences is now great enough that the three domains Archaea, Bacteria and Eucarya (Woese et al., 1990) can be reliably defined in terms of their sequence "signatures". Approximately 50 homologous positions (or nucleotide pairs) in the small subunit rRNA characterize and distinguish among the three. In addition, the three can be recognized by a variety of nonhomologous rRNA characters, either individual positions and/or higher-order structural features. The Crenarchaeota and the Euryarchaeota, the two archaeal kingdoms, can also be defined and distinguished by their characteristic compositions at approximately fifteen positions in the small subunit rRNA molecule.
Novel End-to-End Molecular Biology Approach for Direct Nanopore 1D cDNA Sequencing of Reverse Transcribed mRNAs Purified from Cell Cultures by the NASA ISS WetLab2 SPM
Continued space bioscience research onboard the International Space Station (ISS) and future long-duration flight missions to the Moon or Mars will require the ability to conduct on-orbit molecular analysis of biological samples independently from Earth. In the last year two new molecular analytic technologies have been installed and the technologies demonstrated onboard the ISS: The Sample Prep Module (SPM) WetLab-2 (WL2) qRT-PCR toolbox and the Oxford Nanopore MinIon Biomolecule Sequencer. Here we describe protocol development and integration into existing ISS technology for end-to-end on-orbit biological sample processing and molecular analysis with real time results generated utilizing only field offline analytic software. For this experiment we isolated primary cells from bone marrow flushes of wild type B6129SF2 mice (Jackson Labs) long bones. The cell isolate was then processed using the SPM to produce total 147nanograms of RNA. The total RNA was purified to only messenger RNA (mRNA) and transferred to Smartcycler Thermocycle ISS kit consumable tube using Eppendorf gel loading pipette tips for further processing. Complementary first strand cDNA was synthesized using OLIGO dT priming followed by addition of SuperScript II Reverse Transcriptase and thermal cycling as per manufacturers instruction. All thermal cycling was conducted using the ISS WetLab-2 Cephid Smarcycler real time thermal cycler. Our protocol takes advantage of mRNAs native poly(A) tail, synthesized in vivo to protect the mRNA from degradation by endonucleases, to eliminate end-prep for adapter ligation. The adapted library is purified using MyOne C1 Streptavidin beads before elution in buffer. The pre-sequencing library is diluted in the loading buffer and injected into the MinIon sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown software with local basecalling. The sequencing read produced 34.5 million events and local basecalling produced 117,301 successful reads. NCBI Blast of the data for the mouse genome resulted in 2,462 successful nucleotide collection matches (gene sequences) exceeding 70 homology. These results demonstrate the viability of this novel flight ready end-to-end sample analytic methodology and provide a real time homolog for flight experimentation utilizing supply kits and technologies that have already been demonstrated on ISS.
Phylogenetic origins of the plant mitochondrion based on a comparative analysis of 5S ribosomal RNA sequences
The complete nucleotide sequences of 5S ribosomal RNAs from Rhodocyclus gelatinosa, Rhodobacter sphaeroides, and Pseudomonas cepacia were determined. Comparisons of these 5S RNA sequences show that rather than being phylogenetically related to one another, the two photosynthetic bacterial 5S RNAs share more sequence and signature homology with the RNAs of two nonphotosynthetic strains. Rhodobacter sphaeroides is specifically related to Paracoccus denitrificans and Rc. gelatinosa is related to Ps. cepacia. These results support earlier 16S ribosomal RNA studies and add two important groups to the 5S RNA data base. Unique 5S RNA structural features previously found in P. denitrificans are present also in the 5S RNA of Rb. sphaeroides; these provide the basis for subdivisional signatures. The immediate consequence of obtaining these new sequences is that it is possible to clarify the phylogenetic origins of the plant mitochondrion. In particular, a close phylogenetic relationship is found between the plant mitochondria and members of the alpha subdivision of the purple photosynthetic bacteria, namely, Rb. sphaeroides, P. denitrificans, and Rhodospirillum rubrum.
Functional role of myosin-binding protein H in thick filaments of developing vertebrate fast-twitch skeletal muscle
Myosin-binding protein H (MyBP-H) is a component of the vertebrate skeletal muscle sarcomere with sequence and domain homology to myosin-binding protein C (MyBP-C). Whereas skeletal muscle isoforms of MyBP-C (fMyBP-C, sMyBP-C) modulate muscle contractility via interactions with actin thin filaments and myosin motors within the muscle sarcomere “C-zone,” MyBP-H has no known function. This is in part due to MyBP-H having limited expression in adult fast-twitch muscle and no known involvement in muscle disease. Quantitative proteomics reported here reveal that MyBP-H is highly expressed in prenatal rat fast-twitch muscles and larval zebrafish, suggesting a conserved role in muscle development and prompting studies to define its function. We take advantage of the genetic control of the zebrafish model and a combination of structural, functional, and biophysical techniques to interrogate the role of MyBP-H. Transgenic, FLAG-tagged MyBP-H or fMyBP-C both localize to the C-zones in larval myofibers, whereas genetic depletion of endogenous MyBP-H or fMyBP-C leads to increased accumulation of the other, suggesting competition for C-zone binding sites. Does MyBP-H modulate contractility in the C-zone? Globular domains critical to MyBP-C’s modulatory functions are absent from MyBP-H, suggesting that MyBP-H may be functionally silent. However, our results suggest an active role. In vitro motility experiments indicate MyBP-H shares MyBP-C’s capacity as a molecular “brake.” These results provide new insights and raise questions about the role of the C-zone during muscle development.
Numerical classification of coding sequences
DNA sequences coding for protein may be represented by counts of nucleotides or codons. A complete reading frame may be abbreviated by its base count, e.g. A76C158G121T74, or with the corresponding codon table, e.g. (AAA)0(AAC)1(AAG)9 ... (TTT)0. We propose that these numerical designations be used to augment current methods of sequence annotation. Because base counts and codon tables do not require revision as knowledge of function evolves, they are well-suited to act as cross-references, for example to identify redundant GenBank entries. These descriptors may be compared, in place of DNA sequences, to extract homologous genes from large databases. This approach permits rapid searching with good selectivity.
Transcription factor IID in the Archaea: sequences in the Thermococcus celer genome would encode a product closely related to the TATA-binding protein of eukaryotes
The first step in transcription initiation in eukaryotes is mediated by the TATA-binding protein, a subunit of the transcription factor IID complex. We have cloned and sequenced the gene for a presumptive homolog of this eukaryotic protein from Thermococcus celer, a member of the Archaea (formerly archaebacteria). The protein encoded by the archaeal gene is a tandem repeat of a conserved domain, corresponding to the repeated domain in its eukaryotic counterparts. Molecular phylogenetic analyses of the two halves of the repeat are consistent with the duplication occurring before the divergence of the archael and eukaryotic domains. In conjunction with previous observations of similarity in RNA polymerase subunit composition and sequences and the finding of a transcription factor IIB-like sequence in Pyrococcus woesei (a relative of T. celer) it appears that major features of the eukaryotic transcription apparatus were well-established before the origin of eukaryotic cellular organization. The divergence between the two halves of the archael protein is less than that between the halves of the individual eukaryotic sequences, indicating that the average rate of sequence change in the archael protein has been less than in its eukaryotic counterparts. To the extent that this lower rate applies to the genome as a whole, a clearer picture of the early genes (and gene families) that gave rise to present-day genomes is more apt to emerge from the study of sequences from the Archaea than from the corresponding sequences from eukaryotes.
NEAR: Neural Embeddings for Amino acid Relationships
Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.
Evolution of EF-hand calcium-modulated proteins. IV. Exon shuffling did not determine the domain compositions of EF-hand proteins
In the previous three reports in this series we demonstrated that the EF-hand family of proteins evolved by a complex pattern of gene duplication, transposition, and splicing. The dendrograms based on exon sequences are nearly identical to those based on protein sequences for troponin C, the essential light chain myosin, the regulatory light chain, and calpain. This validates both the computational methods and the dendrograms for these subfamilies. The proposal of congruence for calmodulin, troponin C, essential light chain, and regulatory light chain was confirmed. There are, however, significant differences in the calmodulin dendrograms computed from DNA and from protein sequences. In this study we find that introns are distributed throughout the EF-hand domain and the interdomain regions. Further, dendrograms based on intron type and distribution bear little resemblance to those based on protein or on DNA sequences. We conclude that introns are inserted, and probably deleted, with relatively high frequency. Further, in the EF-hand family exons do not correspond to structural domains and exon shuffling played little if any role in the evolution of this widely distributed homolog family. Calmodulin has had a turbulent evolution. Its dendrograms based on protein sequence, exon sequence, 3'-tail sequence, intron sequences, and intron positions all show significant differences.
Cloning of the cDNA for U1 small nuclear ribonucleoprotein particle 70K protein from Arabidopsis thaliana
We cloned and sequenced a plant cDNA that encodes U1 small nuclear ribonucleoprotein (snRNP) 70K protein. The plant U1 snRNP 70K protein cDNA is not full length and lacks the coding region for 68 amino acids in the amino-terminal region as compared to human U1 snRNP 70K protein. Comparison of the deduced amino acid sequence of the plant U1 snRNP 70K protein with the amino acid sequence of animal and yeast U1 snRNP 70K protein showed a high degree of homology. The plant U1 snRNP 70K protein is more closely related to the human counter part than to the yeast 70K protein. The carboxy-terminal half is less well conserved but, like the vertebrate 70K proteins, is rich in charged amino acids. Northern analysis with the RNA isolated from different parts of the plant indicates that the snRNP 70K gene is expressed in all of the parts tested. Southern blotting of genomic DNA using the cDNA indicates that the U1 snRNP 70K protein is coded by a single gene.
Isolation and characterization of a novel calmodulin-binding protein from potato
Tuberization in potato is controlled by hormonal and environmental signals. Ca(2+), an important intracellular messenger, and calmodulin (CaM), one of the primary Ca(2+) sensors, have been implicated in controlling diverse cellular processes in plants including tuberization. The regulation of cellular processes by CaM involves its interaction with other proteins. To understand the role of Ca(2+)/CaM in tuberization, we have screened an expression library prepared from developing tubers with biotinylated CaM. This screening resulted in isolation of a cDNA encoding a novel CaM-binding protein (potato calmodulin-binding protein (PCBP)). Ca(2+)-dependent binding of the cDNA-encoded protein to CaM is confirmed by (35)S-labeled CaM. The full-length cDNA is 5 kb long and encodes a protein of 1309 amino acids. The deduced amino acid sequence showed significant similarity with a hypothetical protein from another plant, Arabidopsis. However, no homologs of PCBP are found in nonplant systems, suggesting that it is likely to be specific to plants. Using truncated versions of the protein and a synthetic peptide in CaM binding assays we mapped the CaM-binding region to a 20-amino acid stretch (residues 1216-1237). The bacterially expressed protein containing the CaM-binding domain interacted with three CaM isoforms (CaM2, CaM4, and CaM6). PCBP is encoded by a single gene and is expressed differentially in the tissues tested. The expression of CaM, PCBP, and another CaM-binding protein is similar in different tissues and organs. The predicted protein contained seven putative nuclear localization signals and several strong PEST motifs. Fusion of the N-terminal region of the protein containing six of the seven nuclear localization signals to the reporter gene beta-glucuronidase targeted the reporter gene to the nucleus, suggesting a nuclear role for PCBP.
Transcription in archaea
Using the sequences of all the known transcription-associated proteins from Bacteria and Eucarya (a total of 4,147), we have identified their homologous counterparts in the four complete archaeal genomes. Through extensive sequence comparisons, we establish the presence of 280 predicted transcription factors or transcription-associated proteins in the four archaeal genomes, of which 168 have homologs only in Bacteria, 51 have homologs only in Eucarya, and the remaining 61 have homologs in both phylogenetic domains. Although bacterial and eukaryotic transcription have very few factors in common, each exclusively shares a significantly greater number with the Archaea, especially the Bacteria. This last fact contrasts with the obvious close relationship between the archaeal and eukaryotic transcription mechanisms per se, and in particular, basic transcription initiation. We interpret these results to mean that the archaeal transcription system has retained more ancestral characteristics than have the transcription mechanisms in either of the other two domains.
An intron within the 16S ribosomal RNA gene of the archaeon Pyrobaculum aerophilum
The 16S rRNA genes of Pyrobaculum aerophilum and Pyrobaculum islandicum were amplified by the polymerase chain reaction, and the resulting products were sequenced directly. The two organisms are closely related by this measure (over 98% similar). However, they differ in that the (lone) 16S rRNA gene of Pyrobaculum aerophilum contains a 713-bp intron not seen in the corresponding gene of Pyrobaculum islandicum. To our knowledge, this is the only intron so far reported in the small subunit rRNA gene of a prokaryote. Upon excision the intron is circularized. A secondary structure model of the intron-containing rRNA suggests a splicing mechanism of the same type as that invoked for the tRNA introns of the Archaea and Eucarya and 23S rRNAs of the Archaea. The intron contains an open reading frame whose protein translation shows no certain homology with any known protein sequence.
DNA Recombinase Proteins, their Function and Structure in the Active Form, a Computational Study
Homologous recombination is a crucial sequence of reactions in all cells for the repair of double strand DNA (dsDNA) breaks. While it was traditionally considered as a means for generating genetic diversity, it is now known to be essential for restart of collapsed replication forks that have met a lesion on the DNA template (Cox et al., 2000). The central stage of this process requires the presence of the DNA recombinase protein, RecA in bacteria, RadA in archaea, or Rad51 in eukaryotes, which leads to an ATP-mediated DNA strand-exchange process. Despite many years of intense study, some aspects of the biochemical mechanism, and structure of the active form of recombinase proteins are not well understood. Our theoretical study is an attempt to shed light on the main structural and mechanistic issues encountered on the RecA of the e-coli, the RecA of the extremely radio resistant Deinococcus Radiodurans (promoting an inverse DNA strand-exchange repair), and the homolog human Rad51. The conformational changes are analyzed for the naked enzymes, and when they are linked to ATP and ADP. The average structures are determined over 2ns time scale of Langevian dynamics using a collision frequency of 1.0 ps(sup -1). The systems are inserted in an octahedron periodic box with a 10 Angstrom buffer of water molecules explicitly described by the TIP3P model. The corresponding binding free energies are calculated in an implicit solvent using the Poisson-Boltzmann solvent accessible surface area, MM-PBSA model. The role of the ATP is not only in stabilizing the interaction RecA-DNA, but its hydrolysis is required to allow the DNA strand-exchange to proceed. Furthermore, we extended our study, using the hybrid QM/MM method, on the mechanism of this chemical process. All the calculations were performed using the commercial code Amber 9.