Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Sequence Similarity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Quantifying Structural Relationships of Metal-Binding Sites Suggests Origins of Biological Electron Transfer

Biological redox reactions drive planetary biogeochemical cycles. Using a novel, structure-guided sequence analysis of proteins, we explored the patterns of evolution of enzymes responsible for these reactions. Our analysis reveals that the folds that bind transition metal–containing ligands have similar structural geometry and amino acid sequences across the full diversity of proteins. Similarity across folds reflects the availability of key transition metals over geological time and strongly suggests that transition metal–ligand binding had a small number of common peptide origins. We observe that structures central to our similarity network come primarily from oxidoreductases, suggesting that ancestral peptides may have also facilitated electron transfer reactions. Last, our results reveal that the earliest biologically functional peptides were likely available before the assembly of fully functional protein domains over 3.8 billion years ago. Thus, life is a special, very complex form of motion of matter, but this form did not always exist, and it is not separated from inorganic nature by an impassable abyss; rather, it arose from inorganic nature as a new property in the process of evolution of the world. We must study the history of this evolution if we want to solve the problem of the origin of life.

Yana Bromberg↗

Classification and evolution of EF-hand proteins

Forty-five distinct subfamilies of EF-hand proteins have been identified. They contain from two to eight EF-hands that are recognizable by amino acid sequence as being statistically similar to other EF-hand domains. All proteins within one subfamily are congruent to one another, i.e. the dendrogram computed from one of the EF-hand domains is similar, within statistical error, to the dendrogram computed from another(s) domain. Thirteen subfamilies--including Calmodulin, Troponin C, Essential light chain, Regulatory light chain--referred to collectively as CTER, are congruent with one another. They appear to have evolved from a single ur-domain by two cycles of gene duplication and fusion. The subfamilies of CTER subsequently evolved by gene duplications and speciations. The remaining 32 subfamilies do not show such general patterns of congruence; however, some--such as S100, intestinal calcium binding protein (calbindin 9 kd), and trichohylin--do not form congruent clusters of subfamilies. Nearly all of the domains 1, 3, 5, and 7 are most similar to other ODD domains. Correspondingly the EVEN numbered domains of all 45 subfamilies most closely resemble EVEN domains of other subfamilies. Many sequence and chemical characteristics do not show systemic trends by subfamily or species of host organisms; such homoplasy is widespread. Eighteen of the subfamilies are heterochimeric; in addition to multiple EF-hands they contain domains of other evolutionary origins.

Non-NASA Center↗

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES↗

Matrix Metalloproteinases as Candidate Antigenic Determinants for Anti‐Tumor Autoantibodies in Human Ovarian Cancer: A Post Hoc Analysis

Circulating antibodies in patients with cancer can facilitate the identification of accessible epitopes on autoantigens expressed by tumors. To identify previously unrecognized protein targets in ovarian cancer, we computationally assessed a heptapeptide consensus motif (VPELGHE, flanked by two cysteine residues yielding a cyclic nonapeptide under oxidizing conditions) previously discovered via phage display-based epitope mapping of autoantibodies in patients. Eight proteins associated with ovarian cancer encompass amino acid sequences similar to the consensus motif and were, therefore, considered as candidate native autoantigens. Among these candidate targets, however, matrix metalloproteinase 14 (MMP14) demonstrates gene expression that is both high and negatively correlated with survival in ovarian cancer patient cohorts. MMP14 protein levels are also stable in tumor versus non-tumor tissues. Moreover, the corresponding heptapeptide mimic in MMP14 occurs within an α-helical secondary structural element observed in its catalytic domain. These findings demonstrate that a subset of patient-derived autoantibodies may interact with a previously unknown antigenic epitope found in MMP14 and other MMPs, thereby providing opportunities for the development of new targeted agents.

Biochemistry & Molecular Biology↗

Characterization and distribution of a maize cDNA encoding a peptide similar to the catalytic region of second messenger dependent protein kinases

Maize (Zea mays) roots respond to a variety of environmental stimuli which are perceived by a specialized group of cells, the root cap. We are studying the transduction of extracellular signals by roots, particularly the role of protein kinases. Protein phosphorylation by kinases is an important step in many eukaryotic signal transduction pathways. As a first phase of this research we have isolated a cDNA encoding a maize protein similar to fungal and animal protein kinases known to be involved in the transduction of extracellular signals. The deduced sequence of this cDNA encodes a polypeptide containing amino acids corresponding to 33 out of 34 invariant or nearly invariant sequence features characteristic of protein kinase catalytic domains. The maize cDNA gene product is more closely related to the branch of serine/threonine protein kinase catalytic domains composed of the cyclic-nucleotide- and calcium-phospholipid-dependent subfamilies than to other protein kinases. Sequence identity is 35% or more between the deduced maize polypeptide and all members of this branch. The high structural similarity strongly suggests that catalytic activity of the encoded maize protein kinase may be regulated by second messengers, like that of all members of this branch whose regulation has been characterized. Northern hybridization with the maize cDNA clone shows a single 2400 base transcript at roughly similar levels in maize coleoptiles, root meristems, and the zone of root elongation, but the transcript is less abundant in mature leaves. In situ hybridization confirms the presence of the transcript in all regions of primary maize root tissue.

NASA Program Space Biology↗

Characterization of the proteins comprising the integral matrix of Strongylocentrotus purpuratus embryonic spicules

In the present study, we enumerate and characterize the proteins that comprise the integral spicule matrix of the Strongylocentrotus purpuratus embryo. Two-dimensional gel electrophoresis of [35S]methionine radiolabeled spicule matrix proteins reveals that there are 12 strongly radiolabeled spicule matrix proteins and approximately three dozen less strongly radiolabeled spicule matrix proteins. The majority of the proteins have acidic isoelectric points; however, there are several spicule matrix proteins that have more alkaline isoelectric points. Western blotting analysis indicates that SM50 is the spicule matrix protein with the most alkaline isoelectric point. In addition, two distinct SM30 proteins are identified in embryonic spicules, and they have apparent molecular masses of approximately 43 and 46 kDa. Comparisons between embryonic spicule matrix proteins and adult spine integral matrix proteins suggest that the embryonic 43-kDa SM30 protein is an embryonic isoform of SM30. An adult 49-kDa spine matrix protein is also identified as a possible adult isoform of SM30. Analysis of the SM30 amino acid sequences indicates that a portion of SM30 proteins is very similar to the carbohydrate recognition domain of C-type lectin proteins.

Non-NASA Center↗

A novel beta-glucosidase from the cell wall of maize (Zea mays L.): rapid purification and partial characterization

Plants have a variety of glycosidic conjugates of hormones, defense compounds, and other molecules that are hydrolyzed by beta-glucosidases (beta-D-glucoside glucohydrolases, E.C. 3.2.1.21). Workers have reported several beta-glucosidases from maize (Zea mays L.; Poaceae), but have localized them mostly by indirect means. We have purified and partly characterized a 58-Ku beta-glucosidase from maize, which we conclude from a partial sequence analysis, from kinetic data, and from its localization is not identical to any of those already reported. A monoclonal antibody, mWP 19, binds this enzyme, and localizes it in the cell walls of maize coleoptiles. An earlier report showed that mWP19 inhibits peroxidase activity in crude cell wall extracts and can immunoprecipitate peroxidase activity from these extracts, yet purified preparations of the 58 Ku protein had little or no peroxidase activity. The level of sequence similarity between beta-glucosidases and peroxidases makes it unlikely that these enzymes share epitopes in common. Contrary to a previous conclusion, these results suggest that the enzyme recognized by mWP19 is not a peroxidase, but there is a wall peroxidase closely associated with the 58 Ku beta-glucosidase in crude preparations. Other workers also have co-purified distinct proteins with beta-glucosidases. We found no significant charge in the level of immunodetectable beta-glucosidase in mesocotyls or coleoptiles that precedes the red light-induced changes in the growth rate of these tissues.

Non-NASA Center↗

Genomic analysis and identification of a novel superantigen, SargEY, in Staphylococcus argenteus isolated from atopic dermatitis lesions

During surveillance of Staphylococcus aureus in lesions from patients with atopic dermatitis (AD), we isolated Staphylococcus argenteus, a species registered in 2011 as a new member of the genus Staphylococcus and previously considered a lineage of S. aureus. Genome sequence comparisons between S. argenteus isolates and representative S. aureus clinical isolates from various origins revealed that the S. argenteus genome from AD patients closely resembles that of S. aureus causing skin infections. We previously reported that 17%–22% of S. aureus isolated from skin infections produce staphylococcal enterotoxin Y (SEY), which predominantly induces T-cell proliferation via the T-cell receptor (TCR) Vα pathway. Complete genome sequencing of S. argenteus isolates revealed a gene encoding a protein similar to superantigen SEY, designated as SargEY, on its chromosome. Population structure analysis of S. argenteus revealed that these isolates are ST2250 lineage, which was the only lineage positive for the SEY-like gene among S. argenteus. Recombinant SargEY demonstrated immunological cross-reactivity with anti-SEY serum. SargEY could induce proliferation of human CD4 + and CD8 + T cells, as well as production of TNF-α and IFN-γ. SargEY showed emetic activity in a marmoset monkey model. S arg EY and SET (a phylogenetically close but uncharacterized SE) revealed their dependency on TCR Vα in inducing human T-cell proliferation. Additionally, TCR sequencing revealed other previously undescribed Vα repertoires induced by SEH. S arg EY and SEY may play roles in exacerbating the respective toxin-producing strains in AD.

59 BASIC BIOLOGICAL SCIENCES↗

Isolation, characterization, and amino acid sequences of auracyanins, blue copper proteins from the green photosynthetic bacterium Chloroflexus aurantiacus

Three small blue copper proteins designated auracyanin A, auracyanin B-1, and auracyanin B-2 have been isolated from the thermophilic green gliding photosynthetic bacterium Chloroflexus aurantiacus. All three auracyanins are peripheral membrane proteins. Auracyanin A was described previously (Trost, J. T., McManus, J. D., Freeman, J. C., Ramakrishna, B. L., and Blankenship, R. E. (1988) Biochemistry 27, 7858-7863) and is not glycosylated. The two B forms are glycoproteins and have almost identical properties to each other, but are distinct from the A form. The sodium dodecyl sulfate-polyacrylamide gel electrophoresis apparent monomer molecular masses are 14 (A), 18 (B-2), and 22 (B-1) kDa. The amino acid sequences of the B forms are presented. All three proteins have similar absorbance, circular dichroism, and resonance Raman spectra, but the electron spin resonance signals are quite different. Laser flash photolysis kinetic analysis of the reactions of the three forms of auracyanin with lumiflavin and flavin mononucleotide semiquinones indicates that the site of electron transfer is negatively charged and has an accessibility similar to that found in other blue copper proteins. Copper analysis indicates that all three proteins contain 1 mol of copper per mol of protein. All three auracyanins exhibit a midpoint redox potential of +240 mV. Light-induced absorbance changes and electron spin resonance signals suggest that auracyanin A may play a role in photosynthetic electron transfer. Kinetic data indicate that all three proteins can donate electrons to cytochrome c-554, the electron donor to the photosynthetic reaction center.

NASA Discipline Exobiology↗

Use of Split‐Intein Proteins to Design a Small Molecule Biosensor in Plants

Understanding how plants perceive their environment is fundamental to advancing agricultural productivity and sustainability. Many biological small molecules, including those involved in microbial recognition, act rapidly at the plant cell surface, but the absence of tools to visualise these dynamics has limited our ability to dissect plant–microbe communication. To address this gap, we sought to create a genetically encoded biosensor that couples ligand-induced protein dimerization with the production of a fluorescent reporter. Inteins are peptide regions that excise themselves from precursor proteins and ligate the flanking chains (exteins). When each half of a split intein is fused to one of two dimerizing proteins, ligand binding brings them into proximity, inducing intein splicing and ligation of flanking extein sequences (Kang et al. 2022). Similar to previous studies, we split the yeast vacuolar ATPase subunit 1 (VMA1) intein, creating a protein biosensor that produces eGFP upon protein dimerization after ligand binding (Figure 1A) (Mootz et al. 2003). Specifically, eGFP halves (i.e., non-functional N- and C-terminal GFP fragments) were fused to the intein halves, resulting in two fusion proteins: N-terminal GFP::N-terminal intein and C-terminal intein::C-terminal GFP (Figure 1A).

Boone, Brandon A. [Oak Ridge National Laboratory (↗

Two genes with similarity to bacterial response regulators are rapidly and specifically induced by cytokinin in Arabidopsis

Cytokinins are central regulators of plant growth and development, but little is known about their mode of action. By using differential display, we identified a gene, IBC6 (for induced by cytokinin), from etiolated Arabidopsis seedlings, that is induced rapidly by cytokinin. The steady state level of IBC6 mRNA was elevated within 10 min by the exogenous application of cytokinin, and this induction did not require de novo protein synthesis. IBC6 was not induced by other plant hormones or by light. A second Arabidopsis gene with a sequence highly similar to IBC6 was identified. This IBC7 gene also was induced by cytokinin, although with somewhat slower kinetics and to a lesser extent. The pattern of expression of the two genes was similar, with higher expression in leaves, rachises, and flowers and lower transcript levels in roots and siliques. Sequence analysis revealed that IBC6 and IBC7 are similar to the receiver domain of bacterial two-component response regulators. This homology, coupled with previously published work on the CKI1 histidine kinase homolog, suggests that these proteins may play a role in early cytokinin signaling.

Non-NASA Center↗

Zuotin, a putative Z-DNA binding protein in Saccharomyces cerevisiae

A putative Z-DNA binding protein, named zuotin, was purified from a yeast nuclear extract by means of a Z-DNA binding assay using [32P]poly(dG-m5dC) and [32P]oligo(dG-Br5dC)22 in the presence of B-DNA competitor. Poly(dG-Br5dC) in the Z-form competed well for the binding of a zuotin containing fraction, but salmon sperm DNA, poly(dG-dC) and poly(dA-dT) were not effective. Negatively supercoiled plasmid pUC19 did not compete, whereas an otherwise identical plasmid pUC19(CG), which contained a (dG-dC)7 segment in the Z-form was an excellent competitor. A Southwestern blot using [32P]poly(dG-m5dC) as a probe in the presence of MgCl2 identified a protein having a molecular weight of 51 kDa. The 51 kDa zuotin was partially sequenced at the N-terminal and the gene, ZUO1, was cloned, sequenced and expressed in Escherichia coli; the expressed zuotin showed similar Z-DNA binding activity, but with lower affinity than zuotin that had been partially purified from yeast. Zuotin was deduced to have a number of potential phosphorylation sites including two CDC28 (homologous to the human and Schizosaccharomyces pombe cdc2) phosphorylation sites. The hexapeptide motif KYHPDK was found in zuotin as well as in several yeast proteins, DnaJ of E.coli, csp29 and csp32 proteins of Drosophila and the small t and large T antigens of the polyoma virus. A 60 amino acid segment of zuotin has similarity to several histone H1 sequences. Disruption of ZUO1 in yeast resulted in a slow growth phenotype.

NASA Discipline Exobiology↗

Enhanced polymorph metastability drives glycine nucleation in aqueous salt solutions

Crystal nucleation from aqueous solutions influences countless geological, biochemical, astrophysical, environmental, and materials science–related phenomena, including ice formation, the manufacturing of active pharmaceutical ingredients, development of diseases such as Alzheimer’s and the origin of life itself. Understanding and controlling nucleation is essential for designing materials with specific properties, developing strategies to inhibit or promote crystallization in various contexts and preventing pathological aggregation in neurodegenerative diseases. Similar to the protein structure prediction problem—where a single amino acid sequence can in theory adopt one most stable conformation but in practice may sample multiple competing conformations—crystal nucleation faces a parallel challenge: the same chemical species can form diverse polymorphs under different environmental conditions (e.g., temperature, pressure, solvent). Each polymorph presents its own set of physical and chemical properties, highlighting the importance of understanding and controlling polymorph selection in fields ranging from pharmaceuticals to materials design. Despite advances in experimental and computational methods for studying phase transitions and polymorph stability, nucleation remains challenging due to its nanoscale nature. Furthermore, in practical settings, salts and impurities can further influence crystal nucleation in diverse contexts, from scaling in pipelines and desalination plants to the durability of concrete and the efficiency of battery materials. This can lead to the formation of polymorphs that may differ from the most stable phase in pure solutions. Or, even though the final structure might appear same irrespective of whether the environment contained impurities or not, the mechanism through which it was formed might be completely different and not intuitive.

Wang, Ruiyu [University of Maryland, College Park,↗

Native Chemical Ligation of Peptoid Oligomers

Bioorganic chemists are inspired by natural biopolymers to design peptidomimetic oligomers that can exhibit sequence-structure-function relationships. Biomimetic polymers can be synthesized to incorporate a specific sequence of nonbiological monomer units using a variety of iterative solution-phase or solid-phase reaction schemes. These protocols generally provide access to a vast diversity of oligomeric compounds but are limited with respect to their ability to attain protein-like chain lengths. This constraint can preclude access to sequence-defined synthetic macromolecules with sufficient sizes required to exhibit tertiary structure and other protein-mimetic attributes. In contrast, peptide chemists have overcome this limitation by developing convergent synthetic methods, such as native chemical ligation, to join individual, smaller peptide chains together to make larger peptides or full proteins. A similar convergent approach is needed to establish efficient synthetic routes to non-natural sequence-defined macromolecules. Herein, we adapt the peptide native chemical ligation method to peptoid oligomers, demonstrating how short chains can be conjoined to create sequence-defined peptoid macromolecules. Nanosheet-forming peptoid polymers with distinct surface loop display domains were generated by sequential ligation of several discrete fragments. This method provides a reliable convergent ligation route for sequence-defined polypeptoids that results in a native amide bond joining the fragments. We envision that this strategy will be useful in synthesizing peptoid-based proteomimetics that incorporate diverse chemical features.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sequence information signal processor for local and global string comparisons

A sequence information signal processing integrated circuit chip designed to perform high speed calculation of a dynamic programming algorithm based upon the algorithm defined by Waterman and Smith. The signal processing chip of the present invention is designed to be a building block of a linear systolic array, the performance of which can be increased by connecting additional sequence information signal processing chips to the array. The chip provides a high speed, low cost linear array processor that can locate highly similar global sequences or segments thereof such as contiguous subsequences from two different DNA or protein sequences. The chip is implemented in a preferred embodiment using CMOS VLSI technology to provide the equivalent of about 400,000 transistors or 100,000 gates. Each chip provides 16 processing elements, and is designed to provide 16 bit, two's compliment operation for maximum score precision of between -32,768 and +32,767. It is designed to provide a comparison between sequences as long as 4,194,304 elements without external software and between sequences of unlimited numbers of elements with the aid of external software. Each sequence can be assigned different deletion and insertion weight functions. Each processor is provided with a similarity measure device which is independently variable. Thus, each processor can contribute to maximum value score calculation using a different similarity measure.

Peterson, John C.↗

Evolution and folding of repeat proteins

Repeat proteins are made with tandem copies of similar amino acid stretches that fold into elongated architectures. These proteins constitute excellent model systems to investigate how evolution relates to structure, folding, and function. Here, we propose a scheme to map evolutionary information at the sequence level to a coarse-grained model for repeat-protein folding and use it to investigate the folding of thousands of repeat proteins. We model the energetics by a combination of an inverse Potts-model scheme with an explicit mechanistic model of duplications and deletions of repeats to calculate the evolutionary parameters of the system at the single-residue level. These parameters are used to inform an Ising-like model that allows for the generation of folding curves, apparent domain emergence, and occupation of intermediate states that are highly compatible with experimental data in specific case studies. We analyzed the folding of thousands of natural Ankyrin repeat proteins and found that a multiplicity of folding mechanisms are possible. Fully cooperative all-or-none transitions are obtained for arrays with enough sequence-similar elements and strong interactions between them, while noncooperative element-by-element intermittent folding arose if the elements are dissimilar and the interactions between them are energetically weak. Additionally, we characterized nucleation-propagation and multidomain folding mechanisms. We show that the global stability and cooperativity of the repeating arrays can be predicted from simple sequence scores.

Ezequiel A. Galpern↗

Uncovering Sequence and Structural Characteristics of Fungal Expansin‐Related Proteins With Potential to Drive Substrate Targeting

Expansins loosen plant cell wall networks through disrupting non-covalent bonds between cellulose microfibrils and matrix polysaccharides. Whereas expansins were first discovered in plants, expansin-related proteins have since been identified in bacteria and fungi. The biological function of microbial expansins remains unclear; however, several studies have shown distinct binding preferences toward different structural polysaccharides. Earlier studies of bacterial expansin-related proteins uncovered sequence and structural features that correlate to substrate binding. Herein, 20 fungal expansin-related sequences were recombinantly produced in Komagataella phaffii, and the purified proteins were compared in terms of substrate binding to cellulosic and chitinous substrates. The impact of pH on the zeta potential of prioritized substrates was also measured, and Principal Component Analysis was performed to uncover correlations between protein characteristics (e.g., pI, hydrophobicity, surface charge distribution) and measured substrate binding preferences. Whereas acidic proteins with a predicted pI less than 5.0 preferentially bound to chitin, basic proteins with pI greater than 8.0 preferentially bound to xylan and xylan-containing fiber. Similar to many cellulases, binding to cellulose was correlated to relatively high aromatic amino acid content in the protein sequence and presence of a carbohydrate binding module (CBM), which in the case of expansins is a C-terminal CBM63. Whereas overall sequence characteristics could be correlated to substrate binding preference, the identity of amino acids occupying conserved positions that impact protein activity was better correlated with loosenin versus expansin classifications.

chitin↗

Alpha-amylase from the Hyperthermophilic Archaeon Thermococcus thioreducens

Extremophiles are microorganisms that thrive in, from an anthropocentric view, extreme environments such as hot springs. The ability of survival at extreme conditions has rendered enzymes from extremophiles to be of interest in industrial applications. One approach to producing these extremozymes entails the expression of the enzyme-encoding gene in a mesophilic host such as E.coli. This method has been employed in the effort to produce an alpha-amylase from a hyperthermophile (an organism that displays optimal growth above 80 C) isolated from a hydrothermal vent at the Rainbow vent site in the Atlantic Ocean. alpha-amylases catalyze the hydrolysis of starch to produce smaller sugars and constitute a class of industrial enzymes having approximately 25% of the enzyme market. One application for thermostable alpha-amylases is the starch liquefaction process in which starch is converted into fructose and glucose syrups. The a-amylase encoding gene from the hyperthermophile Thermococcus thioreducens was cloned and sequenced, revealing high similarity with other archaeal hyperthermophilic a-amylases. The gene encoding the mature protein was expressed in E.coli. Initial characterization of this enzyme has revealed an optimal amylolytic activity between 85-90 C and around pH 5.3-6.0.

Bernhardsdotter, E. C. M. J.↗