Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Epitopes recognition of SARS-CoV-2 nucleocapsid RNA binding domain by human monoclonal antibodies

Coronavirus nucleocapsid protein (NP) of SARS-CoV-2 plays a central role in many functions important for virus proliferation including packaging and protecting genomic RNA. The protein shares sequence, structure, and architecture with nucleocapsid proteins from betacoronaviruses. The N-terminal domain (NP RBD ) binds RNA and the C-terminal domain is responsible for dimerization. After infection, NP is highly expressed and triggers robust host immune response. The anti-NP antibodies are not protective and not neutralizing but can effectively detect viral proliferation soon after infection. Two structures of SARS-CoV-2 NP RBD were determined providing a continuous model from residue 48 to 173, including RNA binding region and key epitopes. Five structures of NP RBD complexes with human mAbs were isolated using an antigen-bait sorting. Complexes revealed a distinct complement-determining regions and unique sets of epitope recognition. This may assist in the early detection of pathogens and designing peptide-based vaccines. Mutations that significantly increase viral load were mapped on developed, full length NP model, likely impacting interactions with host proteins and viral RNA.

59 BASIC BIOLOGICAL SCIENCES↗

Putting Humpty Dumpty Back Together Again: What Does Protein Quantification Mean in Bottom-Up Proteomics?

Bottom-up proteomics provides peptide measurements and has been invaluable for moving proteomics into large-scale analyses. Commonly, a single quantitative value is reported for each protein-coding gene by aggregating peptide quantities into protein groups following protein inference or parsimony. However, given the complexity of both RNA splicing and post-translational protein modification, it is overly simplistic to assume that all peptides that map to a singular protein-coding gene will demonstrate the same quantitative response. Here, by assuming that all peptides from a protein-coding sequence are representative of the same protein, we may miss the discovery of important biological differences. To capture the contributions of existing proteoforms, we need to reconsider the practice of aggregating protein values to a single quantity per protein-coding gene.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Identification of a nuclear localization sequence in the polyomavirus capsid protein VP2

A nuclear localization signal (NLS) has been identified in the C-terminal (Glu307-Glu-Asp-Gly-Pro-Gln-Lys-Lys-Lys-Arg-Arg-Leu318) amino acid sequence of the polyomavirus minor capsid protein VP2. The importance of this amino acid sequence for nuclear transport of newly synthesized VP2 was demonstrated by a genetic "subtractive" study using the constructs pSG5VP2 (expressing full-length VP2) and pSG5 delta 3VP2 (expressing truncated VP2, lacking amino acids Glu307-Leu318). These constructs were transfected into COS-7 cells, and the intracellular localization of the VP2 protein was determined by indirect immunofluorescence. These studies revealed that the full-length VP2 was localized in the nucleus, while the truncated VP2 protein was localized in the cytoplasm and not transported to the nucleus. A biochemical "additive" approach was also used to determine whether this sequence could target nonnuclear proteins to the nucleus. A synthetic peptide identical to VP2 amino acids Glu307-Leu318 was cross-linked to the nonnuclear proteins bovine serum albumin (BSA) or immunoglobulin G (IgG). The conjugates were then labeled with fluorescein isothiocyanate and microinjected into the cytoplasm of NIH 3T6 cells. Both conjugates localized in the nucleus of the microinjected cells, whereas unconjugated BSA and IgG remained in the cytoplasm. Taken together, these genetic subtractive and biochemical additive approaches have identified the C-terminal sequence of polyoma-virus VP2 (containing amino acids Glu307-Leu318) as the NLS of this protein.

Non-NASA Center↗

Aspartate Residues in a Forisome-Forming SEO Protein Are Critical for Protein Body Assembly and Ca 2+ Responsiveness

Forisomes are protein bodies known exclusively from sieve elements of legumes. Forisomes contribute to the regulation of phloem transport due to their unique Ca 2+ -controlled, reversible swelling. The assembly of forisomes from sieve element occlusion (SEO) protein monomers in developing sieve elements and the mechanism(s) of Ca 2+ -dependent forisome contractility are poorly understood because the amino acid sequences of SEO proteins lack conventional protein–protein interaction and Ca 2+ -binding motifs. Here we selected amino acids potentially responsible for forisome-specific functions by analyzing SEO protein sequences in comparison to those of the widely distributed SEO-related (SEOR), or SEOR proteins. SEOR proteins resemble SEO proteins closely but lack any Ca 2+ responsiveness. We exchanged identified candidate residues by directed mutagenesis of the Medicago truncatula SEO1 gene, expressed the mutated genes in yeast ( Saccharomyces cerevisiae ) and studied the structural and functional phenotypes of the forisome-like bodies that formed in the transgenic cells. We identified three aspartate residues critical for Ca 2+ responsiveness and two more that were required for forisome-like bodies to assemble. The phenotypes observed further suggested that Ca 2+ -controlled and pH-inducible swelling effects in forisome-like bodies proceeded by different yet interacting mechanisms. Finally, we observed a previously unknown surface striation in native forisomes and in recombinant forisome-like bodies that could serve as an indicator of successful forisome assembly. To conclude, this study defines a promising path to the elucidation of the so-far elusive molecular mechanisms of forisome assembly and Ca 2+ -dependent contractility.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid evolution of cis-regulatory sequences via local point mutations

Although the evolution of protein-coding sequences within genomes is well understood, the same cannot be said of the cis-regulatory regions that control transcription. Yet, changes in gene expression are likely to constitute an important component of phenotypic evolution. We simulated the evolution of new transcription factor binding sites via local point mutations. The results indicate that new binding sites appear and become fixed within populations on microevolutionary timescales under an assumption of neutral evolution. Even combinations of two new binding sites evolve very quickly. We predict that local point mutations continually generate considerable genetic variation that is capable of altering gene expression.

Non-NASA Center↗

Evolution of early life inferred from protein and ribonucleic acid sequences

The chemical structures of ferredoxin, 5S ribosomal RNA, and c-type cytochrome sequences have been employed to construct a phylogenetic tree which connects all major photosynthesizing organisms: the three types of bacteria, blue-green algae, and chloroplasts. Anaerobic and aerobic bacteria, eukaryotic cytoplasmic components and mitochondria are also included in the phylogenetic tree. Anaerobic nonphotosynthesizing bacteria similar to Clostridium were the earliest organisms, arising more than 3.2 billion years ago. Bacterial photosynthesis evolved nearly 3.0 billion years ago, while oxygen-evolving photosynthesis, originating in the blue-green algal line, came into being about 2.0 billion years ago. The phylogenetic tree supports the symbiotic theory of the origin of eukaryotes.

Dayhoff, M. O.↗

Evolution of EF-hand calcium-modulated proteins. IV. Exon shuffling did not determine the domain compositions of EF-hand proteins

In the previous three reports in this series we demonstrated that the EF-hand family of proteins evolved by a complex pattern of gene duplication, transposition, and splicing. The dendrograms based on exon sequences are nearly identical to those based on protein sequences for troponin C, the essential light chain myosin, the regulatory light chain, and calpain. This validates both the computational methods and the dendrograms for these subfamilies. The proposal of congruence for calmodulin, troponin C, essential light chain, and regulatory light chain was confirmed. There are, however, significant differences in the calmodulin dendrograms computed from DNA and from protein sequences. In this study we find that introns are distributed throughout the EF-hand domain and the interdomain regions. Further, dendrograms based on intron type and distribution bear little resemblance to those based on protein or on DNA sequences. We conclude that introns are inserted, and probably deleted, with relatively high frequency. Further, in the EF-hand family exons do not correspond to structural domains and exon shuffling played little if any role in the evolution of this widely distributed homolog family. Calmodulin has had a turbulent evolution. Its dendrograms based on protein sequence, exon sequence, 3'-tail sequence, intron sequences, and intron positions all show significant differences.

NASA Discipline Exobiology↗

Identification of amino acid sequences in the polyomavirus capsid proteins that serve as nuclear localization signals

The molecular mechanism participating in the transport of newly synthesized proteins from the cytoplasm to the nucleus in mammalian cells is poorly understood. Recently, the nuclear localization signal sequences (NLS) of many nuclear proteins have been identified, and most have been found to be composed of a highly basic amino acid stretch. A genetic "subtractive" and a biochemical "additive" approach were used in our studies to identify the NLS's of the polyomavirus structural capsid proteins. An NLS was identified at the N-terminus (Ala1-Pro-Lys-Arg-Lys-Ser-Gly-Val-Ser-Lys-Cys11) of the major capsid protein VP1 and at the C-terminus (Glu307 -Glu-Asp-Gly-Pro-Glu-Lys-Lys-Lys-Arg-Arg-Leu318) of the VP2/VP3 minor capsid proteins.

NASA Discipline Number 93-10↗

Proteome-scale Deployment of Protein Structure Prediction Workflows on the Summit Supercomputer

Deep learning has contributed to major advances in the prediction of protein structure from sequence, a fundamental problem in structural bioinformatics. With predictions now approaching the accuracy of crystallographic experiments, and with accelerators like GPUs and TPUs making inference using large models rapid, genome-level structure prediction becomes an obvious aim. Leadership-class computing resources can be used to perform genome-scale protein structure prediction using state-of-the-art deep learning models, providing a wealth of new data for systems biology applications. Here we describe our efforts to efficiently deploy the AlphaFold v.2 program, for full-proteome structure prediction, at scale on the Oak Ridge Leadership Computing Facility's resources, including the Summit supercomputer. We performed inference to produce the predicted structures for 40,526 protein sequences, corresponding to four prokaryotic proteomes and one plant proteome, using under 4,400 total Summit node hours, equivalent to using the majority of the supercomputer for a little over one hour. We also designed an optimized structure refinement that reduced the time for the relaxation stage of the AlphaFold pipeline by over 10X for longer sequences. We demonstrate the types of analyses that can be performed on proteome-scale collections of sequences, including a search for novel quaternary structures and implications for functional annotation.

Gao, Mu↗

luoyunan/ECNet: First release

ECNet (evolutionary context-integrated neural network) is a deep learning model that guides protein engineering by predicting protein fitness from the sequence. It integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe.

Luo, Yunan↗

Desulfovulcanus ferrireducens gen. nov., sp. nov., A Thermophilic Autotrophic Iron and Sulfate-reducing Bacterium from Subseafloor Basalt That Grows on Akaganeite and Lepidocrocite

A deep-sea thermophilic bacterium, strain Ax17T, was isolated from 25 °C hydrothermal fluid at Axial Seamount. It was obligately anaerobic and autotrophic, oxidized molecular hydrogen and formate, and reduced synthetic nanophase Fe(III) (oxyhydr)oxide minerals, sulfate, sulfite, thiosulfate, and elemental sulfur for growth. It produced up to 20 mM Fe2+ when grown on ferrihydrite but < 5 mM Fe2+ when grown on akaganéite, lepidocrocite, hematite, and goethite. It was a straight to curved rod that grew at temperatures ranging from 35 to 70 °C (optimum 65 °C) and a minimum doubling time of 7.1 h, in the presence of 1.5–6% NaCl (optimum 3%) and pH 5–9 (optimum 8.0). Phylogenetic analysis based on 16S rRNA gene sequences indicated that the strain was 90–92% identical to other genera of the family Desulfonauticaceae in the phylum Pseudomonadota. The genome of Ax17T was sequenced, which yielded 2,585,834 bp and contained 2407 protein-coding sequences. Based on overall genome relatedness index analyses and its unique phenotypic characteristics, strain Ax17T is suggested to represent a novel genus and species, for which the name Desulfovulcanus ferrireducens is proposed. The type strain is Ax17T (= DSM 111878T = ATCC TSD-233T).

Srishti Kashyap↗

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Information contained in protein shapes

The sequence of local conformations at C-alpha atoms of a protein has been considered as an informational message string. The total self-information contents and self-information per letter have been evaluated for 83 globular proteins whose structures are known from X-ray crystallography. The derived information contents provide a method of quantitating structural specificity of proteins. This method of analysis enables repeating, intricate structural features to be recognized. Among the globular proteins whose structures have been solved, high potential iron protein stands out with the largest three-letter dependence.

Sundaram, K.↗