Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Validation and functional characterization of transcription factors in wheat using cell-free protein expression and high-throughput sequencing technologies

Transcription factors (TFs) are critical biomolecules that control and regulate gene expression in every organism. In wheat, there are 3,606 genes annotated as putative TFs which are primarily uncharacterized. In this work, using a DAP-seq approach, we have validated and characterized a select number of TFs (5bl and 5dl) belonging to GeBP family of proteins, responsible for plant cellular growth, development, and differentiation. We found that top DNA binding motifs for 5bl and 5dl are (G/T)N(T/G)GTGGT and (C/G)AA(C/G)AA respectively. These motifs/peaks are enriched in the promoter regions of the genes, which are associated with rRNA biogenesis and protein maturation pathways. This study is also a first demonstration that DAP-seq can be applied for such a complex and large genome such as wheat.

59 BASIC BIOLOGICAL SCIENCES↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Archaebacterial rhodopsin sequences: Implications for evolution

It was proposed over 10 years ago that the archaebacteria represent a separate kingdom which diverged very early from the eubacteria and eukaryotes. It follows that investigations of archaebacterial characteristics might reveal features of early evolution. So far, two genes, one for bacteriorhodopsin and another for halorhodopsin, both from Halobacterium halobium, have been sequenced. We cloned and sequenced the gene coding for the polypeptide of another one of these rhodopsins, a halorhodopsin in Natronobacterium pharaonis. Peptide sequencing of cyanogen bromide fragments, and immuno-reactions of the protein and synthetic peptides derived from the C-terminal gene sequence, confirmed that the open reading frame was the structural gene for the pharaonis halorhodopsin polypeptide. The flanking DNA sequences of this gene, as well as those of other bacterial rhodopsins, were compared to previously proposed archaebacterial consensus sequences. In pairwise comparisons of the open reading frame with DNA sequences for bacterio-opsin and halo-opsin from Halobacterium halobium, silent divergences were calculated. These indicate very considerable evolutionary distance between each pair of genes, even in the dame organism. In spite of this, three protein sequences show extensive similarities, indicating strong selective pressures.

Lanyi, J. K.↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation

Preliminary: This final report corresponds to a grant (DE-SC0021216) that was awarded to the University of Montana. Mid-way through the grant period, I relocated from the University of Montana to the University of Arizona. The grant was ended at University of Montana in late 2022, with all efforts concluding on 08/26/22; the remaining funds supporting the project were relinquished by University of Montana, and were later awarded to University of Arizona under a new grant, with start date 04/01/23. This report focuses on results of research efforts at UMontana through 08/26/22. Results: We made progress in each of the three aims of the proposal. We released software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes. We made substantial progress in developing software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we made notable progress in developing AI methods (specifically: a neural embedding model) for identifying similarities between protein sequences based on amino-wise latent vectors. These efforts were supplemented by development of methods for protein modeling in support of predicting protein-drug binding activity, and by my leadership of a team in the NIH/DOE 2021 Petabyte-Scale Sequence Search hack-a-thon.

59 BASIC BIOLOGICAL SCIENCES↗

Cross-reactive immunogenicity of group A streptococcal vaccines designed using a recurrent neural network to identify conserved M protein linear epitopes

The M protein of group A streptococci (Strep A) is a major virulence determinant and protective antigen. The N-terminal sequence of the protein defines the more than 200 M types of Strep A and also contains epitopes that elicit opsonic antibodies, some of which cross-react with heterologous M types. Current efforts to develop broadly protective M protein-based vaccines are directed at identifying potential cross-protective epitopes located in the N-terminal regions of cluster-related M proteins for use as vaccine antigens. In this study, we have used a comprehensive approach using the recurrent neural network ABCpred and IEDB epitope conservancy analysis tools to predict 16 residue linear B-cell epitopes from 117 clinically relevant M types of Strep A (~88% of global Strep A infections). Furthermore, to examine the immunogenicity of these epitope-based vaccines, nine peptides that together shared ≥60% sequence identity with 37 heterologous M proteins were incorporated into two recombinant hybrid protein vaccines, in which the epitopes were repeated 2 or 3 times, respectively. The combined immune responses of immunized rabbits showed that the vaccines elicited significant levels of antibodies against all nine vaccine epitopes present in homologous N-terminal 1–50 amino acid synthetic M peptides, as well as cross-reactive antibodies against 16 of 37 heterologous M peptides predicted to contain similar epitopes. The epitope-specificity of the cross-reactive antibodies was confirmed by ELISA inhibition assays and functional opsonic activity was assayed in HL-60-based bactericidal assays. The results provide important information for the future design of broadly protective M protein-based Strep A vaccines.

60 APPLIED LIFE SCIENCES↗

Frameshifting Stimulatory Sequence Induces Large Structural Change of Ribosomal Proteins When Bound to E. coli Ribosomes

Biological macromolecular machines occupy a continuum of structural conformations to perform cellular tasks. Mapping this conformational space provides an insight into its functionality. While the cryo-electron microscopy resolution revolution has expanded our ability to characterize the conformational continuums, there are obstacles in structurally characterizing regions of high flexibility. These technical barriers have impeded characterization of flexible ribosomal proteins when the ribosome is interacting with mRNA stem-loop structures such as a frameshifting stimulatory sequence (FSS). Small-angle neutron/X-ray scattering and electron microscopy were used to study ribosomal samples and compared structural differences between a ribosome that is bound to an FSS stem-loop compared to a ribosome bound to linear mRNA. This comparison shows that a large protein stalk elongates by 22% when the 70S interacts with an mRNA stem-loop. Finally, our results suggest that ribosomal proteins have extensive flexibility and may influence important ribosomal mechanisms, such as those that involve FSS.

36 MATERIALS SCIENCE↗

Origins of the protein synthesis cycle

Largely derived from experiments in molecular evolution, a theory of protein synthesis cycles has been constructed. The sequence begins with ordered thermal proteins resulting from the self-sequencing of mixed amino acids. Ordered thermal proteins then aggregate to cell-like structures. When they contained proteinoids sufficiently rich in lysine, the structures were able to synthesize offspring peptides. Since lysine-rich proteinoid (LRP) also catalyzes the polymerization of nucleoside triphosphate to polynucleotides, the same microspheres containing LRP could have synthesized both original cellular proteins and cellular nucleic acids. The LRP within protocells would have provided proximity advantageous for the origin and evolution of the genetic code.

Fox, S. W.↗

A pollen-specific novel calmodulin-binding protein with tetratricopeptide repeats

Calcium is essential for pollen germination and pollen tube growth. A large body of information has established a link between elevation of cytosolic Ca(2+) at the pollen tube tip and its growth. Since the action of Ca(2+) is primarily mediated by Ca(2+)-binding proteins such as calmodulin (CaM), identification of CaM-binding proteins in pollen should provide insights into the mechanisms by which Ca(2+) regulates pollen germination and tube growth. In this study, a CaM-binding protein from maize pollen (maize pollen calmodulin-binding protein, MPCBP) was isolated in a protein-protein interaction-based screening using (35)S-labeled CaM as a probe. MPCBP has a molecular mass of about 72 kDa and contains three tetratricopeptide repeats (TPR) suggesting that it is a member of the TPR family of proteins. MPCBP protein shares a high sequence identity with two hypothetical TPR-containing proteins from Arabidopsis. Using gel overlay assays and CaM-Sepharose binding, we show that the bacterially expressed MPCBP binds to bovine CaM and three CaM isoforms from Arabidopsis in a Ca(2+)-dependent manner. To map the CaM-binding domain several truncated versions of the MPCBP were expressed in bacteria and tested for their ability to bind CaM. Based on these studies, the CaM-binding domain was mapped to an 18-amino acid stretch between the first and second TPR regions. Gel and fluorescence shift assays performed with CaM and a CaM-binding synthetic peptide further confirmed MPCBP binding to CaM. Western, Northern, and reverse transcriptase-polymerase chain reaction analysis have shown that MPCBP expression is specific to pollen. MPCBP was detected in both soluble and microsomal proteins. Immunoblots showed the presence of MPCBP in mature and germinating pollen. Pollen-specific expression of MPCBP, its CaM-binding properties, and the presence of TPR motifs suggest a role for this protein in Ca(2+)-regulated events during pollen germination and growth.

NASA Discipline Plant Biology↗

Gate-based quantum computing for protein design

Protein design is a technique to engineer proteins by permuting amino acids in the sequence to obtain novel functionalities. However, exploring all possible combinations of amino acids is generally impossible due to the exponential growth of possibilities with the number of designable sites. The present work introduces circuits implementing a pure quantum approach, Grover’s algorithm, to solve protein design problems. Our algorithms can adjust to implement any custom pair-wise energy tables and protein structure models. Moreover, the algorithm’s oracle is designed to consist of only adder functions. Quantum computer simulators validate the practicality of our circuits, containing up to 234 qubits. However, a smaller circuit is implemented on real quantum devices. Our results show that using iterations, the circuits find the correct results among all N possibilities, providing the expected quadratic speed up of Grover’s algorithm over classical methods (i.e.,).

59 BASIC BIOLOGICAL SCIENCES↗

A Re-Evaluation of African Swine Fever Genotypes Based on p72 Sequences Reveals the Existence of Only Six Distinct p72 Groups

The African swine fever virus (ASFV) is currently causing a world-wide pandemic of a highly lethal disease in domestic swine and wild boar. Currently, recombinant ASF live-attenuated vaccines based on a genotype II virus strain are commercially available in Vietnam. With 25 reported ASFV genotypes in the literature, it is important to understand the molecular basis and usefulness of ASFV genotyping, as well as the true significance of genotypes in the epidemiology, transmission, evolution, control, and prevention of ASFV. Historically, genotyping of ASFV was used for the epidemiological tracking of the disease and was based on the analysis of small fragments that represent less than 1% of the viral genome. The predominant method for genotyping ASFV relies on the sequencing of a fragment within the gene encoding the structural p72 protein. Genotype assignment has been accomplished through automated phylogenetic trees or by comparing the target sequence to the most closely related genotyped p72 gene. To evaluate its appropriateness for the classification of genotypes by p72, we reanalyzed all available genomic data for ASFV. We conclude that the majority of p72-based genotypes, when initially created, were neither identified under any specific methodological criteria nor correctly compared with the already existing ASFV genotypes. Based on our analysis of the p72 protein sequences, we propose that the current twenty-five genotypes, created exclusively based on the p72 sequence, should be reduced to only six genotypes. To help differentiate between the new and old genotype classification systems, we propose that Arabic numerals (1, 2, 8, 9, 15, and 23) be used instead of the previously used Roman numerals. Furthermore, we discuss the usefulness of genotyping ASFV isolates based only on the p72 gene sequence.

59 BASIC BIOLOGICAL SCIENCES↗

Self-Assembly of Repetitive Segment and Random Segment Polymer Architectures

Recent advances in chemical synthesis have created new methodologies for synthesizing sequence-controlled synthetic polymers, but rational design of monomer sequence for desired properties remains challenging. In this work, we synthesize periodic polymers with repetitive segments using a sequence-controlled ring-opening metathesis polymerization (ROMP) method, which draws inspiration from proteins containing repetitive sequence motifs. The repetitive segment architecture is shown to dramatically affect the self-assembly behavior of these materials. In this work, our results show that polymers with identical repetitive sequences assemble into uniform spherical nanoparticles after thermal annealing, whereas copolymers with random placement of segments with different sequences exhibit disordered assemblies without a well-defined morphology. Overall, these results bring a new understanding to the role of periodic repetitive sequences in polymer assembly.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Blueprinting extendable nanomaterials with standardized protein blocks

A wooden house frame consists of many different lumber pieces, but because of the regularity of these building blocks, the structure can be designed using straightforward geometrical principles. The design of multicomponent protein assemblies, in comparison, has been much more complex, largely owing to the irregular shapes of protein structures. Here we describe extendable linear, curved and angled protein building blocks, as well as inter-block interactions, that conform to specified geometric standards; assemblies designed using these blocks inherit their extendability and regular interaction surfaces, enabling them to be expanded or contracted by varying the number of modules, and reinforced with secondary struts. Using X-ray crystallography and electron microscopy, we validate nanomaterial designs ranging from simple polygonal and circular oligomers that can be concentrically nested, up to large polyhedral nanocages and unbounded straight ‘train track’ assemblies with reconfigurable sizes and geometries that can be readily blueprinted. Because of the complexity of protein structures and sequence–structure relationships, it has not previously been possible to build up large protein assemblies by deliberate placement of protein backbones onto a blank three-dimensional canvas; the simplicity and geometric regularity of our design platform now enables construction of protein nanomaterials according to ‘back of an envelope’ architectural blueprints.

36 MATERIALS SCIENCE↗

A small-angle neutron scattering study of the physical mechanism that drives the action of a viral fusion peptide

Viruses have evolved a variety of ways for delivering their genetic cargo to a target cell. One mechanism relies on a short sequence from a protein of the virus that is referred to as a fusion peptide. In some cases, the isolated fusion peptide is also capable of causing membranes to fuse. Infection by HIV-1 involves the 23 amino acid N-terminal sequence of its gp41 envelope protein, which is capable of causing membranes to fuse by itself, but the mechanism by which it does so is not fully understood. In this study, a variant of the gp41 fusion peptide that does not strongly promote fusion was studied in the presence of vesicles composed of a mixture of unsaturated lipids and cholesterol by small-angle neutron scattering and circular dichroism spectroscopy to improve the understanding of the mechanism that drives vesicle fusion. The peptide concentration and cholesterol content govern both the peptide conformation and its impact on the bilayer structure. The results indicate that the mechanism that drives vesicle fusion by the peptide is a strong distortion of the bilayer structure by the peptide when it adopts the β-sheet conformation.

60 APPLIED LIFE SCIENCES↗

Design of Broadly Cross-Reactive M Protein–Based Group A Streptococcal Vaccines

Group A streptococcal infections are a significant cause of global morbidity and mortality. A leading vaccine candidate is the surface M protein, a major virulence determinant and protective Ag. An obstacle to the development of M protein–based vaccines is the >200 different M types defined by the N-terminal sequences that contain protective epitopes. Despite sequence variability, M proteins share coiled-coil structural motifs that bind host proteins required for virulence. In this study, we exploit this potential Achilles heel of conserved structure to predict cross-reactive M peptides that could serve as broadly protective vaccine Ags. Combining sequences with structural predictions, six heterologous M peptides in a sequence-related cluster were predicted to elicit cross-reactive Abs with the remaining five nonvaccine M types in the cluster. The six-valent vaccine elicited Abs in rabbits that reacted with all 11 M peptides in the cluster and functional opsonic Abs against vaccine and nonvaccine M types in the cluster. We next immunized mice with four sequence-unrelated M peptides predicted to contain different coiled-coil propensities and tested the antisera for cross-reactivity against 41 heterologous M peptides. Based on these results, we developed an improved algorithm to select cross-reactive peptide pairs using additional parameters of coiled-coil length and propensity. The revised algorithm accurately predicted cross-reactive Ab binding, improving the Matthews correlation coefficient from 0.42 to 0.74. These results form the basis for selecting the minimum number of N-terminal M peptides to include in potentially broadly efficacious multivalent vaccines that could impact the overall global burden of group A streptococcal diseases.

60 APPLIED LIFE SCIENCES↗

Quantifying Structural Relationships of Metal-Binding Sites Suggests Origins of Biological Electron Transfer

Biological redox reactions drive planetary biogeochemical cycles. Using a novel, structure-guided sequence analysis of proteins, we explored the patterns of evolution of enzymes responsible for these reactions. Our analysis reveals that the folds that bind transition metal–containing ligands have similar structural geometry and amino acid sequences across the full diversity of proteins. Similarity across folds reflects the availability of key transition metals over geological time and strongly suggests that transition metal–ligand binding had a small number of common peptide origins. We observe that structures central to our similarity network come primarily from oxidoreductases, suggesting that ancestral peptides may have also facilitated electron transfer reactions. Last, our results reveal that the earliest biologically functional peptides were likely available before the assembly of fully functional protein domains over 3.8 billion years ago. Thus, life is a special, very complex form of motion of matter, but this form did not always exist, and it is not separated from inorganic nature by an impassable abyss; rather, it arose from inorganic nature as a new property in the process of evolution of the world. We must study the history of this evolution if we want to solve the problem of the origin of life.

Yana Bromberg↗

Structural and biochemical analyses of selectivity determinants in chimeric Streptococcus Class A sortase enzymes

Abstract Sequence variation in related proteins is an important characteristic that modulates activity and selectivity. An example of a protein family with a large degree of sequence variation is that of bacterial sortases, which are cysteine transpeptidases on the surface of gram‐positive bacteria. Class A sortases are responsible for attachment of diverse proteins to the cell wall to facilitate environmental adaption and interaction. These enzymes are also used in protein engineering applications for sortase‐mediated ligations (SML) or sortagging of protein targets. We previously investigated SrtA from Streptococcus pneumoniae , identifying a number of putative β7–β8 loop‐mediated interactions that affected in vitro enzyme function. We identified residues that contributed to the ability of S. pneumoniae SrtA to recognize several amino acids at the P1′ position of the substrate motif, underlined in LPXT G , in contrast to the strict P1′ Gly recognition of SrtA from Staphylococcus aureus . However, motivated by the lack of a structural model for the active, monomeric form of S. pneumoniae SrtA, here, we expanded our studies to other Streptococcus SrtA proteins. We solved the first monomeric structure of S. agalactiae SrtA which includes the C‐terminus, and three others of β7–β8 loop chimeras from S. pyogenes and S. agalactiae SrtA. These structures and accompanying biochemical data support our previously identified β7–β8 loop‐mediated interactions and provide additional insight into their role in Class A sortase substrate selectivity. A greater understanding of individual SrtA sequence and structural determinants of target selectivity may also facilitate the design or discovery of improved sortagging tools.

Gao, Melody↗