Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Population-based heteropolymer design to mimic protein mixtures

Biological fluids, the most complex blends, have compositions that constantly vary and cannot be molecularly defined. Despite these uncertainties, proteins fluctuate, fold, function and evolve as programmed. We propose that in addition to the known monomeric sequence requirements, protein sequences encode multi-pair interactions at the segmental level to navigate random encounters; synthetic heteropolymers capable of emulating such interactions can replicate how proteins behave in biological fluids individually and collectively. Here, we extracted the chemical characteristics and sequential arrangement along a protein chain at the segmental level from natural protein libraries and used the information to design heteropolymer ensembles as mixtures of disordered, partially folded and folded proteins. For each heteropolymer ensemble, the level of segmental similarity to that of natural proteins determines its ability to replicate many functions of biological fluids including assisting protein folding during translation, preserving the viability of fetal bovine serum without refrigeration, enhancing the thermal stability of proteins and behaving like synthetic cytosol under biologically relevant conditions. Molecular studies further translated protein sequence information at the segmental level into intermolecular interactions with a defined range, degree of diversity and temporal and spatial availability. This framework provides valuable guiding principles to synthetically realize protein properties, engineer bio/abiotic hybrid materials and, ultimately, realize matter-to-life transformations.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaFold Protein Structure Database for Sequence-Independent Molecular Replacement

Crystallographic phasing recovers the phase information that is lost during a diffraction experiment. Molecular replacement is a commonly used phasing method for crystal structures in the protein data bank. In one form it uses a protein sequence to search a structure database to find suitable templates for phasing. However, sequence information is not always available, such as when proteins are crystallized with unknown binding partner proteins or when the crystal is of a contaminant. The recent development of AlphaFold published the predicted protein structures for every protein from twenty distinct species. In this work, we tested whether AlphaFold-predicted E. coli protein structures were accurate enough to enable sequence-independent phasing of diffraction data from two crystallization contaminants of unknown sequence. Using each of more than 4000 predicted structures as a search model, robust molecular replacement solutions were obtained, which allowed the identification and structure determination of YncE and YadF. Our results demonstrate the general utility of the AlphaFold-predicted structure database with respect to sequence-independent crystallographic phasing.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

TwoFold: Highly accurate structure and affinity prediction for protein-ligand complexes from sequences

We describe our development of ab initio protein-ligand binding pose prediction models based on transformers and binding affinity prediction models based on the neural tangent kernel (NTK). Folding both protein and ligand, the TwoFold models achieve efficient and quality predictions matching state-of-the-art implementations while additionally reconstructing protein structures. In conclusion, solving NTK models points to a new use case for highly optimized linear solver benchmarking codes on HPC.

60 APPLIED LIFE SCIENCES↗

DNCON2_Inter: predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning

Deep learning methods that achieved great success in predicting intrachain residue-residue contacts have been applied to predict interchain contacts between proteins. However, these methods require multiple sequence alignments (MSAs) of a pair of interacting proteins (dimers) as input, which are often difficult to obtain because there are not many known protein complexes available to generate MSAs of sufficient depth for a pair of proteins. In recognizing that multiple sequence alignments of a monomer that forms homomultimers contain the co-evolutionary signals of both intrachain and interchain residue pairs in contact, we applied DNCON2 (a deep learning-based protein intrachain residue-residue contact predictor) to predict both intrachain and interchain contacts for homomultimers using multiple sequence alignment (MSA) and other co-evolutionary features of a single monomer followed by discrimination of interchain and intrachain contacts according to the tertiary structure of the monomer. We name this tool DNCON2_Inter. Allowing true-positive predictions within two residue shifts, the best average precision was obtained for the Top-L/10 predictions of 22.9% for homodimers and 17.0% for higher-order homomultimers. In some instances, especially where interchain contact densities are high, DNCON2_Inter predicted interchain contacts with 100% precision. We also developed Con_Complex, a complex structure reconstruction tool that uses predicted contacts to produce the structure of the complex. Using Con_Complex, we show that the predicted contacts can be used to accurately construct the structure of some complexes. Our experiment demonstrates that monomeric multiple sequence alignments can be used with deep learning to predict interchain contacts of homomeric proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Library Screening, In Vivo Confirmation, and Structural and Bioinformatic Analysis of Pentapeptide Sequences as Substrates for Protein Farnesyltransferase

Protein farnesylation is a post-translational modification where a 15-carbon farnesyl isoprenoid is appended to the C-terminal end of a protein by farnesyltransferase (FTase). This process often causes proteins to associate with the membrane and participate in signal transduction pathways. The most common substrates of FTase are proteins that have C-terminal tetrapeptide CaaX box sequences where the cysteine is the site of modification. However, recent work has shown that five amino acid sequences can also be recognized, including the pentapeptides CMIIM and CSLMQ. In this work, peptide libraries were initially used to systematically vary the residues in those two parental sequences using an assay based on Matrix Assisted Laser Desorption Ionization–Mass Spectrometry (MALDI-MS). In addition, 192 pentapeptide sequences from the human proteome were screened using that assay to discover additional extended CaaaX-box motifs. Selected hits from that screening effort were rescreened using an in vivo yeast reporter protein assay. The X-ray crystal structure of CMIIM bound to FTase was also solved, showing that the C-terminal tripeptide of that sequence interacted with the enzyme in a similar manner as the C-terminal tripeptide of CVVM, suggesting that the tripeptide comprises a common structural element for substrate recognition in both tetrapeptide and pentapeptide sequences. Molecular dynamics simulation of CMIIM bound to FTase further shed light on the molecular interactions involved, showing that a putative catalytically competent Zn(II)-thiolate species was able to form. Bioinformatic predictions of tetrapeptide (CaaX-box) reactivity correlated well with the reactivity of pentapeptides obtained from in vivo analysis, reinforcing the importance of the C-terminal tripeptide motif. This analysis provides a structural framework for understanding the reactivity of extended CaaaX-box motifs and a method that may be useful for predicting the reactivity of additional FTase substrates bearing CaaaX-box sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Generative design of de novo proteins based on secondary-structure constraints using an attention-based diffusion model

We report two generative deep-learning models that predict amino acid sequences and 3D protein structures on the basis of secondary-structure design objectives via either the overall content or the per-residue structure. Both models are robust regarding imperfect inputs and offer de novo design capacity because they can discover new protein sequences not yet discovered from natural mechanisms or systems. The residue-level secondary-structure design model generally yields higher accuracy and more diverse sequences. These findings suggest unexplored opportunities for protein designs and functional outcomes within the vast amino acid sequences beyond known proteins. Our models, based on an attention-based diffusion model and trained on a dataset extracted from experimentally known 3D protein structures, offer numerous downstream applications in the conditional generative design of various biological or engineering systems. Future work could include additional conditioning and an exploration of other functional properties of the generated proteins for various properties beyond structural objectives.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Comparative Virulence and Genomic Analysis of Streptococcus suis Isolates

Streptococcus suis is a zoonotic bacterial swine pathogen causing substantial economic and health burdens to the pork industry. Mechanisms used by S. suis to colonize and cause disease remain unknown and vaccines and/or intervention strategies currently do not exist. Studies addressing virulence mechanisms used by S. suis have been complicated because different isolates can cause a spectrum of disease outcomes ranging from lethal systemic disease to asymptomatic carriage. The objectives of this study were to evaluate the virulence capacity of nine United States S. suis isolates following intranasal challenge in swine and then perform comparative genomic analyses to identify genomic attributes associated with swine-virulent phenotypes. No correlation was found between the capacity to cause disease in swine and the functional characteristics of genome size, serotype, sequence type (ST), or in vitro virulence-associated phenotypes. A search for orthologs found in highly virulent isolates and not found in non-virulent isolates revealed numerous predicted protein coding sequences specific to each category. While none of these predicted protein coding sequences have been previously characterized as potential virulence factors, this analysis does provide a reliable one-to-one assignment of specific genes of interest that could prove useful in future allelic replacement and/or functional genomic studies. Collectively, this report provides a framework for future allelic replacement and/or functional genomic studies investigating genetic characteristics underlying the spectrum of disease outcomes caused by S. suis isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Convergent behavior of extended stalk regions from staphylococcal surface proteins with widely divergent sequence patterns

Staphylococcus epidermidis and Staphylococcus aureus are highly problematic bacteria in hospital settings. A major challenge is their ability to form biofilms on abiotic or biotic surfaces. Biofilms are well-organized, multicellular bacterial aggregates that resist antibiotic treatment and often lead to recurrent infections. Bacterial cell wall-anchored (CWA) proteins are important players in biofilm formation and infection. Many have putative stalk-like regions or regions of low complexity near the cell wall-anchoring motif. Recent work demonstrated the strong propensity of the stalk region of S. epidermidis accumulation-associated protein (Aap) to remain highly extended under solution conditions that typically induce compaction. This behavior is consistent with the expected function of a stalk-like region that is covalently attached to the cell wall peptidoglycan and projects the adhesive domains of Aap away from the cell surface. In this study, we evaluate whether the ability to resist compaction is a common theme among stalk regions from various staphylococcal CWA proteins. Circular dichroism spectroscopy was used to examine secondary structure changes as a function of temperature and cosolvents along with sedimentation velocity analytical ultracentrifugation, size-exclusion chromatography, and SAXS to characterize structural characteristics in solution. All stalk regions tested are intrinsically disordered, lacking secondary structure beyond random coil and polyproline type II helix, and they all sample highly extended conformations. Remarkably, the Ser-Asp dipeptide repeat region of SdrC exhibited nearly identical behavior in solution when compared to the Aap Pro/Gly-rich region, despite highly divergent sequence patterns, indicating conservation of function by various distinct staphylococcal CWA protein stalk regions.

59 BASIC BIOLOGICAL SCIENCES↗

Signal sequences target enzymes and structural proteins to bacterial microcompartments and are critical for microcompartment formation

ABSTRACT Spatial organization of pathway enzymes has emerged as a promising tool to address several challenges in metabolic engineering, such as flux imbalances and off-target product formation. Bacterial microcompartments (MCPs) are a spatial organization strategy used natively by many bacteria to encapsulate metabolic pathways that produce toxic, volatile intermediates. Several recent studies have focused on engineering MCPs to encapsulate heterologous pathways of interest, but how this engineering affects MCP assembly and function is poorly understood. In this study, we investigated the role of signal sequences, short domains that target proteins to the MCP core, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterized two novel Pdu signal sequences on the structural proteins PduM and PduB, which constitute the first report of metabolosome signal sequences on structural proteins rather than enzymes. We then explored the role of enzymatic and structural Pdu signal sequences on MCP assembly by deleting their encoding sequences from the genome alone and in combination. Deleting enzymatic signal sequences decreased the MCP formation, but this defect could be recovered in some cases by overexpressing genes encoding the knocked-out signal sequence fused to a heterologous protein. By contrast, deleting structural signal sequences caused similar defects to knocking out the genes encoding the full-length PduM and PduB proteins. Our results contribute to a growing understanding of how MCPs form and function in bacteria and provide strategies to mitigate assembly disruption when encapsulating heterologous pathways in MCPs. IMPORTANCE Spatially organizing biosynthetic pathway enzymes is a promising strategy to increase pathway throughput and yield. Bacterial microcompartments (MCPs) are proteinaceous organelles that many bacteria natively use as a spatial organization strategy to encapsulate niche metabolic pathways, providing significant metabolic benefits. Encapsulating heterologous pathways of interest in MCPs could confer these benefits to industrially relevant pathways. Here, we investigate the role of signal sequences, short domains that target proteins for encapsulation in MCPs, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterize two novel signal sequences on structural proteins, constituting the first Pdu signal sequences found on structural proteins rather than enzymes, and perform knockout studies to compare the impacts of enzymatic and structural signal sequences on MCP assembly. Our results demonstrate that enzymatic and structural signal sequences play critical but distinct roles in Pdu MCP assembly and provide design rules for engineering MCPs while minimizing disruption to MCP assembly.

Johnson, Elizabeth R. (ORCID:0000000179236881)↗

Epitopes recognition of SARS-CoV-2 nucleocapsid RNA binding domain by human monoclonal antibodies

Coronavirus nucleocapsid protein (NP) of SARS-CoV-2 plays a central role in many functions important for virus proliferation including packaging and protecting genomic RNA. The protein shares sequence, structure, and architecture with nucleocapsid proteins from betacoronaviruses. The N-terminal domain (NP RBD ) binds RNA and the C-terminal domain is responsible for dimerization. After infection, NP is highly expressed and triggers robust host immune response. The anti-NP antibodies are not protective and not neutralizing but can effectively detect viral proliferation soon after infection. Two structures of SARS-CoV-2 NP RBD were determined providing a continuous model from residue 48 to 173, including RNA binding region and key epitopes. Five structures of NP RBD complexes with human mAbs were isolated using an antigen-bait sorting. Complexes revealed a distinct complement-determining regions and unique sets of epitope recognition. This may assist in the early detection of pathogens and designing peptide-based vaccines. Mutations that significantly increase viral load were mapped on developed, full length NP model, likely impacting interactions with host proteins and viral RNA.

59 BASIC BIOLOGICAL SCIENCES↗

Putting Humpty Dumpty Back Together Again: What Does Protein Quantification Mean in Bottom-Up Proteomics?

Bottom-up proteomics provides peptide measurements and has been invaluable for moving proteomics into large-scale analyses. Commonly, a single quantitative value is reported for each protein-coding gene by aggregating peptide quantities into protein groups following protein inference or parsimony. However, given the complexity of both RNA splicing and post-translational protein modification, it is overly simplistic to assume that all peptides that map to a singular protein-coding gene will demonstrate the same quantitative response. Here, by assuming that all peptides from a protein-coding sequence are representative of the same protein, we may miss the discovery of important biological differences. To capture the contributions of existing proteoforms, we need to reconsider the practice of aggregating protein values to a single quantity per protein-coding gene.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Aspartate Residues in a Forisome-Forming SEO Protein Are Critical for Protein Body Assembly and Ca 2+ Responsiveness

Forisomes are protein bodies known exclusively from sieve elements of legumes. Forisomes contribute to the regulation of phloem transport due to their unique Ca 2+ -controlled, reversible swelling. The assembly of forisomes from sieve element occlusion (SEO) protein monomers in developing sieve elements and the mechanism(s) of Ca 2+ -dependent forisome contractility are poorly understood because the amino acid sequences of SEO proteins lack conventional protein–protein interaction and Ca 2+ -binding motifs. Here we selected amino acids potentially responsible for forisome-specific functions by analyzing SEO protein sequences in comparison to those of the widely distributed SEO-related (SEOR), or SEOR proteins. SEOR proteins resemble SEO proteins closely but lack any Ca 2+ responsiveness. We exchanged identified candidate residues by directed mutagenesis of the Medicago truncatula SEO1 gene, expressed the mutated genes in yeast ( Saccharomyces cerevisiae ) and studied the structural and functional phenotypes of the forisome-like bodies that formed in the transgenic cells. We identified three aspartate residues critical for Ca 2+ responsiveness and two more that were required for forisome-like bodies to assemble. The phenotypes observed further suggested that Ca 2+ -controlled and pH-inducible swelling effects in forisome-like bodies proceeded by different yet interacting mechanisms. Finally, we observed a previously unknown surface striation in native forisomes and in recombinant forisome-like bodies that could serve as an indicator of successful forisome assembly. To conclude, this study defines a promising path to the elucidation of the so-far elusive molecular mechanisms of forisome assembly and Ca 2+ -dependent contractility.

59 BASIC BIOLOGICAL SCIENCES↗