Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Evolution of EF-hand calcium-modulated proteins. III. Exon sequences confirm most dendrograms based on protein sequences: calmodulin dendrograms show significant lack of parallelism

In the first report in this series we presented dendrograms based on 152 individual proteins of the EF-hand family. In the second we used sequences from 228 proteins, containing 835 domains, and showed that eight of the 29 subfamilies are congruent and that the EF-hand domains of the remaining 21 subfamilies have diverse evolutionary histories. In this study we have computed dendrograms within and among the EF-hand subfamilies using the encoding DNA sequences. In most instances the dendrograms based on protein and on DNA sequences are very similar. Significant differences between protein and DNA trees for calmodulin remain unexplained. In our fourth report we evaluate the sequences and the distribution of introns within the EF-hand family and conclude that exon shuffling did not play a significant role in its evolution.

Non-NASA Center↗

Conservation of Shannon's redundancy for proteins

Concepts of information theory are applied to examine various proteins in terms of their redundancy in natural originators such as animals and plants. The Monte Carlo method is used to derive information parameters for random protein sequences. Real protein sequence parameters are compared with the standard parameters of protein sequences having a specific length. The tendency of a chain to contain some amino acids more frequently than others and the tendency of a chain to contain certain amino acid pairs more frequently than other pairs are used as randomness measures of individual protein sequences. Non-periodic proteins are generally found to have random Shannon redundancies except in cases of constraints due to short chain length and genetic codes. Redundant characteristics of highly periodic proteins are discussed. A degree of periodicity parameter is derived.

Gatlin, L. L.↗

Size and Structure of the Sequence Space of Repeat Proteins

The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family— the total number of sequences in that family—can be estimated using models of maximum entropy trained on multiple sequence alignments of naturally occurring amino acid sequences. We analyzed and calculated the size of three abundant repeat proteins families, whose members are large proteins made of many repetitions of conserved portions of *30 amino acids. While amino acid conservation at each position of the alignment explains most of the reduction of diversity relative to completely random sequences, we found that correlations between amino acid usage at different positions significantly impact that diversity. We quantified the impact of different types of correlations, functional and evolutionary, on sequence diversity. Analysis of the detailed structure of the coding space of the families revealed a rugged landscape, with many local energy minima of varying sizes with a hierarchical structure, reminiscent of frustrated energy landscapes of spin glass in physics. This clustered structure indicates a multiplicity of subtypes within each family and suggests new strategies for protein design.

Jacopo Marchi↗

ForceGen: End-to-end de novo protein generation based on nonlinear mechanical unfolding responses using a language diffusion model

Through evolution, nature has presented a set of remarkable protein materials, including elastins, silks, keratins and collagens with superior mechanical performances that play crucial roles in mechanobiology. However, going beyond natural designs to discover proteins that meet specified mechanical properties remains challenging. Here, we report a generative model that predicts protein designs to meet complex nonlinear mechanical property-design objectives. Our model leverages deep knowledge on protein sequences from a pretrained protein language model and maps mechanical unfolding responses to create proteins. Via full-atom molecular simulations for direct validation, we demonstrate that the designed proteins are de novo, and fulfill the targeted mechanical properties, including unfolding energy and mechanical strength, as well as the detailed unfolding force-separation curves. Our model offers rapid pathways to explore the enormous mechanobiological protein sequence space unconstrained by biological synthesis, using mechanical features as the target to enable the discovery of protein materials with superior mechanical properties.

59 BASIC BIOLOGICAL SCIENCES↗

Opportunities and Challenges for Machine Learning-Assisted Enzyme Engineering

Enzymes can be engineered at the level of their amino acid sequences to optimize key properties such as expression, stability, substrate range, and catalytic efficiency or even to unlock new catalytic activities not found in nature. Because the search space of possible proteins is vast, enzyme engineering usually involves discovering an enzyme starting point that has some level of the desired activity followed by directed evolution to improve its “fitness” for a desired application. Recently, machine learning (ML) has emerged as a powerful tool to complement this empirical process. ML models can contribute to (1) starting point discovery by functional annotation of known protein sequences or generating novel protein sequences with desired functions and (2) navigating protein fitness landscapes for fitness optimization by learning mappings between protein sequences and their associated fitness values. In this Outlook, we explain how ML complements enzyme engineering and discuss its future potential to unlock improved engineering outcomes.

60 APPLIED LIFE SCIENCES↗

Optimizing metaproteomics database construction: lessons from a study of the vaginal microbiome

Metaproteomics, a method for untargeted, high-throughput identification of proteins in complex samples, provides functional information about microbial communities and can tie functions to specific taxa. Metaproteomics often generates less data than other omics techniques, but analytical workflows can be improved to increase usable data in metaproteomic outputs. Identification of peptides in the metaproteomic analysis is performed by comparing mass spectra of sample peptides to a reference database of protein sequences. Although these protein databases are an integral part of the metaproteomic analysis, few studies have explored how database composition impacts peptide identification. Here, we used cervicovaginal lavage (CVL) samples from a study of bacterial vaginosis (BV) to compare the performance of databases built using six different strategies. We evaluated broad versus sample-matched databases, as well as databases populated with proteins translated from metagenomic sequencing of the same samples versus sequences from public repositories. Smaller sample-matched databases performed significantly better, driven by the statistical constraints on large databases. Additionally, large databases attributed up to 34% of significant bacterial hits to taxa absent from the sample, as determined orthogonally by 16S rRNA gene sequencing. We also tested a set of hybrid databases which included bacterial proteins from NCBI RefSeq and translated bacterial genes from the samples. These hybrid databases had the best overall performance, identifying 1,068 unique human and 1,418 unique bacterial proteins, ~30% more than a database populated with proteins from typical vaginal bacteria and fungi. Our findings can help guide the optimal identification of proteins while maintaining statistical power for reaching biological conclusions.

59 BASIC BIOLOGICAL SCIENCES↗

Reclassification of ASFV into 7 Biotypes Using Unsupervised Machine Learning

In 2007, an outbreak of African swine fever (ASF), a deadly disease of domestic swine and wild boar caused by the African swine fever virus (ASFV), occurred in Georgia and has since spread globally. Historically, ASFV was classified into 25 different genotypes. However, a newly proposed system recategorized all ASFV isolates into 6 genotypes exclusively using the predicted protein sequences of p72. However, ASFV has a large genome that encodes between 150–200 genes, and classifications using a single gene are insufficient and misleading, as strains encoding an identical p72 often have significant mutations in other areas of the genome. We present here a new classification of ASFV based on comparisons performed considering the entire encoded proteome. A curated database consisting of the protein sequences predicted to be encoded by 220 reannotated ASFV genomes was analyzed for similarity between homologous protein sequences. Weights were applied to the protein identity matrices and averaged to generate a genome-genome identity matrix that was then analyzed by an unsupervised machine learning algorithm, DBSCAN, to separate the genomes into distinct clusters. We conclude that all available ASFV genomes can be classified into 7 distinct biotypes.

59 BASIC BIOLOGICAL SCIENCES↗

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Dissecting neurofilament tail sequence-phosphorylation-structure relationships with multicomponent reconstituted protein brushes

Neurofilaments (NFs) are multisubunit, bottlebrush-shaped intermediate filaments abundant in the axonal cytoskeleton. Each NF subunit contains a long intrinsically disordered tail domain, which protrudes from the NF core to form a “brush” surrounding each NF. Precisely how the tails’ variable charge patterns and repetitive phosphorylation sites mediate their conformation within the brush remains an open question in axonal biology. We address this problem by grafting recombinant NF tail protein constructs NF-Light, -Medium, and -Heavy (NFL, NFM, and NFH) to surfaces, yielding protein brushes of defined stoichiometry that can be phosphorylated in vitro. Atomic force microscopy measurements reveal that brush height depends on composition monotonically but not always linearly for binary NFL:NFM or NFL:NFH systems, and that NFM-based brushes are highly extended, while brushes incorporating the much larger NFH are surprisingly compact even after multisite phosphorylation. Complementary self-consistent field theory (SCFT) predicts multilayer brush morphologies for NFM and phosphorylated NFH brushes. Further experiments and SCFT analysis with designed mutants reveal that N-terminal negative charges in the NFH tail repel phosphorylated residues to generate the multilayer morphology, while the C-terminal charge-neutral region contributes to multilayer brush morphology but not total brush height. Charge-shuffled NFM variants show that charge segregation promotes brush collapse near physiological ionic strengths. Collectively, this study supports a role for NFM in establishing a dynamic range for NF brush conformation, lending insight into previous in vitro and in vivo findings. More broadly, this work establishes a platform for dissecting contributions of disordered protein sequence to conformation at interfaces.

Science & Technology - Other Topics↗

Fast and versatile sequence-independent protein docking for nanomaterials design using RPXDock

Computationally designed multi-subunit assemblies have shown considerable promise for a variety of applications, including a new generation of potent vaccines. One of the major routes to such materials is rigid body sequence-independent docking of cyclic oligomers into architectures with point group or lattice symmetries. Current methods for docking and designing such assemblies are tailored to specific classes of symmetry and are difficult to modify for novel applications. Here we describe RPXDock, a fast, flexible, and modular software package for sequence-independent rigid-body protein docking across a wide range of symmetric architectures that is easily customizable for further development. RPXDock uses an efficient hierarchical search and a residue-pair transform ( RPX ) scoring method to rapidly search through multidimensional docking space. We describe the structure of the software, provide practical guidelines for its use, and describe the available functionalities including a variety of score functions and filtering tools that can be used to guide and refine docking results towards desired configurations.

59 BASIC BIOLOGICAL SCIENCES↗

Controllable Phycobilin Modification: An Alternative Photoacclimation Response in Cryptophyte Algae

Cryptophyte algae are well-known for their ability to survive under low light conditions using their auxiliary light harvesting antennas, phycobiliproteins. Mainly acting to absorb light where chlorophyll cannot (500-650 nm), phycobiliproteins also play an instrumental role in helping cryptophyte algae respond to changes in light intensity through the process of photoacclimation. Until recently, photoacclimation in cryptophyte algae was only observed as a change in the cellular concentration of phycobiliproteins; however, an additional photoacclimation response was recently discovered that causes shifts in the phycobiliprotein absorbance peaks following growth under red, blue, or green light. Here, we reproduce this newly identified photoacclimation response in two species of cryptophyte algae and elucidate the origin of the response on the protein level. We compare isolated native and photoacclimated phycobiliproteins for these two species using spectroscopy and mass spectrometry, and we report the X-ray structures of each phycobiliprotein and the corresponding photoacclimated complex. We find that neither the protein sequences nor the protein structures are modified by photoacclimation. We conclude that cryptophyte algae change one chromophore in the phycobiliprotein β subunits in response to changes in the spectral quality of light. Ultrafast pump-probe spectroscopy shows that the energy transfer is weakly affected by photoacclimation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto↗

Build Optimization Software Tools (BOOST) v2.0.0

Build Optimization Software Tools (BOOST) accelerate the design of DNA, RNA and protein sequences for their synthesis and assembly in an automated and scalable fashion. The BOOST Juggler automates the following two design tasks: - Reverse-Translation of protein sequences into DNA sequences - Codon Juggling of DNA sequences The BOOST Polisher provides the following design tasks: - The Verification of DNA sequences against DNA synthesis constraints - The Modification of DNA sequences in case they violate any DNA synthesis constraint In addition, the BOOST Polisher can be instructed to verify and modify DNA sequences against sequence patterns, such as restriction sites. The BOOST Partitioner supports both Gibson/chewback and Yeast assembly (also known as Transformation-Associated Recombination, or TAR) methods when performing the decomposition of large sequences that exceed the maximum length of DNA synthesis into synthesizable building blocks. The Partitioner can be instructed to find building blocks with appropriate overlap sequences, making their assembly into the complete construct more efficient. BOOST Workflows enable, in a customized fashion, the execution of the Juggler, Polisher, and Partitioner tools on a batch of sequences.

Sharkey, Michael↗

Sequence-defined structural transitions by calcium-responsive proteins

Biopolymer sequences dictate their functions, and protein-based polymers are a promising platform to establish sequence–function relationships for novel biopolymers. To efficiently explore vast sequence spaces of natural proteins, sequence repetition is a common strategy to tune and amplify specific functions. This strategy is applied to repeats-in-toxin (RTX) proteins with calcium-responsive folding behavior, which stems from tandem repeats of the nonapeptide GGXGXDXUX in which X can be any amino acid and U is a hydrophobic amino acid. To determine the functional range of this nonapeptide, we modified a naturally occurring RTX protein that forms β-roll structures in the presence of calcium. Sequence modifications focused on calcium-binding turns within the repetitive region, including either global substitution of nonconserved residues or complete replacement with tandem repeats of a consensus nonapeptide GGAGXDTLY. Some sequence modifications disrupted the typical transition from intrinsically disordered random coils to folded β rolls, despite conservation of the underlying nonapeptide sequence. Proteins enriched with smaller, hydrophobic amino acids adopted secondary structures in the absence of calcium and underwent structural rearrangements in calcium-rich environments. In contrast, proteins with bulkier, hydrophilic amino acids maintained intrinsic disorder in the absence of calcium. In conclusion, these results indicate a significant role of nonconserved amino acids in calcium-responsive folding, thereby revealing a strategy to leverage sequences in the design of tunable, calcium-responsive biopolymers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Non-Genomic Origins of Proteins and Metabolism

It is proposed that evolution of inanimate matter to cells endowed with a nucleic acid- based coding of genetic information was preceded by an evolutionary phase, in which peptides not coded by nucleic acids were able to self-organize into networks capable of evolution towards increasing metabolic complexity. Recent findings that truly different, simple peptides (Keefe and Szostak, 2001) can perform the same function (such as ATP binding) provide experimental support for this mechanism of early protobiological evolution. The central concept underlying this mechanism is that the reproduction of cellular functions alone was sufficient for self-maintenance of protocells, and that self- replication of macromolecules was not required at this stage of evolution. The precise transfer of information between successive generations of the earliest protocells was unnecessary and, possibly, undesirable. The key requirement in the initial stage of protocellular evolution was an ability to rapidly explore a large number of protein sequences in order to discover a set of molecules capable of supporting self- maintenance and growth of protocells. Undoubtedly, the essential protocellular functions were carried out by molecules not nearly as efficient or as specific as contemporary proteins. Many, potentially unrelated sequences could have performed each of these functions at an evolutionarily acceptable level. As evolution progressed, however proteins must have performed their functions with increasing efficiency and specificity. This, in turn, put additional constraints on protein sequences and the fraction of proteins capable of performing their functions at the required level decreased. At some point, the likelihood of generating a sufficiently efficient set of proteins through a non-coded synthesis was so small that further evolution was not possible without storing information about the sequences of these proteins. Beyond this point, further evolution required coupling between proteins and informational polymers that is characteristic to all known forms of life. The emergence of such coupling must be postulated in any scenario of the origin of life, no matter whether it starts with RNA or proteins. To examine the evolutionary potential of non-genomic systems, a simple, computationally tractable model, which is still capable of capturing the essential features of the real system, has been studied computationally. Both constructive and destructive processes have been introduced into the model in a stochastic manner. Instead of assuming random reaction sets, only a suite of protobiologically plausible reactions has been considered. Peptides have been explicitly considered as protoenzymes and their catalytic efficiencies have been assigned on the basis of biochemical principles and experimental estimates. Simulations have been carried out using a novel approach (The Next Reaction Method) that is appropriate even for very low concentrations of reactants. Studies have focused on global autocatalytic processes and their diversity.

Pohorille, Andrew↗

Sequence Design of Random Heteropolymers as Protein Mimics

Random heteropolymers (RHPs) have been computationally designed and experimentally shown to recapitulate protein-like phase behavior and function. However, unlike proteins, RHP sequences are only statistically defined and cannot be sequenced. Recent developments in reversible-deactivation radical polymerization allowed simulated polymer sequences based on the well-established Mayo–Lewis equation to more accurately reflect ground-truth sequences that are experimentally synthesized. This led to opportunities to perform bioinformatics-inspired analysis on simulated sequences to guide the design, synthesis, and interpretation of RHPs. We compared batches on the order of 10000 simulated RHP sequences that vary by synthetically controllable and measurable RHP characteristics such as chemical heterogeneity and average degree of polymerization. Our analysis spans across 3 levels: segments along a single chain, sequences within a batch, and batch-averaged statistics. We discuss simulator fidelity and highlight the importance of robust segment definition. Examples are presented that demonstrate the use of simulated sequence analysis for in-silico iterative design to mimic protein hydrophobic/hydrophilic segment distributions in RHPs and compare RHP and protein sequence segments to explain experimental results of RHPs that mimic protein function. To facilitate the community use of this workflow, the simulator and analysis modules have been made available through an open source toolkit, the RHPapp.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Self-driving laboratories to autonomously navigate the protein fitness landscape

Abstract Protein engineering has nearly limitless applications across chemistry, energy and medicine, but creating new proteins with improved or novel functions remains slow, labor-intensive and inefficient. Here we present the Self-driving Autonomous Machines for Protein Landscape Exploration (SAMPLE) platform for fully autonomous protein engineering. SAMPLE is driven by an intelligent agent that learns protein sequence–function relationships, designs new proteins and sends designs to a fully automated robotic system that experimentally tests the designed proteins and provides feedback to improve the agent’s understanding of the system. We deploy four SAMPLE agents with the goal of engineering glycoside hydrolase enzymes with enhanced thermal tolerance. Despite showing individual differences in their search behavior, all four agents quickly converge on thermostable enzymes. Self-driving laboratories automate and accelerate the scientific discovery process and hold great potential for the fields of protein engineering and synthetic biology.

Rapp, Jacob T. (ORCID:0000000226835204)↗

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,↗