Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Chapter 10: Advances in Protein Engineering and its Application in Synthetic Biology

Protein engineering has been used successfully in fields ranging from medicine to food science to biofuels. Applications of protein engineering include developing antiviral peptides or other protein therapeutics, antibody engineering, designing protein-based logic circuits, engineering enzymes to be more specific or to function under industrially relevant conditions such as at higher temperatures or high/low pH, modifying cell signaling or regulatory functions, and so on. Advances in recombinant DNA, "omics," and CRISPR-Cas (clustered regularly interspaced short palindromic repeats and its associated proteins) technologies, combined with high-throughput screening facilities, will lead to improved methods for protein engineering, enabling easy modification of more proteins/enzymes for new specific applications. New methods for rational design, directed evolution, and computer-aided protein design will further accelerate the speed of protein evolution and expand the scope for protein engineering. In this chapter we discuss general protein engineering strategies and advances in engineering proteins with desired functions, focusing on the "design" and "build" part of the design-build-test-learn cycle.

BIOMASS FUELS↗

Effect of polyphenols on the rheology, microstructure and in vitro digestion of pea protein gels at various pH

Polyphenols exist widely in plants and interact with plant proteins distinctly depending on the environmental pH. This could potentially affect the gelling property of plant proteins and their digestion. In the present study, pea protein suspensions containing 0 %, 0.5 % and 1 % green tea polyphenols (GTP) were heated to form gels at pH5, pH7, and pH8.5. A strain amplitude sweep showed that the storage modulus (G’) and critical strain of pea protein gels decreased with increased GTP concentration at all examined pH. Gelation dynamic showed gelling of pea protein-GTP was delayed at pH7 and pH8.5 compared to the control. Differential scanning calorimetry showed that pea proteins became less heat resistant in the presence of GTP at pH5 and pH7, but not at pH8.5. Ultra-small and small-angle X-ray scattering showed that the radius of gyration of small- and medium-sized aggregates in pea protein gels was increased ~10–16 % and 22–30 % at pH7 and pH8.5, respectively when 1 % GTP was present. Finally, under such pH conditions, the proportion of structure with a radius of ~10–200 nm was increased in pea protein-GTP gels based on the volume size distribution. In vitro digestion found the soluble protein content of digesta of pea protein-GTP gels had a 5–8.5 % decrement with the presence of larger peptides when compared to pea protein gels.

59 BASIC BIOLOGICAL SCIENCES↗

Internal Fragments Generated by Electron Ionization Dissociation Enhance Protein Top-Down Mass Spectrometry

Top-down proteomics by mass spectrometry (MS) involves the mass measurement of an intact protein followed by subsequent activation of the protein to generate product ions. Electron-based fragmentation methods like electron capture dissociation and electron transfer dissociation are widely used for these types of analyses. Recently, electron ionization dissociation (EID), which utilizes higher energy electrons (>20 eV) has been suggested to be more efficient for top-down protein fragmentation compared to other electron-based dissociation methods. In this paper, we demonstrate that the use of EID enhances protein fragmentation and subsequent detection of protein fragments. Protein product ions can form by either single cleavage events, resulting in terminal fragments containing the C-terminus or N-terminus of the protein, or by multiple cleavage events to give rise to internal fragments that include neither the C-terminus nor the N-terminus of the protein. Conventionally, internal fragments have been disregarded, as reliable assignments of these fragments were limited. Here, we demonstrate that internal fragments generated by EID can account for ~20–40% of the mass spectral signals detected by top-down EID-MS experiments. By including internal fragments, the extent of the protein sequence that can be explained from a single tandem mass spectrum increases from ~50 to ~99% for 29 kDa carbonic anhydrase II and 8.6 kDa ubiquitin. When searching for internal fragments during data analysis, previously unassigned peaks can be readily and accurately assigned to confirm a given protein sequence and to enhance the utility of top-down protein sequencing experiments.

59 BASIC BIOLOGICAL SCIENCES↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Properties of protein unfolded states suggest broad selection for expanded conformational ensembles

Much attention is being paid to conformational biases in the ensembles of intrinsically disordered proteins. However, it is currently unknown whether or how conformational biases within the disordered ensembles of foldable proteins affect function in vivo. Recently, we demonstrated that water can be a good solvent for unfolded polypeptide chains, even those with a hydrophobic and charged sequence composition typical of folded proteins. These results run counter to the generally accepted model that protein folding begins with hydrophobicity-driven chain collapse. Here we investigate what other features, beyond amino acid composition, govern chain collapse. We found that local clustering of hydrophobic and/or charged residues leads to significant collapse of the unfolded ensemble of pertactin, a secreted autotransporter virulence protein from Bordetella pertussis , as measured by small angle X-ray scattering (SAXS). Sequence patterns that lead to collapse also correlate with increased intermolecular polypeptide chain association and aggregation. Crucially, sequence patterns that support an expanded conformational ensemble enhance pertactin secretion to the bacterial cell surface. Similar sequence pattern features are enriched across the large and diverse family of autotransporter virulence proteins, suggesting sequence patterns that favor an expanded conformational ensemble are under selection for efficient autotransporter protein secretion, a necessary prerequisite for virulence. More broadly, we found that sequence patterns that lead to more expanded conformational ensembles are enriched across water-soluble proteins in general, suggesting protein sequences are under selection to regulate collapse and minimize protein aggregation, in addition to their roles in stabilizing folded protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted protein degradation: from small molecules to complex organelles—a Keystone Symposia report

Targeted protein degradation is critical for proper cellular function and development. Protein degradation pathways, such as the ubiquitin proteasomes system, autophagy, and endosome–lysosome pathway, must be tightly regulated to ensure proper elimination of misfolded and aggregated proteins and regulate changing protein levels during cellular differentiation, while ensuring that normal proteins remain unscathed. Protein degradation pathways have also garnered interest as a means to selectively eliminate target proteins that may be difficult to inhibit via other mechanisms. Additionally, on June 7 and 8, 2021, several experts in protein degradation pathways met virtually for the Keystone eSymposium “Targeting protein degradation: from small molecules to complex organelles.” The event brought together researchers working in different protein degradation pathways in an effort to begin to develop a holistic, integrated vision of protein degradation that incorporates all the major pathways to understand how changes in them can lead to disease pathology and, alternatively, how they can be leveraged for novel therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Bacterial hemophilin homologs and their specific type eleven secretor proteins have conserved roles in heme capture and are diversifying as a family

Cellular life relies on enzymes that require metals, which must be acquired from extracellular sources. Bacteria utilize surface and secreted proteins to acquire such valuable nutrients from their environment. These include the cargo proteins of the type eleven secretion system (T11SS), which have been connected to host specificity, metal homeostasis, and nutritional immunity evasion. This Sec-dependent, Gram-negative secretion system is encoded by organisms throughout the phylum Proteobacteria, including human pathogens Neisseria meningitidis, Proteus mirabilis, Acinetobacter baumannii, and Haemophilus influenzae. Experimentally verified T11SS-dependent cargo include transferrin-binding protein B (TbpB), the hemophilin homologs heme receptor protein C (HrpC), hemophilin A (HphA), the immune evasion protein factor-H binding protein (fHbp), and the host symbiosis factor nematode intestinal localization protein C (NilC). Here, we examined the specificity of T11SS systems for their cognate cargo proteins using taxonomically distributed homolog pairs of T11SS and hemophilin cargo and explored the ligand binding ability of those hemophilin cargo homologs. In vivo expression in Escherichia coli of hemophilin homologs revealed that each is secreted in a specific manner by its cognate T11SS protein. Sequence analysis and structural modeling suggest that all hemophilin homologs share an N-terminal ligand-binding domain with the same topology as the ligand-binding domains of the Haemophilus haemolyticus heme binding protein (Hpl) and HphA. We term this signature feature of this group of proteins the hemophilin ligand-binding domain. Network analysis of hemophilin homologs revealed five subclusters and representatives from four of these showed variable heme-binding activities, which, combined with sequence-structure variation, suggests that hemophilins are diversifying in function.

59 BASIC BIOLOGICAL SCIENCES↗

Heterologous expression of a fully active Azotobacter vinelandii nitrogenase Fe protein in Escherichia coli

ABSTRACT The functional versatility of the Fe protein, the reductase component of nitrogenase, makes it an appealing target for heterologous expression, which could facilitate future biotechnological adaptations of nitrogenase-based production of valuable chemical commodities. Yet, the heterologous synthesis of a fully active Fe protein of Azotobacter vinelandii ( Av NifH) in Escherichia coli has proven to be a challenging task. Here, we report the successful synthesis of a fully active Av NifH protein upon co-expression of this protein with Av IscS/U and Av NifM in E. coli . Our metal, activity, electron paramagnetic resonance, and X-ray absorption spectroscopy/extended X-ray absorption fine structure (EXAFS) data demonstrate that the heterologously expressed Av NifH protein has a high [Fe 4 S 4 ] cluster content and is fully functional in nitrogenase catalysis and assembly. Moreover, our phylogenetic analyses and structural predictions suggest that Av NifM could serve as a chaperone and assist the maturation of a cluster-replete Av NifH protein. Given the crucial importance of the Fe protein for the functionality of nitrogenase, this work establishes an effective framework for developing a heterologous expression system of the complete, two-component nitrogenase system; additionally, it provides a useful tool for further exploring the intricate biosynthetic mechanism of this structurally unique and functionally important metalloenzyme. IMPORTANCE The heterologous expression of a fully active Azotobacter vinelandii Fe protein (AvNifH) has never been accomplished. Given the functional importance of this protein in nitrogenase catalysis and assembly, the successful expression of AvNifH in Escherichia coli as reported herein supplies a key element for the further development of heterologous expression systems that explore the catalytic versatility of the Fe protein, either on its own or as a key component of nitrogenase, for nitrogenase-based biotechnological applications in the future. Moreover, the “clean” genetic background of the heterologous expression host allows for an unambiguous assessment of the effect of certain nif-encoded protein factors, such as AvNifM described in this work, in the maturation of AvNifH, highlighting the utility of this heterologous expression system in further advancing our understanding of the complex biosynthetic mechanism of nitrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

Scalable production of recombinant three-finger proteins: from inclusion bodies to high quality molecular probes

The three-finger proteins are a collection of disulfide bond rich proteins of great biomedical interests. Scalable recombinant expression and purification of bioactive three-finger proteins is quite difficult. We introduce a working pipeline for expression, purification and validation of disulfide-bond rich three-finger proteins using E. coli as the expression host. With this pipeline, we have successfully obtained highly purified and bioactive recombinant α-Βungarotoxin, k-Bungarotoxin, Hannalgesin, Mambalgin-1, α-Cobratoxin, MTα, Slurp1, Pate B etc. Milligrams to hundreds of milligrams of recombinant three finger proteins were obtained within weeks in the lab. The recombinant proteins showed specificity in binding assay and six of them were crystallized and structurally validated using X-ray diffraction protein crystallography. Our pipeline allows refolding and purifying recombinant three finger proteins under optimized conditions and can be scaled up for massive production of three finger proteins. As many three finger proteins have attractive therapeutic or research interests and due to the extremely high quality of the recombinant three finger proteins we obtained, our method provides a competitive alternative to either their native counterparts or chemically synthetic ones and should facilitate related research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

Investigation of design principles for metal-binding and conductive protein assemblies

Throughout the lifetime of this initiative, including renewals, we focused on understanding the fundamental principles of protein-protein interface design that enable predictable and modular spatial and kinetic control of multi-component protein self-assembly in 1D, 2D, and 3D, including the interface with inorganic materials, small molecules, and metal ions. We designed individual protein components that bind specific metal ions, including REEs and transport ions across lipid membranes. We created helical 1D filaments of repeating units with programmed periodicity, pitch, and multi-component environmentally responsive self-assembling protein fibers. We showed that these filaments reversibly assemble and disassemble under specific pH conditions and created end-specific caps that independently tune the balance of attachment and detachment rates at each terminus of the filament. Using similar filaments, we succeeded in binding arrays of heme and chlorophyll molecules and assembling patterned helical coatings around carbon nanotubes in efforts to create de novo conductive nanowires. By arraying REE binding sites in a large circular tandem array with a repeat protein-based cyclic oligomer, we created a molecular scaffold for superradiance and paramagnetic quantum sensing. We created a range of one-component and two-component self-assembling 2D arrays and showed that when designed to engage cell receptors, these arrays can control cell behavior from outside the cell signal to inside the cell. We designed helical repeat proteins with variable lengths displaying charged residues in a pattern matched to the cation lattice of mica. achieved a range of ordered states with an epitaxial match to the underlying crystal lattice. We further applied the learned principles of protein-induced biomineralization to design proteins with an interface lattice matching CaCO 3 and guide the formation of specific crystal forms of CaCO 3 from solution, a significant advance toward the global need to manage carbon. In all cases of mineral lattice matching and biomineralization, we followed assembly using molecularly resolved in situ AFM imaging and extracted information about assembly pathways and energetics, applying deep learning to quantify the dynamics of protein self-organization. We developed techniques for using dynamic metal-dependent interfaces on protein nanopores for discriminatively sensing dilute REEs in solution and demonstrated the use of strong metal-binding interfaces to drive nanocage disassembly for conditional nanocompartmentalization applications. This grant supported 11 people, including Asim Bera, Evans Brackenbrough, Andrew Borst, Nikita Hanikel, Timothy Huddy, Emily Joyce, Alex Young-Seug Kang, Ryan Kibler, Joshua Morris Lubner, Harley Pyles, and Shuai Zhang. The research effort culminated in the production of published papers and theses. Electronic Thesis/Dissertation are distributed by ProQuest/UMI Dissertation Publishing and made available on an open access basis through UW Libraries ResearchWorks Service.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Protein Extraction, Precipitation, and Recovery from Chlorella sorokiniana Using Mechanochemical Methods

Protein extraction, precipitation, and recovery methods were evaluated by this study using a green alga—Chlorella sorokiniana. A mechanochemical cell disruption process was applied to facilitate protein extraction from microalgal biomass. Optimization of the mechanochemical process resulted in milling conditions that achieved a protein extraction of 52.7 ± 6.45%. The consequent acid precipitation method was optimized to recover 98.7% of proteins from the microalgal slurry. The measured protein content of the protein isolate was 41.4% w/w. These results indicate that the precipitation method is successful at recovering the extracted proteins in the algal slurry; however, the removal of non-protein solids during centrifugation and pH adjustment is not complete. The energy balance analysis elucidated that the energy demand of the protein extraction and recovery operation, at 0.83 MJ/kg dry algal biomass, is much lower than previous studies using high-pressure homogenization and membrane filtration. This study concludes that mechanochemical protein extraction and recovery is an effective, low-energy processing method, which could be used by algal biorefineries to prepare algal proteins for value-added chemical production as well as to make algal carbohydrates and lipids in the residual biomass more accessible for biofuel production.

59 BASIC BIOLOGICAL SCIENCES↗

Isolation and characterization of a novel calmodulin-binding protein from potato

Tuberization in potato is controlled by hormonal and environmental signals. Ca(2+), an important intracellular messenger, and calmodulin (CaM), one of the primary Ca(2+) sensors, have been implicated in controlling diverse cellular processes in plants including tuberization. The regulation of cellular processes by CaM involves its interaction with other proteins. To understand the role of Ca(2+)/CaM in tuberization, we have screened an expression library prepared from developing tubers with biotinylated CaM. This screening resulted in isolation of a cDNA encoding a novel CaM-binding protein (potato calmodulin-binding protein (PCBP)). Ca(2+)-dependent binding of the cDNA-encoded protein to CaM is confirmed by (35)S-labeled CaM. The full-length cDNA is 5 kb long and encodes a protein of 1309 amino acids. The deduced amino acid sequence showed significant similarity with a hypothetical protein from another plant, Arabidopsis. However, no homologs of PCBP are found in nonplant systems, suggesting that it is likely to be specific to plants. Using truncated versions of the protein and a synthetic peptide in CaM binding assays we mapped the CaM-binding region to a 20-amino acid stretch (residues 1216-1237). The bacterially expressed protein containing the CaM-binding domain interacted with three CaM isoforms (CaM2, CaM4, and CaM6). PCBP is encoded by a single gene and is expressed differentially in the tissues tested. The expression of CaM, PCBP, and another CaM-binding protein is similar in different tissues and organs. The predicted protein contained seven putative nuclear localization signals and several strong PEST motifs. Fusion of the N-terminal region of the protein containing six of the seven nuclear localization signals to the reporter gene beta-glucuronidase targeted the reporter gene to the nucleus, suggesting a nuclear role for PCBP.

NASA Discipline Plant Biology↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

Dissecting the structural heterogeneity of proteins by native mass spectrometry

Abstract A single gene yields many forms of proteins via combinations of posttranscriptional/posttranslational modifications. Proteins also fold into higher‐order structures and interact with other molecules. The combined molecular diversity leads to the heterogeneity of proteins that manifests as distinct phenotypes. Structural biology has generated vast amounts of data, effectively enabling accurate structural prediction by computational methods. However, structures are often obtained heterologously under homogeneous states in vitro. The lack of native heterogeneity under cellular context creates challenges in precisely connecting the structural data to phenotypes. Mass spectrometry (MS) based proteomics methods can profile proteome composition of complex biological samples. Most MS methods follow the “bottom‐up” approach, which denatures and digests proteins into short peptide fragments for ease of detection. Coupled with chemical biology approaches, higher‐order structures can be probed via incorporation of covalent labels on native proteins that are maintained at the peptide level. Alternatively, native MS follows the “top‐down” approach and directly analyzes intact proteins under nondenaturing conditions. Various tandem MS activation methods can dissect the intact proteins for in‐depth structural elucidation. Herein, we review recent native MS applications for characterizing heterogeneous samples, including proteins binding to mixtures of ligands, homo/hetero‐complexes with varying stoichiometry, intrinsically disordered proteins with dynamic conformations, glycoprotein complexes with mixed modification states, and active membrane protein complexes in near‐native membrane environments. We summarize the benefits, challenges, and ongoing developments in native MS, with the hope to demonstrate an emerging technology that complements other tools by filling the knowledge gaps in understanding the molecular heterogeneity of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing Structural, Thermal, and Functional Characteristics of Marigold Flower Protein as a Sustainable Food Ingredient

The demand for sustainable and alternative protein sources has been on the rise, driving interest in the valorization of underutilized plants. This study evaluated Calendula officinalis (marigold), a common floral waste, as a sustainable alternative protein source for the food industry. The primary objective of this study was to investigate the physicochemical properties of protein fractions from Calendula officinalis flower to evaluate their potential as a novel protein ingredient. Extraction of the Calendula officinalis flower yielded 92.17% of the crude protein. A sequential extraction of albumin, globulin, glutelin, and prolamin from marigold flower revealed albumin as the dominant fraction (65.47%) and exhibited the highest protein functionality, including water-holding capacity (2.37 g/g), oil-holding capacity (2.49 g/g), and emulsifying capacity (65.22 mL/g). Compared with other protein fractions, glutelin showed a relatively high emulsifying and foaming capacity (EC: 59.13 mL/g; FC: 16.23%). Differential scanning calorimetry revealed high thermal stability for albumin (T p = 105.28 °C) and glutelin (T p = 97.6 °C). Sodium Dodecyl Sulfate–Polyacrylamide Gel Electrophoresis (SDS-PAGE) and Liquid Chromatography–Mass Spectrometry (LC-MS) confirmed the presence of abundant low-molecular-weight polypeptides (<37 kDa), which enhanced emulsification, while scanning electron microscopy revealed porous structures aligned with hydration properties. Antioxidant activity was higher in albumin and glutelin, linked to surface hydrophobicity. LC-MS/MS identified 33 short-chain proteins, including oxidoreductase proteins and lipid-transfer proteins. Findings highlight marigold flower proteins as a sustainable, functional ingredient for a diverse range of food applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Population-based heteropolymer design to mimic protein mixtures

Biological fluids, the most complex blends, have compositions that constantly vary and cannot be molecularly defined. Despite these uncertainties, proteins fluctuate, fold, function and evolve as programmed. We propose that in addition to the known monomeric sequence requirements, protein sequences encode multi-pair interactions at the segmental level to navigate random encounters; synthetic heteropolymers capable of emulating such interactions can replicate how proteins behave in biological fluids individually and collectively. Here, we extracted the chemical characteristics and sequential arrangement along a protein chain at the segmental level from natural protein libraries and used the information to design heteropolymer ensembles as mixtures of disordered, partially folded and folded proteins. For each heteropolymer ensemble, the level of segmental similarity to that of natural proteins determines its ability to replicate many functions of biological fluids including assisting protein folding during translation, preserving the viability of fetal bovine serum without refrigeration, enhancing the thermal stability of proteins and behaving like synthetic cytosol under biologically relevant conditions. Molecular studies further translated protein sequence information at the segmental level into intermolecular interactions with a defined range, degree of diversity and temporal and spatial availability. This framework provides valuable guiding principles to synthetically realize protein properties, engineer bio/abiotic hybrid materials and, ultimately, realize matter-to-life transformations.

59 BASIC BIOLOGICAL SCIENCES↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗