Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Structural studies of intrinsically disordered MLL -fusion protein AF9 in complex with peptidomimetic inhibitors

AF9 (MLLT3) and its paralog ENL(MLLT1) are members of the YEATS family of proteins with important role in transcriptional and epigenetic regulatory complexes. These proteins are two common MLL fusion partners in MLL -rearranged leukemias. The oncofusion proteins MLL-AF9/ENL recruit multiple binding partners, including the histone methyltransferase DOT1L, leading to aberrant transcriptional activation and enhancing the expression of a characteristic set of genes that drive leukemogenesis. The interaction between AF9 and DOT1L is mediated by an intrinsically disordered C-terminal ANC1 homology domain (AHD) in AF9, which undergoes folding upon binding of DOT1L and other partner proteins. We have recently reported peptidomimetics that disrupt the recruitment of DOT1L by AF9 and ENL, providing a proof-of-concept for targeting AHD and assessing its druggability. Intrinsically disordered proteins, such as AF9 AHD, are difficult to study and characterize experimentally on a structural level. In this study, we present a successful protein engineering strategy to facilitate structural investigation of the intrinsically disordered AF9 AHD domain in complex with peptidomimetic inhibitors by using maltose binding protein (MBP) as a crystallization chaperone connected with linkers of varying flexibility and length. The strategic incorporation of disulfide bonds provided diffraction-quality crystals of the two disulfide-bridged MBP–AF9 AHD fusion proteins in complex with the peptidomimetics. These successfully determined first series of 2.1–2.6 Å crystal complex structures provide high-resolution insights into the interactions between AHD and its inhibitors, shedding light on the role of AHD in recruiting various binding partner proteins. We show that the overall complex structures closely resemble the reported NMR structure of AF9 AHD/DOT1L with notable difference in the conformation of the β-hairpin region, stabilized through conserved hydrogen bonds network. These first series of AF9 AHD/peptidomimetics complex structures are providing insights of the protein–inhibitor interactions and will facilitate further development of novel inhibitors targeting the AF9/ENL AHD domain.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Skeletides: A Modular, Simplified Physical Model of Protein Secondary Structure

Three-dimensional (3D) models are essential for visualization and conceptual understanding of complex architectures such as protein structure. Although there is a plethora of software platforms that allow digital depictions of protein structure at an atomic level in silico, physical models are needed to convey an intuitive understanding of biomolecular architecture. However, it is a challenge to represent all the relevant features of proteins in a single physical model due to their sheer structural complexity. Here, we describe a modular protein model that focuses only on representation of the secondary structure—the underlying structural skeleton. The simplified model consists of amino acid units, which can be linked together to reproduce the two most fundamental structural features of protein secondary structure: the relative positions of the amino acid alpha carbon atoms, and the intra-main chain hydrogen bonding pattern. We use 3D printing and magnets to create a set of three modular amino acid building blocks, which when linked together into a chain, can faithfully represent alpha helices, beta sheets, and turns. These simple to make models can be used to quickly assemble the alpha carbon trace of an entire protein domain, which conveys a tactile experience of the complexity of protein skeletal architecture. These models highlight the modularity of protein structure: where a single structural unit, the amino acid, can be linked together to form a larger, regular secondary structure. These models also have a propensity to spontaneously organize into alpha helices and beta sheets, as demonstrated by their ability to autonomously assemble when placed in a circulating water tank. These models provide a missing educational tool to expand knowledge of protein structure, foster deeper insight into protein folding, and inspire greater interest in biomacromolecular architecture.

36 MATERIALS SCIENCE↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

SARS-CoV-2 spike protein variant binding affinity to an angiotensin-converting enzyme 2 fusion glycoproteins

Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2), the causative agent of the Coronavirus disease 2019 (Covid-19) pandemic, continues to evolve and circulate globally. Current prophylactic and therapeutic countermeasures against Covid-19 infection include vaccines, small molecule drugs, and neutralizing monoclonal antibodies. SARS-CoV-2 infection is mainly mediated by the viral spike glycoprotein binding to angiotensin converting enzyme 2 (ACE2) on host cells for viral entry. As emerging mutations in the spike protein evade efficacy of spike-targeted countermeasures, a potential strategy to counter SARS-CoV-2 infection is to competitively block the spike protein from binding to the host ACE2 using a soluble recombinant fusion protein that contains a human ACE2 and an IgG1-Fc domain (ACE2-Fc). Here, we have established Chinese Hamster Ovary (CHO) cell lines that stably express ACE2-Fc proteins in which the ACE2 domain either has or has no catalytic activity. The fusion proteins were produced and purified to partially characterize physicochemical properties and spike protein binding. Our results demonstrate the ACE2-Fc fusion proteins are heavily N-glycosylated, sensitive to thermal stress, and actively bind to five spike protein variants (parental, alpha, beta, delta, and omicron) with different affinity. Our data demonstrates a proof-of-concept production strategy for ACE2-Fc fusion glycoproteins that can bind to different spike protein variants to support the manufacture of potential alternative countermeasures for emerging SARS-CoV-2 variants.

60 APPLIED LIFE SCIENCES↗

3D Visualization of Proteins within Metal–Organic Frameworks via Ferritin‐Enabled Electron Microscopy

Abstract Electron tomography holds great promise as a tool for investigating the 3D morphologies and internal structures of metal‐organic framework‐based protein biocomposites (protein@MOFs). Understanding the 3D spatial arrangement of proteins within protein@MOFs is paramount for developing synthetic methods to control their spatial localization and distribution patterns within the biocomposite crystals. In this study, the naturally occurring iron oxide mineral core of the protein horse spleen ferritin (Fn) is leveraged as a contrast agent to directly observe individual proteins once encapsulated into MOFs by electron microscopy techniques. This methodology couples scanning electron microscopy, transmission electron microscopy, and electron tomography to garner detailed 2D and 3D structural interpretations of where proteins spatially lie in Fn@MOF crystals, addressing the significant gaps in understanding how synthetic conditions relate to overall protein spatial localization and aggregation. These findings collectively reveal that adjusting the ligand‐to‐metal ratios, protein concentration, and the use of denaturing agents alters how proteins are arranged, localized, and aggregated within MOF crystals.

Chemistry↗

Engineering an efficient and bright split Corynactis californica green fluorescent protein

Split green fluorescent protein (GFP) has been used in a panoply of cellular biology applications to study protein translocation, monitor protein solubility and aggregation, detect protein–protein interactions, enhance protein crystallization, and even map neuron contacts. Recent work shows the utility of split fluorescent proteins for large scale labeling of proteins in cells using CRISPR, but sets of efficient split fluorescent proteins that do not cross-react are needed for multiplexing experiments. We present a new monomeric split green fluorescent protein (ccGFP) engineered from a tetrameric GFP found in Corynactis californica, a bright red colonial anthozoan similar to sea anemones and scleractinian stony corals. Split ccGFP from C. californica complements up to threefold faster compared to the original Aequorea victoria split GFP and enable multiplexed labeling with existing A. victoria split YFP and CFP.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

3D-equivariant graph neural networks for protein model quality assessment

Quality assessment (QA) of predicted protein tertiary structure models plays an important role in ranking and using them. With the recent development of deep learning end-to-end protein structure prediction techniques for generating highly confident tertiary structures for most proteins, it is important to explore corresponding QA strategies to evaluate and select the structural models predicted by them since these models have better quality and different properties than the models predicted by traditional tertiary structure prediction methods. We develop EnQA, a novel graph-based 3D-equivariant neural network method that is equivariant to rotation and translation of 3D objects to estimate the accuracy of protein structural models by leveraging the structural features acquired from the state-of-the-art tertiary structure prediction method—AlphaFold2. We train and test the method on both traditional model datasets (e.g. the datasets of the Critical Assessment of Techniques for Protein Structure Prediction) and a new dataset of high-quality structural models predicted only by AlphaFold2 for the proteins whose experimental structures were released recently. Our approach achieves state-of-the-art performance on protein structural models predicted by both traditional protein structure prediction methods and the latest end-to-end deep learning method—AlphaFold2. It performs even better than the model QA scores provided by AlphaFold2 itself. The results illustrate that the 3D-equivariant graph neural network is a promising approach to the evaluation of protein structural models. Integrating AlphaFold2 features with other complementary sequence and structural features is important for improving protein model QA.

59 BASIC BIOLOGICAL SCIENCES↗

Putting Humpty Dumpty Back Together Again: What Does Protein Quantification Mean in Bottom-Up Proteomics?

Bottom-up proteomics provides peptide measurements and has been invaluable for moving proteomics into large-scale analyses. Commonly, a single quantitative value is reported for each protein-coding gene by aggregating peptide quantities into protein groups following protein inference or parsimony. However, given the complexity of both RNA splicing and post-translational protein modification, it is overly simplistic to assume that all peptides that map to a singular protein-coding gene will demonstrate the same quantitative response. Here, by assuming that all peptides from a protein-coding sequence are representative of the same protein, we may miss the discovery of important biological differences. To capture the contributions of existing proteoforms, we need to reconsider the practice of aggregating protein values to a single quantity per protein-coding gene.

59 BASIC BIOLOGICAL SCIENCES↗

evSeq: Cost-Effective Amplicon Sequencing of Every Variant in a Protein Library

Widespread availability of protein sequence-fitness data would revolutionize both our biochemical understanding of proteins and our ability to engineer them. Unfortunately, even though thousands of protein variants are generated and evaluated for fitness during a typical protein engineering campaign, most are never sequenced, leaving a wealth of potential sequence-fitness information untapped. Primarily, this is because sequencing is unnecessary for many protein engineering strategies; the added cost and effort of sequencing is thus unjustified. It also results from the fact that, even though many lower cost sequencing strategies have been developed, they often require at least some sequencing or computational resources, both of which can be barriers to access. In this work, we present every variant sequencing (evSeq), a method and collection of tools/standardized components for sequencing a variable region within every variant gene produced during a protein engineering campaign at a cost of cents per variant. evSeq was designed to democratize low-cost sequencing for protein engineers and, indeed, anyone interested in engineering biological systems. Execution of its wet-lab component is simple, requires no sequencing experience to perform, relies only on resources and services typically available to biology labs, and slots neatly into existing protein engineering workflows. Analysis of evSeq data is likewise made simple by its accompanying software (found at github.com/fhalab/evSeq, documentation at fhalab.github.io/evSeq), which can be run on a personal laptop and was designed to be accessible to users with no computational experience. Here, low-cost and easy to use, evSeq makes collection of extensive protein variant sequence-fitness data practical.

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of knotted tandem repeat proteins

De novo protein design methods can create proteins with folds not yet seen in nature. These methods largely focus on optimizing the compatibility between the designed sequence and the intended conformation, without explicit consideration of protein folding pathways. Deeply knotted proteins, whose topologies may introduce substantial barriers to folding, thus represent an interesting test case for protein design. Here we report our attempts to design proteins with trefoil (3 1 ) and pentafoil (5 1 ) knotted topologies. We extended previously described algorithms for tandem repeat protein design in order to construct deeply knotted backbones and matching designed repeat sequences (N = 3 repeats for the trefoil and N = 5 for the pentafoil). We confirmed the intended conformation for the trefoil design by X ray crystallography, and we report here on this protein’s structure, stability, and folding behaviour. The pentafoil design misfolded into an asymmetric structure (despite a 5-fold symmetric sequence); two of the four repeat-repeat units matched the designed backbone while the other two diverged to form local contacts, leading to a trefoil rather than pentafoil knotted topology. Our results also provide insights into the folding of knotted proteins.

59 BASIC BIOLOGICAL SCIENCES↗

A General Framework to Learn Tertiary Structure for Protein Sequence Characterization

During the past five years, deep-learning algorithms have enabled ground-breaking progress towards the prediction of tertiary structure from a protein sequence. Very recently, we developed SAdLSA, a new computational algorithm for protein sequence comparison via deep-learning of protein structural alignments. SAdLSA shows significant improvement over established sequence alignment methods. In this contribution, we show that SAdLSA provides a general machine-learning framework for structurally characterizing protein sequences. By aligning a protein sequence against itself, SAdLSA generates a fold distogram for the input sequence, including challenging cases whose structural folds were not present in the training set. About 70% of the predicted distograms are statistically significant. Although at present the accuracy of the intra-sequence distogram predicted by SAdLSA self-alignment is not as good as deep-learning algorithms specifically trained for distogram prediction, it is remarkable that the prediction of single protein structures is encoded by an algorithm that learns ensembles of pairwise structural comparisons, without being explicitly trained to recognize individual structural folds. As such, SAdLSA can not only predict protein folds for individual sequences, but also detects subtle, yet significant, structural relationships between multiple protein sequences using the same deep-learning neural network. The former reduces to a special case in this general framework for protein sequence annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme-Directed Functionalization of Designed, Two-Dimensional Protein Lattices

The design and construction of crystalline protein arrays to selectively assemble ordered nanoscale materials has potential applications in sensing, catalysis and medicine. Whereas numerous designs have been implemented for the bottom-up construction of novel protein assemblies, the generation of artificial functional materials has been relatively unexplored. Enzyme-directed post-translational modifications are responsible for the functional diversity of the proteome and thus, could be harnessed to selectively modify artificial protein assemblies. In this study, we describe the use of phosphopantetheinyl transferases (PPTases), a class of enzymes that covalently modify proteins using coenzyme A (CoA), to site-selectively tailor the surface of designed, two-dimensional (2D) protein crystals. We demonstrate that a short peptide (ybbR) or a molecular tag (CoA) can be covalently tethered to 2D arrays to enable enzymatic functionalization using Sfp PPTase. Here, the site-specific modification of two different protein array platforms is facilitated by PPTases to afford both small-molecule- and protein-functionalized surfaces with no loss in crystalline order. This work highlights the potential for chemoenzymatic modification of large protein surfaces towards the generation of sophisticated protein platforms reminiscent of the complex landscape of cell surfaces.

36 MATERIALS SCIENCE↗

De Novo Design of Proteins That Bind Naphthalenediimides, Powerful Photooxidants with Tunable Photophysical Properties

De novo protein design provides a framework to test our understanding of protein function and build proteins with cofactors and functions not found in nature. Here, we report the design of proteins designed to bind powerful photooxidants and the evaluation of the use of these proteins to generate diffusible small-molecule reactive species. Because excited-state dynamics are influenced by the dynamics and hydration of a photooxidant’s environment, it was important to not only design a binding site but also to evaluate its dynamic properties. Thus, we used computational design in conjunction with molecular dynamics (MD) simulations to design a protein, designated NBP (NDI Binding Protein), that held a naphthalenediimide (NDI), a powerful photooxidant, in a programmable molecular environment. Solution NMR confirmed the structure of the complex. We evaluated two NDI cofactors in this de novo protein using ultrafast pump–probe spectroscopy to evaluate light-triggered intra- and intermolecular electron transfer function. Moreover, we demonstrated the utility of this platform to activate multiple molecular probes for protein labeling.

carbonyls↗

Rapid and automated design of two-component protein nanomaterials using ProteinMPNN

The design of protein–protein interfaces using physics-based design methods such as Rosetta requires substantial computational resources and manual refinement by expert structural biologists. Deep learning methods promise to simplify protein–protein interface design and enable its application to a wide variety of problems by researchers from various scientific disciplines. Here, we test the ability of a deep learning method for protein sequence design, ProteinMPNN, to design two-component tetrahedral protein nanomaterials and benchmark its performance against Rosetta. ProteinMPNN had a similar success rate to Rosetta, yielding 13 new experimentally confirmed assemblies, but required orders of magnitude less computation and no manual refinement. The interfaces designed by ProteinMPNN were substantially more polar than those designed by Rosetta, which facilitated in vitro assembly of the designed nanomaterials from independently purified components. Crystal structures of several of the assemblies confirmed the accuracy of the design method at high resolution. Our results showcase the potential of deep learning–based methods to unlock the widespread application of designed protein–protein interfaces and self-assembling protein nanomaterials in biotechnology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Pulse Protein Isolates as Competitive Food Ingredients: Origin, Composition, Functionalities, and the State-of-the-Art Manufacturing

The ever-increasing world population and environmental stress are leading to surging demand for nutrient-rich food products with cleaner labeling and improved sustainability. Plant proteins, accordingly, are gaining enormous popularity compared with counterpart animal proteins in the food industry. While conventional plant protein sources, such as wheat and soy, cause concerns about their allergenicity, peas, beans, chickpeas, lentils, and other pulses are becoming important staples owing to their agronomic and nutritional benefits. However, the utilization of pulse proteins is still limited due to unclear pulse protein characteristics and the challenges of characterizing them from extensively diverse varieties within pulse crops. To address these challenges, the origins and compositions of pulse crops were first introduced, while an overarching description of pulse protein physiochemical properties, e.g., interfacial properties, aggregation behavior, solubility, etc., are presented. For further enhanced functionalities, appropriate modifications (including chemical, physical, and enzymatic treatment) are necessary. Among them, non-covalent complexation and enzymatic strategies are especially preferable during the value-added processing of clean-label pulse proteins for specific focus. This comprehensive review aims to provide an in-depth understanding of the interrelationships between the composition, structure, functional characteristics, and advanced modification strategies of pulse proteins, which is a pillar of high-performance pulse protein in future food manufacturing.

59 BASIC BIOLOGICAL SCIENCES↗