Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Dissecting the structural heterogeneity of proteins by native mass spectrometry

Abstract A single gene yields many forms of proteins via combinations of posttranscriptional/posttranslational modifications. Proteins also fold into higher‐order structures and interact with other molecules. The combined molecular diversity leads to the heterogeneity of proteins that manifests as distinct phenotypes. Structural biology has generated vast amounts of data, effectively enabling accurate structural prediction by computational methods. However, structures are often obtained heterologously under homogeneous states in vitro. The lack of native heterogeneity under cellular context creates challenges in precisely connecting the structural data to phenotypes. Mass spectrometry (MS) based proteomics methods can profile proteome composition of complex biological samples. Most MS methods follow the “bottom‐up” approach, which denatures and digests proteins into short peptide fragments for ease of detection. Coupled with chemical biology approaches, higher‐order structures can be probed via incorporation of covalent labels on native proteins that are maintained at the peptide level. Alternatively, native MS follows the “top‐down” approach and directly analyzes intact proteins under nondenaturing conditions. Various tandem MS activation methods can dissect the intact proteins for in‐depth structural elucidation. Herein, we review recent native MS applications for characterizing heterogeneous samples, including proteins binding to mixtures of ligands, homo/hetero‐complexes with varying stoichiometry, intrinsically disordered proteins with dynamic conformations, glycoprotein complexes with mixed modification states, and active membrane protein complexes in near‐native membrane environments. We summarize the benefits, challenges, and ongoing developments in native MS, with the hope to demonstrate an emerging technology that complements other tools by filling the knowledge gaps in understanding the molecular heterogeneity of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing Structural, Thermal, and Functional Characteristics of Marigold Flower Protein as a Sustainable Food Ingredient

The demand for sustainable and alternative protein sources has been on the rise, driving interest in the valorization of underutilized plants. This study evaluated Calendula officinalis (marigold), a common floral waste, as a sustainable alternative protein source for the food industry. The primary objective of this study was to investigate the physicochemical properties of protein fractions from Calendula officinalis flower to evaluate their potential as a novel protein ingredient. Extraction of the Calendula officinalis flower yielded 92.17% of the crude protein. A sequential extraction of albumin, globulin, glutelin, and prolamin from marigold flower revealed albumin as the dominant fraction (65.47%) and exhibited the highest protein functionality, including water-holding capacity (2.37 g/g), oil-holding capacity (2.49 g/g), and emulsifying capacity (65.22 mL/g). Compared with other protein fractions, glutelin showed a relatively high emulsifying and foaming capacity (EC: 59.13 mL/g; FC: 16.23%). Differential scanning calorimetry revealed high thermal stability for albumin (T p = 105.28 °C) and glutelin (T p = 97.6 °C). Sodium Dodecyl Sulfate–Polyacrylamide Gel Electrophoresis (SDS-PAGE) and Liquid Chromatography–Mass Spectrometry (LC-MS) confirmed the presence of abundant low-molecular-weight polypeptides (<37 kDa), which enhanced emulsification, while scanning electron microscopy revealed porous structures aligned with hydration properties. Antioxidant activity was higher in albumin and glutelin, linked to surface hydrophobicity. LC-MS/MS identified 33 short-chain proteins, including oxidoreductase proteins and lipid-transfer proteins. Findings highlight marigold flower proteins as a sustainable, functional ingredient for a diverse range of food applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Population-based heteropolymer design to mimic protein mixtures

Biological fluids, the most complex blends, have compositions that constantly vary and cannot be molecularly defined. Despite these uncertainties, proteins fluctuate, fold, function and evolve as programmed. We propose that in addition to the known monomeric sequence requirements, protein sequences encode multi-pair interactions at the segmental level to navigate random encounters; synthetic heteropolymers capable of emulating such interactions can replicate how proteins behave in biological fluids individually and collectively. Here, we extracted the chemical characteristics and sequential arrangement along a protein chain at the segmental level from natural protein libraries and used the information to design heteropolymer ensembles as mixtures of disordered, partially folded and folded proteins. For each heteropolymer ensemble, the level of segmental similarity to that of natural proteins determines its ability to replicate many functions of biological fluids including assisting protein folding during translation, preserving the viability of fetal bovine serum without refrigeration, enhancing the thermal stability of proteins and behaving like synthetic cytosol under biologically relevant conditions. Molecular studies further translated protein sequence information at the segmental level into intermolecular interactions with a defined range, degree of diversity and temporal and spatial availability. This framework provides valuable guiding principles to synthetically realize protein properties, engineer bio/abiotic hybrid materials and, ultimately, realize matter-to-life transformations.

59 BASIC BIOLOGICAL SCIENCES↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

A dynamic protein interactome drives energy conservation and electron flux in Thermococcus kodakarensis

ABSTRACT Life is supported by energy gains fueled by catabolism of a wide range of substrates, each reliant on the selective partitioning of electrons through redox ( red uction and ox idation) reactions. Electron flux through tunable and regulated protein interactions provides dynamic routes for energy conservation, but how electron flux is regulated in vivo , particularly for archaeal metabolisms that support rapid growth at the thermodynamic limits of life, is poorly understood. Identification of bona fide in vivo protein assemblies and how such assemblies dictate the totality of electron flux is critical to our understanding of the regulation imposed on metabolism, energy production, and energy conservation. Here, 25 key proteins in central metabolic redox pathways in the model, genetically accessible, hyperthermophilic archaeon Thermococcus kodakarensis , were purified to reveal an extensive, dynamic, and tightly interconnected network of protein interactions that responds to environmental cues (such as the availability of various reductive sinks) to direct electron flux to maximize energetic gains. Interactions connecting disparate functions suggest many catabolic and anabolic activities occur in spatial proximity in vivo , and while protein complexes have been historically defined under optimal conditions, many of these complexes appear to maintain alternative partnerships in changing conditions. The totality of the results obtained redefines our understanding of in vivo assemblies driving ancient metabolic strategies supporting the growth of modern Archaea. IMPORTANCE Given the potential for rational genetic manipulations of biofuel- and biotech-promising archaea to yield transformative results for major markets, it is a priority to define how the metabolisms of such species are controlled, at least in part, by in vivo protein assemblies, and from such, define routes of energy flux that can be most efficiently altered toward biofuel or biotechnological gains. Proteinaceous electron carriers (PECs, such as ferredoxins) offer the potential for specific protein–protein interactions to coordinate selective reductive flow. Employing the model, genetically accessible, hyperthermophilic archaeon, Thermococcus kodakarensis , we establish the metabolic protein interactome of 25 key redox proteins, revealing that each redox active protein has a dynamic partnership profile, suggesting catabolic and anabolic activities may occur in concert and in temporal and spatial proximity in vivo . These results reveal critical importance in evaluating the newly identified partnerships and their role and utility in providing regulated redox flux in T. kodakarensis .

Williams, Sere A. (ORCID:0000000235509590)↗

Exploring the fragmentation efficiency of proteins analyzed by MALDI-TOF-TOF tandem mass spectrometry using computational and statistical analyses

Matrix-assisted laser desorption/ionization time-of-flight-time-of-flight (MALDI-TOF-TOF) tandem mass spectrometry (MS/MS) is a rapid technique for identifying intact proteins from unfractionated mixtures by top-down proteomic analysis. MS/MS allows isolation of specific intact protein ions prior to fragmentation, allowing fragment ion attribution to a specific precursor ion. However, the fragmentation efficiency of mature, intact protein ions by MS/MS post-source decay (PSD) varies widely, and the biochemical and structural factors of the protein that contribute to it are poorly understood. With the advent of protein structure prediction algorithms such as Alphafold2, we have wider access to protein structures for which no crystal structure exists. In this work, we use a statistical approach to explore the properties of bacterial proteins that can affect their gas phase dissociation via PSD. We extract various protein properties from Alphafold2 predictions and analyze their effect on fragmentation efficiency. Our results show that the fragmentation efficiency from cleavage of the polypeptide backbone on the C-terminal side of glutamic acid (E) and asparagine (N) residues were nearly equal. In addition, we found that the rearrangement and cleavage on the C-terminal side of aspartic acid (D) residues that result from the aspartic acid effect (AAE) were higher than for E- and N-residues. From residue interaction network analysis, we identified several local centrality measures and discussed their implications regarding the AAE. We also confirmed the selective cleavage of the backbone at D-proline bonds in proteins and further extend it to N-proline bonds. Finally, we note an enhancement of the AAE mechanism when the residue on the C-terminal side of D-, E- and N-residues is glycine. To the best of our knowledge, this is the first report of this phenomenon. Our study demonstrates the value of using statistical analyses of protein sequences and their predicted structures to better understand the fragmentation of the intact protein ions in the gas phase.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Learning Prediction of Protein Complex Structures

Proteins interact to form protein complex to carry out biological functions such as catalytic chemical reaction. Therefore, it is important to develop computational methods to predict protein-protein interaction and the structures of protein complexes to study and enhance protein function. In this project, we successfully developed several deep learning methods to predict inter-protein contacts and the reinforcement learning and optimization methods to reconstruct protein complex structures from predicted inter-chain contacts. The methods were integrated with the MULTICOM protein complex structure prediction system and applied to predict the complex structures of biomass production-related proteins of green algae. During the two and a half years of research and development, all the specific milestones of the project were achieved successfully. 16 publications/manuscripts were produced. 10 software tools were developed. A patent application was submitted. Our MULTICOM predictors leveraging some tools developed in this project were ranked among the top predictors in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022.

59 BASIC BIOLOGICAL SCIENCES↗

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Structural studies of intrinsically disordered MLL -fusion protein AF9 in complex with peptidomimetic inhibitors

AF9 (MLLT3) and its paralog ENL(MLLT1) are members of the YEATS family of proteins with important role in transcriptional and epigenetic regulatory complexes. These proteins are two common MLL fusion partners in MLL -rearranged leukemias. The oncofusion proteins MLL-AF9/ENL recruit multiple binding partners, including the histone methyltransferase DOT1L, leading to aberrant transcriptional activation and enhancing the expression of a characteristic set of genes that drive leukemogenesis. The interaction between AF9 and DOT1L is mediated by an intrinsically disordered C-terminal ANC1 homology domain (AHD) in AF9, which undergoes folding upon binding of DOT1L and other partner proteins. We have recently reported peptidomimetics that disrupt the recruitment of DOT1L by AF9 and ENL, providing a proof-of-concept for targeting AHD and assessing its druggability. Intrinsically disordered proteins, such as AF9 AHD, are difficult to study and characterize experimentally on a structural level. In this study, we present a successful protein engineering strategy to facilitate structural investigation of the intrinsically disordered AF9 AHD domain in complex with peptidomimetic inhibitors by using maltose binding protein (MBP) as a crystallization chaperone connected with linkers of varying flexibility and length. The strategic incorporation of disulfide bonds provided diffraction-quality crystals of the two disulfide-bridged MBP–AF9 AHD fusion proteins in complex with the peptidomimetics. These successfully determined first series of 2.1–2.6 Å crystal complex structures provide high-resolution insights into the interactions between AHD and its inhibitors, shedding light on the role of AHD in recruiting various binding partner proteins. We show that the overall complex structures closely resemble the reported NMR structure of AF9 AHD/DOT1L with notable difference in the conformation of the β-hairpin region, stabilized through conserved hydrogen bonds network. These first series of AF9 AHD/peptidomimetics complex structures are providing insights of the protein–inhibitor interactions and will facilitate further development of novel inhibitors targeting the AF9/ENL AHD domain.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

SARS-CoV-2 spike protein variant binding affinity to an angiotensin-converting enzyme 2 fusion glycoproteins

Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2), the causative agent of the Coronavirus disease 2019 (Covid-19) pandemic, continues to evolve and circulate globally. Current prophylactic and therapeutic countermeasures against Covid-19 infection include vaccines, small molecule drugs, and neutralizing monoclonal antibodies. SARS-CoV-2 infection is mainly mediated by the viral spike glycoprotein binding to angiotensin converting enzyme 2 (ACE2) on host cells for viral entry. As emerging mutations in the spike protein evade efficacy of spike-targeted countermeasures, a potential strategy to counter SARS-CoV-2 infection is to competitively block the spike protein from binding to the host ACE2 using a soluble recombinant fusion protein that contains a human ACE2 and an IgG1-Fc domain (ACE2-Fc). Here, we have established Chinese Hamster Ovary (CHO) cell lines that stably express ACE2-Fc proteins in which the ACE2 domain either has or has no catalytic activity. The fusion proteins were produced and purified to partially characterize physicochemical properties and spike protein binding. Our results demonstrate the ACE2-Fc fusion proteins are heavily N-glycosylated, sensitive to thermal stress, and actively bind to five spike protein variants (parental, alpha, beta, delta, and omicron) with different affinity. Our data demonstrates a proof-of-concept production strategy for ACE2-Fc fusion glycoproteins that can bind to different spike protein variants to support the manufacture of potential alternative countermeasures for emerging SARS-CoV-2 variants.

60 APPLIED LIFE SCIENCES↗

3D Visualization of Proteins within Metal–Organic Frameworks via Ferritin‐Enabled Electron Microscopy

Abstract Electron tomography holds great promise as a tool for investigating the 3D morphologies and internal structures of metal‐organic framework‐based protein biocomposites (protein@MOFs). Understanding the 3D spatial arrangement of proteins within protein@MOFs is paramount for developing synthetic methods to control their spatial localization and distribution patterns within the biocomposite crystals. In this study, the naturally occurring iron oxide mineral core of the protein horse spleen ferritin (Fn) is leveraged as a contrast agent to directly observe individual proteins once encapsulated into MOFs by electron microscopy techniques. This methodology couples scanning electron microscopy, transmission electron microscopy, and electron tomography to garner detailed 2D and 3D structural interpretations of where proteins spatially lie in Fn@MOF crystals, addressing the significant gaps in understanding how synthetic conditions relate to overall protein spatial localization and aggregation. These findings collectively reveal that adjusting the ligand‐to‐metal ratios, protein concentration, and the use of denaturing agents alters how proteins are arranged, localized, and aggregated within MOF crystals.

Chemistry↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

3D-equivariant graph neural networks for protein model quality assessment

Quality assessment (QA) of predicted protein tertiary structure models plays an important role in ranking and using them. With the recent development of deep learning end-to-end protein structure prediction techniques for generating highly confident tertiary structures for most proteins, it is important to explore corresponding QA strategies to evaluate and select the structural models predicted by them since these models have better quality and different properties than the models predicted by traditional tertiary structure prediction methods. We develop EnQA, a novel graph-based 3D-equivariant neural network method that is equivariant to rotation and translation of 3D objects to estimate the accuracy of protein structural models by leveraging the structural features acquired from the state-of-the-art tertiary structure prediction method—AlphaFold2. We train and test the method on both traditional model datasets (e.g. the datasets of the Critical Assessment of Techniques for Protein Structure Prediction) and a new dataset of high-quality structural models predicted only by AlphaFold2 for the proteins whose experimental structures were released recently. Our approach achieves state-of-the-art performance on protein structural models predicted by both traditional protein structure prediction methods and the latest end-to-end deep learning method—AlphaFold2. It performs even better than the model QA scores provided by AlphaFold2 itself. The results illustrate that the 3D-equivariant graph neural network is a promising approach to the evaluation of protein structural models. Integrating AlphaFold2 features with other complementary sequence and structural features is important for improving protein model QA.

59 BASIC BIOLOGICAL SCIENCES↗

Putting Humpty Dumpty Back Together Again: What Does Protein Quantification Mean in Bottom-Up Proteomics?

Bottom-up proteomics provides peptide measurements and has been invaluable for moving proteomics into large-scale analyses. Commonly, a single quantitative value is reported for each protein-coding gene by aggregating peptide quantities into protein groups following protein inference or parsimony. However, given the complexity of both RNA splicing and post-translational protein modification, it is overly simplistic to assume that all peptides that map to a singular protein-coding gene will demonstrate the same quantitative response. Here, by assuming that all peptides from a protein-coding sequence are representative of the same protein, we may miss the discovery of important biological differences. To capture the contributions of existing proteoforms, we need to reconsider the practice of aggregating protein values to a single quantity per protein-coding gene.

59 BASIC BIOLOGICAL SCIENCES↗

evSeq: Cost-Effective Amplicon Sequencing of Every Variant in a Protein Library

Widespread availability of protein sequence-fitness data would revolutionize both our biochemical understanding of proteins and our ability to engineer them. Unfortunately, even though thousands of protein variants are generated and evaluated for fitness during a typical protein engineering campaign, most are never sequenced, leaving a wealth of potential sequence-fitness information untapped. Primarily, this is because sequencing is unnecessary for many protein engineering strategies; the added cost and effort of sequencing is thus unjustified. It also results from the fact that, even though many lower cost sequencing strategies have been developed, they often require at least some sequencing or computational resources, both of which can be barriers to access. In this work, we present every variant sequencing (evSeq), a method and collection of tools/standardized components for sequencing a variable region within every variant gene produced during a protein engineering campaign at a cost of cents per variant. evSeq was designed to democratize low-cost sequencing for protein engineers and, indeed, anyone interested in engineering biological systems. Execution of its wet-lab component is simple, requires no sequencing experience to perform, relies only on resources and services typically available to biology labs, and slots neatly into existing protein engineering workflows. Analysis of evSeq data is likewise made simple by its accompanying software (found at github.com/fhalab/evSeq, documentation at fhalab.github.io/evSeq), which can be run on a personal laptop and was designed to be accessible to users with no computational experience. Here, low-cost and easy to use, evSeq makes collection of extensive protein variant sequence-fitness data practical.

59 BASIC BIOLOGICAL SCIENCES↗