Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Identification and characterization of proteins of unknown function (PUFs) in Clostridium thermocellum DSM 1313 strains as potential genetic engineering targets

Abstract Background Mass spectrometry-based proteomics can identify and quantify thousands of proteins from individual microbial species, but a significant percentage of these proteins are unannotated and hence classified as proteins of unknown function (PUFs). Due to the difficulty in extracting meaningful metabolic information, PUFs are often overlooked or discarded during data analysis, even though they might be critically important in functional activities, in particular for metabolic engineering research. Results We optimized and employed a pipeline integrating various “guilt-by-association” (GBA) metrics, including differential expression and co-expression analyses of high-throughput mass spectrometry proteome data and phylogenetic coevolution analysis, and sequence homology-based approaches to determine putative functions for PUFs in Clostridium thermocellum . Our various analyses provided putative functional information for over 95% of the PUFs detected by mass spectrometry in a wild-type and/or an engineered strain of C. thermocellum . In particular, we validated a predicted acyltransferase PUF (WP_003519433.1) with functional activity towards 2-phenylethyl alcohol, consistent with our GBA and sequence homology-based predictions. Conclusions This work demonstrates the value of leveraging sequence homology-based annotations with empirical evidence based on the concept of GBA to broadly predict putative functions for PUFs, opening avenues to further interrogation via targeted experiments.

09 BIOMASS FUELS↗

Direct selection of functional fluorescent-protein antibody fusions by yeast display

Antibodies are important reagents for research, diagnostics, and therapeutics. Many examples of chimeric proteins combining the specific target recognition of antibodies with complementing functionalities such as fluorescence, toxicity or enzymatic activity have been described. However, antibodies selected solely on the basis of their binding specificities are not necessarily ideal candidates for the construction of chimeras. Here, we describe a high throughput method based on yeast display to directly select antibodies most suitable for conversion to fluorescent chimera. A library of scFv binders was converted to a fluorescent chimeric form, by cloning thermal green protein into the linker between VH and VL, and directly selecting for both binding and fluorescent functionality. This allowed us to directly identify antibodies functional in the single chain TGP format, that manifest higher protein expression, easier protein purification, and one-step binding assays.

59 BASIC BIOLOGICAL SCIENCES↗

AlphaFold -assisted structure determination of a bacterial protein of unknown function using X-ray and electron crystallography

Macromolecular crystallography generally requires the recovery of missing phase information from diffraction data to reconstruct an electron-density map of the crystallized molecule. Most recent structures have been solved using molecular replacement as a phasing method, requiring an a priori structure that is closely related to the target protein to serve as a search model; when no such search model exists, molecular replacement is not possible. New advances in computational machine-learning methods, however, have resulted in major advances in protein structure predictions from sequence information. Methods that generate predicted structural models of sufficient accuracy provide a powerful approach to molecular replacement. Taking advantage of these advances, AlphaFold predictions were applied to enable structure determination of a bacterial protein of unknown function (UniProtKB Q63NT7, NCBI locus BPSS0212) based on diffraction data that had evaded phasing attempts using MIR and anomalous scattering methods. Using both X-ray and micro-electron (microED) diffraction data, it was possible to solve the structure of the main fragment of the protein using a predicted model of that domain as a starting point. The use of predicted structural models importantly expands the promise of electron diffraction, where structure determination relies critically on molecular replacement.

molecular replacement↗

Design of metal-mediated protein assemblies via hydroxamic acid functionalities

The self-assembly of proteins into sophisticated multicomponent assemblies is a hallmark of all living systems and has spawned extensive efforts in the construction of novel synthetic protein architectures with emergent functional properties. Protein assemblies in nature are formed via selective association of multiple protein surfaces through intricate noncovalent protein-protein interactions, a challenging task to accurately replicate in the de novo design of multiprotein systems. In this protocol, we describe the application of metal-coordinating hydroxamate (HA) motifs to direct the metal-mediated assembly of polyhedral protein architectures and 3D crystalline protein frameworks (protein-MOFs). This strategy has been implemented using an asymmetric cytochrome cb562 monomer through selective, concurrent association of Fe 3+ and Zn 2+ ions to form polyhedral cages. Furthermore, the use of ditopic HA linkers as bridging ligands with metal-binding protein nodes has allowed the construction of crystalline 3D protein-MOF lattices. The protocol is divided into two major sections: (1) the development of a Cys-reactive HA molecule for protein derivatization and self-assembly of protein-HA conjugates into polyhedral cages and (2) the synthesis of ditopic HA bridging ligands for the construction of ferritin-based protein-MOFs using symmetric metal-binding protein nodes. Furthermore, protein cages can be analyzed using analytical ultracentrifugation (AUC), transmission electron microscopy (TEM) and single-crystal X-ray diffraction (sc-XRD) techniques. HA-mediated protein-MOFs are formed in sitting-drop vapor diffusion crystallization trays and are probed via sc-XRD and multi-crystal small-angle X-ray scattering (SAXS) measurements. Ligand synthesis, construction of HA-mediated assemblies, and post-assembly analysis as described in this protocol can be performed by a graduate-level researcher within six weeks.

36 MATERIALS SCIENCE↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide characterization of the soybean DOMAIN OF UNKNOWN FUNCTION 679 membrane protein gene family highlights their potential involvement in growth and stress response

The DMP (DUF679 membrane proteins) family is a plant-specific gene family that encodes membrane proteins. The DMP family genes are suggested to be involved in various programmed cell death processes and gamete fusion during double fertilization in Arabidopsis. However, their functional relevance in other crops remains unknown. This study identified 14 genes from the DMP family in soybean (Glycine max) and characterized their physiochemical properties, subcellular location, gene structure, and promoter regions using bioinformatics tools. Additionally, their tissue-specific and stress-responsive expressions were analyzed using publicly available transcriptome data. Phylogenetic analysis of 198 DMPs from monocots and dicots revealed six clades, with clade-I encoding senescence-related AtDMP1/2 orthologues and clade-II including pollen-specific AtDMP8/9 orthologues. The largest clade, clade-III, predominantly included monocot DMPs, while monocot- and dicot-specific DMPs were assembled in clade-IV and clade-VI, respectively. Evolutionary analysis suggests that soybean GmDMPs underwent purifying selection during evolution. Using 68 transcriptome datasets, expression profiling revealed expression in diverse tissues and distinct responses to abiotic and biotic stresses. The genes Glyma.09G237500 and Glyma.18G098300 showed pistil-abundant expression by qPCR, suggesting they could be potential targets for female organ-mediated haploid induction. Furthermore, cis-acting regulatory elements primarily related to stress-, hormone-, and light-induced pathways regulate GmDMPs, which is consistent with their divergent expression and suggests involvement in growth and stress responses. Overall, our study provides a comprehensive report on the soybean GmDMP family and a framework for further biological functional analysis of DMP genes in soybean or other crops.

59 BASIC BIOLOGICAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗