Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Structural- and Functional-Informed Machine Learning for Protein Function Prediction

In this project we aimed to extend methods for protein function prediction to include structural prediction data, and benchmark methods against existing tools. We proposed to apply the method to large metagenome datasets, and develop approaches to examine activity-based protein profiling results for protein function-structure patterns. Nitrogen cycle protein families were previously identified and are used here to provide a proof-of-principle for use of structure prediction in protein function classification.

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of buttressed loops for sculpting protein functions

In natural proteins, structured loops have central roles in molecular recognition, signal transduction and enzyme catalysis. However, because of the intrinsic flexibility and irregularity of loop regions, organizing multiple structured loops at protein functional sites has been very difficult to achieve by de novo protein design. Here we describe a solution to this problem that designs tandem repeat proteins with structured loops (9–14 residues) buttressed by extensive hydrogen bonding interactions. Experimental characterization shows that the designs are monodisperse, highly soluble, folded and thermally stable. Crystal structures are in close agreement with the design models, with the loops structured and buttressed as designed. We demonstrate the functionality afforded by loop buttressing by designing and characterizing binders for extended peptides in which the loops form one side of an extended binding pocket. The ability to design multiple structured loops should contribute generally to efforts to design new protein functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function

Abstract Motivation Millions of protein sequences have been generated by numerous genome and transcriptome sequencing projects. However, experimentally determining the function of the proteins is still a time consuming, low-throughput, and expensive process, leading to a large protein sequence-function gap. Therefore, it is important to develop computational methods to accurately predict protein function to fill the gap. Even though many methods have been developed to use protein sequences as input to predict function, much fewer methods leverage protein structures in protein function prediction because there was lack of accurate protein structures for most proteins until recently. Results We developed TransFun—a method using a transformer-based protein language model and 3D-equivariant graph neural networks to distill information from both protein sequences and structures to predict protein function. It extracts feature embeddings from protein sequences using a pre-trained protein language model (ESM) via transfer learning and combines them with 3D structures of proteins predicted by AlphaFold2 through equivariant graph neural networks. Benchmarked on the CAFA3 test dataset and a new test dataset, TransFun outperforms several state-of-the-art methods, indicating that the language model and 3D-equivariant graph neural networks are effective methods to leverage protein sequences and structures to improve protein function prediction. Combining TransFun predictions and sequence similarity-based predictions can further increase prediction accuracy. Availability and implementation The source code of TransFun is available at https://github.com/jianlin-cheng/TransFun.

59 BASIC BIOLOGICAL SCIENCES↗

Large language models generate functional protein sequences across diverse families

Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here, in this paper, we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

Watching a signaling protein function: What has been learned over four decades of time-resolved studies of photoactive yellow protein

Photoactive yellow protein (PYP) is a signaling protein whose internal p-coumaric acid chromophore undergoes reversible, light-induced trans-to-cis isomerization, which triggers a sequence of structural changes that ultimately lead to a signaling state. Since its discovery nearly 40 years ago, PYP has attracted much interest and has become one of the most extensively studied proteins found in nature. The method of time-resolved crystallography, pioneered by Keith Moffat, has successfully characterized intermediates in the PYP photocycle at near atomic resolution over 12 decades of time down to the sub-picosecond time scale, allowing one to stitch together a movie and literally watch a protein as it functions. But how close to reality is this movie? To address this question, results from numerous complementary time-resolved techniques including x-ray crystallography, x-ray scattering, and spectroscopy are discussed. Emerging from spectroscopic studies is a general consensus that three time constants are required to model the excited state relaxation, with a highly strained ground-state cis intermediate formed in less than 2.4 ps. Persistent strain drives the sequence of structural transitions that ultimately produce the signaling state. Crystal packing forces produce a restoring force that slows somewhat the rates of interconversion between the intermediates. Moreover, the solvent composition surrounding PYP can influence the number and structures of intermediates as well as the rates at which they interconvert. When chloride is present, the PYP photocycle in a crystal closely tracks that in solution, which suggests the epic movie of the PYP photocycle is indeed based in reality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Monomer-scale design of functional protein polymers using consensus repeat sequences

Protein-based polymers possess chemically defined sequences that can encode diverse properties and functions into a new class of biopolymeric materials. However, sequence variation that emerges from evolution can obscure the sequence–function relationships of naturally derived polymers. One strategy to clarify these relationships is to identify common sequences between proteins with similar functions. These conserved sequences often emerge from repeat proteins, and “consensus repeat sequences” provide a convenient platform for systematic investigations of biopolymer sequence–property relationships. In this review, we highlight recent approaches to engineer tunable polymeric materials using monomer-scale design of consensus repeat proteins. Here, we explore established and emerging protein-based materials with mechanical resilience, thermodynamic phase behavior, chemical responsiveness, biomolecular transport, and hierarchical structure. Overall, recent advances in the monomer-scale design of repetitive protein polymers present exciting fundamental and translational opportunities for polymer scientists and engineers.

36 MATERIALS SCIENCE↗

Controlling Mineralization with Protein–Functionalized Peptoid Nanotubes

Sequence-defined foldamers that self-assemble into well-defined architectures are promising scaffolds to template inorganic mineralization. However, it has been challenging to achieve robust control of nucleation and growth without sequence redesign or extensive experimentation. Here, peptoid nanotubes functionalized with a panel of solid-binding proteins are used to mineralize homogeneously distributed and monodisperse anatase nanocrystals from the water-soluble TiBALDH precursor. Crystallite size is systematically tuned between 1.4 and 4.4 nm by changing protein coverage and the identity and valency of the genetically engineered solid-binding segments. The approach is extended to the synthesis of gold nanoparticles and, using a protein encoding both material-binding specificities, to the fabrication of titania/gold nanocomposites capable of photocatalysis under visible-light illumination. Here, beyond uncovering critical roles for hierarchical organization and denticity on solid-binding protein mineralization outcomes, the strategy described herein should prove valuable for the fabrication of hierarchical hybrid materials incorporating a broad range of inorganic components.

36 MATERIALS SCIENCE↗

ER-associated VAP27-1 and VAP27-3 proteins functionally link the lipid-binding ORP2A at the ER-chloroplast contact sites

Abstract The plant endoplasmic reticulum (ER) contacts heterotypic membranes at membrane contact sites (MCSs) through largely undefined mechanisms. For instance, despite the well-established and essential role of the plant ER-chloroplast interactions for lipid biosynthesis, and the reported existence of physical contacts between these organelles, almost nothing is known about the ER-chloroplast MCS identity. Here we show that the Arabidopsis ER membrane-associated VAP27 proteins and the lipid-binding protein ORP2A define a functional complex at the ER-chloroplast MCSs. Specifically, through in vivo and in vitro association assays, we found that VAP27 proteins interact with the outer envelope membrane (OEM) of chloroplasts, where they bind to ORP2A. Through lipidomic analyses, we established that VAP27 proteins and ORP2A directly interact with the chloroplast OEM monogalactosyldiacylglycerol (MGDG), and we demonstrated that the loss of the VAP27-ORP2A complex is accompanied by subtle changes in the acyl composition of MGDG and PG. We also found that ORP2A interacts with phytosterols and established that the loss of the VAP27-ORP2A complex alters sterol levels in chloroplasts. We propose that, by interacting directly with OEM lipids, the VAP27-ORP2A complex defines plant-unique MCSs that bridge ER and chloroplasts and are involved in chloroplast lipid homeostasis.

59 BASIC BIOLOGICAL SCIENCES↗

A core of cell wall proteins functions in wall integrity responses in Arabidopsis thaliana

Abstract Cell walls surround all plant cells, and their composition and structure are tightly regulated to maintain cellular and organismal homeostasis. In response to wall damage, the cell wall integrity (CWI) system is engaged to ameliorate effects on plant growth. Despite the central role CWI plays in plant development, our current understanding of how this system functions at the molecular level is limited. Here, we investigated the transcriptomes of etiolated seedlings of mutants of Arabidopsis thaliana with defects in three major wall polysaccharides, pectin ( quasimodo2 ), cellulose ( cellulose synthase3 je5 ), and xyloglucan ( xyloglucan xylosyltransferase1 and 2 ), to probe whether changes in the expression of cell wall‐related genes occur and are similar or different when specific wall components are reduced or missing. Many changes occurred in the transcriptomes of pectin‐ and cellulose‐deficient plants, but fewer changes occurred in the transcriptomes of xyloglucan‐deficient plants. We hypothesize that this might be because pectins interact with other wall components and/or integrity sensors, whereas cellulose forms a major load‐bearing component of the wall; defects in either appear to trigger the expression of structural proteins to maintain wall cohesion in the absence of a major polysaccharide. This core set of genes functioning in CWI in plants represents an attractive target for future genetic engineering of robust and resilient cell walls.

59 BASIC BIOLOGICAL SCIENCES↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing Structural, Thermal, and Functional Characteristics of Marigold Flower Protein as a Sustainable Food Ingredient

The demand for sustainable and alternative protein sources has been on the rise, driving interest in the valorization of underutilized plants. This study evaluated Calendula officinalis (marigold), a common floral waste, as a sustainable alternative protein source for the food industry. The primary objective of this study was to investigate the physicochemical properties of protein fractions from Calendula officinalis flower to evaluate their potential as a novel protein ingredient. Extraction of the Calendula officinalis flower yielded 92.17% of the crude protein. A sequential extraction of albumin, globulin, glutelin, and prolamin from marigold flower revealed albumin as the dominant fraction (65.47%) and exhibited the highest protein functionality, including water-holding capacity (2.37 g/g), oil-holding capacity (2.49 g/g), and emulsifying capacity (65.22 mL/g). Compared with other protein fractions, glutelin showed a relatively high emulsifying and foaming capacity (EC: 59.13 mL/g; FC: 16.23%). Differential scanning calorimetry revealed high thermal stability for albumin (T p = 105.28 °C) and glutelin (T p = 97.6 °C). Sodium Dodecyl Sulfate–Polyacrylamide Gel Electrophoresis (SDS-PAGE) and Liquid Chromatography–Mass Spectrometry (LC-MS) confirmed the presence of abundant low-molecular-weight polypeptides (<37 kDa), which enhanced emulsification, while scanning electron microscopy revealed porous structures aligned with hydration properties. Antioxidant activity was higher in albumin and glutelin, linked to surface hydrophobicity. LC-MS/MS identified 33 short-chain proteins, including oxidoreductase proteins and lipid-transfer proteins. Findings highlight marigold flower proteins as a sustainable, functional ingredient for a diverse range of food applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Neural networks to learn protein sequence–function relationships from deep mutational scanning data

Understanding the relationship between protein sequence and function is necessary to design new and useful proteins with applications in bioenergy, medicine, and agriculture. The mapping from sequence to function is tremendously complex because it involves thousands of molecular interactions that are coupled over multiple lengths and timescales. We show that neural networks can learn the sequence–function mapping from large protein datasets. Neural networks are appealing for this task because they can learn complicated relationships from data, make few assumptions about the nature of the sequence–function relationship, and can learn general rules that apply across the length of the protein sequence. We demonstrate that learned models can be applied to design new proteins with properties that exceed natural sequences.

59 BASIC BIOLOGICAL SCIENCES↗

DeepComplex: A Web Server of Predicting Protein Complex Structures by Deep Learning Inter-chain Contact Prediction and Distance-Based Modelling

Proteins interact to form complexes. Predicting the quaternary structure of protein complexes is useful for protein function analysis, protein engineering, and drug design. However, few user-friendly tools leveraging the latest deep learning technology for inter-chain contact prediction and the distance-based modelling to predict protein quaternary structures are available. To address this gap, we develop DeepComplex, a web server for predicting structures of dimeric protein complexes. It uses deep learning to predict inter-chain contacts in a homodimer or heterodimer. The predicted contacts are then used to construct a quaternary structure of the dimer by the distance-based modelling, which can be interactively viewed and analysed. The web server is freely accessible and requires no registration. It can be easily used by providing a job name and an email address along with the tertiary structure for one chain of a homodimer or two chains of a heterodimer. The output webpage provides the multiple sequence alignment, predicted inter-chain residue-residue contact map, and predicted quaternary structure of the dimer.

59 BASIC BIOLOGICAL SCIENCES↗