Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Integrated Proteomic and Glycoproteomic Characterization of Human High-Grade Serous Ovarian Carcinoma

Many gene products exhibit great structural heterogeneity because of an array of modifications. These modifications are not directly encoded in the genomic template but often affect the functionality of proteins. Protein glycosylation plays a vital role in proper protein functions. However, the analysis of glycoproteins has been challenging compared with other protein modifications, such as phosphorylation. Here, we perform an integrated proteomic and glycoproteomic analysis of 83 prospectively collected high-grade serous ovarian carcinoma (HGSC) and 23 non-tumor tissues. Integration of the expression data from global proteomics and glycoproteomics reveals tumor-specific glycosylation, uncovers different glycosylation associated with three tumor clusters, and identifies glycosylation enzymes that were correlated with the altered glycosylation. In addition to providing a valuable resource, these results provide insights into the potentialroles of glycosylation in the pathogenesis of HGSC, with the possibility of distinguishing pathological outcomes of ovarian tumors from non-tumors, as well as classifying tumor clusters.

Hu, Yingwei↗

DeepComplex: A Web Server of Predicting Protein Complex Structures by Deep Learning Inter-chain Contact Prediction and Distance-Based Modelling

Proteins interact to form complexes. Predicting the quaternary structure of protein complexes is useful for protein function analysis, protein engineering, and drug design. However, few user-friendly tools leveraging the latest deep learning technology for inter-chain contact prediction and the distance-based modelling to predict protein quaternary structures are available. To address this gap, we develop DeepComplex, a web server for predicting structures of dimeric protein complexes. It uses deep learning to predict inter-chain contacts in a homodimer or heterodimer. The predicted contacts are then used to construct a quaternary structure of the dimer by the distance-based modelling, which can be interactively viewed and analysed. The web server is freely accessible and requires no registration. It can be easily used by providing a job name and an email address along with the tertiary structure for one chain of a homodimer or two chains of a heterodimer. The output webpage provides the multiple sequence alignment, predicted inter-chain residue-residue contact map, and predicted quaternary structure of the dimer.

59 BASIC BIOLOGICAL SCIENCES↗

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

De Novo Design of Proteins That Bind Naphthalenediimides, Powerful Photooxidants with Tunable Photophysical Properties

De novo protein design provides a framework to test our understanding of protein function and build proteins with cofactors and functions not found in nature. Here, we report the design of proteins designed to bind powerful photooxidants and the evaluation of the use of these proteins to generate diffusible small-molecule reactive species. Because excited-state dynamics are influenced by the dynamics and hydration of a photooxidant’s environment, it was important to not only design a binding site but also to evaluate its dynamic properties. Thus, we used computational design in conjunction with molecular dynamics (MD) simulations to design a protein, designated NBP (NDI Binding Protein), that held a naphthalenediimide (NDI), a powerful photooxidant, in a programmable molecular environment. Solution NMR confirmed the structure of the complex. We evaluated two NDI cofactors in this de novo protein using ultrafast pump–probe spectroscopy to evaluate light-triggered intra- and intermolecular electron transfer function. Moreover, we demonstrated the utility of this platform to activate multiple molecular probes for protein labeling.

carbonyls↗

Escherichia coli amino acid auxotrophic expression host strains for investigating protein structure–function relationships

Abstract A set of C43(DE3) and BL21(DE3) Escherichia coli host strains that are auxotrophic for various amino acids is briefly reviewed. These strains require the addition of a defined set of one or more amino acids in the growth medium, and have been specifically designed for overproduction of membrane or water-soluble proteins selectively labelled with stable isotopes, such as 2H, 13C and 15N. The strains described here are available for use and have been deposited into public strain banks. Although they cannot fully eliminate the possibility of isotope dilution and mixing, metabolic scrambling of the different amino acid types can be minimized through a careful consideration of the bacterial metabolic pathways. The use of a suitable auxotrophic expression host strain with an appropriately isotopically labelled growth medium ensures high levels of isotope labelling efficiency as well as selectivity for providing deeper insight into protein structure–function relationships.

Biochemistry & Molecular Biology↗

Mass spectrometry structural analysis of intrinsically disordered phosphoproteins

Phosphorylation is a ubiquitous protein modification that is known to play important roles in many biological phenomena including cell signaling, the opening and closing of membrane protein channels, and even triggering of amyloid protein aggregation. Despite the effects phosphorylation has on protein function, the impact phosphorylation has on the structure of proteins is not well understood. Here, to determine how phosphorylation affects the structure of proteins, top-down mass spectrometry (TD-MS) and ion mobility-mass spectrometry (IM-MS) were performed on various phosphorylated proteins and their dephosphorylated proteoforms. TD-MS with collision- and electron-based fragmentation techniques was utilized to locate phosphorylation sites on the intrinsically disordered amyloid proteins β-casein and α-synuclein. TD-MS also provided evidence that alkaline phosphatase dephosphorylates β-casein from the N-terminus to the C-terminus. Furthermore, IM-MS of common phosphorylated proteins such as β-casein, α-casein, ovalbumin, and phosvitin indicates that phosphorylation promotes compaction of protein structure in denaturing as well as native conditions. Increases in abundance of more compact conformers are also observed when the disease related amyloid protein α-synuclein is phosphorylated at serine 129. We interpret the increased abundance of more compact conformers when proteins are phosphorylated as evidence that salt bridges form between negatively charged phosphates and positively charged residues, which alters protein structure. Salt bridge formation due to phosphorylation could be a mechanism for regulating protein function and be responsible for many of the phenomena observed in nature.

Amyloid proteins↗

Exact reaction coordinates for flap opening in HIV-1 protease

The primary goal of protein science is to understand how proteins function, which requires understanding the functional dynamics responsible for transitions between different functional structures of a protein. A central concept is the exact reaction coordinates that can determine the value of committor for any protein configuration, which provide the optimal description of functional dynamics. Despite intensive efforts, identifying the exact reaction coordinates (RCs) in complex molecules remains a formidable challenge. Using the recently developed generalized work functional, we report the discovery of the exact RCs for an important functional process—the flap opening of HIV-1 protease. Our results show that this process has six RCs, each one is a linear combination of ~240 backbone dihedrals, providing the precise definition of collectivity and cooperativity in the functional dynamics of a protein. Applying bias potentials along each RC can accelerate flap opening by 10 3 to 10 4 folds. The success in identifying the RCs of a protein with 198 residues represents a significant progress beyond that of the alanine dipeptide, currently the only other complex molecule for which the exact RCs for its conformational changes are known. Our results suggest that the generalized work functional (GWF) might be the fundamental operator of mechanics that controls protein dynamics.

Science & Technology - Other Topics↗

Effect of the abolition of intersubunit salt bridges on allosteric protein structural dynamics

A salt bridge, one of the representative structural factors established by non-covalent interactions, plays a crucial role in stabilizing the structure and regulating the protein function, but its role in dynamic processes has been elusive. Here, to scrutinize the structural and functional roles of the salt bridge in the process of performing the protein function, we investigated the effects of salt bridges on the allosteric structural transition of homodimeric hemoglobin (HbI) by applying time-resolved X-ray solution scattering (TRXSS) to the K30D mutant, in which the interfacial salt bridges of the wild type (WT) are abolished. The TRXSS data of K30D are consistent with the kinetic model that requires one monomer intermediate in addition to three structurally distinct dimer intermediates (I 1 , I 2 , and I 3 ) observed in WT and other mutants. The kinetic and structural analyses show that K30D has an accelerated biphasic transition from I 2 to I 3 by more than nine times compared to WT and lacks significant structural changes in the transition from R-like I 2 to T-like I 3 observed in WT, unveiling that the loss of the salt bridges interrupts the R–T allosteric transition of HbI. Besides, the correlation between the bimolecular CO recombination rates in K30D, WT, and other mutants reveals that the bimolecular CO recombination is abnormally decelerated in K30D, indicating that the salt bridges also affect the cooperative ligand binding in HbI. These comparisons of the structural dynamics and kinetics of K30D and WT show that the interfacial salt bridges not only assist the physical connection of two subunits but also play a critical role in the global structural signal transduction of one subunit to the other subunit via a series of well-organized structural transitions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep representation learning improves prediction of LacI-mediated transcriptional repression

Significance The understanding of protein function increases with new experimental and evolutionary datasets. A major challenge is to apply machine learning to these datasets to capture essential features of protein function. Here, we analyze the experimentally determined repression function for tens of thousands of mutants of the LacI protein. This study provides a continuous, noncategorical repression value across a majority of all single mutations and for thousands of higher-order mutations. To develop a top-performing model for the prediction of repression by LacI, we compare several leading variant effect prediction algorithms. A deep representation learning paradigm, first trained across millions of proteins from all known protein families and then fine-tuned using LacI experimental data, offers the highest predictive performance of repression function.

42 ENGINEERING↗

Designed and biologically active protein lattices

Versatile methods to organize proteins in space are required to enable complex biomaterials, engineered biomolecular scaffolds, cell-free biology, and hybrid nanoscale systems. Here, we demonstrate how the tailored encapsulation of proteins in DNA-based voxels can be combined with programmable assembly that directs these voxels into biologically functional protein arrays with prescribed and ordered two-dimensional (2D) and three-dimensional (3D) organizations. We apply the presented concept to ferritin, an iron storage protein, and its iron-free analog, apoferritin, in order to form single-layers, double-layers, as well as several types of 3D protein lattices. Our study demonstrates that internal voxel design and inter-voxel encoding can be effectively employed to create protein lattices with designed organization, as confirmed by in situ X-ray scattering and cryo-electron microscopy 3D imaging. The assembled protein arrays maintain structural stability and biological activity in environments relevant for protein functionality. The framework design of the arrays then allows small molecules to access the ferritins and their iron cores and convert them into apoferritin arrays through the release of iron ions. The presented study introduces a platform approach for creating bio-active protein-containing ordered nanomaterials with desired 2D and 3D organizations.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Local Topological Free Energy of the SARS-CoV-2 Spike Protein

The novel coronavirus SARS-CoV-2 infects human cells using a mechanism that involves binding and structural rearrangement of its Spike protein. Understanding protein rearrangement and identifying specific amino acids where mutations affect protein rearrangement has attracted much attention for drug development. In this manuscript, we use a mathematical method to characterize the local topology/geometry of the SARS-CoV-2 Spike protein backbone. Our results show that local conformational changes in the FP, HR1, and CH domains are associated with global conformational changes in the RBD domain. The SARS-CoV-2 variants analyzed in this manuscript (alpha, beta, gamma, delta Mink, G614, N501) show differences in the local conformations of the FP, HR1, and CH domains as well. Finally, most mutations of concern are either in or in the vicinity of high local topological free energy conformations, suggesting that high local topological free energy conformations could be targets for mutations with significant impact of protein function. Namely, the residues 484, 570, 614, 796, and 969, which are present in variants of concern and are targeted as important in protein function, are predicted as such from our model.

60 APPLIED LIFE SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing chemical bioproduction with rational control of bacterial post-translational modifications

Efficient conversion of inexpensive feedstocks to valuable chemicals by microbes is critical for a robust bioeconomy, but the ability to rationally design bacteria is hampered by insufficient knowledge of how post translational modifications (PTMs) control bacterial protein function and thus bioproduction phenotypes. Our study will focus on the lysine acetylation, a ubiquitous bacterial PTM that can affect the function of enzymes in central metabolism that are often critical for bioproduction processes, disrupt transcriptional regulation, and reduce translation. However, most lysine acetylation data is observational, which means that we do not know when, how, and what specific acetylated residues affect protein function and bacterial physiology. For our model host, we will use a Pseudomonas putida strain that we previously engineered to convert lignocellulosic feedstocks into chemicals such as itaconic acid (ITA). With this strain, we use a dynamic two-stage bioproduction process in which ITA is produced during a non-growth associated production phase. Production is highest during growth stages when lysine acetylation is low in other organisms (early stationary phase) and stalls in conditions where acetylation is highest (late stationary phase). The switch from high to stalled ITA production is also correlated with an unexpected increase in acetate levels – the precursor to non-enzymatic lysine acetylation. As such, we predict that lysine acetylation plays a substantial role in regulating the metabolic pathways required for ITA production. We will develop a generalizable approach that combines high-throughput genetic screens and cutting-edge genome engineering with state-of-the-art proteomics, metabolomics, and genetic code expansion methods to identify and modulate lysine acetylation patterns in bacteria. Ultimately, these strategies aim to manipulate protein expression and acetylation patterns to enhance bioproduction phenotypes (e.g., sustained ITA production in late stationary phase).

60 APPLIED LIFE SCIENCES↗

Integrating multimodal data through interpretable heterogeneous ensembles

Motivation: Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate approaches that rely on a uniform integrated representation reinforce the consensus among the modalities but may lose exclusive local information. The alternative late integration approach that can address this challenge has not been systematically studied for biomedical problems. Results: We propose Ensemble Integration (EI) as a novel systematic implementation of the late integration approach. EI infers local predictive models from the individual data modalities using appropriate algorithms and uses heterogeneous ensemble algorithms to integrate these local models into a global predictive model. We also propose a novel interpretation method for EI models. We tested EI on the problems of predicting protein function from multimodal STRING data and mortality due to coronavirus disease 2019 (COVID-19) from multimodal data in electronic health records. We found that EI accomplished its goal of producing significantly more accurate predictions than each individual modality. It also performed better than several established early integration methods for each of these problems. The interpretation of a representative EI model for COVID-19 mortality prediction identified several disease-relevant features, such as laboratory test (blood urea nitrogen and calcium) and vital sign measurements (minimum oxygen saturation) and demographics (age). These results demonstrated the effectiveness of the EI framework for biomedical data integration and predictive modeling.

59 BASIC BIOLOGICAL SCIENCES↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗