Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Exact reaction coordinates for flap opening in HIV-1 protease

The primary goal of protein science is to understand how proteins function, which requires understanding the functional dynamics responsible for transitions between different functional structures of a protein. A central concept is the exact reaction coordinates that can determine the value of committor for any protein configuration, which provide the optimal description of functional dynamics. Despite intensive efforts, identifying the exact reaction coordinates (RCs) in complex molecules remains a formidable challenge. Using the recently developed generalized work functional, we report the discovery of the exact RCs for an important functional process—the flap opening of HIV-1 protease. Our results show that this process has six RCs, each one is a linear combination of ~240 backbone dihedrals, providing the precise definition of collectivity and cooperativity in the functional dynamics of a protein. Applying bias potentials along each RC can accelerate flap opening by 10 3 to 10 4 folds. The success in identifying the RCs of a protein with 198 residues represents a significant progress beyond that of the alanine dipeptide, currently the only other complex molecule for which the exact RCs for its conformational changes are known. Our results suggest that the generalized work functional (GWF) might be the fundamental operator of mechanics that controls protein dynamics.

Science & Technology - Other Topics↗

Effect of the abolition of intersubunit salt bridges on allosteric protein structural dynamics

A salt bridge, one of the representative structural factors established by non-covalent interactions, plays a crucial role in stabilizing the structure and regulating the protein function, but its role in dynamic processes has been elusive. Here, to scrutinize the structural and functional roles of the salt bridge in the process of performing the protein function, we investigated the effects of salt bridges on the allosteric structural transition of homodimeric hemoglobin (HbI) by applying time-resolved X-ray solution scattering (TRXSS) to the K30D mutant, in which the interfacial salt bridges of the wild type (WT) are abolished. The TRXSS data of K30D are consistent with the kinetic model that requires one monomer intermediate in addition to three structurally distinct dimer intermediates (I 1 , I 2 , and I 3 ) observed in WT and other mutants. The kinetic and structural analyses show that K30D has an accelerated biphasic transition from I 2 to I 3 by more than nine times compared to WT and lacks significant structural changes in the transition from R-like I 2 to T-like I 3 observed in WT, unveiling that the loss of the salt bridges interrupts the R–T allosteric transition of HbI. Besides, the correlation between the bimolecular CO recombination rates in K30D, WT, and other mutants reveals that the bimolecular CO recombination is abnormally decelerated in K30D, indicating that the salt bridges also affect the cooperative ligand binding in HbI. These comparisons of the structural dynamics and kinetics of K30D and WT show that the interfacial salt bridges not only assist the physical connection of two subunits but also play a critical role in the global structural signal transduction of one subunit to the other subunit via a series of well-organized structural transitions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep representation learning improves prediction of LacI-mediated transcriptional repression

Significance The understanding of protein function increases with new experimental and evolutionary datasets. A major challenge is to apply machine learning to these datasets to capture essential features of protein function. Here, we analyze the experimentally determined repression function for tens of thousands of mutants of the LacI protein. This study provides a continuous, noncategorical repression value across a majority of all single mutations and for thousands of higher-order mutations. To develop a top-performing model for the prediction of repression by LacI, we compare several leading variant effect prediction algorithms. A deep representation learning paradigm, first trained across millions of proteins from all known protein families and then fine-tuned using LacI experimental data, offers the highest predictive performance of repression function.

42 ENGINEERING↗

Designed and biologically active protein lattices

Versatile methods to organize proteins in space are required to enable complex biomaterials, engineered biomolecular scaffolds, cell-free biology, and hybrid nanoscale systems. Here, we demonstrate how the tailored encapsulation of proteins in DNA-based voxels can be combined with programmable assembly that directs these voxels into biologically functional protein arrays with prescribed and ordered two-dimensional (2D) and three-dimensional (3D) organizations. We apply the presented concept to ferritin, an iron storage protein, and its iron-free analog, apoferritin, in order to form single-layers, double-layers, as well as several types of 3D protein lattices. Our study demonstrates that internal voxel design and inter-voxel encoding can be effectively employed to create protein lattices with designed organization, as confirmed by in situ X-ray scattering and cryo-electron microscopy 3D imaging. The assembled protein arrays maintain structural stability and biological activity in environments relevant for protein functionality. The framework design of the arrays then allows small molecules to access the ferritins and their iron cores and convert them into apoferritin arrays through the release of iron ions. The presented study introduces a platform approach for creating bio-active protein-containing ordered nanomaterials with desired 2D and 3D organizations.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Local Topological Free Energy of the SARS-CoV-2 Spike Protein

The novel coronavirus SARS-CoV-2 infects human cells using a mechanism that involves binding and structural rearrangement of its Spike protein. Understanding protein rearrangement and identifying specific amino acids where mutations affect protein rearrangement has attracted much attention for drug development. In this manuscript, we use a mathematical method to characterize the local topology/geometry of the SARS-CoV-2 Spike protein backbone. Our results show that local conformational changes in the FP, HR1, and CH domains are associated with global conformational changes in the RBD domain. The SARS-CoV-2 variants analyzed in this manuscript (alpha, beta, gamma, delta Mink, G614, N501) show differences in the local conformations of the FP, HR1, and CH domains as well. Finally, most mutations of concern are either in or in the vicinity of high local topological free energy conformations, suggesting that high local topological free energy conformations could be targets for mutations with significant impact of protein function. Namely, the residues 484, 570, 614, 796, and 969, which are present in variants of concern and are targeted as important in protein function, are predicted as such from our model.

60 APPLIED LIFE SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Gene Fusion: A Genome Wide Survey

As a well known fact, organisms form larger and complex multimodular (composite or chimeric) and mostly multi-functional proteins through gene fusion of two or more individual genes which have independent evolution histories and functions. We call each of these components a module. The existence of multimodular proteins may improves the efficiency in gene regulation and in cellular functions, and thus may give the host organism advantages in adaptation to environments. Analysis of all gene fusions in present-day organisms should allow us to examine the patterns of gene fusion in context with cellular functions, to trace back the evolution processes from the ancient smaller and uni-functional proteins to the present-day larger and complex multi-functional proteins, and to estimate the minimal number of ancestor proteins that existed in the last common ancestor for all life on earth. Although many multimodular proteins have been experimentally known, identification of gene fusion events systematically at genome scale had not been possible until recently when large number of completed genome sequences have been becoming available. In addition, technical difficulties for such analysis also exist due to the complexity of this biological and evolutionary process. We report from this study a new strategy to computationally identify multimodular proteins using completed genome sequences and the results surveyed from 22 organisms with the data from over 40 organisms to be presented during the meeting. Additional information is contained in the original extended abstract.

Liang, Ping↗

Enhancing chemical bioproduction with rational control of bacterial post-translational modifications

Efficient conversion of inexpensive feedstocks to valuable chemicals by microbes is critical for a robust bioeconomy, but the ability to rationally design bacteria is hampered by insufficient knowledge of how post translational modifications (PTMs) control bacterial protein function and thus bioproduction phenotypes. Our study will focus on the lysine acetylation, a ubiquitous bacterial PTM that can affect the function of enzymes in central metabolism that are often critical for bioproduction processes, disrupt transcriptional regulation, and reduce translation. However, most lysine acetylation data is observational, which means that we do not know when, how, and what specific acetylated residues affect protein function and bacterial physiology. For our model host, we will use a Pseudomonas putida strain that we previously engineered to convert lignocellulosic feedstocks into chemicals such as itaconic acid (ITA). With this strain, we use a dynamic two-stage bioproduction process in which ITA is produced during a non-growth associated production phase. Production is highest during growth stages when lysine acetylation is low in other organisms (early stationary phase) and stalls in conditions where acetylation is highest (late stationary phase). The switch from high to stalled ITA production is also correlated with an unexpected increase in acetate levels – the precursor to non-enzymatic lysine acetylation. As such, we predict that lysine acetylation plays a substantial role in regulating the metabolic pathways required for ITA production. We will develop a generalizable approach that combines high-throughput genetic screens and cutting-edge genome engineering with state-of-the-art proteomics, metabolomics, and genetic code expansion methods to identify and modulate lysine acetylation patterns in bacteria. Ultimately, these strategies aim to manipulate protein expression and acetylation patterns to enhance bioproduction phenotypes (e.g., sustained ITA production in late stationary phase).

60 APPLIED LIFE SCIENCES↗

DNA Recombinase Proteins, their Function and Structure in the Active Form, a Computational Study

Homologous recombination is a crucial sequence of reactions in all cells for the repair of double strand DNA (dsDNA) breaks. While it was traditionally considered as a means for generating genetic diversity, it is now known to be essential for restart of collapsed replication forks that have met a lesion on the DNA template (Cox et al., 2000). The central stage of this process requires the presence of the DNA recombinase protein, RecA in bacteria, RadA in archaea, or Rad51 in eukaryotes, which leads to an ATP-mediated DNA strand-exchange process. Despite many years of intense study, some aspects of the biochemical mechanism, and structure of the active form of recombinase proteins are not well understood. Our theoretical study is an attempt to shed light on the main structural and mechanistic issues encountered on the RecA of the e-coli, the RecA of the extremely radio resistant Deinococcus Radiodurans (promoting an inverse DNA strand-exchange repair), and the homolog human Rad51. The conformational changes are analyzed for the naked enzymes, and when they are linked to ATP and ADP. The average structures are determined over 2ns time scale of Langevian dynamics using a collision frequency of 1.0 ps(sup -1). The systems are inserted in an octahedron periodic box with a 10 Angstrom buffer of water molecules explicitly described by the TIP3P model. The corresponding binding free energies are calculated in an implicit solvent using the Poisson-Boltzmann solvent accessible surface area, MM-PBSA model. The role of the ATP is not only in stabilizing the interaction RecA-DNA, but its hydrolysis is required to allow the DNA strand-exchange to proceed. Furthermore, we extended our study, using the hybrid QM/MM method, on the mechanism of this chemical process. All the calculations were performed using the commercial code Amber 9.

Carra, Claudio↗

Integrating multimodal data through interpretable heterogeneous ensembles

Motivation: Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate approaches that rely on a uniform integrated representation reinforce the consensus among the modalities but may lose exclusive local information. The alternative late integration approach that can address this challenge has not been systematically studied for biomedical problems. Results: We propose Ensemble Integration (EI) as a novel systematic implementation of the late integration approach. EI infers local predictive models from the individual data modalities using appropriate algorithms and uses heterogeneous ensemble algorithms to integrate these local models into a global predictive model. We also propose a novel interpretation method for EI models. We tested EI on the problems of predicting protein function from multimodal STRING data and mortality due to coronavirus disease 2019 (COVID-19) from multimodal data in electronic health records. We found that EI accomplished its goal of producing significantly more accurate predictions than each individual modality. It also performed better than several established early integration methods for each of these problems. The interpretation of a representative EI model for COVID-19 mortality prediction identified several disease-relevant features, such as laboratory test (blood urea nitrogen and calcium) and vital sign measurements (minimum oxygen saturation) and demographics (age). These results demonstrated the effectiveness of the EI framework for biomedical data integration and predictive modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting functional divergence in protein evolution by site-specific rate shifts

Most modern tools that analyze protein evolution allow individual sites to mutate at constant rates over the history of the protein family. However, Walter Fitch observed in the 1970s that, if a protein changes its function, the mutability of individual sites might also change. This observation is captured in the "non-homogeneous gamma model", which extracts functional information from gene families by examining the different rates at which individual sites evolve. This model has recently been coupled with structural and molecular biology to identify sites that are likely to be involved in changing function within the gene family. Applying this to multiple gene families highlights the widespread divergence of functional behavior among proteins to generate paralogs and orthologs.

Review↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Biosensor and optogenetics for systems biology of yeast branched-chain alcohol production and tolerance

In this project we combined synthetic biology, systems biology, protein engineering and metabolic engineering to study, control, and improve the production of branched chain alcohols (BCAs), a class of advanced biofuels preferred by the DOE, in the yeast Saccharomyces cerevisiae. This involved the development and application of optogenetic systems as a new modality of dynamic control of native and engineered metabolic pathways, using light as inducible or repressible agent. The optogenetic systems include gene circuits for light control of gene expression, as well as light-assembled synthetic organelles and photo-switchable protein binders to control metabolic and protein function with light at the protein level. In addition, we developed the first genetically encoded biosensor for BCA production in yeast, which we used to design high throughput assays to identify highly productive strains, pathways, and enzymes. We also showed that this biosensor can be functionally co-expressed with optogenetic circuits in the same strain, raising the possibility of establishing, for the first time, computer-interfaced closed-loop controls of engineered metabolic pathways. The yeast gene deletion library was also utilized to conduct the first systems-level study on BCA toxicity in yeast, which uncovered key fundamental principles of yeast sensitivity and tolerance to these alcohols, allowing us to design highly tolerant strains with increased BCA production. This project, thus comprises the development of several new technologies, which we integrated to make new discoveries on the dynamics of BCA production and their mechanisms of cellular toxicity and tolerance, as well as to establish new paradigms to engineer and control metabolic pathways and microbial fermentations with light, for the production of BCAs and other products of interest to the DOE.

2-metyl-1-butanol↗