Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Accelerating crystal structure determination with iterative AlphaFold prediction

Experimental structure determination can be accelerated with artificial intelligence (AI)-based structure-prediction methods such as AlphaFold . Here, an automatic procedure requiring only sequence information and crystallographic data is presented that uses AlphaFold predictions to produce an electron-density map and a structural model. Iterating through cycles of structure prediction is a key element of this procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. This procedure was applied to X-ray data for 215 structures released by the Protein Data Bank in a recent six-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2 Å. Predictions from the iterative template-guided prediction procedure were more accurate than those obtained without templates. It is concluded that AlphaFold predictions obtained based on sequence information alone are usually accurate enough to solve the crystallographic phase problem with molecular replacement, and a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization is suggested.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling receptor flexibility in the structure-based design of KRAS G12C inhibitors

KRAS has long been referred to as an ‘undruggable’ target due to its high affinity for its cognate ligands (GDP and GTP) and its lack of readily exploited allosteric binding pockets. Recent progress in the development of covalent inhibitors of KRAS G12C has revealed that occupancy of an allosteric binding site located between the α3-helix and switch-II loop of KRAS G12C —sometimes referred to as the ‘switch-II pocket’—holds great potential in the design of direct inhibitors of KRAS G12C . In studying diverse switch-II pocket binders during the development of sotorasib (AMG 510), the first FDA-approved inhibitor of KRAS G12C , we found the dramatic conformational flexibility of the switch-II pocket posing significant challenges toward the structure-based design of inhibitors. Here, we present our computational approaches for dealing with receptor flexibility in the prediction of ligand binding pose and binding affinity. For binding pose prediction, we modified the covalent docking program CovDock to allow for protein conformational mobility. This new docking approach, termed as FlexCovDock, improves success rates from 55 to 89% for binding pose prediction on a dataset of 10 cross-docking cases and has been prospectively validated across diverse ligand chemotypes. For binding affinity prediction, we found standard free energy perturbation (FEP) methods could not adequately handle the significant conformational change of the switch-II loop. We developed a new computational strategy to accelerate conformational transitions through the use of targeted protein mutations. Using this methodology, the mean unsigned error (MUE) of binding affinity prediction were reduced from 1.44 to 0.89 kcal/mol on a set of 14 compounds. These approaches were of significant use in facilitating the structure-based design of KRAS G12C inhibitors and are anticipated to be of further use in the design of covalent (and noncovalent) inhibitors of other conformationally labile protein targets.

59 BASIC BIOLOGICAL SCIENCES↗

PigmentHunter: A point-and-click application for automated chlorophyll-protein simulations

Chlorophyll proteins (CPs) are the workhorses of biological photosynthesis, working together to absorb solar energy, transfer it to chemically active reaction centers, and control the charge-separation process that drives its storage as chemical energy. Yet predicting CP optical and electronic properties remains a serious challenge, driven by the computational difficulty of treating large, electronically coupled molecular pigments embedded in a dynamically structured protein environment. To address this challenge, we introduce here an analysis tool called PigmentHunter, which automates the process of preparing CP structures for molecular dynamics (MD), running short MD simulations on the nanoHUB.org science gateway, and then using electrostatic and steric analysis routines to predict optical absorption, fluorescence, and circular dichroism spectra within a Frenkel exciton model. Inter-pigment couplings are evaluated using point-dipole or transition-charge coupling models, while site energies can be estimated using both electrostatic and ring-deformation approaches. The package is built in a Jupyter Notebook environment, with a point-and-click interface that can be used either to manually prepare individual structures or to batch-process many structures at once. Here, we illustrate PigmentHunter’s capabilities with example simulations on spectral line shapes in the light harvesting 2 complex, site energies in the Fenna–Matthews–Olson protein, and ring deformation in photosystems I and II.

14 SOLAR ENERGY↗

Light-modulated abundance of an mRNA encoding a calmodulin-regulated, chromatin-associated NTPase in pea

A CDNA encoding a 47 kDa nucleoside triphosphatase (NTPase) that is associated with the chromatin of pea nuclei has been cloned and sequenced. The translated sequence of the cDNA includes several domains predicted by known biochemical properties of the enzyme, including five motifs characteristic of the ATP-binding domain of many proteins, several potential casein kinase II phosphorylation sites, a helix-turn-helix region characteristic of DNA-binding proteins, and a potential calmodulin-binding domain. The deduced primary structure also includes an N-terminal sequence that is a predicted signal peptide and an internal sequence that could serve as a bipartite-type nuclear localization signal. Both in situ immunocytochemistry of pea plumules and immunoblots of purified cell fractions indicate that most of the immunodetectable NTPase is within the nucleus, a compartment proteins typically reach through nuclear pores rather than through the endoplasmic reticulum pathway. The translated sequence has some similarity to that of human lamin C, but not high enough to account for the earlier observation that IgG against human lamin C binds to the NTPase in immunoblots. Northern blot analysis shows that the NTPase MRNA is strongly expressed in etiolated plumules, but only poorly or not at all in the leaf and stem tissues of light-grown plants. Accumulation of NTPase mRNA in etiolated seedlings is stimulated by brief treatments with both red and far-red light, as is characteristic of very low-fluence phytochrome responses. Southern blotting with pea genomic DNA indicates the NTPase is likely to be encoded by a single gene.

NASA Discipline Number 40-50↗

UniKP: a unified framework for the prediction of enzyme kinetic parameters

Prediction of enzyme kinetic parameters is essential for designing and optimizing enzymes for various biotechnological and industrial applications, but the limited performance of current prediction tools on diverse tasks hinders their practical applications. Here, we introduce UniKP, a unified framework based on pretrained language models for the prediction of enzyme kinetic parameters, including enzyme turnover number (k cat ), Michaelis constant (K m ), and catalytic efficiency (k cat / K m ), from protein sequences and substrate structures. A two-layer framework derived from UniKP (EF-UniKP) has also been proposed to allow robust k cat prediction in considering environmental factors, including pH and temperature. In addition, four representative re-weighting methods are systematically explored to successfully reduce the prediction error in high-value prediction tasks. We have demonstrated the application of UniKP and EF-UniKP in several enzyme discovery and directed evolution tasks, leading to the identification of new enzymes and enzyme mutants with higher activity. UniKP is a valuable tool for deciphering the mechanisms of enzyme kinetics and enables novel insights into enzyme engineering and their industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Biomineralization Potential of Metal-Sequestering Bacteria for Rare Earth Elements (REE) Material Synthesis

Extremophilic methylotrophs, such as Methylotuvimicrobium alcaliphilum 20ZR are known for their natural ability to carry out REE capture and uptake. It has been predicted that proteins associated with surface layers facilitate transport and homeostasis of essential metals. Indeed, electron microscopic analysis coupled with antibody staining revealed that cell envelop of metal-grown M. alcaliphilum 20ZR contains metal-binding proteins in the base of the cup-shaped structures of S-layers. While less explored, the proteins associated with S-layers are also predicted to contribute to scavenging of other minerals essential for core metabolism. Here, we employed bottom-up proteomic study to compare the protein content of S-layer fractions of M. alcaliphilum 20ZR cultures grown in the presence and absence trace metals to identify putative components REE-scavenging machinery.

36 MATERIALS SCIENCE↗

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

Dissecting the structural heterogeneity of proteins by native mass spectrometry

Abstract A single gene yields many forms of proteins via combinations of posttranscriptional/posttranslational modifications. Proteins also fold into higher‐order structures and interact with other molecules. The combined molecular diversity leads to the heterogeneity of proteins that manifests as distinct phenotypes. Structural biology has generated vast amounts of data, effectively enabling accurate structural prediction by computational methods. However, structures are often obtained heterologously under homogeneous states in vitro. The lack of native heterogeneity under cellular context creates challenges in precisely connecting the structural data to phenotypes. Mass spectrometry (MS) based proteomics methods can profile proteome composition of complex biological samples. Most MS methods follow the “bottom‐up” approach, which denatures and digests proteins into short peptide fragments for ease of detection. Coupled with chemical biology approaches, higher‐order structures can be probed via incorporation of covalent labels on native proteins that are maintained at the peptide level. Alternatively, native MS follows the “top‐down” approach and directly analyzes intact proteins under nondenaturing conditions. Various tandem MS activation methods can dissect the intact proteins for in‐depth structural elucidation. Herein, we review recent native MS applications for characterizing heterogeneous samples, including proteins binding to mixtures of ligands, homo/hetero‐complexes with varying stoichiometry, intrinsically disordered proteins with dynamic conformations, glycoprotein complexes with mixed modification states, and active membrane protein complexes in near‐native membrane environments. We summarize the benefits, challenges, and ongoing developments in native MS, with the hope to demonstrate an emerging technology that complements other tools by filling the knowledge gaps in understanding the molecular heterogeneity of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Mercury methylation by metabolically versatile and cosmopolitan marine bacteria

Microbes transform aqueous mercury (Hg) into methylmercury (MeHg), a potent neurotoxin that accumulates in terrestrial and marine food webs, with potential impacts on human health. This process requires the gene pair hgcAB, which encodes for proteins that actuate Hg methylation, and has been well described for anoxic environments. However, recent studies report potential MeHg formation in suboxic seawater, although the microorganisms involved remain poorly understood. In this study, we conducted large-scale multi-omic analyses to search for putative microbial Hg methylators along defined redox gradients in Saanich Inlet, British Columbia, a model natural ecosystem with previously measured Hg and MeHg concentration profiles. Analysis of gene expression profiles along the redoxcline identified several putative Hg methylating microbial groups, including Calditrichaeota, SAR324 and Marinimicrobia, with the last the most active based on hgc transcription levels. Marinimicrobia hgc genes were identified from multiple publicly available marine metagenomes, consistent with a potential key role in marine Hg methylation. Computational homology modelling predicts that Marinimicrobia HgcAB proteins contain the highly conserved amino acid sites and folding structures required for functional Hg methylation. Furthermore, a number of terminal oxidases from aerobic respiratory chains were associated with several putative novel Hg methylators. Our findings thus reveal potential novel marine Hg-methylating microorganisms with a greater oxygen tolerance and broader habitat range than previously recognized.

59 BASIC BIOLOGICAL SCIENCES↗

Predicted structural proteome of Sphagnum divinum and proteome-scale annotation

Sphagnum-dominated peatlands store a substantial amount of terrestrial carbon. The genus is undersampled and under-studied. No experimental crystal structure from any Sphagnum species exists in the Protein Data Bank and fewer than 200 Sphagnum-related genes have structural models available in the AlphaFold Protein Structure Database. Tools and resources are needed to help bridge these gaps, and to enable the analysis of other structural proteomes now made possible by accurate structure prediction. We present the predicted structural proteome (25,134 primary transcripts) of Sphagnum divinum computed using AlphaFold, structural alignment results of all high-confidence models against an annotated nonredundant crystallographic database of over 90,000 structures, a structure-based classification of putative Enzyme Commission (EC) numbers across this proteome, and the computational method to perform this proteome-scale structure-based annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Structural analyses of macromolecules by solution scattering (CRADA Final Report)

New innovative Small Angle X-ray Scattering (SAXS) methods to visualized macromolecules in solution at low resolution has led the SIBYLS group at Lawrence Berkeley Lab in being a world leader in biological SAXS. This research is particularly important now as LBNL establishes high-priority research topics including molecular to mesoscale analysis. This new expanding area of research relies on and ties in with our efforts towards developing a solution structure modeling tool for RNA structure prediction. Due to the pandemic our efforts focused on the SARS-CoV-2 proteins and RNA interactions. Specifically, we characterized SARS-CoC-2 proteins interaction with RNA (Wilanowski et al. 2021 -Hammel group) and Nucleoprotein interaction with RNA (Schneidman group).

59 BASIC BIOLOGICAL SCIENCES↗

A chloroplast protein atlas reveals punctate structures and spatial organization of biosynthetic pathways

Chloroplasts are eukaryotic photosynthetic organelles that drive the global carbon cycle. Despite their importance, our understanding of their protein composition, function, and spatial organization remains limited. Here, we determined the localizations of 1,034 candidate chloroplast proteins using fluorescent protein tagging in the model alga Chlamydomonas reinhardtii. The localizations provide insights into the functions of poorly characterized proteins; identify novel components of nucleoids, plastoglobules, and the pyrenoid; and reveal widespread protein targeting to multiple compartments. We discovered and further characterized cellular organizational features, including eleven chloroplast punctate structures, cytosolic crescent structures, and unexpected spatial distributions of enzymes within the chloroplast. We also used machine learning to predict the localizations of other nuclear-encoded Chlamydomonas proteins. The strains and localization atlas developed here will serve as a resource to accelerate studies of chloroplast architecture and functions.

59 BASIC BIOLOGICAL SCIENCES↗

Genome sequence, phylogenetic analysis, and structure-based annotation reveal metabolic potential of Chlorella sp. SLA-04

Algae are a broad class of photosynthetic eukaryotes that are phylogenetically and physiologically diverse. Most of the phylogenetic diversity has been inferred from 18S rDNA sequencing since there are only a few complete genomes available in public databases. Here we use ultra-long-read Nanopore sequencing to determine a gapless, telomere-to-telomere complete genome sequence of Chlorella sp. SLA-04, previously described as Chlorella sorokiniana SLA-04. Chlorella sp. SLA-04 is a green alga that grows to high cell density in a wide variety of environments - high and neutral pH, high and low alkalinity, and high and low salinity. SLA-04's ability to grow in high pH and high alkalinity media without external CO 2 supply is favorable for large-scale algal biomass production. Phylogenetic analysis performed using ribosomal DNA and conserved protein sequences consistently reveal that Chlorella sp. SLA-04 forms a distinct lineage from other strains of Chlorella sorokiniana. We complement traditional genome annotation methods with high throughput structural predictions and demonstrate that this approach expands functional prediction of the SLA-04 proteome. Genomic analysis of the SLA-04 genome identifies the genes capable of utilizing TCA cycle intermediates to replenish cytosolic acetyl-CoA pools for lipid production. We also identify a complete metabolic pathway for sphingolipid anabolism that may allow SLA-04 to readily adapt to changing environmental conditions and facilitate robust cultivation in mass production systems. Altogether, this work clarifies the phylogeny of Chlorella sp. SLA-04 within Trebouxiophyceae and demonstrates how structural predictions can be used to improve annotation beyond sequencebased methods.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-guided design of a broadly cross-reactive multivalent group a streptococcal vaccine

The M protein of group A streptococci (Strep A) is a major virulence determinant and protective antigen. The N-terminal region of the M protein is variable in sequence, defines the M/emm type, and contains epitopes that elicit opsonic antibodies that protect animals from challenge infections. Although there are >200 M types of Strep A, there is now evidence that structurally related M proteins can be grouped into clusters and that immunity may be cluster-specific in addition to M type-specific. This observation has led to recent studies of structure-based design of multivalent M peptide vaccines to select peptides predicted to cross-react with heterologous M types to improve vaccine coverage. In the current study, we have applied a refined series of peptide structural algorithms to predict immunological cross-reactivity among 117N-terminal M peptides representing the most prevalent M types of Strep A. Based on the results of the structural analyses, in combination with global M type prevalence data, we constructed a 32-valent vaccine containing 19 cross-reactive vaccine candidates predicted to cross-react with 37 heterologous M peptides to which were added 13 type-specific M peptides. Further, the 4-protein recombinant vaccine was immunogenic in rabbits and elicited significant levels of antibodies against 31/32 (97%) vaccine peptides and 28/37 (76%) peptides predicted to cross-react. The vaccine antisera also promoted opsonophagocytic killing of vaccine and cross-reactive M types of Strep A. Based on a recent analysis of M type prevalence of Strep A, the potential global coverage of the 32-valent vaccine is ~90%, ranging from 68% in Africa to 95% in North America. Our results indicate the utility of structure-based design that may be applied to future studies of broadly protective M peptide vaccines.

60 APPLIED LIFE SCIENCES↗

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗