Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AlphaFold”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Protein folds vs. protein folding: Differing questions, different challenges

We report protein fold prediction using deep-learning artificial intelligence (AI) has transformed the field of protein structure prediction. By combining physical and geometric constraints—and especially patterns extracted from the Protein Data Bank —these machine learning algorithms can predict protein structures at or near atomic resolution and do so in seconds. Today, these computational methods have now solved more than 200 million protein structures, which are accessible from the AlphaFold Protein Structure Database. This accomplishment seems all the more remarkable because few thought it possible or saw it coming. Deservedly, deep-learning AI was named Science magazine’s 2021 “breakthrough of the year”. Clearly, deep-learning AI represents a major advance in protein fold prediction.

54 ENVIRONMENTAL SCIENCES↗

The C2 domain augments Ras GTPase-activating protein catalytic activity

Regulation of Ras GTPases by GTPase-activating proteins (GAPs) is essential for their normal signaling. Nine of the ten GAPs for Ras contain a C2 domain immediately proximal to their canonical GAP domain, and in RasGAP (p120GAP, p120RasGAP;RASA1) mutation of this domain is associated with vascular malformations in humans. Here, we show that the C2 domain of RasGAP is required for full catalytic activity toward Ras. Analyses of the RasGAP C2-GAP crystal structure, AlphaFold models, and sequence conservation reveal direct C2 domain interaction with the Ras allosteric lobe. This is achieved by an evolutionarily conserved surface centered around RasGAP residue R707, point mutation of which impairs the catalytic advantage conferred by the C2 domain in vitro. In mice,R707Cmutation phenocopies the vascular and signaling defects resulting from constitutive disruption of theRASA1gene. In SynGAP, mutation of the equivalent conserved C2 domain surface impairs catalytic activity. Our results indicate that the C2 domain is required to achieve full catalytic activity of GAPs for Ras.

Science & Technology - Other Topics↗

A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers

Abstract Motivation Deep learning has revolutionized protein tertiary structure prediction recently. The cutting-edge deep learning methods such as AlphaFold can predict high-accuracy tertiary structures for most individual protein chains. However, the accuracy of predicting quaternary structures of protein complexes consisting of multiple chains is still relatively low due to lack of advanced deep learning methods in the field. Because interchain residue–residue contacts can be used as distance restraints to guide quaternary structure modeling, here we develop a deep dilated convolutional residual network method (DRCon) to predict interchain residue–residue contacts in homodimers from residue–residue co-evolutionary signals derived from multiple sequence alignments of monomers, intrachain residue–residue contacts of monomers extracted from true/predicted tertiary structures or predicted by deep learning, and other sequence and structural features. Results Tested on three homodimer test datasets (Homo_std dataset, DeepHomo dataset and CASP-CAPRI dataset), the precision of DRCon for top L/5 interchain contact predictions (L: length of monomer in a homodimer) is 43.46%, 47.10% and 33.50% respectively at 6 Å contact threshold, which is substantially better than DeepHomo and DNCON2_inter and similar to Glinter. Moreover, our experiments demonstrate that using predicted tertiary structure or intrachain contacts of monomers in the unbound state as input, DRCon still performs well, even though its accuracy is lower than using true tertiary structures in the bound state are used as input. Finally, our case study shows that good interchain contact predictions can be used to build high-accuracy quaternary structure models of homodimers. Availability and implementation The source code of DRCon is available at https://github.com/jianlin-cheng/DRCon. The datasets are available at https://zenodo.org/record/5998532#.YgF70vXMKsB. Supplementary information Supplementary data are available at Bioinformatics online.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A gated graph transformer for protein complex structure quality assessment and its performance in CASP15

Abstract Motivation Proteins interact to form complexes to carry out essential biological functions. Computational methods such as AlphaFold-multimer have been developed to predict the quaternary structures of protein complexes. An important yet largely unsolved challenge in protein complex structure prediction is to accurately estimate the quality of predicted protein complex structures without any knowledge of the corresponding native structures. Such estimations can then be used to select high-quality predicted complex structures to facilitate biomedical research such as protein function analysis and drug discovery. Results In this work, we introduce a new gated neighborhood-modulating graph transformer to predict the quality of 3D protein complex structures. It incorporates node and edge gates within a graph transformer framework to control information flow during graph message passing. We trained, evaluated and tested the method (called DProQA) on newly-curated protein complex datasets before the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) and then blindly tested it in the 2022 CASP15 experiment. The method was ranked 3rd among the single-model quality assessment methods in CASP15 in terms of the ranking loss of TM-score on 36 complex targets. The rigorous internal and external experiments demonstrate that DProQA is effective in ranking protein complex structures. Availability and implementation The source code, data, and pre-trained models are available at https://github.com/jianlin-cheng/DProQA.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes↗

Structural and biochemical characterization of the mitomycin C repair exonuclease MrfB

Mitomycin C (MMC) repair factor A (mrfA) and factor B (mrfB), encode a conserved helicase and exonuclease that repair DNA damage in the soil-dwelling bacterium Bacillus subtilis. Here we have focused on the characterization of MrfB, a DEDDh exonuclease in the DnaQ superfamily. We solved the structure of the exonuclease core of MrfB to a resolution of 2.1 Å, in what appears to be an inactive state. In this conformation, a predicted α-helix containing the catalytic DEDDh residue Asp172 adopts a random coil, which moves Asp172 away from the active site and results in the occupancy of only one of the two catalytic Mg 2+ ions. We propose that MrfB resides in this inactive state until it interacts with DNA to become activated. By comparing our structure to an AlphaFold prediction as well as other DnaQ-family structures, we located residues hypothesized to be important for exonuclease function. Using exonuclease assays we show that MrfB is a Mg 2+ -dependent 3'–5' DNA exonuclease. We show that Leu113 aids in coordinating the 3' end of the DNA substrate, and that a basic loop is important for substrate binding. This work provides insight into the function of a recently discovered bacterial exonuclease important for the repair of MMC-induced DNA adducts.

59 BASIC BIOLOGICAL SCIENCES↗

Enzymes in 3D: Synthesis, remodelling, and hydrolysis of cell wall (1,3;1,4)-β-glucans

Abstract Recent breakthroughs in structural biology have provided valuable new insights into enzymes involved in plant cell wall metabolism. More specifically, the molecular mechanism of synthesis of (1,3;1,4)-β-glucans, which are widespread in cell walls of commercially important cereals and grasses, has been the topic of debate and intense research activity for decades. However, an inability to purify these integral membrane enzymes or apply transgenic approaches without interpretative problems associated with pleiotropic effects has presented barriers to attempts to define their synthetic mechanisms. Following the demonstration that some members of the CslF sub-family of GT2 family enzymes mediate (1,3;1,4)-β-glucan synthesis, the expression of the corresponding genes in a heterologous system that is free of background complications has now been achieved. Biochemical analyses of the (1,3;1,4)-β-glucan synthesized in vitro, combined with 3-dimensional (3D) cryogenic-electron microscopy and AlphaFold protein structure predictions, have demonstrated how a single CslF6 enzyme, without exogenous primers, can incorporate both (1,3)- and (1,4)-β-linkages into the nascent polysaccharide chain. Similarly, 3D structures of xyloglucan endo-transglycosylases and (1,3;1,4)-β-glucan endo- and exohydrolases have allowed the mechanisms of (1,3;1,4)-β-glucan modification and degradation to be defined. X-ray crystallography and multi-scale modeling of a broad specificity GH3 β-glucan exohydrolase recently revealed a previously unknown and remarkable molecular mechanism with reactant trajectories through which a polysaccharide exohydrolase can act with a processive action pattern. The availability of high-quality protein 3D structural predictions should prove invaluable for defining structures, dynamics, and functions of other enzymes involved in plant cell wall metabolism in the immediate future.

59 BASIC BIOLOGICAL SCIENCES↗

Structure of the Ndc80 complex and its interactions at the yeast kinetochore–microtubule interface

The conserved Ndc80 kinetochore complex, Ndc80c, is the principal link between mitotic spindle microtubules and centromere-associated proteins. We used AlphaFold 2 (AF2) to obtain predictions of the Ndc80 ‘loop’ structure and of the Ndc80 : Nuf2 globular head domains that interact with the Dam1 subunit of the heterodecameric DASH/Dam1 complex (Dam1c). The predictions guided design of crystallizable constructs, with structures close to the predicted ones. The Ndc80 ‘loop’ is a stiff, α-helical ‘switchback’ structure; AF2 predictions and positions of preferential cleavage sites indicate that flexibility within the long Ndc80c rod occurs instead at a hinge closer to the globular head. Conserved stretches of the Dam1 C terminus bind Ndc80c such that phosphorylation of Dam1 serine residues 257, 265 and 292 by the mitotic kinase Ipl1/Aurora B can release this contact during error correction of mis-attached kinetochores. We integrate the structural results presented here into our current molecular model of the kinetochore–microtubule interface. The model illustrates how multiple interactions between Ndc80c, DASH/Dam1c and the microtubule lattice stabilize kinetochore attachments.

59 BASIC BIOLOGICAL SCIENCES↗

The 1.3 Å resolution structure of the truncated group Ia type IV pilin from Pseudomonas aeruginosa strain P1

The type IV pilus is a diverse molecular machine capable of conferring a variety of functions and is produced by a wide range of bacterial species. The ability of the pilus to perform host-cell adherence makes it a viable target for the development of vaccines against infection by human pathogens such as Pseudomonas aeruginosa . Here, the 1.3 Å resolution crystal structure of the N-terminally truncated type IV pilin from P. aeruginosa strain P1 (ΔP1) is reported, the first structure of its phylogenetically linked group (group I) to be discussed in the literature. The structure was solved from X-ray diffraction data that were collected 20 years ago with a molecular-replacement search model generated using AlphaFold ; the effectiveness of other search models was analyzed. Examination of the high-resolution ΔP1 structure revealed a solvent network that aids in maintaining the fold of the protein. On comparing the sequence and structure of P1 with a variety of type IV pilins, it was observed that there are cases of higher structural similarities between the phylogenetic groups of P. aeruginosa than there are between the same phylogenetic group, indicating that a structural grouping of pilins may be necessary in developing antivirulence drugs and vaccines. These analyses also identified the α–β loop as the most structurally diverse domain of the pilins, which could allow it to serve a role in pilus recognition. Studies of ΔP1 in vitro polymerization demonstrate that the optimal hydrophobic catalyst for the oligomerization of the pilus from strain K122 is not conducive for pilus formation of ΔP1; a model of a three-start helical assembly using the ΔP1 structure indicates that the α–β loop and the D-loop prevent in vitro polymerization.

Bragagnolo, Nicholas↗

Elongation Factor Tu Acts as a Chaperone to Activate an Antibacterial RNase Toxin

Many Gram-negative bacterial species use contact-dependent growth inhibition (CDI) systems to deliver toxic proteins into neighboring competitors. CDI + strains deploy CdiA effector proteins, which translocate their C-terminal toxin (CT) domains into target bacteria through a receptor-mediated delivery pathway. To protect against auto-intoxication, CDI + bacteria also produce CdiI immunity proteins that neutralize CT toxin activity. Here, we present the crystal structure of the CT·CdiI O32:H37 complex from Escherichia coli O32:H37. CT O32:H37 adopts the same fold as the tRNase domain of colicin D, and the nucleases share similar catalytic centers. However, unlike colicin D, which cleaves the anticodon loops of tRNA Arg isoacceptors, CT O32:H37 exhibits nonspecific RNase activity. Notably, we find that endogenous elongation factor Tu (EF-Tu) co-purifies with the over-produced CT·CdiI O32:H37 complex. Although EF-Tu does not bind stably to CT O32:H37 in the absence of CdiI O32:H37 , the translation factor is required for toxic RNase activity in vitro. AlphaFold 3 modeling and site-directed mutagenesis indicate that CT O32:H37 interacts with the N-terminal GTPase domain of EF-Tu. EF-Tu appears to stabilize residue Trp52 within the hydrophobic core of the toxin, which in turn supports the RNase active site through an unusual hydrogen-bonding interaction with the catalytic His67 residue. Furthermore, EF-Tu is hijacked as an essential co-factor to organize the toxin's catalytic center.

Nhan, Dinh Quan [University of California, Santa B↗

Increasing thermostability of the key photorespiratory enzyme glycerate 3‐kinase by structure‐based recombination

As global temperatures rise, improving crop yields will require enhancing the thermotolerance of crops. One approach for improving thermotolerance is using bioengineering to increase the thermostability of enzymes catalysing essential biological processes. Photorespiration is an essential recycling process in plants that is integral to photosynthesis and crop growth. The enzymes of photorespiration are targets for enhancing plant thermotolerance as this pathway limits carbon fixation at elevated temperatures. We explored the effects of temperature on the activity of the photorespiratory enzyme glycerate kinase (GLYK) from various organisms and the homologue from the thermophilic alga Cyanidioschyzon merolae was more thermotolerant than those from mesophilic plants, including Arabidopsis thaliana. To understand enzyme features underlying the thermotolerance of C. merolae GLYK (CmGLYK), we performed molecular dynamics simulations using AlphaFold-predicted structures, which revealed greater movement of loop regions of mesophilic plant GLYKs at higher temperatures compared to CmGLYK. Based on these simulations, hybrid proteins were produced and analysed. These hybrid enzymes contained loop regions from CmGLYK replacing the most mobile corresponding loops of AtGLYK. Two of these hybrid enzymes had enhanced thermostability, with melting temperatures increased by 6 °C. One hybrid with three grafted loops maintained higher activity at elevated temperatures. Whilst this hybrid enzyme exhibited enhanced thermostability and a similar Km for ATP compared to AtGLYK, its Km for glycerate increased threefold. This study demonstrates that molecular dynamics simulation-guided structure-based recombination offers a promising strategy for enhancing the thermostability of other plant enzymes with possible application to increasing the thermotolerance of plants under warming climates.

59 BASIC BIOLOGICAL SCIENCES↗

Composition and in situ structure of the Methanospirillum hungatei cell envelope and surface layer

Archaea share genomic similarities with Eukarya and cellular architectural similarities with Bacteria, though archaeal and bacterial surface layers (S-layers) differ. Using cellular cryo–electron tomography, we visualized the S-layer lattice surroundingMethanospirillum hungatei, a methanogenic archaeon. Though more compact than known structures,M. hungatei’s S-layer is a flexible hexagonal lattice of dome-shaped tiles, uniformly spaced from both the overlying cell sheath and the underlying cell membrane. Subtomogram averaging resolved the S-layer hexamer tile at 6.4-angstrom resolution. By fitting an AlphaFold model into hexamer tiles in flat and curved conformations, we uncover intra- and intertile interactions that contribute to the S-layer’s cylindrical and flexible architecture, along with a spacer extension for cell membrane attachment.M. hungateicell’s end plug structure, likely composed of S-layer isoforms, further highlights the uniqueness of this archaeal cell. These structural features offer advantages for methane release and reflect divergent evolutionary adaptations to environmental pressures during early microbial emergence.

Science & Technology - Other Topics↗

Robust deep learning–based protein sequence design using ProteinMPNN

Although deep learning has revolutionized protein structure prediction, almost all experimentally characterized de novo protein designs have been generated using physically based approaches such as Rosetta. Here, we describe a deep learning–based protein sequence design method, ProteinMPNN, that has outstanding performance in both in silico and experimental tests. On native protein backbones, ProteinMPNN has a sequence recovery of 52.4% compared with 32.9% for Rosetta. The amino acid sequence at different positions can be coupled between single or multiple chains, enabling application to a wide range of current protein design challenges. We demonstrate the broad utility and high accuracy of ProteinMPNN using x-ray crystallography, cryo–electron microscopy, and functional studies by rescuing previously failed designs, which were made using Rosetta or AlphaFold, of protein monomers, cyclic homo-oligomers, tetrahedral nanoparticles, and target-binding proteins.

59 BASIC BIOLOGICAL SCIENCES↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

pnnl/PTMPSI

PTM-Psi is a Python 3 package that combines several capabilities to streamline the workflow to interrogate the impact of PTMs on proteins using well-established software packages. The workflow of the PTM-Psi software package includes input files and launch instances from standard packages such as AlphaFold, NWChem, GROMACS, and the Autodock Suite

Mejia-Rodriguez, Daniel↗

Multi-Omics integration can be used to rescue metabolic information for some of the dark region of the Pseudomonas putida proteome

In every omics experiment, genes or their products are identified for which even state of the art tools are unable to assign a function. In the biotechnology chassis organism Pseudomonas putida, these proteins of unknown function make up 14% of the proteome. This missing information can bias analyses since these proteins can carry out functions which impact the engineering of organisms. As a consequence of predicting protein function across all organisms, function prediction tools generally fail to use all of the types of data available for any specific organism, including protein and transcript expression information. Additionally, the release of Alphafold predictions for all Uniprot proteins provides a novel opportunity for leveraging structural information. We constructed a bespoke machine learning model to predict the function of recalcitrant proteins of unknown function in Pseudomonas putida based on these sources of data, which annotated 1079 terms to 213 proteins. Among the predicted functions supplied by the model, we found evidence for a significant overrepresentation of nitrogen metabolism and macromolecule processing proteins. These findings were corroborated by manual analyses of selected proteins which identified, among others, a functionally unannotated operon that likely encodes a branch of the shikimate pathway.

60 APPLIED LIFE SCIENCES↗