Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A single amino acid change led to structural and functional differentiation of PvHd1 to control flowering in switchgrass

Abstract Switchgrass, a forage and bioenergy crop, occurs as two main ecotypes with different but overlapping ranges of adaptation. The two ecotypes differ in a range of characteristics, including flowering time. Flowering time determines the duration of vegetative development and therefore biomass accumulation, a key trait in bioenergy crops. No causal variants for flowering time differences between switchgrass ecotypes have, as yet, been identified. In this study, we mapped a robust flowering time quantitative trait locus (QTL) on chromosome 4K in a biparental F2 population and characterized the flowering-associated transcription factor gene PvHd1, an ortholog of CONSTANS in Arabidopsis and Heading date 1 in rice, as the underlying causal gene. Protein modeling predicted that a serine to glycine substitution at position 35 (p.S35G) in B-Box domain 1 greatly altered the global structure of the PvHd1 protein. The predicted variation in protein compactness was supported in vitro by a 4 °C shift in denaturation temperature. Overexpressing the PvHd1-p.35S allele in a late-flowering CONSTANS-null Arabidopsis mutant rescued earlier flowering, whereas PvHd1-p.35G had a reduced ability to promote flowering, demonstrating that the structural variation led to functional divergence. Our findings provide us with a tool to manipulate the timing of floral transition in switchgrass cultivars and, potentially, expand their cultivation range.

59 BASIC BIOLOGICAL SCIENCES↗

Binding Free Energy Analysis of Colicin D, E3 and E8 to Their Respective Cognate Immunity Proteins Using Computational Simulations

Colicins are antimicrobial proteins produced by bacteria for the purpose of destroying neighboring bacteria. Colicin activity is neutralized by a specific cognate immunity protein in order to protect the host. This study investigates the structural and binding mechanisms underlying the interaction of colicin-D, -E3 and -E8 to their respective immunity proteins (ImD, Im3 and Im8) using structure prediction, molecular dynamics (MD) simulations and MM-PBSA approach of free energy calculations. High-confidence colicin-immunity (Col-Im) complex structures predicted using AlphaFold2 were subjected to MD simulations of 150 ns with GROMACS and were analyzed for the binding free energy calculation using gmx_MMPBSA. Results showed that the complex of Col_E3-Im3 exhibited the most favorable binding free energy, driven by strong van der Waals and electrostatic interactions. Col_D-ImD and Col_E8-Im8 also showed the favorable binding. Electrostatics and hydrogen bonding emerged as a key factor driving binding and stability, while polar solvation acted as a destabilizing factor across all systems. These outcomes provide an understanding of the molecular mechanisms of Col-Im systems, with potential applications for developing natural antimicrobials for food safety.

Biochemistry & Molecular Biology↗

Integrating solvation shell structure in experimentally driven molecular dynamics using x-ray solution scattering data

In the past few decades, prediction of macromolecular structures beyond the native conformation has been aided by the development of molecular dynamics (MD) protocols aimed at exploration of the energetic landscape of proteins. Yet, the computed structures do not always agree with experimental observables, calling for further development of the MD strategies to bring the computations and experiments closer together. Here, we report a scalable, efficient MD simulation approach that incorporates an x-ray solution scattering signal as a driving force for the conformational search of stable structural configurations outside of the native basin. We further demonstrate the importance of inclusion of the hydration layer effect for a precise description of the processes involving large changes in the solvent exposed area, such as unfolding. Utilization of the graphics processing unit allows for an efficient all-atom calculation of scattering patterns on-the-fly, even for large biomolecules, resulting in a speed-up of the calculation of the associated driving force. The utility of the methodology is demonstrated on two model protein systems, the structural transition of lysine-, arginine-, ornithine-binding protein and the folding of deca-alanine. We discuss how the present approach will aid in the interpretation of dynamical scattering experiments on protein folding and association.

Hsu, Darren J.↗

Equivariant Graph Attention Network - 3D Conformers & Feature Fusion

EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.

Shim, Heesung [Lawrence Livermore National Laborat↗

Rapid Computational Identification of Therapeutic Targets for Pathogens

Biological threats continue to persist and evolve as an important challenge to national security. There are multiple ways in which novel viral pathogens could emerge to pose a serious threat to human health. This project developed a pathogen target identification tool that can rapidly respond to a novel or emerging viral biological threat. A set of computational tools were developed that provide detailed information on the newly sequenced genes, their protein products and the drug target sites for the proteins that are best suited for biological countermeasure development. Three key innovations were developed in the project. 1) Development of a new extensive database of protein pocket structures with structure-based search algorithms to rapidly link novel protein targets with the complete collection of previously experimentally solved protein structures. 2) A novel clustering pipeline was introduced to group matching structures and associated small-molecule binding ligands into a consensus protein pocket with the associated small-molecule chemotypes predicted to fit in the pocket site. The matching experimentally solved structures were used to inform the value of different target sites. 3) Where there are viral protein targets with pockets structurally matched to similar human proteins, a biological knowledge graph, which links molecular interactions with human disease, was used to further assess the potential negative impact of a viral protein target with similarities to human proteins that could have important off target side effects. In total, the project produced a new resource for rapid and detailed assessment of promising targets for countermeasures, reflecting the ongoing wet lab, clinical, and computational data being collected. These capabilities will improve the ability to respond to a biological threat in multiple domains.

59 BASIC BIOLOGICAL SCIENCES↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

Rational Design of Lanmodulin Variants for Size-Based Selectivity of Individual Rare Earth Elements

Rare earth elements (REEs) are essential to modern technologies, yet their high physical and chemical similarity makes separation of individual REEs difficult and environmentally taxing. Metalloproteins offer a promising alternative for selective REE binding, as they tend to have high metal ion affinity and specificity. Lanmodulin (LanM), in particular, has arisen as a potential candidate for REE separation as it exhibits picomolar affinity for elements in the REE family. Prior work has shown that the single point mutation D9N can shift LanM’s preference away from lanthanides toward actinides, motivating efforts to tune selectivity of LanM through targeted mutagenesis. Here, we tested the hypothesis that introducing selective aspartic acid to glutamic acid substitutions in the metal coordinating EF hands of LanM would impose steric constraints that would drive LanM affinity away from larger ions, such as La3+, to smaller ions, such as Y3+. To test this hypothesis, a combination of computational and experimental approaches were employed to evaluate the signal mutations LanM D5E and LanM D3E and the double mutants LanM D1ED5E and LanM D3ED9E. Surprisingly, increasing the number of mutations within the metal center did not enhance affinity for smaller REEs, or decrease affinity for larger ions. Only the single point mutation LanM D5E weakened La3+ binding by one order of magnitude relative to LanM wild type (WT), and pairing it with a second mutation to produce LanM D1ED5E drove La3+ affinity to be stronger than that seen for LanM WT. The D3E mutation alone prevented proper expression and folding, but paring it with D9E to produce LanM D3ED9E rescued expression and yielded La3+ affinities comparable to LanM WT. All variants that expressed (LanM D5E, LanM D1ED5E, LanM D3ED9E) displayed Y3+ affinities comparable to LanM WT. Overall, these results highlight the tunability of LanM’s metal-binding environment but also expose current limitations in predicting structural responses to point mutations within a protein sequence. This work establishes a foundation that can be used for refining computational and experimental strategies to engineer metalloproteins with tailored REE selectivity.

Close, Emily [Pacific Northwest National Laborator↗

Hierarchical organization and assembly of the archaeal cell sheath from an amyloid-like protein

Abstract Certain archaeal cells possess external proteinaceous sheath, whose structure and organization are both unknown. By cellular cryogenic electron tomography (cryoET), here we have determined sheath organization of the prototypical archaeon, Methanospirillum hungatei . Fitting of Alphafold-predicted model of the sheath protein (SH) monomer into the 7.9 Å-resolution structure reveals that the sheath cylinder consists of axially stacked β-hoops, each of which is comprised of two to six 400 nm-diameter rings of β-strand arches (β-rings). With both similarities to and differences from amyloid cross-β fibril architecture, each β-ring contains two giant β-sheets contributed by ~ 450 SH monomers that entirely encircle the outer circumference of the cell. Tomograms of immature cells suggest models of sheath biogenesis: oligomerization of SH monomers into β-ring precursors after their membrane-proximal cytoplasmic synthesis, followed by translocation through the unplugged end of a dividing cell, and insertion of nascent β-hoops into the immature sheath cylinder at the junction of two daughter cells.

59 BASIC BIOLOGICAL SCIENCES↗

Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator

Gene activator proteins comprise distinct DNA-binding and transcriptional activation domains (ADs). Because few ADs have been described, we tested domains tiling all yeast transcription factors for activation in vivo and identified 150 ADs. By mRNA display, we showed that 73% of ADs bound the Med15 subunit of Mediator, and that binding strength was correlated with activation. AD-Mediator interaction in vitro was unaffected by a large excess of free activator protein, pointing to a dynamic mechanism of interaction. Structural modeling showed that ADs interact with Med15 without shape complementarity (‘fuzzy’ binding). ADs shared no sequence motifs, but mutagenesis revealed biochemical and structural constraints. Finally, a neural network trained on AD sequences accurately predicted ADs in human proteins and in other yeast proteins, including chromosomal proteins and chromatin remodeling complexes. These findings solve the longstanding enigma of AD structure and function and provide a rationale for their role in biology.

60 APPLIED LIFE SCIENCES↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Putting AlphaFold models to work with phenix.process_predicted_model and ISOLDE

AlphaFold has recently become an important tool in providing models for experimental structure determination by X-ray crystallography and cryo-EM. Large parts of the predicted models typically approach the accuracy of experimentally determined structures, although there are frequently local errors and errors in the relative orientations of domains. Importantly, residues in the model of a protein predicted by AlphaFold are tagged with a predicted local distance difference test score, informing users about which regions of the structure are predicted with less confidence. AlphaFold also produces a predicted aligned error matrix indicating its confidence in the relative positions of each pair of residues in the predicted model. The phenix.process_predicted_model tool downweights or removes low-confidence residues and can break a model into confidently predicted domains in preparation for molecular replacement or cryo-EM docking. These confidence metrics are further used in ISOLDE to weight torsion and atom–atom distance restraints, allowing the complete AlphaFold model to be interactively rearranged to match the docked fragments and reducing the need for the rebuilding of connecting regions.

59 BASIC BIOLOGICAL SCIENCES↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

Human XPG nuclease structure, assembly, and activities with insights for neurodegeneration and cancer from pathogenic mutations

Xeroderma pigmentosum group G (XPG) protein is both a functional partner in multiple DNA damage responses (DDR) and a pathway coordinator and structure-specific endonuclease in nucleotide excision repair (NER). Different mutations in the XPG gene ERCC5 lead to either of two distinct human diseases: Cancer-prone xeroderma pigmentosum (XP-G) or the fatal neurodevelopmental disorder Cockayne syndrome (XP-G/CS). To address the enigmatic structural mechanism for these differing disease phenotypes and for XPG’s role in multiple DDRs, here we determined the crystal structure of human XPG catalytic domain (XPGcat), revealing XPG-specific features for its activities and regulation. Furthermore, XPG DNA binding elements conserved with FEN1 superfamily members enable insights on DNA interactions. Notably, all but one of the known pathogenic point mutations map to XPGcat, and both XP-G and XP-G/CS mutations destabilize XPG and reduce its cellular protein levels. Mapping the distinct mutation classes provides structure-based predictions for disease phenotypes: Residues mutated in XP-G are positioned to reduce local stability and NER activity, whereas residues mutated in XP-G/CS have implied long-range structural defects that would likely disrupt stability of the whole protein, and thus interfere with its functional interactions. Combined data from crystallography, biochemistry, small angle X-ray scattering, and electron microscopy unveil an XPG homodimer that binds, unstacks, and sculpts duplex DNA at internal unpaired regions (bubbles) into strongly bent structures, and suggest how XPG complexes may bind both NER bubble junctions and replication forks. Collective results support XPG scaffolding and DNA sculpting functions in multiple DDR processes to maintain genome stability.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of α $IIb$ $β$ 3 integrin structures along its minimum free energy activation pathway

The adhesion protein integrin is a transmembrane heterodimer that plays a pivotal role in cellular processes such as cell signaling and cell migration. To execute its function, integrin undergoes extensive conformational changes from a bent-closed to an extended-open state. Resolving the structures across these changes remains a challenge with both experimental and computational methods, but it is crucial for understanding the activation mechanism of integrin. We address this challenge for the platelet integrin α IIb β 3 by employing finite temperature string method with structures of the images along the initial guess path generated by a multiscale data-driven framework. The full-length all-atom structures along the resulting minimum free energy path between the inactive bent-closed and active extended-open states of α IIb β 3 integrin are consistent with a variety of experimentally resolved structures. Changes in these predicted structures along the path show that the extension and separation of the α and β subunits from the bent-closed to the extended-open state require correlated movements between the subdomain pairs in α IIb β 3 . Furthermore, these results provide new insights into integrin activation mechanism, and the predicted structures have potential applications in guiding the design of integrin-targeting therapeutics.

Dasetty, Siva [University of Chicago, IL (United S↗

RNA target highlights in CASP15 : Evaluation of predicted models by structure providers

Abstract The first RNA category of the Critical Assessment of Techniques for Structure Prediction competition was only made possible because of the scientists who provided experimental structures to challenge the predictors. In this article, these scientists offer a unique and valuable analysis of both the successes and areas for improvement in the predicted models. All 10 RNA‐only targets yielded predictions topologically similar to experimentally determined structures. For one target, experimentalists were able to phase their x‐ray diffraction data by molecular replacement, showing a potential application of structure predictions for RNA structural biologists. Recommended areas for improvement include: enhancing the accuracy in local interaction predictions and increased consideration of the experimental conditions such as multimerization, structure determination method, and time along folding pathways. The prediction of RNA–protein complexes remains the most significant challenge. Finally, given the intrinsic flexibility of many RNAs, we propose the consideration of ensemble models.

59 BASIC BIOLOGICAL SCIENCES↗

Heterologous expression of a fully active Azotobacter vinelandii nitrogenase Fe protein in Escherichia coli

ABSTRACT The functional versatility of the Fe protein, the reductase component of nitrogenase, makes it an appealing target for heterologous expression, which could facilitate future biotechnological adaptations of nitrogenase-based production of valuable chemical commodities. Yet, the heterologous synthesis of a fully active Fe protein of Azotobacter vinelandii ( Av NifH) in Escherichia coli has proven to be a challenging task. Here, we report the successful synthesis of a fully active Av NifH protein upon co-expression of this protein with Av IscS/U and Av NifM in E. coli . Our metal, activity, electron paramagnetic resonance, and X-ray absorption spectroscopy/extended X-ray absorption fine structure (EXAFS) data demonstrate that the heterologously expressed Av NifH protein has a high [Fe 4 S 4 ] cluster content and is fully functional in nitrogenase catalysis and assembly. Moreover, our phylogenetic analyses and structural predictions suggest that Av NifM could serve as a chaperone and assist the maturation of a cluster-replete Av NifH protein. Given the crucial importance of the Fe protein for the functionality of nitrogenase, this work establishes an effective framework for developing a heterologous expression system of the complete, two-component nitrogenase system; additionally, it provides a useful tool for further exploring the intricate biosynthetic mechanism of this structurally unique and functionally important metalloenzyme. IMPORTANCE The heterologous expression of a fully active Azotobacter vinelandii Fe protein (AvNifH) has never been accomplished. Given the functional importance of this protein in nitrogenase catalysis and assembly, the successful expression of AvNifH in Escherichia coli as reported herein supplies a key element for the further development of heterologous expression systems that explore the catalytic versatility of the Fe protein, either on its own or as a key component of nitrogenase, for nitrogenase-based biotechnological applications in the future. Moreover, the “clean” genetic background of the heterologous expression host allows for an unambiguous assessment of the effect of certain nif-encoded protein factors, such as AvNifM described in this work, in the maturation of AvNifH, highlighting the utility of this heterologous expression system in further advancing our understanding of the complex biosynthetic mechanism of nitrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis of the amidase ClbL central to the biosynthesis of the genotoxin colibactin

Colibactin is a genotoxic natural product produced by select commensal bacteria in the human gut microbiota. The compound is a bis-electrophile that is predicted to form interstrand DNA cross-links in target cells, leading to double-strand DNA breaks. The biosynthesis of colibactin is carried out by a mixed NRPS–PKS assembly line with several noncanonical features. An amidase, ClbL, plays a key role in the pathway, catalyzing the final step in the formation of the pseudodimeric scaffold. ClbL couples α-aminoketone and β-ketothioester intermediates attached to separate carrier domains on the NRPS–PKS assembly. Here, the 1.9 Å resolution structure of ClbL is reported, providing a structural basis for this key step in the colibactin biosynthetic pathway. The structure reveals an open hydrophobic active site surrounded by flexible loops, and comparison with homologous amidases supports its unusual function and predicts macromolecular interactions with pathway carrier-protein substrates. Modeling protein–protein interactions supports a predicted molecular basis for enzyme–carrier domain interactions. Overall, the work provides structural insight into this unique enzyme that is central to the biosynthesis of colibactin.

59 BASIC BIOLOGICAL SCIENCES↗

ppdx : Automated modeling of protein–protein interaction descriptors for use with machine learning

This paper describes ppdx, a python workflow tool that combines protein sequence alignment, homology modeling, and structural refinement, to compute a broad array of descriptors for characterizing protein–protein interactions. The descriptors can be used to predict various properties of interest, such as protein–protein binding affinities, or inhibitory concentrations (IC 50 ), using approaches that range from simple regression to more complex machine learning models. The software is highly modular. It supports different protocols for generating structures, and 95 descriptors can be currently computed. More protocols and descriptors can be easily added. The implementation is highly parallel and can fully exploit the available cores in a single workstation, or multiple nodes on a supercomputer, allowing many systems to be analyzed simultaneously. As an illustrative application, ppdx is used to parametrize a model that predicts the IC 50 of a set of antigens and a class of antibodies directed to the influenza hemagglutinin stalk.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗