Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Unraveling the Molecular Origin of Prey-Wrapping Spider Silk's Unique Mechanical Properties and Assembly Process Using NMR

Prey wrapping spider silk's unique mechanical properties are investigated confirming the silk's high degree of extensibility and superior toughness compared to other types of spider silk. For the first time, the pre-spinning dope phase is studied in isotope-enriched intact aciniform (AC) silk glands using solution NMR that reveals a combination of α-helical domains linked by disordered random coil chains consistent with previously proposed “beads-on-a-string” models. The model is further refined through the AlphaFold2 protein structure prediction tool. Finally, extensive magic angle spinning (MAS) solid-state (SS) NMR data for isotopically-enriched fibers is used to refine the structural model for AC silk from two species, A. aurantia and A. argentata. The SSNMR data shows that the AC silk fibers are highly α-helical, coiled-coil in structure but, also exhibit significant β-sheet components that can be traced back to the Gly-rich disordered linker regions in the pre-spinning dope phase that are converted to β-sheet structures during fiber formation. This combination of mechanical and structural characterization enhances the understanding of AC silk's liquid-to-solid transition and structure-mechanics relationship. In conclusion, these prey wrap silk results and models will provide the basis for the design of biomimetic materials inspired by the AC spider silk system.

36 MATERIALS SCIENCE↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Universally Accessible Structural Data on Macromolecular Conformation, Assembly, and Dynamics by Small Angle X-Ray Scattering for DNA Repair Insights.

Structures provide a critical breakthrough step for biological analyses, and small angle X-ray scattering (SAXS) is a powerful structural technique to study dynamic DNA repair proteins. As toxic and mutagenic repair intermediates need to be prevented from inadvertently harming the cell, DNA repair proteins often chaperone these intermediates through dynamic conformations, coordinated assemblies, and allosteric regulation. By measuring structural conformations in solution for both proteins, DNA, RNA, and their complexes, SAXS provides insight into initial DNA damage recognition, mechanisms for validation of their substrate, and pathway regulation. Here, we describe exemplary SAXS analyses of a DNA damage response protein spanning from what can be derived directly from the data to obtaining super resolution through the use of SAXS selection of atomic models. We outline strategies and tactics for practical SAXS data collection and analysis. Making these structural experiments in reach of any basic and clinical researchers who have protein, SAXS data can readily be collected at government-funded synchrotrons, typically at no cost for academic researchers. In addition to discussing how SAXS complements and enhances cryo-electron microscopy, X-ray crystallography, NMR, and computational modeling, we furthermore discuss taking advantage of recent advances in protein structure prediction in combination with SAXS analysis.

Chinnam, Naga Babu↗

Characterization of aromatic acid/proton symporters in Pseudomonas putida KT2440 toward efficient microbial conversion of lignin-related aromatics

Pseudomonas putida KT2440 (hereafter KT2440) is a well-studied platform bacterium for the production of industrially valuable chemicals from heterogeneous mixtures of aromatic compounds obtained from lignin depolymerization. KT2440 can grow on lignin-related monomers, such as ferulate (FA), 4-coumarate (4CA), vanillate (VA), 4-hydroxybenzoate (4HBA), and protocatechuate (PCA). Genes associated with their catabolism are known, but knowledge about the uptake systems remains limited. In this work, we studied the KT2440 transporters of lignin-related monomers and their substrate selectivity. Based on the inhibition by protonophores, we focused on five genes encoding aromatic acid/H+ symporter family transporters categorized into major facilitator superfamily that uses the proton motive force. Furthermore, the mutants of PP_1376 (pcaK) and PP_3349 (hcnK) exhibited significantly reduced growth on PCA/4HBA and FA/4CA, respectively, while no change was observed on VA for any of the five gene mutants. At pH 9.0, the conversion of these compounds by hcnK mutant (FA/4CA) and vanK mutant (VA) was dramatically reduced, revealing that these transporters are crucial for the uptake of the anionic substrates at high pH. Uptake assays using 14 C-labeled substrates in Escherichia coli and biosensor-based assays confirmed that PcaK, HcnK, and VanK have ability to take up PCA, FA/4CA, and VA/PCA, respectively. Additionally, analyses of the predicted protein structures suggest that the size and hydropathic properties of the substrate-binding sites of these transporters determine their substrate preferences. Overall, this study reveals that at physiological pH, PcaK and HcnK have a major role in the uptake of PCA/4HBA and FA/4CA, respectively, and VanK is a VA/PCA transporter. This information can contribute to the engineering of strains for the efficient conversion of lignin-related monomers to value-added chemicals.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting the structural basis of targeted protein degradation by integrating molecular dynamics simulations with structural mass spectrometry

Targeted protein degradation (TPD) is a promising approach in drug discovery for degrading proteins implicated in diseases. A key step in this process is the formation of a ternary complex where a heterobifunctional molecule induces proximity of an E3 ligase to a protein of interest (POI), thus facilitating ubiquitin transfer to the POI. In this work, we characterize 3 steps in the TPD process. (1) We simulate the ternary complex formation of SMARCA2 bromodomain and VHL E3 ligase by combining hydrogen-deuterium exchange mass spectrometry with weighted ensemble molecular dynamics (MD). (2) We characterize the conformational heterogeneity of the ternary complex using Hamiltonian replica exchange simulations and small-angle X-ray scattering. (3) We assess the ubiquitination of the POI in the context of the full Cullin-RING Ligase, confirming experimental ubiquitinomics results. Differences in degradation efficiency can be explained by the proximity of lysine residues on the POI relative to ubiquitin.

59 BASIC BIOLOGICAL SCIENCES↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced polymorph metastability drives glycine nucleation in aqueous salt solutions

Crystal nucleation from aqueous solutions influences countless geological, biochemical, astrophysical, environmental, and materials science–related phenomena, including ice formation, the manufacturing of active pharmaceutical ingredients, development of diseases such as Alzheimer’s and the origin of life itself. Understanding and controlling nucleation is essential for designing materials with specific properties, developing strategies to inhibit or promote crystallization in various contexts and preventing pathological aggregation in neurodegenerative diseases. Similar to the protein structure prediction problem—where a single amino acid sequence can in theory adopt one most stable conformation but in practice may sample multiple competing conformations—crystal nucleation faces a parallel challenge: the same chemical species can form diverse polymorphs under different environmental conditions (e.g., temperature, pressure, solvent). Each polymorph presents its own set of physical and chemical properties, highlighting the importance of understanding and controlling polymorph selection in fields ranging from pharmaceuticals to materials design. Despite advances in experimental and computational methods for studying phase transitions and polymorph stability, nucleation remains challenging due to its nanoscale nature. Furthermore, in practical settings, salts and impurities can further influence crystal nucleation in diverse contexts, from scaling in pipelines and desalination plants to the durability of concrete and the efficiency of battery materials. This can lead to the formation of polymorphs that may differ from the most stable phase in pure solutions. Or, even though the final structure might appear same irrespective of whether the environment contained impurities or not, the mechanism through which it was formed might be completely different and not intuitive.

Wang, Ruiyu [University of Maryland, College Park,↗

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity↗

SFold v0.1

This is a scientific software package to integrate Small Angle X-ray Scattering (SAXS) experimental data into OpenFold deep learning models to improve protein structure prediction.

Prince, Stephanie [Lawrence Berkeley National Labo↗

Hierarchical, rotation‐equivariant neural networks to select structural models of protein complexes

Abstract Predicting the structure of multi‐protein complexes is a grand challenge in biochemistry, with major implications for basic science and drug discovery. Computational structure prediction methods generally leverage predefined structural features to distinguish accurate structural models from less accurate ones. This raises the question of whether it is possible to learn characteristics of accurate models directly from atomic coordinates of protein complexes, with no prior assumptions. Here we introduce a machine learning method that learns directly from the 3D positions of all atoms to identify accurate models of protein complexes, without using any precomputed physics‐inspired or statistical terms. Our neural network architecture combines multiple ingredients that together enable end‐to‐end learning from molecular structures containing tens of thousands of atoms: a point‐based representation of atoms, equivariance with respect to rotation and translation, local convolutions, and hierarchical subsampling operations. When used in combination with previously developed scoring functions, our network substantially improves the identification of accurate structural models among a large set of possible models. Our network can also be used to predict the accuracy of a given structural model in absolute terms. The architecture we present is readily applicable to other tasks involving learning on 3D structures of large atomic systems.

Eismann, Stephan↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Foldy: An open-source web application for interactive protein structure analysis

Foldy is a cloud-based application that allows non-computational biologists to easily utilize advanced AI-based structural biology tools, including AlphaFold and DiffDock. With many deployment options, it can be employed by individuals, labs, universities, and companies in the cloud without requiring hardware resources, but it can also be configured to utilize locally available computers. Foldy enables scientists to predict the structure of proteins and complexes up to 6000 amino acids with AlphaFold, visualize Pfam annotations, and dock ligands with AutoDock Vina and DiffDock. In our manuscript, we detail Foldy’s interface design, deployment strategies, and optimization for various user scenarios. We demonstrate its application through case studies including rational enzyme design and analyzing proteins with domains of unknown function. Furthermore, we compare Foldy’s interface and management capabilities with other open and closed source tools in the field, illustrating its practicality in managing complex data and computation tasks. Our manuscript underlines the benefits of Foldy as a day-to-day tool for life science researchers, and shows how Foldy can make modern tools more accessible and efficient.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of protein structure and function with RFdiffusion

Abstract There has been considerable recent progress in designing new proteins using deep-learning methods 1–9 . Despite this progress, a general deep-learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher-order symmetric architectures, has yet to be described. Diffusion models 10,11 have had considerable success in image and language generative modelling but limited success when applied to protein modelling, probably due to the complexity of protein backbone geometry and sequence–structure relationships. Here we show that by fine-tuning the RoseTTAFold structure prediction network on protein structure denoising tasks, we obtain a generative model of protein backbones that achieves outstanding performance on unconditional and topology-constrained protein monomer design, protein binder design, symmetric oligomer design, enzyme active site scaffolding and symmetric motif scaffolding for therapeutic and metal-binding protein design. We demonstrate the power and generality of the method, called RoseTTAFold diffusion (RFdiffusion), by experimentally characterizing the structures and functions of hundreds of designed symmetric assemblies, metal-binding proteins and protein binders. The accuracy of RFdiffusion is confirmed by the cryogenic electron microscopy structure of a designed binder in complex with influenza haemagglutinin that is nearly identical to the design model. In a manner analogous to networks that produce images from user-specified inputs, RFdiffusion enables the design of diverse functional proteins from simple molecular specifications.

Science & Technology - Other Topics↗

Nonsynonymous amino acid changes in the α-chain of complement component 5 influence longitudinal susceptibility to Plasmodium falciparum infections and severe malarial anemia in kenyan children

Background: Severe malarial anemia (SMA; Hb < 5.0 g/dl) is a leading cause of childhood morbidity and mortality in holoendemic Plasmodium falciparum transmission regions such as western Kenya. Methods: We investigated the relationship between two novel complement component 5 (C5) missense mutations [rs17216529:C>T, p(Val145Ile) and rs17610:C>T, p(Ser1310Asn)] and longitudinal outcomes of malaria in a cohort of Kenyan children (under 60 mos, n = 1,546). Molecular modeling was used to investigate the impact of the amino acid transitions on the C5 protein structure. Results: Prediction of the wild-type and mutant C5 protein structures did not reveal major changes to the overall structure. However, based on the position of the variants, subtle differences could impact on the stability of C5b. The influence of the C5 genotypes/haplotypes on the number of malaria and SMA episodes over 36 months was determined by Poisson regression modeling. Genotypic analyses revealed that inheritance of the homozygous mutant (TT) for rs17216529:C>T enhanced the risk for both malaria (incidence rate ratio, IRR = 1.144, 95%CI: 1.059–1.236, p = 0.001) and SMA (IRR = 1.627, 95%CI: 1.201–2.204, p = 0.002). In the haplotypic model, carriers of TC had increased risk of malaria (IRR = 1.068, 95%CI: 1.017–1.122, p = 0.009), while carriers of both wild-type alleles (CC) were protected against SMA (IRR = 0.679, 95%CI: 0.542–0.850, p = 0.001). Conclusion: Collectively, these findings show that the selected C5 missense mutations influence the longitudinal risk of malaria and SMA in immune-naïve children exposed to holoendemic P. falciparum transmission through a mechanism that remains to be defined.

60 APPLIED LIFE SCIENCES↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗