Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Predicting the structural basis of targeted protein degradation by integrating molecular dynamics simulations with structural mass spectrometry

Targeted protein degradation (TPD) is a promising approach in drug discovery for degrading proteins implicated in diseases. A key step in this process is the formation of a ternary complex where a heterobifunctional molecule induces proximity of an E3 ligase to a protein of interest (POI), thus facilitating ubiquitin transfer to the POI. In this work, we characterize 3 steps in the TPD process. (1) We simulate the ternary complex formation of SMARCA2 bromodomain and VHL E3 ligase by combining hydrogen-deuterium exchange mass spectrometry with weighted ensemble molecular dynamics (MD). (2) We characterize the conformational heterogeneity of the ternary complex using Hamiltonian replica exchange simulations and small-angle X-ray scattering. (3) We assess the ubiquitination of the POI in the context of the full Cullin-RING Ligase, confirming experimental ubiquitinomics results. Differences in degradation efficiency can be explained by the proximity of lysine residues on the POI relative to ubiquitin.

59 BASIC BIOLOGICAL SCIENCES↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced polymorph metastability drives glycine nucleation in aqueous salt solutions

Crystal nucleation from aqueous solutions influences countless geological, biochemical, astrophysical, environmental, and materials science–related phenomena, including ice formation, the manufacturing of active pharmaceutical ingredients, development of diseases such as Alzheimer’s and the origin of life itself. Understanding and controlling nucleation is essential for designing materials with specific properties, developing strategies to inhibit or promote crystallization in various contexts and preventing pathological aggregation in neurodegenerative diseases. Similar to the protein structure prediction problem—where a single amino acid sequence can in theory adopt one most stable conformation but in practice may sample multiple competing conformations—crystal nucleation faces a parallel challenge: the same chemical species can form diverse polymorphs under different environmental conditions (e.g., temperature, pressure, solvent). Each polymorph presents its own set of physical and chemical properties, highlighting the importance of understanding and controlling polymorph selection in fields ranging from pharmaceuticals to materials design. Despite advances in experimental and computational methods for studying phase transitions and polymorph stability, nucleation remains challenging due to its nanoscale nature. Furthermore, in practical settings, salts and impurities can further influence crystal nucleation in diverse contexts, from scaling in pipelines and desalination plants to the durability of concrete and the efficiency of battery materials. This can lead to the formation of polymorphs that may differ from the most stable phase in pure solutions. Or, even though the final structure might appear same irrespective of whether the environment contained impurities or not, the mechanism through which it was formed might be completely different and not intuitive.

Wang, Ruiyu [University of Maryland, College Park,↗

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity↗

SFold v0.1

This is a scientific software package to integrate Small Angle X-ray Scattering (SAXS) experimental data into OpenFold deep learning models to improve protein structure prediction.

Prince, Stephanie [Lawrence Berkeley National Labo↗

Neurosymbolic Hybrid Approach to Driver Collision Warning

There are two main algorithmic approaches to autonomous driving systems: (1) An end-to-end system in which a single deep neural network learns to map sensory input directly into appropriate warning and driving responses. (2) A mediated hybrid recognition system in which a system is created by combining independent modules that detect each semantic feature. While some researchers believe that deep learning can solve any problem, others believe that a more engineered and symbolic approach is needed to cope with complex environments with less data. Deep learning alone has achieved state-of-the-art results in many areas, from complex gameplay to predicting protein structures. In particular, in image classification and recognition, deep learning models have achieved accuracies as high as humans. But sometimes it can be very difficult to debug if the deep learning model doesn't work. Deep learning models can be vulnerable and are very sensitive to changes in data distribution. Generalization can be problematic. It's usually hard to prove why it works or doesn't. Deep learning models can also be vulnerable to adversarial attacks. Here, we combine deep learning-based object recognition and tracking with an adaptive neurosymbolic network agent, called the Non-Axiomatic Reasoning System (NARS), that can adapt to its environment by building concepts based on perceptual sequences. We achieved an improved intersection-over-union (IOU) object recognition performance of 0.65 in the adaptive retraining model compared to IOU 0.31 in the COCO data pre-trained model. We improved the object detection limits using RADAR sensors in a simulated environment, and demonstrated the weaving car detection capability by combining deep learning-based object detection and tracking with a neurosymbolic model.

Wang, Pei↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Geometrical analysis of Cys-Cys bridges in proteins and their prediction from incomplete structural information

Analysis of C-alpha atom positions from cysteines involved in disulphide bridges in protein crystals shows that their geometric characteristics are unique with respect to other Cys-Cys, non-bridging pairs. They may be used for predicting disulphide connections in incompletely determined protein structures, such as low resolution crystallography or theoretical folding experiments. The basic unit for analysis and prediction is the 3 x 3 distance matrix for Cx positions of residues (i - 1), Cys(i), (i +1) with (j - 1), Cys(j), (j + 1). In each of its columns, row and diagonal vector--outer distances are larger than the central distance. This analysis is compared with some analytical models.

NASA Discipline Exobiology↗

Foldy: An open-source web application for interactive protein structure analysis

Foldy is a cloud-based application that allows non-computational biologists to easily utilize advanced AI-based structural biology tools, including AlphaFold and DiffDock. With many deployment options, it can be employed by individuals, labs, universities, and companies in the cloud without requiring hardware resources, but it can also be configured to utilize locally available computers. Foldy enables scientists to predict the structure of proteins and complexes up to 6000 amino acids with AlphaFold, visualize Pfam annotations, and dock ligands with AutoDock Vina and DiffDock. In our manuscript, we detail Foldy’s interface design, deployment strategies, and optimization for various user scenarios. We demonstrate its application through case studies including rational enzyme design and analyzing proteins with domains of unknown function. Furthermore, we compare Foldy’s interface and management capabilities with other open and closed source tools in the field, illustrating its practicality in managing complex data and computation tasks. Our manuscript underlines the benefits of Foldy as a day-to-day tool for life science researchers, and shows how Foldy can make modern tools more accessible and efficient.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of protein structure and function with RFdiffusion

Abstract There has been considerable recent progress in designing new proteins using deep-learning methods 1–9 . Despite this progress, a general deep-learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher-order symmetric architectures, has yet to be described. Diffusion models 10,11 have had considerable success in image and language generative modelling but limited success when applied to protein modelling, probably due to the complexity of protein backbone geometry and sequence–structure relationships. Here we show that by fine-tuning the RoseTTAFold structure prediction network on protein structure denoising tasks, we obtain a generative model of protein backbones that achieves outstanding performance on unconditional and topology-constrained protein monomer design, protein binder design, symmetric oligomer design, enzyme active site scaffolding and symmetric motif scaffolding for therapeutic and metal-binding protein design. We demonstrate the power and generality of the method, called RoseTTAFold diffusion (RFdiffusion), by experimentally characterizing the structures and functions of hundreds of designed symmetric assemblies, metal-binding proteins and protein binders. The accuracy of RFdiffusion is confirmed by the cryogenic electron microscopy structure of a designed binder in complex with influenza haemagglutinin that is nearly identical to the design model. In a manner analogous to networks that produce images from user-specified inputs, RFdiffusion enables the design of diverse functional proteins from simple molecular specifications.

Science & Technology - Other Topics↗

Nonsynonymous amino acid changes in the α-chain of complement component 5 influence longitudinal susceptibility to Plasmodium falciparum infections and severe malarial anemia in kenyan children

Background: Severe malarial anemia (SMA; Hb < 5.0 g/dl) is a leading cause of childhood morbidity and mortality in holoendemic Plasmodium falciparum transmission regions such as western Kenya. Methods: We investigated the relationship between two novel complement component 5 (C5) missense mutations [rs17216529:C>T, p(Val145Ile) and rs17610:C>T, p(Ser1310Asn)] and longitudinal outcomes of malaria in a cohort of Kenyan children (under 60 mos, n = 1,546). Molecular modeling was used to investigate the impact of the amino acid transitions on the C5 protein structure. Results: Prediction of the wild-type and mutant C5 protein structures did not reveal major changes to the overall structure. However, based on the position of the variants, subtle differences could impact on the stability of C5b. The influence of the C5 genotypes/haplotypes on the number of malaria and SMA episodes over 36 months was determined by Poisson regression modeling. Genotypic analyses revealed that inheritance of the homozygous mutant (TT) for rs17216529:C>T enhanced the risk for both malaria (incidence rate ratio, IRR = 1.144, 95%CI: 1.059–1.236, p = 0.001) and SMA (IRR = 1.627, 95%CI: 1.201–2.204, p = 0.002). In the haplotypic model, carriers of TC had increased risk of malaria (IRR = 1.068, 95%CI: 1.017–1.122, p = 0.009), while carriers of both wild-type alleles (CC) were protected against SMA (IRR = 0.679, 95%CI: 0.542–0.850, p = 0.001). Conclusion: Collectively, these findings show that the selected C5 missense mutations influence the longitudinal risk of malaria and SMA in immune-naïve children exposed to holoendemic P. falciparum transmission through a mechanism that remains to be defined.

60 APPLIED LIFE SCIENCES↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

Predicting receptor-ligand pairing preferences in plant-microbe interfaces via molecular dynamics and machine learning

Microbiome assembly, structure, and dynamics significantly influence plant health. Secreted microbial signaling molecules initiate and mediate symbiosis by binding to structurally compatible plant receptors. For example, lipo-chitooligosaccharides (LCOs), produced by nitrogen-fixing rhizobial bacteria and various fungi, are recognized by plant lysin motif receptor-like kinases (LysM-RLKs), which activate the common symbiotic pathway. Accurately predicting these molecular interactions could reveal complementary signatures underlying the initial stages of endosymbiosis. Despite the breakthrough in protein-ligand structure prediction with deep learning-based tools, such as AlphaFold3, the large size and highly flexible nature of signaling compounds like LCOs present major challenges for detailed structural characterization and binding-affinity prediction. Typical structure-/physics-based methods of ligand virtual screening are designed for small, drug-like molecules, often rely on high-resolution, experimentally determined structures of the protein receptors, and rarely achieve sufficient sampling to obtain converged thermodynamic quantities with large ligands. In this study, we developed a hybrid molecular dynamics/machine learning (MD/ML) approach capable of predicting binding affinity rankings with high accuracy in systems involving large, flexible ligands, despite limited experimental structural information. Using coarse initial structural models, the predictions using the MD/ML workflow achieved strong alignment with experimental trends, particularly in the top-affinity tier for four legume LysM-RLKs (LYR3) binding to LCOs and a chitooligosaccharide. Furthermore, the MD-based conformation selection protocol provided critical structural insights into substrate specificity and binding mechanisms. This study demonstrates a powerful method to screen for challenging cognate ligand-receptors and advance our understanding of the molecular basis of microbial colonization in plants.

Lipo-chitooligosaccharides↗

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗