Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AlphaFold”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Selective deuteration of an RNA:RNA complex for structural analysis using small-angle scattering

The structures of RNA:RNA complexes regulate many biological processes. Despite their importance, protein-free RNA:RNA complexes represent a tiny fraction of experimentally determined structures. Here, we describe a joint small-angle X-ray and neutron scattering (SAXS/SANS) approach to structurally interrogate conformational changes in a model RNA:RNA complex. Using SAXS, we measured the solution structures of the individual RNAs and of the overall RNA:RNA complex. With SANS, we demonstrate, as a proof of principle, that isotope labeling and contrast matching (CM) can be combined to probe the bound state structure of an RNA within a selectively deuterated RNA:RNA complex. Furthermore, we show that experimental scattering data can validate and improve predicted AlphaFold 3 RNA:RNA complex structures to reflect its solution structure. In conclusion, our work demonstrates that in silico modeling, SAXS, and CM-SANS can be used in concert to directly analyze conformational changes within RNAs when in complex, enhancing our understanding of RNA structure in functional assemblies.

HIV-1 dimerization initiation site↗

Phosphoserine Charge State Drives Ion Condensation and Spatial Polyamine Presentation in Multirepeat Silaffin

Diatom silaffins direct silica biomineralization through heavily post-translationally modified repeat domains, yet how these modifications reshape the multirepeat conformational ensemble remains unknown. We report all-atom MD simulations of a 195-residue construct spanning repeats R1–R7 of Sil1p from Cylindrotheca fusiformis, carrying the full complement of native PTMs: phosphoserine (pSer), long-chain polyamines (LCPA), dimethyllysine (MLY), and trimethylhydroxylysine phosphate (TPL). We simulated three variants (Native, P1/singly deprotonated phosphate, and P2/doubly deprotonated phosphate) at two NaCl concentrations in triplicate for 500 ns each. All systems disorder from the AlphaFold 3 starting structure. Phosphate charge state, not ionic strength, is the dominant control of ensemble compaction and ion organization. Doubly deprotonated phosphate organizes an extensive Na + condensation shell (∼100 ions, 20% of box Na + ) and a heterogeneous bridging network that integrates both pSer and TPL phosphate groups. The resulting ensemble is compact with LCPA side chains exhibiting above-median solvent accessibility in 84% of simulation frames in P2 at 300 mM. This is higher than any other condition we simulated and indicates that polyamine groups are preferentially surface-presented in the most compact, ion-organized state. A charge-neutralization control confirms that this compact state is a structured intermediate maintained by the bridging network, not a simple collapsed globule. Here, this repeat-scale spatial organization is not captured by single-repeat peptide studies. Understanding the dynamics, mechanism, and spatial organization of PTM-rich silaffin at the repeat scale is a step closer to hierarchical biomimetic materials beyond simple silica morphologies.

Amines↗

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES↗

Thermophilic Chassis-Enabled High-Throughput Selection of a Thermostable Fluorogenic Reporter

Thermostable proteins show increased shelf life and performance at elevated temperatures and under harsh conditions, resulting in lower costs for various industrial and biotechnological applications. However, due to a limited understanding of the relationship between stability and function, protein stabilization remains primarily a trial-and-error approach. Therefore, building a combinatorial library of mutations predicted to improve stability, followed by experimental testing, represents a markedly improved methodology. However, the lack of high-throughput approaches to screen even a moderately sized library presents a major bottleneck in the field. Here, in this study, we use a thermophile, Parageobacillus thermoglucosidasius (Ptherm) to rapidly screen combinatorial libraries consisting of rationally designed thermostabilizing mutations (∼10 3 –10 4 ) of a mesophilic fluorescent reporter, Y-FAST. On a Petri dish, microbial growth at an elevated temperature and exposure to fluorogen yielded several colonies of Ptherm that showed distinct fluorescence at 55 and 68 °C in our two sequentially generated libraries using Rosetta and ProteinMPNN, respectively. The Y-FAST variants isolated from fluorescent colonies were brighter than Y-FAST and showed higher resistance to thermal and chemical denaturation. AlphaFold-predicted structures and MD simulations revealed stability-enhancing salt bridges and hydrogen bond networks in the isolated FAST variants. The moderately thermostable FAST (tsFAST) and hyperstable FAST (hsFAST) were then demonstrated as translation reporters for protein expression and folding at elevated temperatures, such as 55 and 68 °C. Our approach of combinatorial library generation and high-throughput screening in a thermophilic chassis could, in principle, be extended to other proteins fused to these translation reporters. Furthermore, the hsFAST protein is small─half the size of the green fluorescent protein─and does not require oxygen for maturation, making it ideal for engineering extremophilic anaerobes for biosensing and bioconversion.

59 BASIC BIOLOGICAL SCIENCES↗

Cryo-EM Structure of the Mnx Protein Complex Reveals a Tunnel Framework for the Mechanism of Manganese Biomineralization

The global manganese cycle relies on microbes to oxidize soluble Mn(II) to insoluble Mn(IV) oxides. Some microbes require peroxide or superoxide as oxidants, but others can use O 2 directly, via multicopper oxidase (MCO) enzymes. One of these, MnxG from Bacillus sp. strain PL-12, was isolated in tight association with small accessory proteins, MnxE and MnxF. The protein complex, called Mnx, has eluded crystallization efforts, but we now report the 3D structure of a point mutant using cryo-EM single particle analysis, cross-linking mass spectrometry, and AlphaFold Multimer prediction. The ß-sheet–rich complex features MnxG enzyme, capped by a heterohexameric ring of alternating MnxE and MnxF subunits, and a tunnel that runs through MnxG and its MnxE 3 F 3 cap. The tunnel dimensions and charges can accommodate the mechanistically inferred binuclear manganese intermediates. Furthermore, comparison with the Fe(II)-oxidizing MCO, ceruloplasmin, identifies likely coordinating groups for the Mn(II) substrate, at the entrance to the tunnel. Thus, the 3D structure provides a rationale for the established manganese oxidase mechanism, and a platform for further experiments to elucidate mechanistic details of manganese biomineralization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AF2Complex predicts direct physical interactions in multimeric proteins with deep learning

Abstract Accurate descriptions of protein-protein interactions are essential for understanding biological systems. Remarkably accurate atomic structures have been recently computed for individual proteins by AlphaFold2 (AF2). Here, we demonstrate that the same neural network models from AF2 developed for single protein sequences can be adapted to predict the structures of multimeric protein complexes without retraining. In contrast to common approaches, our method, AF2Complex, does not require paired multiple sequence alignments. It achieves higher accuracy than some complex protein-protein docking strategies and provides a significant improvement over AF-Multimer, a development of AlphaFold for multimeric proteins. Moreover, we introduce metrics for predicting direct protein-protein interactions between arbitrary protein pairs and validate AF2Complex on some challenging benchmark sets and the E. coli proteome. Lastly, using the cytochrome c biogenesis system I as an example, we present high-confidence models of three sought-after assemblies formed by eight members of this system.

59 BASIC BIOLOGICAL SCIENCES↗

Structural characterization of a soil viral auxiliary metabolic gene product – a functional chitosanase

Metagenomics is unearthing the previously hidden world of soil viruses. Many soil viral sequences in metagenomes contain putative auxiliary metabolic genes (AMGs) that are not associated with viral replication. Here, we establish that AMGs on soil viruses actually produce functional, active proteins. We focus on AMGs that potentially encode chitosanase enzymes that metabolize chitin – a common carbon polymer. We express and functionally screen several chitosanase genes identified from environmental metagenomes. One expressed protein showing endo-chitosanase activity (V-Csn) is crystalized and structurally characterized at ultra-high resolution, thus representing the structure of a soil viral AMG product. This structure provides details about the active site, and together with structure models determined using AlphaFold, facilitates understanding of substrate specificity and enzyme mechanism. Our findings support the hypothesis that soil viruses contribute auxiliary functions to their hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical organization and assembly of the archaeal cell sheath from an amyloid-like protein

Abstract Certain archaeal cells possess external proteinaceous sheath, whose structure and organization are both unknown. By cellular cryogenic electron tomography (cryoET), here we have determined sheath organization of the prototypical archaeon, Methanospirillum hungatei . Fitting of Alphafold-predicted model of the sheath protein (SH) monomer into the 7.9 Å-resolution structure reveals that the sheath cylinder consists of axially stacked β-hoops, each of which is comprised of two to six 400 nm-diameter rings of β-strand arches (β-rings). With both similarities to and differences from amyloid cross-β fibril architecture, each β-ring contains two giant β-sheets contributed by ~ 450 SH monomers that entirely encircle the outer circumference of the cell. Tomograms of immature cells suggest models of sheath biogenesis: oligomerization of SH monomers into β-ring precursors after their membrane-proximal cytoplasmic synthesis, followed by translocation through the unplugged end of a dividing cell, and insertion of nascent β-hoops into the immature sheath cylinder at the junction of two daughter cells.

59 BASIC BIOLOGICAL SCIENCES↗

Structural diversity and clustering of bacterial flagellar outer domains

Supercoiled flagellar filaments function as mechanical propellers within the bacterial flagellum complex, playing a crucial role in motility. Flagellin, the building block of the filament, features a conserved inner D0/D1 core domain across different bacterial species. In contrast, approximately half of the flagellins possess additional, highly divergent outer domain(s), suggesting varied functional potential. In this study, we report atomic structures of flagellar filaments from three distinct bacterial species: Cupriavidus gilardii , Stenotrophomonas maltophilia , and Geovibrio thiophilus . Our findings reveal that the flagella from the facultative anaerobic G. thiophilus possesses a significantly more negatively charged surface, potentially enabling adhesion to positively charged minerals. Furthermore, we analyze all AlphaFold predicted structures for annotated bacterial flagellins, categorizing the flagellin outer domains into 682 structural clusters. This classification provides insights into the prevalence and experimental verification of these outer domains. Remarkably, two of the flagellar structures reported herein belong to a distinct cluster, indicating additional opportunities on the study of the functional diversity of flagellar outer domains. Our findings underscore the complexity of bacterial flagellins and open up possibilities for future studies into their varied roles beyond motility.

Science & Technology - Other Topics↗

Molecular model of TFIIH recruitment to the transcription-coupled repair machinery

Transcription-coupled repair (TCR) is a vital nucleotide excision repair sub-pathway that removes DNA lesions from actively transcribed DNA strands. Binding of CSB to lesion-stalled RNA Polymerase II (Pol II) initiates TCR by triggering the recruitment of downstream repair factors. Yet it remains unknown how transcription factor IIH (TFIIH) is recruited to the intact TCR complex. Combining existing structural data with AlphaFold predictions, we build an integrative model of the initial TFIIH-bound TCR complex. We show how TFIIH can be first recruited in an open repair-inhibited conformation, which requires subsequent CAK module removal and conformational closure to process damaged DNA. In our model, CSB, CSA, UVSSA, elongation factor 1 (ELOF1), and specific Pol II and UVSSA-bound ubiquitin moieties come together to provide interaction interfaces needed for TFIIH recruitment. STK19 acts as a linchpin of the assembly, orienting the incoming TFIIH and bridging Pol II to core TCR factors and DNA. Molecular simulations of the TCR-associated CRL4CSA ubiquitin ligase complex unveil the interplay of segmental DDB1 flexibility, continuous Cullin4A flexibility, and the key role of ELOF1 for Pol II ubiquitination that enables TCR. Collectively, these findings elucidate the coordinated assembly of repair proteins in early TCR.

Paul, Tanmoy↗

Unlocking expanded flagellin perception through rational receptor engineering

Abstract The surface-localized receptor kinase FLS2 detects the flg22 epitope from bacterial flagella. FLS2 is conserved across land plants, but bacterial pathogens exhibit polymorphic flg22 epitopes. Most FLS2 homologues possess narrow perception ranges, but four with expanded perception have been identified. Using diversity analyses, AlphaFold modelling and amino acid properties, key residues enabling expanded recognition were mapped to FLS2’s concave surface, interacting with the co-receptor and polymorphic flg22 residues. Synthetic biology enabled engineering of expanded recognition from QvFLS2 (Quercus variabilis) into a homologue with canonical perception. A similar approach enabled transfer ofAgrobacteriumperception from FLS2 XL (Vitis riparia) into VrFLS2. Evolutionary analyses across three plant orders showed residues under positive selection aligning with those binding the co-receptor and flg22’s C terminus, suggesting more alleles with expanded perception exist. Our experimental data enabled the identification of specific receptor amino acid properties and AlphaFold3 metrics that facilitate predicting FLS2–flg22 recognition. This study provides a framework for rational receptor engineering to enhance pathogen restriction.

Plant Sciences↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of domain of unknown function 507 (DUF507) reveals a new protein fold

The crystal structure of the domain of unknown function family 507 protein from Aquifex aeolicus is reported (AaDUF507, UniProt O67633, 183 residues). The structure was determined in two space groups (C222 1 and P3 2 21) at 1.9 Å resolution. The phase problem was solved by molecular replacement using an AlphaFold model as the search model. AaDUF507 is a Y-shaped α-helical protein consisting of an anti-parallel 4-helix bundle base and two helical arms that extend 30-Å from the base. The two crystal structures differ by a 25° rigid body rotation of the C-terminal arm. The tertiary structure exhibits pseudo-twofold symmetry. The structural symmetry mirrors internal sequence similarity: residues 11–57 and 102–148 are 30% identical and 53% similar with an E-value of 0.002. In one of the structures, electron density for an unknown ligand, consistent with nicotinamide or similar molecule, may indicate a functional site. Docking calculations suggest potential ligand binding hot spots in the region between the helical arms. Structure-based query of the Protein Data Bank revealed no other protein with a similar tertiary structure, leading us to propose that AaDUF507 represents a new protein fold.

36 MATERIALS SCIENCE↗

A multi-ancestry GWAS of Fuchs corneal dystrophy highlights the contributions of laminins, collagen, and endothelial cell regulation

Fuchs endothelial corneal dystrophy (FECD) is a leading indication for corneal transplantation, but its molecular etiology remains poorly understood. We performed genome-wide association studies (GWAS) of FECD in the Million Veteran Program followed by multi-ancestry meta-analysis with the previous largest FECD GWAS, for a total of 3970 cases and 333,794 controls. We confirm the previous four loci, and identify eight novel loci: SSBP3, THSD7A, LAMB1, PIDD1, RORA, HS3ST3B1, LAMA5, and COL18A1. We further confirm the TCF4 locus in GWAS for admixed African and Hispanic/Latino ancestries and show an enrichment of European-ancestry haplotypes at TCF4 in FECD cases. Among the novel associations are low frequency missense variants in laminin genes LAMA5 and LAMB1 which, together with previously reported LAMC1, form laminin-511 (LM511). AlphaFold 2 protein modeling, validated through homology, suggests that mutations at LAMA5 and LAMB1 may destabilize LM511 by altering inter-domain interactions or extracellular matrix binding. Finally, phenome-wide association scans and colocalization analyses suggest that the TCF4 CTG18.1 trinucleotide repeat expansion leads to dysregulation of ion transport in the corneal endothelium and has pleiotropic effects on renal function.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Genetic control of morphological transitions in a coacervating protein template

Nature routinely exploits liquid–liquid phase separation (LLPS) of proteins to control the assembly and mineralization of hybrid materials. Here, we show that fusion of the Car9 silica-binding peptide to an elastin-like polypeptide (ELP) yields temperature- and sequence-programmable soft matter templates for the synthesis of silicified architectures ranging in size from nanometers to micrometers. Specifically, we demonstrate unprecedented control over the diameter of silica nanoparticles (SiNP) in the 30–60 nm range with 4 nm precision, show that a single arginine residue (R4) in the Car9 sequence underpins the transition from micelles to proteinosomes, and find that substitutions in other basic residues modulate electrostatic repulsion and solvation to enable access to kinetically trapped species. These structures, which include interconnected micelles, small (∼200 nm) and large (>5 µm) vesicles, are readily visualized by SEM imaging following silicification. Molecular dynamics (MD) simulations and AlphaFold predictions reveal that mutations in positively charged residues alter interfacial packing, hydration, and conformational freedom of the silica-binding segments. Overall, our results establish sequence and thermal energy as synergistic levers for morphological control across length scales using solid-binding ELPs and establish mineralization as a powerful tool to visualize the structure of dynamic soft matter assemblies.

hierarchy↗