Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein distance map”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks

Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

A Highly Ordered Nitroxide Side Chain for Distance Mapping and Monitoring Slow Structural Fluctuations in Proteins

Abstract Site-directed spin labeling electron paramagnetic resonance (SDSL-EPR) is an established tool for exploring protein structure and dynamics. Although nitroxide side chains attached to a single cysteine via a disulfide linkage are commonly employed in SDSL-EPR, their internal flexibility complicates applications to monitor slow internal motions in proteins and to structure determination by distance mapping. Moreover, the labile disulfide linkage prohibits the use of reducing agents often needed for protein stability. To enable the application of SDSL-EPR to the measurement of slow internal dynamics, new spin labels with hindered internal motion are desired. Here, we introduce a highly ordered nitroxide side chain, designated R9, attached at a single cysteine residue via a non-reducible thioether linkage. The reaction to introduce R9 is highly selective for solvent-exposed cysteine residues. Structures of R9 at two helical sites in T4 Lysozyme were determined by X-ray crystallography and the mobility in helical sequences was characterized by EPR spectral lineshape analysis, Saturation Transfer EPR, and Saturation Recovery EPR. In addition, interspin distance measurements between pairs of R9 residues are reported. Collectively, all data indicate that R9 will be useful for monitoring slow internal structural fluctuations, and applications to distance mapping via dipolar spectroscopy and relaxation enhancement methods are anticipated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Protein model quality assessment using rotation–equivariant transformations on point clouds

Machine learning research concerning protein structure has seen a surge in popularity over the last years with promising advances for basic science and drug discovery. Working with macromolecular structure in a machine learning context requires an adequate numerical representation, and researchers have extensively studied representations such as graphs, discretized 3D grids, and distance maps. As part of CASP14, we explored a new and conceptually simple representation in a blind experiment: atoms as points in 3D, each with associated features. These features—initially just the basic element type of each atom—are updated through a series of neural network layers featuring rotation-equivariant convolutions. Starting from all atoms, we further aggregate information at the level of alpha carbons before making a prediction at the level of the entire protein structure. We find that this approach yields competitive results in protein model quality assessment despite its simplicity and despite the fact that it incorporates minimal prior information and is trained on relatively little data. As a result, its performance and generality are particularly noteworthy in an era where highly complex, customized machine learning methods such as AlphaFold 2 have come to dominate protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗

Dispersal, habitat filtering, and eco-evolutionary dynamics as drivers of local and global wetland viral biogeography

Abstract Wetlands store 20–30% of the world’s soil carbon, and identifying the microbial controls on these carbon reserves is essential to predicting feedbacks to climate change. Although viral infections likely play important roles in wetland ecosystem dynamics, we lack a basic understanding of wetland viral ecology. Here 63 viral size-fraction metagenomes (viromes) and paired total metagenomes were generated from three time points in 2021 at seven fresh- and saltwater wetlands in the California Bodega Marine Reserve. We recovered 12,826 viral population genomic sequences (vOTUs), only 4.4% of which were detected at the same field site two years prior, indicating a small degree of population stability or recurrence. Viral communities differed most significantly among the seven wetland sites and were also structured by habitat (plant community composition and salinity). Read mapping to a new version of our reference database, PIGEONv2.0 (515,763 vOTUs), revealed 196 vOTUs present over large geographic distances, often reflecting shared habitat characteristics. Wetland vOTU microdiversity was significantly lower locally than globally and lower within than between time points, indicating greater divergence with increasing spatiotemporal distance. Viruses tended to have broad predicted host ranges via CRISPR spacer linkages to metagenome-assembled genomes, and increased SNP frequencies in CRISPR-targeted major tail protein genes suggest potential viral eco-evolutionary dynamics in response to both immune targeting and changes in host cell receptors involved in viral attachment. Together, these results highlight the importance of dispersal, environmental selection, and eco-evolutionary dynamics as drivers of local and global wetland viral biogeography.

Environmental Sciences & Ecology↗

Radius measurement via super-resolution microscopy enables the development of a variable radii proximity labeling platform

The elucidation of protein interaction networks is critical to understanding fundamental biology as well as developing new therapeutics. Proximity labeling platforms (PLPs) are state-of-the-art technologies that enable the discovery and delineation of biomolecular networks through the identification of protein-protein interactions. These platforms work via catalytic generation of reactive probes at a biological region of interest; these probes then diffuse through solution and covalently “tag” proximal biomolecules. The physical distance that the probes diffuse determines the effective labeling radius of the PLP and is a critical parameter that influences the scale and resolution of interactome mapping. As such, by expanding the degrees of labeling resolution offered by PLPs, it is possible to better capture the various size scales of interactomes. At present, however, there is little quantitative understanding of the labeling radii of different PLPs. Here, we report the development of a superresolution microscopy-based assay for the direct quantification of PLP labeling radii. Using this assay, we provide direct extracellular measurements of the labeling radii of state-of-the-art antibody-targeted PLPs, including the peroxidase-based phenoxy radical platform (269 ± 41 nm) and the high-resolution iridium-catalyzed µMap technology (54 ± 12 nm). Last, we apply these insights to the development of a molecular diffusion-based approach to tuning PLP resolution and introduce a new aryl-azide-based µMap platform with an intermediate labeling radius (80 ± 28 nm).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AlignOT: An Optimal Transport Based Algorithm for Fast 3D Alignment With Applications to Cryogenic Electron Microscopy Density Maps

Aligning electron density maps from Cryogenic electron microscopy (cryo-EM) is a first key step for studying multiple conformations of a biomolecule. As this step remains costly and challenging, with standard alignment tools being potentially stuck in local minima, we propose here a new procedure, called AlignOT, which relies on the use of computational optimal transport (OT) to align EM maps in 3D space. By embedding a fast estimation of OT maps within a stochastic gradient descent algorithm, our method searches for a rotation that minimizes the Wasserstein distance between two maps, represented as point clouds. Here, we quantify the impact of various parameters on the precision and accuracy of the alignment, and show that AlignOT can outperform the standard local alignment methods, with an increased range of rotation angles leading to proper alignment. We further benchmark AlignOT on various pairs of experimental maps, which account for different types of conformational heterogeneities and geometric properties. As our experiments show good performance, we anticipate that our method can be broadly applied to align 3D EM maps.

3D alignment↗

A Compact X-Ray System for Macromolecular Crystallography

We describe the design and performance of a high flux x-ray system for macromolecular crystallography that combines a microfocus x-ray generator (40 gm FWHM spot size at a power level of 46.5Watts) and a 5.5 mm focal distance polycapillary optic. The Cu K(sub alpha) X-ray flux produced by this optimized system is 7.0 times above the X-ray flux previously reported. The X-ray flux from the microfocus system is also 3.2 times higher than that produced by the rotating anode generator equipped with a long focal distance graded multilayer monochromator (Green optic; CMF24-48-Cu6) and 30% less than that produced by the rotating anode generator with the newest design of graded multilayer monochromator (Blue optic; CMF12-38-Cu6). Both rotating anode generators operate at a power level of 5000 Watts, dissipating more than 100 times the power of our microfocus x-ray system. Diffraction data collected from small test crystals are of high quality. For example, 42,540 reflections collected at ambient temperature from a lysozyme crystal yielded R(sub sym) 5.0% for the data extending to 1.7A, and 4.8% for the complete set of data to 1.85A. The amplitudes of the reflections were used to calculate difference electron density maps that revealed positions of structurally important ions and water molecules in the crystal of lysozyme using the phases calculated from the protein model.

Gubarev, Mikhail↗

Intracellular stress tomography reveals stress focusing and structural anisotropy in cytoskeleton of living cells

We describe a novel synchronous detection approach to map the transmission of mechanical stresses within the cytoplasm of an adherent cell. Using fluorescent protein-labeled mitochondria or cytoskeletal components as fiducial markers, we measured displacements and computed stresses in the cytoskeleton of a living cell plated on extracellular matrix molecules that arise in response to a small, external localized oscillatory load applied to transmembrane receptors on the apical cell surface. Induced synchronous displacements, stresses, and phase lags were found to be concentrated at sites quite remote from the localized load and were modulated by the preexisting tensile stress (prestress) in the cytoskeleton. Stresses applied at the apical surface also resulted in displacements of focal adhesion sites at the cell base. Cytoskeletal anisotropy was revealed by differential phase lags in X vs. Y directions. Displacements and stresses in the cytoskeleton of a cell plated on poly-L-lysine decayed quickly and were not concentrated at remote sites. These data indicate that mechanical forces are transferred across discrete cytoskeletal elements over long distances through the cytoplasm in the living adherent cell.

NASA Discipline Cell Biology↗

Locations of Halide Ions in Tetragonal Lysozyme Crystals

Anions play an important role in the crystallization of lysozyme, and are known to bind to the crystalline protein. Previous studies employing X-ray crystallography had found one chloride ion binding site in the tetragonal crystal form of the protein and four nitrate ion binding sites in the monoclinic form. Studies using other approaches have reported more chloride ion binding sites, but their locations were not known. Knowing the precise location of these anions is also useful in determining the correct electrostatic fields surrounding the protein. In the first part of this study the anion positions in the tetragonal form were determined from the difference Fourier map obtained from the lysozyme crystals grown in bromide and chloride solutions under identical conditions. The anion locations were then obtained from standard crystallographic methods and five possible anion binding sites were found in this manner. The sole chloride ion binding site found in previous studies was confirmed. The remaining four sites were new ones for tetragonal lysozyme crystals. However, three of these new sites and the previously found one corresponded to the four unique binding sites found for nitrate ions in monoclinic crystals. This suggests that most of the anion binding sites in lysozyme remain unchanged, even when different anions and different crystal forms of lysozyme are employed. It is unlikely that there are many more anions in the tetragonal lysozyme crystal structure. Assuming osmotic equilibrium it can be shown that there are at most three more anions in the crystal channels. Some of the new anion binding sites found in this study were, as expected, in pockets containing basic residues. However, some of them were near neutral, but polar, residues. Thus, the study also showed the importance of uncharged, but polar groups, on the protein surface in determining its electrostatic field. This was important for the second part of this study where the electrostatic field surrounding the protein was accurately determined. This was achieved by solving the linearized version of the Poisson-Boltzmann equation for the protein in solution. The solution was computed employing the commercial code Delphi which uses a finite difference technique. This has recently become available as a module in the general protein visualization code Insight II. Partial charges were assigned to the polar groups of lysozyme for the calculations done here. The calculations showed the complexity of the electrostatic field surrounding the protein. Although most of the region near the protein surface had a positive field strength, the active site cleft was negatively charged and this was projected a considerable distance. This might explain the occurrence of "head-to-side" interactions in the formation of lysozyme aggregates in solution. Pockets of high positive field strength were also found in the vicinity of the anion locations obtained from the crystallographic part of this study, confirming the validity of these calculations. This study clearly shows not only the importance of determining the counterion locations in protein crystals and the electrostatic fields surrounding the protein, but also the advantage of performing them together.

Lim, Kap↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗