Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL-CompBio/pf-gnn_pli

We have developed a deep learning based method to decode the protein-ligand interactions and predict their probability of binding. We use 3-Dimensional structure of proteins and ligand molecules with the ligand docked to the receptor and use the information at the interaction region to assess the binding of the protein-ligand complex.

Bontha, Mridula↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

The leucine-rich repeats in allelic barley MLA immune receptors define specificity towards sequence-unrelated powdery mildew avirulence effectors with a predicted common RNase-like fold

Nucleotide-binding domain leucine-rich repeat-containing receptors (NLRs) in plants can detect avirulence (AVR) effectors of pathogenic microbes. The Mildew locus a (Mla) NLR gene has been shown to confer resistance against diverse fungal pathogens in cereal crops. In barley, Mla has undergone allelic diversification in the host population and confers isolate-specific immunity against the powdery mildew-causing fungal pathogen Blumeria graminis forma specialis hordei (Bgh). We previously isolated the Bgh effectors AVR A1 , AVR A7 , AVR A9 , AVR A13 , and allelic AVR A10 /AVR A22 , which are recognized by matching MLA1, MLA7, MLA9, MLA13, MLA10 and MLA22, respectively. Here, we extend our knowledge of the Bgh effector repertoire by isolating the AVR A6 effector, which belongs to the family of catalytically inactive RNase-Like Proteins expressed in Haustoria (RALPHs). Using structural prediction, we also identified RNase-like folds in AVR A1 , AVR A7 , AVR A10 /AVR A22 , and AVR A13 , suggesting that allelic MLA recognition specificities could detect structurally related avirulence effectors. To better understand the mechanism underlying the recognition of effectors by MLAs, we deployed chimeric MLA1 and MLA6, as well as chimeric MLA10 and MLA22 receptors in plant co-expression assays, which showed that the recognition specificity for AVR A1 and AVR A6 as well as allelic AVR A10 and AVR A22 is largely determined by the receptors’ C-terminal leucine-rich repeats (LRRs). The design of avirulence effector hybrids allowed us to identify four specific AVR A10 and five specific AVR A22 aa residues that are necessary to confer MLA10- and MLA22-specific recognition, respectively. This suggests that the MLA LRR mediates isolate-specific recognition of structurally related AVR A effectors. Thus, functional diversification of multi-allelic MLA receptors may be driven by a common structural effector scaffold, which could be facilitated by proliferation of the RALPH effector family in the pathogen genome.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence and structural implications of a bovine corneal keratan sulfate proteoglycan core protein. Protein 37B represents bovine lumican and proteins 37A and 25 are unique

Amino acid sequence from tryptic peptides of three different bovine corneal keratan sulfate proteoglycan (KSPG) core proteins (designated 37A, 37B, and 25) showed similarities to the sequence of a chicken KSPG core protein lumican. Bovine lumican cDNA was isolated from a bovine corneal expression library by screening with chicken lumican cDNA. The bovine cDNA codes for a 342-amino acid protein, M(r) 38,712, containing amino acid sequences identified in the 37B KSPG core protein. The bovine lumican is 68% identical to chicken lumican, with an 83% identity excluding the N-terminal 40 amino acids. Location of 6 cysteine and 4 consensus N-glycosylation sites in the bovine sequence were identical to those in chicken lumican. Bovine lumican had about 50% identity to bovine fibromodulin and 20% identity to bovine decorin and biglycan. About two-thirds of the lumican protein consists of a series of 10 amino acid leucine-rich repeats that occur in regions of calculated high beta-hydrophobic moment, suggesting that the leucine-rich repeats contribute to beta-sheet formation in these proteins. Sequences obtained from 37A and 25 core proteins were absent in bovine lumican, thus predicting a unique primary structure and separate mRNA for each of the three bovine KSPG core proteins.

NASA Discipline Cell Biology↗

Binding Free Energy Analysis of Colicin D, E3 and E8 to Their Respective Cognate Immunity Proteins Using Computational Simulations

Colicins are antimicrobial proteins produced by bacteria for the purpose of destroying neighboring bacteria. Colicin activity is neutralized by a specific cognate immunity protein in order to protect the host. This study investigates the structural and binding mechanisms underlying the interaction of colicin-D, -E3 and -E8 to their respective immunity proteins (ImD, Im3 and Im8) using structure prediction, molecular dynamics (MD) simulations and MM-PBSA approach of free energy calculations. High-confidence colicin-immunity (Col-Im) complex structures predicted using AlphaFold2 were subjected to MD simulations of 150 ns with GROMACS and were analyzed for the binding free energy calculation using gmx_MMPBSA. Results showed that the complex of Col_E3-Im3 exhibited the most favorable binding free energy, driven by strong van der Waals and electrostatic interactions. Col_D-ImD and Col_E8-Im8 also showed the favorable binding. Electrostatics and hydrogen bonding emerged as a key factor driving binding and stability, while polar solvation acted as a destabilizing factor across all systems. These outcomes provide an understanding of the molecular mechanisms of Col-Im systems, with potential applications for developing natural antimicrobials for food safety.

Biochemistry & Molecular Biology↗

ECUT (Energy Conversion and Utilization Technologies) program: Biocatalysis project

The Annual Report presents the fiscal year (FY) 1988 research activities and accomplishments, for the Biocatalysis Project of the U.S. Department of Energy, Energy Conversion and Utilization Technologies (ECUT) Division. The ECUT Biocatalysis Project is managed by the Jet Propulsion Laboratory, California Institute of Technology. The Biocatalysis Project is a mission-oriented, applied research and exploratory development activity directed toward resolution of the major generic technical barriers that impede the development of biologically catalyzed commercial chemical production. The approach toward achieving project objectives involves an integrated participation of universities, industrial companies and government research laboratories. The Project's technical activities were organized into three work elements: (1) The Molecular Modeling and Applied Genetics work element includes research on modeling of biological systems, developing rigorous methods for the prediction of three-dimensional (tertiary) protein structure from the amino acid sequence (primary structure) for designing new biocatalysis, defining kinetic models of biocatalyst reactivity, and developing genetically engineered solutions to the generic technical barriers that preclude widespread application of biocatalysis. (2) The Bioprocess Engineering work element supports efforts in novel bioreactor concepts that are likely to lead to substantially higher levels of reactor productivity, product yields and lower separation energetics. Results of work within this work element will be used to establish the technical feasibility of critical bioprocess monitoring and control subsystems. (3) The Bioprocess Design and Assessment work element attempts to develop procedures (via user-friendly computer software) for assessing the energy-economics of biocatalyzed chemical production processes, and initiation of technology transfer for advanced bioprocesses.

Baresi, Larry↗

Understanding amyloids to prevent biofilm formation in space

There is a pressing need to search for novel approaches to combat biofilm formation, both in space and in medical applications. Many proteins have the ability to form ordered aggregates called amyloids. Amyloids are known to be an important part of biofilms. The use of anti-amyloid drugs is a novel venue for the development of antimicrobial agents. The ultrastructure of the amyloid aggregate shows a high packing of proteins, the second-order structure of which is dominated by β-sheets. The ability to form an amyloid aggregate is especially typical for proteins containing domains (protein fragments) with sufficient lability to arrange themselves in a tight β-sheet structure. Bioinformatics tools allow the prediction of such behavior of proteins in genomic data. We use GeneLab data of microbial populations identified aboard the International Space Station and other spacecraft to look for bacterial species that utilize amyloid aggregation in biofilm formation. We use a combined bioinformatic approach with a relatively high throughput molecular biology assay and biophysical assays to evaluate the anti-amyloid anti-biofilm approach. The significance of the research extends from understanding basic microbial community responses to spaceflight, to biofouling of the built environments in space as well as the long-term health of astronauts. Bioinformatics shows that onboard the ISS, bacterial species produce far more amyloid and prion proteins than are currently verified, hence their role in bacterial ecosystems is largely unknown. As we propose there is a link between amyloid formation in space and biofilm production, this research should lead to new paths for biofilm remediation in space.

Tomasz Zajkowski↗

Integrating solvation shell structure in experimentally driven molecular dynamics using x-ray solution scattering data

In the past few decades, prediction of macromolecular structures beyond the native conformation has been aided by the development of molecular dynamics (MD) protocols aimed at exploration of the energetic landscape of proteins. Yet, the computed structures do not always agree with experimental observables, calling for further development of the MD strategies to bring the computations and experiments closer together. Here, we report a scalable, efficient MD simulation approach that incorporates an x-ray solution scattering signal as a driving force for the conformational search of stable structural configurations outside of the native basin. We further demonstrate the importance of inclusion of the hydration layer effect for a precise description of the processes involving large changes in the solvent exposed area, such as unfolding. Utilization of the graphics processing unit allows for an efficient all-atom calculation of scattering patterns on-the-fly, even for large biomolecules, resulting in a speed-up of the calculation of the associated driving force. The utility of the methodology is demonstrated on two model protein systems, the structural transition of lysine-, arginine-, ornithine-binding protein and the folding of deca-alanine. We discuss how the present approach will aid in the interpretation of dynamical scattering experiments on protein folding and association.

Hsu, Darren J.↗

Equivariant Graph Attention Network - 3D Conformers & Feature Fusion

EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.

Shim, Heesung [Lawrence Livermore National Laborat↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

Rational Design of Lanmodulin Variants for Size-Based Selectivity of Individual Rare Earth Elements

Rare earth elements (REEs) are essential to modern technologies, yet their high physical and chemical similarity makes separation of individual REEs difficult and environmentally taxing. Metalloproteins offer a promising alternative for selective REE binding, as they tend to have high metal ion affinity and specificity. Lanmodulin (LanM), in particular, has arisen as a potential candidate for REE separation as it exhibits picomolar affinity for elements in the REE family. Prior work has shown that the single point mutation D9N can shift LanM’s preference away from lanthanides toward actinides, motivating efforts to tune selectivity of LanM through targeted mutagenesis. Here, we tested the hypothesis that introducing selective aspartic acid to glutamic acid substitutions in the metal coordinating EF hands of LanM would impose steric constraints that would drive LanM affinity away from larger ions, such as La3+, to smaller ions, such as Y3+. To test this hypothesis, a combination of computational and experimental approaches were employed to evaluate the signal mutations LanM D5E and LanM D3E and the double mutants LanM D1ED5E and LanM D3ED9E. Surprisingly, increasing the number of mutations within the metal center did not enhance affinity for smaller REEs, or decrease affinity for larger ions. Only the single point mutation LanM D5E weakened La3+ binding by one order of magnitude relative to LanM wild type (WT), and pairing it with a second mutation to produce LanM D1ED5E drove La3+ affinity to be stronger than that seen for LanM WT. The D3E mutation alone prevented proper expression and folding, but paring it with D9E to produce LanM D3ED9E rescued expression and yielded La3+ affinities comparable to LanM WT. All variants that expressed (LanM D5E, LanM D1ED5E, LanM D3ED9E) displayed Y3+ affinities comparable to LanM WT. Overall, these results highlight the tunability of LanM’s metal-binding environment but also expose current limitations in predicting structural responses to point mutations within a protein sequence. This work establishes a foundation that can be used for refining computational and experimental strategies to engineer metalloproteins with tailored REE selectivity.

Close, Emily [Pacific Northwest National Laborator↗

Hierarchical organization and assembly of the archaeal cell sheath from an amyloid-like protein

Abstract Certain archaeal cells possess external proteinaceous sheath, whose structure and organization are both unknown. By cellular cryogenic electron tomography (cryoET), here we have determined sheath organization of the prototypical archaeon, Methanospirillum hungatei . Fitting of Alphafold-predicted model of the sheath protein (SH) monomer into the 7.9 Å-resolution structure reveals that the sheath cylinder consists of axially stacked β-hoops, each of which is comprised of two to six 400 nm-diameter rings of β-strand arches (β-rings). With both similarities to and differences from amyloid cross-β fibril architecture, each β-ring contains two giant β-sheets contributed by ~ 450 SH monomers that entirely encircle the outer circumference of the cell. Tomograms of immature cells suggest models of sheath biogenesis: oligomerization of SH monomers into β-ring precursors after their membrane-proximal cytoplasmic synthesis, followed by translocation through the unplugged end of a dividing cell, and insertion of nascent β-hoops into the immature sheath cylinder at the junction of two daughter cells.

59 BASIC BIOLOGICAL SCIENCES↗

Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator

Gene activator proteins comprise distinct DNA-binding and transcriptional activation domains (ADs). Because few ADs have been described, we tested domains tiling all yeast transcription factors for activation in vivo and identified 150 ADs. By mRNA display, we showed that 73% of ADs bound the Med15 subunit of Mediator, and that binding strength was correlated with activation. AD-Mediator interaction in vitro was unaffected by a large excess of free activator protein, pointing to a dynamic mechanism of interaction. Structural modeling showed that ADs interact with Med15 without shape complementarity (‘fuzzy’ binding). ADs shared no sequence motifs, but mutagenesis revealed biochemical and structural constraints. Finally, a neural network trained on AD sequences accurately predicted ADs in human proteins and in other yeast proteins, including chromosomal proteins and chromatin remodeling complexes. These findings solve the longstanding enigma of AD structure and function and provide a rationale for their role in biology.

60 APPLIED LIFE SCIENCES↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Putting AlphaFold models to work with phenix.process_predicted_model and ISOLDE

AlphaFold has recently become an important tool in providing models for experimental structure determination by X-ray crystallography and cryo-EM. Large parts of the predicted models typically approach the accuracy of experimentally determined structures, although there are frequently local errors and errors in the relative orientations of domains. Importantly, residues in the model of a protein predicted by AlphaFold are tagged with a predicted local distance difference test score, informing users about which regions of the structure are predicted with less confidence. AlphaFold also produces a predicted aligned error matrix indicating its confidence in the relative positions of each pair of residues in the predicted model. The phenix.process_predicted_model tool downweights or removes low-confidence residues and can break a model into confidently predicted domains in preparation for molecular replacement or cryo-EM docking. These confidence metrics are further used in ISOLDE to weight torsion and atom–atom distance restraints, allowing the complete AlphaFold model to be interactively rearranged to match the docked fragments and reducing the need for the rebuilding of connecting regions.

59 BASIC BIOLOGICAL SCIENCES↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

Human XPG nuclease structure, assembly, and activities with insights for neurodegeneration and cancer from pathogenic mutations

Xeroderma pigmentosum group G (XPG) protein is both a functional partner in multiple DNA damage responses (DDR) and a pathway coordinator and structure-specific endonuclease in nucleotide excision repair (NER). Different mutations in the XPG gene ERCC5 lead to either of two distinct human diseases: Cancer-prone xeroderma pigmentosum (XP-G) or the fatal neurodevelopmental disorder Cockayne syndrome (XP-G/CS). To address the enigmatic structural mechanism for these differing disease phenotypes and for XPG’s role in multiple DDRs, here we determined the crystal structure of human XPG catalytic domain (XPGcat), revealing XPG-specific features for its activities and regulation. Furthermore, XPG DNA binding elements conserved with FEN1 superfamily members enable insights on DNA interactions. Notably, all but one of the known pathogenic point mutations map to XPGcat, and both XP-G and XP-G/CS mutations destabilize XPG and reduce its cellular protein levels. Mapping the distinct mutation classes provides structure-based predictions for disease phenotypes: Residues mutated in XP-G are positioned to reduce local stability and NER activity, whereas residues mutated in XP-G/CS have implied long-range structural defects that would likely disrupt stability of the whole protein, and thus interfere with its functional interactions. Combined data from crystallography, biochemistry, small angle X-ray scattering, and electron microscopy unveil an XPG homodimer that binds, unstacks, and sculpts duplex DNA at internal unpaired regions (bubbles) into strongly bent structures, and suggest how XPG complexes may bind both NER bubble junctions and replication forks. Collective results support XPG scaffolding and DNA sculpting functions in multiple DDR processes to maintain genome stability.

59 BASIC BIOLOGICAL SCIENCES↗