Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Structure-guided design of a broadly cross-reactive multivalent group a streptococcal vaccine

The M protein of group A streptococci (Strep A) is a major virulence determinant and protective antigen. The N-terminal region of the M protein is variable in sequence, defines the M/emm type, and contains epitopes that elicit opsonic antibodies that protect animals from challenge infections. Although there are >200 M types of Strep A, there is now evidence that structurally related M proteins can be grouped into clusters and that immunity may be cluster-specific in addition to M type-specific. This observation has led to recent studies of structure-based design of multivalent M peptide vaccines to select peptides predicted to cross-react with heterologous M types to improve vaccine coverage. In the current study, we have applied a refined series of peptide structural algorithms to predict immunological cross-reactivity among 117N-terminal M peptides representing the most prevalent M types of Strep A. Based on the results of the structural analyses, in combination with global M type prevalence data, we constructed a 32-valent vaccine containing 19 cross-reactive vaccine candidates predicted to cross-react with 37 heterologous M peptides to which were added 13 type-specific M peptides. Further, the 4-protein recombinant vaccine was immunogenic in rabbits and elicited significant levels of antibodies against 31/32 (97%) vaccine peptides and 28/37 (76%) peptides predicted to cross-react. The vaccine antisera also promoted opsonophagocytic killing of vaccine and cross-reactive M types of Strep A. Based on a recent analysis of M type prevalence of Strep A, the potential global coverage of the 32-valent vaccine is ~90%, ranging from 68% in Africa to 95% in North America. Our results indicate the utility of structure-based design that may be applied to future studies of broadly protective M peptide vaccines.

60 APPLIED LIFE SCIENCES↗

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Fast myosin binding protein C knockout in skeletal muscle alters length-dependent activation and myofilament structure

In striated muscle, the sarcomeric protein myosin-binding protein-C (MyBP-C) is bound to the myosin thick filament and is predicted to stabilize myosin heads in a docked position against the thick filament, which limits crossbridge formation. Here, we use the homozygous Mybpc2 knockout (C2 -/- ) mouse line to remove the fast-isoform MyBP-C from fast skeletal muscle and then conduct mechanical functional studies in parallel with small-angle X-ray diffraction to evaluate the myofilament structure. We report that C2 -/- fibers present deficits in force production and calcium sensitivity. Structurally, passive C2 -/- fibers present altered sarcomere length-independent and -dependent regulation of myosin head conformations, with a shift of myosin heads towards actin. At shorter sarcomere lengths, the thin filament is axially extended in C2 -/- , which we hypothesize is due to increased numbers of low-level crossbridges. These findings provide testable mechanisms to explain the etiology of debilitating diseases associated with MyBP-C.

59 BASIC BIOLOGICAL SCIENCES↗

Persistent Protein Motions in a Rugged Energy Landscape Revealed by Normal Mode Ensemble Analysis

Proteins are allosteric machines that couple motions at distinct, often distant, sites to control biological function. Low-frequency structural vibrations are a mechanism of this long-distance connection and are often used computationally to predict correlations, but experimentally identifying the vibrations associated with specific motions has proved challenging. Spectroscopy is an ideal tool to explore these excitations, but measurements have been largely unable to identify important frequency bands. The result is at odds with some previous calculations and raises the question what methods could successfully characterize protein structural vibrations. Here we show the lack of spectral structure arises in part from the variations in protein structure as the protein samples the energy landscape. However, by averaging over the energy landscape as sampled using an aggregate 18.5 μs of all-atom molecular dynamics simulation of hen egg white lysozyme and normal-mode analyses, we find vibrations with large overlap with functional displacements are surprisingly concentrated in narrow frequency bands. These bands are not apparent in either the ensemble averaged vibrational density of states or isotropic absorption. However, in the case of the ensemble averaged anisotropic absorption, there is persistent spectral structure and overlap between this structure and the functional displacement frequency bands. We systematically lay out heuristics for calculating the spectra robustly, including the need for statistical sampling of the protein and inclusion of adequate water in the spectral calculation. The results show the congested spectrum of these complex molecules obscures important frequency bands associated with function and reveal a method to overcome this congestion by combining structurally sensitive spectroscopy with robust normal mode ensemble analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Strategies to identify and edit improvements in synthetic genome segments episomally

Genome engineering projects often utilize bacterial artificial chromosomes (BACs) to carry multi-kilobase DNA segments at low copy number. However, all stages of whole-genome engineering have the potential to impose mutations on the synthetic genome that can reduce or eliminate the fitness of the final strain. Here, we describe improvements to a multiplex automated genome engineering (MAGE) protocol to improve recombineering frequency and multiplexability. This protocol was applied to recoding an Escherichia coli strain to replace seven codons with synonymous alternatives genome wide. Ten 44 402–47 179 bp de novo synthesized DNA segments contained in a BAC from the recoded strain were unable to complement deletion of the corresponding 33–61 wild-type genes using a single antibiotic resistance marker. Next-generation sequencing (NGS) was used to identify 1–7 non-recoding mutations in essential genes per segment, and MAGE in turn proved a useful strategy to repair these mutations on the recoded segment contained in the BAC when both the recoded and wild-type copies of the mutated genes had to exist by necessity during the repair process. Finally, two web-based tools were used to predict the impact of a subset of non-recoding missense mutations on strain fitness using protein structure and function calls.

59 BASIC BIOLOGICAL SCIENCES↗

Design of Broadly Cross-Reactive M Protein–Based Group A Streptococcal Vaccines

Group A streptococcal infections are a significant cause of global morbidity and mortality. A leading vaccine candidate is the surface M protein, a major virulence determinant and protective Ag. An obstacle to the development of M protein–based vaccines is the >200 different M types defined by the N-terminal sequences that contain protective epitopes. Despite sequence variability, M proteins share coiled-coil structural motifs that bind host proteins required for virulence. In this study, we exploit this potential Achilles heel of conserved structure to predict cross-reactive M peptides that could serve as broadly protective vaccine Ags. Combining sequences with structural predictions, six heterologous M peptides in a sequence-related cluster were predicted to elicit cross-reactive Abs with the remaining five nonvaccine M types in the cluster. The six-valent vaccine elicited Abs in rabbits that reacted with all 11 M peptides in the cluster and functional opsonic Abs against vaccine and nonvaccine M types in the cluster. We next immunized mice with four sequence-unrelated M peptides predicted to contain different coiled-coil propensities and tested the antisera for cross-reactivity against 41 heterologous M peptides. Based on these results, we developed an improved algorithm to select cross-reactive peptide pairs using additional parameters of coiled-coil length and propensity. The revised algorithm accurately predicted cross-reactive Ab binding, improving the Matthews correlation coefficient from 0.42 to 0.74. These results form the basis for selecting the minimum number of N-terminal M peptides to include in potentially broadly efficacious multivalent vaccines that could impact the overall global burden of group A streptococcal diseases.

60 APPLIED LIFE SCIENCES↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of a secretory heme‐binding protein from Nocardia seriolae involved in cell apoptosis

Abstract According to the whole‐genome bioinformatics analysis, a heme‐binding protein from Nocardia seriolae (HBP) was found. HBP was predicted to be a bacterial secretory protein, located at mitochondrial membrane in eukaryotic cells and have a similar protein structure with the heme‐binding protein of Mycobacterium tuberculosis , Rv0203. In this study, HBP was found to be a secretory protein and co‐localized with mitochondria in FHM cells. Quantitative analysis of mitochondrial membrane potential value, caspase‐3 activity, and transcription level of apoptosis‐related genes suggested that overexpression of HBP protein can induce cell apoptosis. In conclusion, HBP was a secretory protein which may target to mitochondria and involve in cell apoptosis in host cells. This research will promote the function study of HBP and deepen the comprehension of the virulence factors and pathogenic mechanisms of N. seriolae .

Wen, Yiming↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary, structural and biochemical evidence for a new interaction site of the leptin obesity protein

The Leptin protein is central to the regulation of energy metabolism in mammals. By integrating evolutionary, structural, and biochemical information, a surface segment, outside of its known receptor contacts, is predicted as a second interaction site that may help to further define its roles in energy balance and its functional differences between humans and other mammals.

Evolution, Molecular↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sas20 is a highly flexible starch-binding protein in the Ruminococcus bromii cell-surface amylosome

Ruminococcus bromii is a keystone species in the human gut that has the rare ability to degrade dietary resistant starch (RS). This bacterium secretes a suite of starch-active proteins that work together within larger complexes called amylosomes that allow R. bromii to bind and degrade RS. Starch adherence system protein 20 (Sas20) is one of the more abundant proteins assembled within amylosomes, but little could be predicted about its molecular features based on amino acid sequence. Here, we performed a structure–function analysis of Sas20 and determined that it features two discrete starch-binding domains separated by a flexible linker. We show that Sas20 domain 1 contains an N-terminal β-sandwich followed by a cluster of α-helices, and the nonreducing end of maltooligosaccharides can be captured between these structural features. Furthermore, the crystal structure of a close homolog of Sas20 domain 2 revealed a unique bilobed starch-binding groove that targets the helical α1,4-linked glycan chains found in amorphous regions of amylopectin and crystalline regions of amylose. Affinity PAGE and isothermal titration calorimetry demonstrated that both domains bind maltoheptaose and soluble starch with relatively high affinity (K d ≤ 20 μM) but exhibit limited or no binding to cyclodextrins. Finally, small-angle X-ray scattering analysis of the individual and combined domains support that these structures are highly flexible, which may allow the protein to adopt conformations that enhance its starch-targeting efficiency. Taken together, we conclude that Sas20 binds distinct features within the starch granule, facilitating the ability of R. bromii to hydrolyze dietary RS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

tinyIFD: A High-Throughput Binding Pose Refinement Workflow Through Induced-Fit Ligand Docking

A critical step in structure-based drug discovery is predicting whether and how a candidate molecule binds to a model of a therapeutic target. However, substantial protein side chain movements prevent current screening methods, such as docking, from accurately predicting the ligand conformations and require expensive refinements to produce viable candidates. Here, we present the development of a high-throughput and flexible ligand pose refinement workflow, called “tinyIFD”. The main features of the workflow include the use of specialized high-throughput, small-system MD simulation code mdgx.cuda and an actively learning model zoo approach. We show the application of this workflow on a large test set of diverse protein targets, achieving 66% and 76% success rates for finding a crystal-like pose within the top-2 and top-5 poses, respectively. We also applied this workflow to the SARS-CoV-2 main protease (M pro ) inhibitors, where we demonstrate the benefit of the active learning aspect in this workflow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of structural transitions in bacterial fatty acid binding proteins that permit ligand entry and exit at membranes

Fatty acid (FA) transfer proteins extract FA from membranes and sequester them to facilitate their movement through the cytosol. Detailed structural information is available for these soluble protein–FA complexes, but the structure of the protein conformation responsible for FA exchange at the membrane is unknown. Staphylococcus aureus FakB1 is a prototypical bacterial FA transfer protein that binds palmitate within a narrow, buried tunnel. Here, we define the conformational change from a “closed” FakB1 state to an “open” state that associates with the membrane and provides a path for entry and egress of the FA. Using NMR spectroscopy, we identified a conformationally flexible dynamic region in FakB1, and X-ray crystallography of FakB1 mutants captured the conformation of the open state. In addition, molecular dynamics simulations show that the new amphipathic α-helix formed in the open state inserts below the phosphate plane of the bilayer to create a diffusion channel for the hydrophobic FA tail to access the hydrocarbon core and place the carboxyl group at the phosphate layer. The membrane binding and catalytic properties of site-directed mutants were consistent with the proposed membrane docked structure predicted by our molecular dynamics simulations. Finally, the structure of the bilayer-associated conformation of FakB1 has local similarities with mammalian FA binding proteins and provides a conceptual framework for how these proteins interact with the membrane to create a diffusion channel from the FA location in the bilayer to the protein interior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The CRISPR effector Cam1 mediates membrane depolarization for phage defence

Prokaryotic type III CRISPR–Cas systems provide immunity against viruses and plasmids using CRISPR-associated Rossman fold (CARF) protein effectors. Recognition of transcripts of these invaders with sequences that are complementary to CRISPR RNA guides leads to the production of cyclic oligoadenylate second messengers, which bind CARF domains and trigger the activity of an effector domain. Whereas most effectors degrade host and invader nucleic acids, some are predicted to contain transmembrane helices without an enzymatic function. Whether and how these CARF–transmembrane helix fusion proteins facilitate the type III CRISPR–Cas immune response remains unknown. Here we investigate the role of cyclic oligoadenylate-activated membrane protein 1 (Cam1) during type III CRISPR immunity. Structural and biochemical analyses reveal that the CARF domains of a Cam1 dimer bind cyclic tetra-adenylate second messengers. In vivo, Cam1 localizes to the membrane, is predicted to form a tetrameric transmembrane pore, and provides defence against viral infection through the induction of membrane depolarization and growth arrest. These results reveal that CRISPR immunity does not always operate through the degradation of nucleic acids, but is instead mediated via a wider range of cellular responses.

59 BASIC BIOLOGICAL SCIENCES↗