Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein Structure Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of a secretory heme‐binding protein from Nocardia seriolae involved in cell apoptosis

Abstract According to the whole‐genome bioinformatics analysis, a heme‐binding protein from Nocardia seriolae (HBP) was found. HBP was predicted to be a bacterial secretory protein, located at mitochondrial membrane in eukaryotic cells and have a similar protein structure with the heme‐binding protein of Mycobacterium tuberculosis , Rv0203. In this study, HBP was found to be a secretory protein and co‐localized with mitochondria in FHM cells. Quantitative analysis of mitochondrial membrane potential value, caspase‐3 activity, and transcription level of apoptosis‐related genes suggested that overexpression of HBP protein can induce cell apoptosis. In conclusion, HBP was a secretory protein which may target to mitochondria and involve in cell apoptosis in host cells. This research will promote the function study of HBP and deepen the comprehension of the virulence factors and pathogenic mechanisms of N. seriolae .

Wen, Yiming↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sas20 is a highly flexible starch-binding protein in the Ruminococcus bromii cell-surface amylosome

Ruminococcus bromii is a keystone species in the human gut that has the rare ability to degrade dietary resistant starch (RS). This bacterium secretes a suite of starch-active proteins that work together within larger complexes called amylosomes that allow R. bromii to bind and degrade RS. Starch adherence system protein 20 (Sas20) is one of the more abundant proteins assembled within amylosomes, but little could be predicted about its molecular features based on amino acid sequence. Here, we performed a structure–function analysis of Sas20 and determined that it features two discrete starch-binding domains separated by a flexible linker. We show that Sas20 domain 1 contains an N-terminal β-sandwich followed by a cluster of α-helices, and the nonreducing end of maltooligosaccharides can be captured between these structural features. Furthermore, the crystal structure of a close homolog of Sas20 domain 2 revealed a unique bilobed starch-binding groove that targets the helical α1,4-linked glycan chains found in amorphous regions of amylopectin and crystalline regions of amylose. Affinity PAGE and isothermal titration calorimetry demonstrated that both domains bind maltoheptaose and soluble starch with relatively high affinity (K d ≤ 20 μM) but exhibit limited or no binding to cyclodextrins. Finally, small-angle X-ray scattering analysis of the individual and combined domains support that these structures are highly flexible, which may allow the protein to adopt conformations that enhance its starch-targeting efficiency. Taken together, we conclude that Sas20 binds distinct features within the starch granule, facilitating the ability of R. bromii to hydrolyze dietary RS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

tinyIFD: A High-Throughput Binding Pose Refinement Workflow Through Induced-Fit Ligand Docking

A critical step in structure-based drug discovery is predicting whether and how a candidate molecule binds to a model of a therapeutic target. However, substantial protein side chain movements prevent current screening methods, such as docking, from accurately predicting the ligand conformations and require expensive refinements to produce viable candidates. Here, we present the development of a high-throughput and flexible ligand pose refinement workflow, called “tinyIFD”. The main features of the workflow include the use of specialized high-throughput, small-system MD simulation code mdgx.cuda and an actively learning model zoo approach. We show the application of this workflow on a large test set of diverse protein targets, achieving 66% and 76% success rates for finding a crystal-like pose within the top-2 and top-5 poses, respectively. We also applied this workflow to the SARS-CoV-2 main protease (M pro ) inhibitors, where we demonstrate the benefit of the active learning aspect in this workflow.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of structural transitions in bacterial fatty acid binding proteins that permit ligand entry and exit at membranes

Fatty acid (FA) transfer proteins extract FA from membranes and sequester them to facilitate their movement through the cytosol. Detailed structural information is available for these soluble protein–FA complexes, but the structure of the protein conformation responsible for FA exchange at the membrane is unknown. Staphylococcus aureus FakB1 is a prototypical bacterial FA transfer protein that binds palmitate within a narrow, buried tunnel. Here, we define the conformational change from a “closed” FakB1 state to an “open” state that associates with the membrane and provides a path for entry and egress of the FA. Using NMR spectroscopy, we identified a conformationally flexible dynamic region in FakB1, and X-ray crystallography of FakB1 mutants captured the conformation of the open state. In addition, molecular dynamics simulations show that the new amphipathic α-helix formed in the open state inserts below the phosphate plane of the bilayer to create a diffusion channel for the hydrophobic FA tail to access the hydrocarbon core and place the carboxyl group at the phosphate layer. The membrane binding and catalytic properties of site-directed mutants were consistent with the proposed membrane docked structure predicted by our molecular dynamics simulations. Finally, the structure of the bilayer-associated conformation of FakB1 has local similarities with mammalian FA binding proteins and provides a conceptual framework for how these proteins interact with the membrane to create a diffusion channel from the FA location in the bilayer to the protein interior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The CRISPR effector Cam1 mediates membrane depolarization for phage defence

Prokaryotic type III CRISPR–Cas systems provide immunity against viruses and plasmids using CRISPR-associated Rossman fold (CARF) protein effectors. Recognition of transcripts of these invaders with sequences that are complementary to CRISPR RNA guides leads to the production of cyclic oligoadenylate second messengers, which bind CARF domains and trigger the activity of an effector domain. Whereas most effectors degrade host and invader nucleic acids, some are predicted to contain transmembrane helices without an enzymatic function. Whether and how these CARF–transmembrane helix fusion proteins facilitate the type III CRISPR–Cas immune response remains unknown. Here we investigate the role of cyclic oligoadenylate-activated membrane protein 1 (Cam1) during type III CRISPR immunity. Structural and biochemical analyses reveal that the CARF domains of a Cam1 dimer bind cyclic tetra-adenylate second messengers. In vivo, Cam1 localizes to the membrane, is predicted to form a tetrameric transmembrane pore, and provides defence against viral infection through the induction of membrane depolarization and growth arrest. These results reveal that CRISPR immunity does not always operate through the degradation of nucleic acids, but is instead mediated via a wider range of cellular responses.

59 BASIC BIOLOGICAL SCIENCES↗

A Case Study of the Glycoside Hydrolase Enzyme Mechanism Using an Automated QM-Cluster Model Building Toolkit

Glycoside hydrolase enzymes are important for hydrolyzing the β-1,4 glycosidic bond in polysaccharides for deconstruction of carbohydrates. The two-step retaining reaction mechanism of Glycoside Hydrolase Family 7 (GH7) was explored with different sized QM-cluster models built by the Residue Interaction Network ResidUe Selector (RINRUS) software using both the wild-type protein and its E217Q mutant. The first step is the glycosylation, in which the acidic residue 217 donates a proton to the glycosidic oxygen leading to bond cleavage. In the subsequent deglycosylation step, one water molecule migrates into the active site and attacks the anomeric carbon. Residue interaction-based QM-cluster models lead to reliable structural and energetic results for proposed glycoside hydrolase mechanisms. The free energies of activation for glycosylation in the largest QM-cluster models were predicted to be 19.5 and 31.4 kcal mol −1 for the wild-type protein and its E217Q mutant, which agree with experimental trends that mutation of the acidic residue Glu217 to Gln will slow down the reaction; and are higher in free energy than the deglycosylation transition states (13.8 and 25.5 kcal mol −1 for the wild-type protein and its mutant, respectively). For the mutated protein, glycosylation led to a low-energy product. This thermodynamic sink may correspond to the intermediate state which was isolated in the X-ray crystal structure. Hence, the glycosylation is validated to be the rate-limiting step in both the wild-type and mutated enzyme.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence-structure-function characterization of the emerging tetracycline destructase family of antibiotic resistance enzymes

Tetracycline destructases (TDases) are flavin monooxygenases which can confer resistance to all generations of tetracycline antibiotics. The recent increase in the number and diversity of reported TDase sequences enables a deep investigation of the TDase sequence-structure-function landscape. Here, we evaluate the sequence determinants of TDase function through two complementary approaches: (1) constructing profile hidden Markov models to predict new TDases, and (2) using multiple sequence alignments to identify conserved positions important to protein function. Using the HMM-based approach we screened 50 high-scoring candidate sequences in Escherichia coli, leading to the discovery of 13 new TDases. The X-ray crystal structures of two new enzymes from Legionella species were determined, and the ability of anhydrotetracycline to inhibit their tetracycline-inactivating activity was confirmed. Using the MSA-based approach we identified 31 amino acid positions 100% conserved across all known TDase sequences. The roles of these positions were analyzed by alanine-scanning mutagenesis in two TDases, to study the impact on cell and in vitro activity, structure, and stability. These results expand the diversity of TDase sequences and provide valuable insights into the roles of important residues in TDases, and flavin monooxygenases more broadly.

60 APPLIED LIFE SCIENCES↗

iPNHOT: a knowledge-based approach for identifying protein-nucleic acid interaction hot spots

The interaction between proteins and nucleic acids plays pivotal roles in various biological processes such as transcription, translation, and gene regulation. Hot spots are a small set of residues that contribute most to the binding affinity of a protein-nucleic acid interaction. Compared to the extensive studies of the hot spots on protein-protein interfaces, the hot spot residues within protein-nucleic acids interfaces remain less well-studied, in part because mutagenesis data for protein-nucleic acids interaction are not as abundant as that for protein-protein interactions. In this study, we built a new computational model, iPNHOT, to effectively predict hot spot residues on protein-nucleic acids interfaces. One training data set and an independent test set were collected from dbAMEPNI and some recent literature, respectively. To build our model, we generated 97 different sequential and structural features and used a two-step strategy to select the relevant features. The final model was built based only on 7 features using a support vector machine (SVM). The features include two unique features such as ΔSASsa 1/2 and esp3, which are newly proposed in this study. Based on the cross validation results, our model gave F1 score and AUROC as 0.725 and 0.807 on the subset collected from ProNIT, respectively, compared to 0.407 and 0.670 of mCSM-NA, a state-of-the art model to predict the thermodynamic effects of protein-nucleic acid interaction. The iPNHOT model was further tested on the independent test set, which showed that our model outperformed other methods. Here, by collecting data from a recently published database dbAMEPNI, we proposed a new model, iPNHOT, to predict hotspots on both protein-DNA and protein-RNA interfaces. The results show that our model outperforms the existing state-of-art models. Our model is available for users through a webserver: http://zhulab.ahu.edu.cn/iPNHOT/.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Biochemical Process Modeling and Simulation (BPMS)

The Biochemical Process Modeling and Simulation project aims to reduce the cost and time of research by applying theory, modeling, and simulation to the most relevant bottlenecks in the biochemical process. We use molecular modeling, quantum mechanics, metabolic modeling, fluid dynamics, and reaction-diffusion methods in close collaboration with pretreatment, hydrolysis, upgrading, and TEA. The project's outcomes are increased yields and efficiency of the biochemical process, added value to products, and reduced price of fuels by specifically targeting catalytic efficiency, reactor design, enzyme efficiency, and microbial design. We work closely with experimental projects to identify problems and iterate with experiments to find and refine solutions. By working with experimentalists, we decide on problems that can be solved with simulation that could otherwise not be solved or would take too long with experiment alone to reach BETO's targets. Over the years, we have produced solutions that have resulted in determining the most likely fatty-acid derivative for passive transport out of bacteria that upgrade biomass, and we have also designed enzyme mutations for enhanced lignin upgrading. Metabolic models have been developed to tune the activity of 2,3 butanediol production for the 2030 target. A computational method to deliver understanding of how complex omics data can be interpreted in the metabolic pathways of organisms used in the Agile Biofoundry. We have found methods to overcome specific barriers and continue to develop those methods. Our reactor studies have guided the design of both the microbes and reactors for aerobic and micro-aerobic production at all scales and have been instrumental in improving the accuracy of techno-economic analysis models. This project is essential in the process of selecting the final processes for 2030 SAF production targets. More specifically, recently, we have: 1) Predicted the strength of the basic structural interactions in commodity plastics to provide guidance for plastics upcycling strategies. 2) Developed computational tool to improve the characterization of lignin-derived compounds 3) Developed new methodologies to enable Machine Learning-based Directed Evolution for protein engineering. 4) Developed Machine Learning methods to predict protein promiscuity and mutations to further improve microbial and enzymatic driven processes and demonstrated the utility of ML approaches to engineering proteins from sparse experimental datasets. 5) Developed new methods to enable high-fidelity simulation of aerobic fermentation at industrial scale and resolving mismatch of time scales through subcycling/operator splitting 7) Identified the difficulty in preventing local high-oxygen conditions in industrial bubble columns, which leads to less-desirable acetoin production, suggesting future research directions in alternative reactor configurations (e.g loop reactors, shallow-channel reactors).

BIOMASS FUELS↗

Custom tuning of Rieske oxygenase reactivity

Rieske oxygenases use a Rieske-type [2Fe-2S] cluster and a mononuclear iron center to initiate a range of chemical transformations. However, few details exist regarding how this catalytic scaffold can be predictively tuned to catalyze divergent reactions. Therefore, in this work, using a combination of structural analyses, as well as substrate and rational protein-based engineering campaigns, we elucidate the architectural trends that govern catalytic outcome in the Rieske monooxygenase TsaM. We identify structural features that permit a substrate to be functionalized by TsaM and pinpoint active-site residues that can be targeted to manipulate reactivity. Exploiting these findings allowed for custom tuning of TsaM reactivity: substrates are identified that support divergent TsaM-catalyzed reactions and variants are created that exclusively catalyze dioxygenation or sequential monooxygenation chemistry. Importantly, we further leverage these trends to tune the reactivity of additional monooxygenase and dioxygenase enzymes, and thereby provide strategies to custom tune Rieske oxygenase reaction outcomes.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying amyloid-related diseases by mapping mutations in low-complexity protein domains to pathologies

Proteins including FUS, hnRNPA2, and TDP-43 reversibly aggregate into amyloid-like fibrils through interactions of their low-complexity domains (LCDs). Mutations in LCDs can promote irreversible amyloid aggregation and disease. We introduce a computational approach to identify mutations in LCDs of disease-associated proteins predicted to increase propensity for amyloid aggregation. We identify several disease-related mutations in the intermediate filament protein keratin-8 (KRT8). Atomic structures of wild-type and mutant KRT8 segments confirm the transition to a pleated strand capable of amyloid formation. Biochemical analysis reveals KRT8 forms amyloid aggregates, and the identified mutations promote aggregation. Aggregated KRT8 is found in Mallory–Denk bodies, observed in hepatocytes of livers with alcoholic steatohepatitis (ASH). We demonstrate that ethanol promotes KRT8 aggregation, and KRT8 amyloids co-crystallize with alcohol. Lastly, KRT8 aggregation can be seeded by liver extract from people with ASH, consistent with the amyloid nature of KRT8 aggregates and the classification of ASH as an amyloid-related condition.

59 BASIC BIOLOGICAL SCIENCES↗