Engineering PapersSearch

SEARCH · Engineering Papers

Results for “protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics

Engineering a new tripartite split-ccGFP system from Corynactis californica for detecting protein–protein interactions

Protein-protein interactions (PPIs) are critical to a range of biological processes and, consequently, aberrant interactions are implicated in many disorders. The study of the complex networks of PPIs promises to elucidate undiscovered roles in cellular processes and the mechanisms of disease. To accomplish this, tools to effectively sense PPIs are necessary. Effective PPI sensors must rapidly detect interactions in real-time with high sensitivity without perturbing the proteins of interest (POIs) under study. Split fluorescent proteins have previously been used to successfully monitor PPIs, in part due to the small size of the tags. Here, we developed an optimized tripartite split GFP system based on Corynactis californica GFP (ccGFP) to detect PPIs in vitro. In this sensor system, ccGFP fragments ccGFP10 and ccGFP11 are tagged to two POIs. PPIs can then be detected via fluorescence by complementation to the third fragment, ccGFP1-9, which reconstitutes functional ccGFP. The optimized ccGFP system shows improved detection kinetics and pH and temperature stability compared to a previous system. We then validated the sensor by monitoring PPIs in two model systems: attractive/repulsive coiled-coils and rapamycin-inducible FRB/FKBP heterodimerization. Finally, we developed an anti-tripartite ccGFP single-chain variable fragment (scFv), which could enable versatile detection of identified protein-protein complexes.

59 BASIC BIOLOGICAL SCIENCES

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Predicting compatibility between ferredoxins and the Fe protein of nitrogenase using in silico protein modeling

Biological nitrogen fixation is the process by which certain bacteria and archaea use the enzyme nitrogenase to reduce atmospheric nitrogen into bioavailable ammonium. Engineering non‐nitrogen‐fixing organisms, like plants, to use nitrogenase could reduce dependency on synthetic fertilizer and mitigate the environmental impacts of industrial fertilizer production. However, nitrogenase activity requires delivery of reducing power by small electron carrying proteins known as ferredoxins and flavodoxins, and successfully engineering nitrogenase into new systems will require a mechanistic understanding of electron delivery by these proteins. Most organisms often have multiple ferredoxins, raising the question of which ferredoxin can support nitrogenase activity. The purpose of this study is to gain insight into how we can predict which ferredoxin is compatible with the Fe protein, the component of nitrogenase that interacts with ferredoxin or flavodoxin. Our in silico protein–protein docking simulations reveal that most ferredoxins and flavodoxins involved in nitrogen fixation have the shortest distance (≤10 Å) between their redox cofactor and the [4Fe‐4S] cluster of the Fe protein. We found shorter cofactor distance contributes to faster intermolecular electron tunneling rates. Bacterial ferredoxins that play a role in nitrogen fixation also exhibit more complementary interactions with the Fe protein than bacterial and plant ferredoxins not involved in this process. Heterologous expression of a set of ferredoxins from both nitrogen‐fixing and non‐nitrogen‐fixing bacteria in the diazotroph Rhodopseudomonas palustris supports our model‐derived prediction that shorter distances between the electron‐carrying cofactors favor nitrogenase compatibility. These findings offer a framework to predict and potentially enhance ferredoxin–nitrogenase compatibility, which will help to improve our ability to engineer nitrogen fixation into non‐nitrogen‐fixing organisms like plants.

59 BASIC BIOLOGICAL SCIENCES

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML

Photoactivable analogs for labeling 25-hydroxyvitamin D3 serum binding protein and for 1,25-dihydroxyvitamin D3 intestinal receptor protein

3-Azidobenzoates and 3-azidonitrobenzoates of 25-hydroxyvitamin D3 as well as 3-deoxy-3-azido-25-hydroxyvitamin D3 and 3-deoxy-3-azido-1,25-dihydroxyvitamin D3 were prepared as photoaffinity labels for vitamin D serum binding protein and 1,25-dihydroxyvitamin D3 intestinal receptor protein. The compounds prepared were easily activated by short- or long-wavelength uv light, as monitored by uv and ir spectrometry. The efficacy of the compounds to compete with 25-hydroxyvitamin D3 or 1,25-dihydroxyvitamin D3 for the binding site of serum binding protein and receptor, respectively, was studied to evaluate the vitamin D label with the highest affinity for the protein. The presence of an azidobenzoate or azidonitrobenzoate substituent at the C-3 position of 25-OH-D3 significantly decreased (10(4)- to 10(6)-fold) the binding activity. However, the labels containing the azido substituent attached directly to the vitamin D skeleton at the C-3 position showed a high affinity, only 20- to 150-fold lower than that of the parent compounds with their respective proteins. Therefore, 3-deoxy-3-azidovitamins present potential ligands for photolabeling of vitamin D proteins and for studying the structures of the protein active sites.

NASA Discipline Musculoskeletal

An Integral Activity-Based Protein Profiling Method for Higher Throughput Determination of Protein Target Sensitivity to Small Molecules

Activity-based protein profiling (ABPP) is a chemoproteomic technique that uses small molecule probes to label active enzymes selectively and covalently in complex proteomes. Competitive ABPP, which involves treatment of the active proteome with an analyte of interest, is especially powerful for profiling how small molecules impact specific protein activities. Advances in higher throughput workflows have made it possible to generate extensive competitive ABPP data across diverse biological samples, making this approach highly appealing for characterizing shared and unique proteins affected by perturbations such as drug or chemical exposures. To use the competitive ABPP approach effectively to understand potential adverse effects of chemicals of concern (CoC), a wide range of concentrations may be needed, particularly for chemicals that lack potency or toxicity data. In this work, we present an integral competitive ABPP method that enables target sensitivity determination for different organophosphate (OP) pesticides as model toxicants. Using previously developed OP-ABPs, we optimized conditions for tandem mass tag (TMT) multiplexing of ABPP samples and compared conventional competitive ABPP involving samples at discrete paraoxon concentrations to pooled samples across that same concentration range. We then expanded our approach to compare protein target sensitivities toward two additional OP pesticides, chlorpyrifos oxon and malaoxon. The results showed that differences in integral intensities for the pooled competition sample can be used to evaluate the relative sensitivity of specific proteins without increasing the overall number of samples. For 8 CoC concentrations of interest, this strategy reduced the number of TMT plexes and the corresponding number of LC–MS/MS analyses 3-fold. In conclusion, we envision the integral ABPP (IABPP) method will provide a means to screen diverse chemicals more rapidly to identify both high and low sensitivity protein targets.

activity-based probes

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES

Pooled PPIseq: Screening the SARS-CoV-2 and human interface with a scalable multiplexed protein-protein interaction assay platform

Protein-Protein Interactions (PPIs) are a key interface between virus and host, and these interactions are important to both viral reprogramming of the host and to host restriction of viral infection. In particular, viral-host PPI networks can be used to further our understanding of the molecular mechanisms of tissue specificity, host range, and virulence. At higher scales, viral-host PPI screening could also be used to screen for small-molecule antivirals that interfere with essential viral-host interactions, or to explore how the PPI networks between interacting viral and host genomes co-evolve. Current high-throughput PPI assays have screened entire viral-host PPI networks. However, these studies are time consuming, often require specialized equipment, and are difficult to further scale. Here, we develop methods that make larger-scale viral-host PPI screening more accessible. This approach combines the mDHFR split-tag reporter with the iSeq2 interaction-barcoding system to permit massively-multiplexed PPI quantification by simple pooled engineering of barcoded constructs, integration of these constructs into budding yeast, and fitness measurements by pooled cell competitions and barcode-sequencing. We applied this method to screen for PPIs between SARS-CoV-2 proteins and human proteins, screening in triplicate >180,000 ORF-ORF combinations represented by >1,000,000 barcoded lineages. Our results complement previous screens by identifying 74 putative PPIs, including interactions between ORF7A with the taste receptors TAS2R41 and TAS2R7, and between NSP4 with the transmembrane KDELR2 and KDELR3. We show that this PPI screening method is highly scalable, enabling larger studies aimed at generating a broad understanding of how viral effector proteins converge on cellular targets to effect replication.

60 APPLIED LIFE SCIENCES

Graph Identification of Proteins in Tomograms (GRIP-Tomo) 2.0: Topologically aware classification for proteins

Cryo-electron tomography (cryo-ET) enables structural characterization of biomolecules under near-native conditions. Existing approaches for interpreting the resulting three-dimensional volumes are computationally expensive and have difficulty interpreting density associated with small proteins/complexes. To explore alternate approaches for identifying proteins in cryo-ET data we pursued a Graph Network and topologically invariant approach. Here, we report on a fast algorithm that classifies particles by searching for nuances of evolutionarily conversed motifs and the geometrical characteristics of protein structure. GRIP-Tomo 2.0 is a machine-learning pipeline that extracts interpretable topological features of protein structures within noisy experimental backgrounds. Compared to version 1.0, the new pipeline includes three upgrades that significantly improve performance including synthetic tomogram generation simulating realistic noise, graph-based persistent feature extraction as protein fingerprints, and high-performance computing acceleration. GRIP-Tomo 2.0 achieves over 90% accuracy in classifying between proteins and noise using both real and synthetic datasets which represents a foundational step toward advancing cryo-ET workflows and empowering automated visual proteomics.

Li, Chengxuan

Protein folding, protein structure and the origin of life: Theoretical methods and solutions of dynamical problems

Theoretical methods and solutions of the dynamics of protein folding, protein aggregation, protein structure, and the origin of life are discussed. The elements of a dynamic model representing the initial stages of protein folding are presented. The calculation and experimental determination of the model parameters are discussed. The use of computer simulation for modeling protein folding is considered.

Weaver, D. L.

Meal composition and plasma amino acid ratios: Effect of various proteins or carbohydrates, and of various protein concentrations

The effects of meals containing various proteins and carbohydrates, and of those containing various proportions of protein (0 percent to 20 percent of a meal, by weight) or of carbohydrate (0 percent to 75 percent), on plasma levels of certain large neutral amino acids (LNAA) in rats previously fasted for 19 hours were examined. Also the plasma tryptophan ratios (the ratio of the plasma trytophan concentration to the summed concentrations of the other large neutral amino acids) and other plasma amino acid ratios were calculated. (The plasma tryptophan ratio has been shown to determine brain tryptophan levels and, thereby, to affect the synthesis and release of the neurotransmitter serotonin). A meal containing 70 percent to 75 percent of an insulin-secreting carbohydrate (dextrose or dextrin) increased plasma insulin levels and the tryptophan ratio; those containing 0 percent or 25 percent carbohydrate failed to do so. Addition of as little as 5 percent casein to a 70 percent carbohydrate meal fully blocked the increase in the plasma tryptophan ratio without affecting the secretion of insulin - probably by contributing much larger quantities of the other LNAA than of tryptophan to the blood. Dietary proteins differed in their ability to suppress the carbohydrate-induced rise in the plasma tryptophan ratio. Addition of 10 percent casein, peanut meal, or gelatin fully blocked this increase, but lactalbumin failed to do so, and egg white did so only partially. (Consumption of the 10 percent gelatin meal also produced a major reduction in the plasma tyrosine ratio, and may thereby have affected brain tyrosine levels and catecholamine synthesis.) These observations suggest that serotonin-releasing neurons in brains of fasted rats are capable of distinguishing (by their metabolic effects) between meals poor in protein but rich in carbohydrates that elicit insulin secretion, and all other meals. The changes in brain serotonin caused by carbohydrate-rich, protein-poor meals may affect subsequent food choice and various serotonin-mediated behaviors.

Yokogoshi, Hidehiko

Mesoscale fractal whey protein particles derived from microscale linear-shaped protein assemblies (Part 1): Manufacturing method and particle characteristics

Whey protein isolates (WPI) are widely used in processed foods for their versatile functional properties. Modifying the structural properties of proteins by assembling them into mesoscale or microscale particles may improve their functionality and broaden their applications. This study aims to manufacture and characterize mesoscale whey protein particles (WPP) derived from WPI. Two types of WPP, WPP1 (0.05 mL/min) and WPP2 (0.25 mL/min), were prepared through a multistep approach involving liquid antisolvent (LAS) precipitation, heat treatment, and microfluidization. Liquid antisolvent precipitation was performed by injecting a 20% (wt/vol) WPI dispersion (pH 7) into an ethanol-glycerol mixture (75:25, vol/vol) under laminar flow, followed by heat treatment at 80°C for 20 min as a particle hardening step. This process produced stable fiber- and ribbon-shaped whey protein assemblies (WPA), which served as precursors to WPP. Subsequent microfluidization (150 MPa, 6 passages) reduced the size of WPA, yielding mesoscale WPP with irregular morphologies and a more uniform size distribution, as revealed by microscopy and dynamic light scattering. ζ-Potential and fluorescence labeling indicated higher surface charge and surface hydrophobicity of WPP compared with untreated WPI. The WPP showed internal mass fractal and surface fractal structures at larger length scales, analyzed using small-angle X-ray scattering. Fourier transform infrared spectroscopy demonstrated an increased fraction of intermolecular β-sheets in WPP, suggesting that hydrogen bonding contributed to their formation. Gel electrophoresis confirmed that disulfide bonds served as the primary cross-links stabilizing the WPP structure. Furthermore, turbidity measurements showed that WPP exhibited superior colloidal phase stability compared with untreated WPI and maintained high colloidal stability under both acidic and neutral pH conditions.

Antisolvent precipitation

Lassa virus protein–protein interactions as mediators of Lassa fever pathogenesis

Viral hemorrhagic Lassa fever (LF), caused by Lassa virus (LASV), is a significant public health concern endemic in West Africa with high morbidity and mortality rates, limited treatment options, and potential for international spread. Despite advances in interrogating its epidemiology and clinical manifestations, the molecular mechanisms driving pathogenesis of LASV and other arenaviruses remain incompletely understood. This review synthesizes current knowledge regarding the role of LASV host-virus interactions in mediating the pathogenesis of LF, with emphasis on interactions between viral and host proteins. Through investigation of these critical protein–protein interactions, we identify potential therapeutic targets and discuss their implications for development of medical countermeasures including antiviral drugs. This review provides an update in recent literature of significant LASV host-virus interactions important in informing the development of targeted therapies and improving clinical outcomes for LF patients. Knowledge gaps are highlighted as opportunities for future research efforts that would advance the field of LASV and arenavirus pathogenesis.

60 APPLIED LIFE SCIENCES

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

This dataset accompanies the publication "ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation". This paper introduces a new AI model for protein sequence generation. This dataset contains data related to experiments discussed in the publication. This includes generated sequences and evaluation metrics supporting all unconditional and bias-controlled experiments in the ProtNHF paper.

60 APPLIED LIFE SCIENCES

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES