Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

De novo design of protein structure and function with RFdiffusion

Abstract There has been considerable recent progress in designing new proteins using deep-learning methods 1–9 . Despite this progress, a general deep-learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher-order symmetric architectures, has yet to be described. Diffusion models 10,11 have had considerable success in image and language generative modelling but limited success when applied to protein modelling, probably due to the complexity of protein backbone geometry and sequence–structure relationships. Here we show that by fine-tuning the RoseTTAFold structure prediction network on protein structure denoising tasks, we obtain a generative model of protein backbones that achieves outstanding performance on unconditional and topology-constrained protein monomer design, protein binder design, symmetric oligomer design, enzyme active site scaffolding and symmetric motif scaffolding for therapeutic and metal-binding protein design. We demonstrate the power and generality of the method, called RoseTTAFold diffusion (RFdiffusion), by experimentally characterizing the structures and functions of hundreds of designed symmetric assemblies, metal-binding proteins and protein binders. The accuracy of RFdiffusion is confirmed by the cryogenic electron microscopy structure of a designed binder in complex with influenza haemagglutinin that is nearly identical to the design model. In a manner analogous to networks that produce images from user-specified inputs, RFdiffusion enables the design of diverse functional proteins from simple molecular specifications.

Science & Technology - Other Topics↗

Discovery of TAK-981, a First-in-Class Inhibitor of SUMO-Activating Enzyme for the Treatment of Cancer

SUMOylation is a reversible post-translational modification that regulates protein function through covalent attachment of small ubiquitin-like modifier (SUMO) proteins. The process of SUMOylating proteins involves an enzymatic cascade, the first step of which entails the activation of a SUMO protein through an ATP-dependent process catalyzed by SUMO-activating enzyme (SAE). Here, we describe the identification of TAK-981, a mechanism-based inhibitor of SAE which forms a SUMO–TAK-981 adduct as the inhibitory species within the enzyme catalytic site. Optimization of selectivity against related enzymes as well as enhancement of mean residence time of the adduct were critical to the identification of compounds with potent cellular pathway inhibition and ultimately a prolonged pharmacodynamic effect and efficacy in preclinical tumor models, culminating in the identification of the clinical molecule TAK-981.

60 APPLIED LIFE SCIENCES↗

Quantifying shifts in natural selection on codon usage between protein regions: a population genetics approach

Codon usage bias (CUB), the non-uniform usage of synonymous codons, occurs across all domains of life. Adaptive CUB is hypothesized to result from various selective pressures, including selection for efficient ribosome elongation, accurate translation, mRNA secondary structure, and/or protein folding. Given the critical link between protein folding and protein function, numerous studies have analyzed the relationship between codon usage and protein structure. The results from these studies have often been contradictory, likely reflecting the differing methods used for measuring codon usage and the failure to appropriately control for confounding factors, such as differences in amino acid usage between protein structures and changes in the frequency of different structures with gene expression. Here we take an explicit population genetics approach to quantify codon-specific shifts in natural selection related to protein structure in S. cerevisiae and E. coli. Unlike other metrics of codon usage, our approach explicitly separates the effects of natural selection, scaled by gene expression, and mutation bias while naturally accounting for a region’s amino acid usage. Bayesian model comparisons suggest selection on codon usage varies only slightly between helix, sheet, and coil secondary structures and, similarly, between structured and intrinsically-disordered regions. Similarly, in contrast to previous findings, we find selection on codon usage only varies slightly at the termini of helices in E. coli. Using simulated data, we show this previous work indicating “non-optimal” codons are enriched at the beginning of helices in S. cerevisiae was due to failure to control for various confounding factors (e.g. amino acid biases, gene expression, etc.), and rather than selection to modulate cotranslational folding. Our results reveal a weak relationship between codon usage and protein structure, indicating that differences in selection on codon usage between structures are slight. In addition to the magnitude of differences in selection between protein structures being slight, the observed shifts appear to be idiosyncratic and largely codon-specific rather than systematic reversals in the nature of selection. Overall, our work demonstrates the statistical power and benefits of studying selective shifts on codon usage or other genomic features from an explicitly evolutionary approach. Limitations of this approach and future potential research avenues are discussed.

59 BASIC BIOLOGICAL SCIENCES↗

Opportunities and challenges for assigning cofactors in cryo-EM density maps of chlorophyll-containing proteins

Abstract The accurate assignment of cofactors in cryo-electron microscopy maps is crucial in determining protein function. This is particularly true for chlorophylls (Chls), for which small structural differences lead to important functional differences. Recent cryo-electron microscopy structures of Chl-containing protein complexes exemplify the difficulties in distinguishing Chl b and Chl f from Chl a . We use these structures as examples to discuss general issues arising from local resolution differences, properties of electrostatic potential maps, and the chemical environment which must be considered to make accurate assignments. We offer suggestions for how to improve the reliability of such assignments.

59 BASIC BIOLOGICAL SCIENCES↗

Challenges in solving structures from radiation-damaged tomograms of protein nanocrystals assessed by simulation

Structure-determination methods are needed to resolve the atomic details that underlie protein function. X-ray crystallography has provided most of our knowledge of protein structure, but is constrained by the need for large, well ordered crystals and the loss of phase information. The rapidly developing methods of serial femtosecond crystallography, micro-electron diffraction and single-particle reconstruction circumvent the first of these limitations by enabling data collection from nanocrystals or purified proteins. However, the first two methods also suffer from the phase problem, while many proteins fall below the molecular-weight threshold required for single-particle reconstruction. Cryo-electron tomography of protein nanocrystals has the potential to overcome these obstacles of mainstream structure-determination methods. Here, a data-processing scheme is presented that combines routines from X-ray crystallography and new algorithms that have been developed to solve structures from tomograms of nanocrystals. This pipeline handles image-processing challenges specific to tomographic sampling of periodic specimens and is validated using simulated crystals. The tolerance of this workflow to the effects of radiation damage is also assessed. The simulations indicate a trade-off between a wider tilt range to facilitate merging data from multiple tomograms and a smaller tilt increment to improve phase accuracy. Since phase errors, but not merging errors, can be overcome with additional data sets, these results recommend distributing the dose over a wide angular range rather than using a finer sampling interval to solve the protein structure.

59 BASIC BIOLOGICAL SCIENCES↗

DiMER

SAND2025-04145O DiMER is a Python based tool that helps researchers understand the functions of genes by searching through multiple biological databases. It takes user-provided data and scans various databases to find the best matches for gene functions, generating a clear summary of results. DiMER identifies the most relevant functional annotations and improves upon previous annotations by replacing instances of "unknown protein function" with more accurate descriptions. DiMER requires minimal setup. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mageeney, Catherine [Sandia National Lab. (SNL-CA)↗

Genome-wide identification and functional prediction of silicon (Si) transporters in poplar (Populus trichocarpa)

Abstract Silicon (Si) enhances plant tolerance to various biotic and abiotic stressors such as salinity, drought, and heat. In addition, Si can be biomineralized within plants to form organic carbon-containing phytoliths that can have ecosystem-level consequences by contributing to long-term carbon sequestration. Si is taken up and transported in plants via different transporter proteins such as influx transporters (e.g., Lsi1, Lsi6) and efflux transporters (e.g., Lsi2). Additionally, the imported Si can be deposited in plant leaves via silicification process using the Siliplant 1 (e.g., Slp1) protein. Functional homologs of these proteins have been reported in different food crops. Here, we performed a genome-wide analysis to identify different Si transporters and Slp1 homologs in the bioenergy crop poplar ( Populus trichocarpa Torr. and A. Gray ex W. Hook). We identified one channel-type Si influx transporter (PtLsi1; Potri.017G083300), one Si efflux transporter (PtLsi2; Potri.012G144000) and two proteins like Slp1 (PtSlp1a; Potri.004G168600 and PtSlp1b; Potri.009G129900 ) in the P. trichocarpa genome. We found a unique sequence (KPKPPVFKPPPVPI) in PtSlp1a which is repeated six times. Repeated presence of this sequence in PtSlp1a indicates that this protein might be important for silicification processes in P. trichocarpa. The mutation profiles of different Si transporters in a P. trichocarpa genome-wide association study population identified significant and impactful mutations in Potri.004G168600 and Potri.009G129900 . Using a publically accessible database ( http://bar.utoronto.ca/eplant_poplar/ ), digital expression analysis of the putative Si transporters in P. trichocarpa found low to moderate expression in the anticipated tissues, such as roots and leaves. Subcellular localization analysis found that PtLsi1/PtLsi2 are localized in the plasma membrane, whereas PtSlp1a/PtSlp1b are found in the extracellular spaces. Protein–Protein interaction analysis of PtLsi1/PtLsi2 identified Delta-1-pyrroline-5-carboxylate synthase (P5CS) as one of the main interacting partners of PtLsi2, which plays a key role in proline biosynthesis. Proline is a well-known participant in biotic and abiotic stress tolerance in plants. These findings will reinforce future efforts to modify Si accumulation for enhancing plant stress tolerance and carbon sequestration in poplar.

59 BASIC BIOLOGICAL SCIENCES↗

Role of backbone strain in de novo design of complex α/β protein structures

Abstract We previously elucidated principles for designing ideal proteins with completely consistent local and non-local interactions which have enabled the design of a wide range of new αβ-proteins with four or fewer β-strands. The principles relate local backbone structures to supersecondary-structure packing arrangements of α-helices and β-strands. Here, we test the generality of the principles by employing them to design larger proteins with five- and six- stranded β-sheets flanked by α-helices. The initial designs were monomeric in solution with high thermal stability, and the nuclear magnetic resonance (NMR) structure of one was close to the design model, but for two others the order of strands in the β-sheet was swapped. Investigation into the origins of this strand swapping suggested that the global structures of the design models were more strained than the NMR structures. We incorporated explicit consideration of global backbone strain into the design methodology, and succeeded in designing proteins with the intended unswapped strand arrangements. These results illustrate the value of experimental structure determination in guiding improvement of de novo design, and the importance of consistency between local, supersecondary, and global tertiary interactions in determining protein topology. The augmented set of principles should inform the design of larger functional proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Core cysteine residues in the Plasminogen-Apple-Nematode (PAN) domain are critical for HGF/c-MET signaling

The Plasminogen-Apple-Nematode (PAN) domain, with a core of four to six cysteine residues, is found in > 28,000 proteins across 959 genera. Still, its role in protein function is not fully understood. The PAN domain was initially characterized in numerous proteins, including HGF. Dysregulation of HGF-mediated signaling results in multiple deadly cancers. The binding of HGF to its cell surface receptor, c-MET, triggers all biological impacts. Here, we show that mutating four core cysteine residues in the HGF PAN domain reduces c-MET interaction, subsequent c-MET autophosphorylation, and phosphorylation of its downstream targets, perinuclear localization, cellular internalization of HGF, and its receptor, c-MET, and c-MET ubiquitination. Furthermore, transcriptional activation of HGF/c-MET signaling-related genes involved in cancer progression, invasion, metastasis, and cell survival were impaired. Thus, targeting the PAN domain of HGF may represent a mechanism for selectively regulating the binding and activation of the c-MET pathway.

59 BASIC BIOLOGICAL SCIENCES↗

Potential pathogenicity determinants identified from structural proteomics of SARS-CoV and SARS-CoV-2

Despite SARS-CoV and SARS-CoV-2 being equipped with highly similar protein arsenals, the corresponding zoonoses have spread among humans at extremely different rates. The specific characteristics of these viruses that led to such distinct outcomes remain unclear. Here, we apply proteome-wide comparative structural analysis aiming to identify the unique molecular elements in the SARS-CoV-2 proteome that may explain the differing consequences. By combining protein modeling and molecular dynamics simulations, we suggest non-conservative substitutions in functional regions of the spike glycoprotein (S), nsp1, and nsp3 that are contributing to differences in virulence. Particularly, we explain why the substitutions at the receptor-binding domain of S affect the structure-dynamics behavior in complexes with putative host receptors. Conservation of functional protein regions within the two taxa is also noteworthy. We suggest that the highly conserved main protease, nsp5, of SARS-CoV and SARS-CoV-2 is part of their mechanism of circumventing the host interferon antiviral response. Overall, most substitutions occur on the protein surfaces and may be modulating their antigenic properties and interactions with other macromolecules. Our results imply that the striking difference in the pervasiveness of SARS-CoV-2 and SARS-CoV among humans seems to significantly derive from molecular features that modulate the efficiency of viral particles in entering the host cells and blocking the host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted Quantification of Protein Phosphorylation and Its Contributions towards Mathematical Modeling of Signaling Pathways

Post-translational modifications (PTMs) are key regulatory mechanisms that can control protein function. Of these, phosphorylation is the most common and widely studied. Because of its importance in regulating cell signaling, precise and accurate measurements of protein phosphorylation across wide dynamic ranges are crucial to understanding how signaling pathways function. Although immunological assays are commonly used to detect phosphoproteins, their lack of sensitivity, specificity, and selectivity often make them unreliable for quantitative measurements of complex biological samples. Recent advances in Mass Spectrometry (MS)-based targeted proteomics have made it a more useful approach than immunoassays for studying the dynamics of protein phosphorylation. Selected reaction monitoring (SRM)—also known as multiple reaction monitoring (MRM)—and parallel reaction monitoring (PRM) can quantify relative and absolute abundances of protein phosphorylation in multiplexed fashions targeting specific pathways. In addition, the refinement of these tools by enrichment and fractionation strategies has improved measurement of phosphorylation of low-abundance proteins. The quantitative data generated are particularly useful for building and parameterizing mathematical models of complex phospho-signaling pathways. Potentially, these models can provide a framework for linking analytical measurements of clinical samples to better diagnosis and treatment of disease.

mathematical modeling↗

De novo design of D-peptide ligands: Application to influenza virus hemagglutinin

D-peptides hold great promise as therapeutics by alleviating the challenges of metabolic stability and immunogenicity in L-peptides. However, current D-peptide discovery methods are severely limited by specific size, structure, and the chemical synthesizability of their protein targets. Here, we describe a computational method for de novo design of D-peptides that bind to an epitope of interest on the target protein using Rosetta’s hotspot-centric approach. The approach comprises identifying hotspot sidechains in a functional protein–protein interaction and grafting these side chains onto much smaller structured peptide scaffolds of opposite chirality. The approach enables more facile design of D-peptides and its applicability is demonstrated by design of D-peptidic binders of influenza A virus hemagglutinin, resulting in identification of multiple D-peptide lead series. The X-ray structure of one of the leads at 2.38 Å resolution verifies the validity of the approach. This method should be generally applicable to targets with detailed structural information, independent of molecular size, and accelerate development of stable, peptide-based therapeutics.

Science & Technology - Other Topics↗

Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15

Abstract Estimating the accuracy of quaternary structural models of protein complexes and assemblies (EMA) is important for predicting quaternary structures and applying them to studying protein function and interaction. The pairwise similarity between structural models is proven useful for estimating the quality of protein tertiary structural models, but it has been rarely applied to predicting the quality of quaternary structural models. Moreover, the pairwise similarity approach often fails when many structural models are of low quality and similar to each other. To address the gap, we developed a hybrid method (MULTICOM_qa) combining a pairwise similarity score (PSS) and an interface contact probability score (ICPS) based on the deep learning inter‐chain contact prediction for estimating protein complex model accuracy. It blindly participated in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 and performed very well in estimating the global structure accuracy of assembly models. The average per‐target correlation coefficient between the model quality scores predicted by MULTICOM_qa and the true quality scores of the models of CASP15 assembly targets is 0.66. The average per‐target ranking loss in using the predicted quality scores to rank the models is 0.14. It was able to select good models for most targets. Moreover, several key factors (i.e., target difficulty, model sampling difficulty, skewness of model quality, and similarity between good/bad models) for EMA are identified and analyzed. The results demonstrate that combining the multi‐model method (PSS) with the complementary single‐model method (ICPS) is a promising approach to EMA.

59 BASIC BIOLOGICAL SCIENCES↗

A glycoprotein B-neutralizing antibody structure at 2.8 Å uncovers a critical domain for herpesvirus fusion initiation

Members of the Herpesviridae, including the medically important alphaherpesvirus varicella-zoster virus (VZV), induce fusion of the virion envelope with cell membranes during entry, and between cells to form polykaryocytes in infected tissues. The conserved glycoproteins, gB, gH and gL, are the core functional proteins of the herpesvirus fusion complex. gB serves as the primary fusogen via its fusion loops, but functions for the remaining gB domains remain unexplained. As a pathway for biological discovery of domain function, our approach used structure-based analysis of the viral fusogen together with a neutralizing antibody. We report here a 2.8 Å cryogenic-electron microscopy structure of native gB recovered from VZV-infected cells, in complex with a human monoclonal antibody, 93k. This high-resolution structure guided targeted mutagenesis at the gB-93k interface, providing compelling evidence that a domain spatially distant from the gB fusion loops is critical for herpesvirus fusion, revealing a potential new target for antiviral therapies. Herpesvirus virions have an outer lipid membrane dotted with glycoproteins that enable fusion with cell membranes to initiate entry and establish infection. Here the authors elucidate the structural mechanism of a neutralizing antibody derived from a patient infected by the herpesvirus varicella-zoster virus and targeted to its fusogen, glycoprotein-B.

59 BASIC BIOLOGICAL SCIENCES↗

Architecture of the human G-protein-methylmalonyl-CoA mutase nanoassembly for B 12 delivery and repair

G-proteins function as molecular switches to power cofactor translocation and confer fidelity in metal trafficking. The G-protein, MMAA, together with MMAB, an adenosyltransferase, orchestrate cofactor delivery and repair of B 12 -dependent human methylmalonyl-CoA mutase (MMUT). The mechanism by which the complex assembles and moves a >1300 Da cargo, or fails in disease, are poorly understood. Herein, we report the crystal structure of the human MMUT-MMAA nano-assembly, which reveals a dramatic 180° rotation of the B 12 domain, exposing it to solvent. The complex, stabilized by MMAA wedging between two MMUT domains, leads to ordering of the switch I and III loops, revealing the molecular basis of mutase-dependent GTPase activation. The structure explains the biochemical penalties incurred by methylmalonic aciduria-causing mutations that reside at the MMAA-MMUT interfaces we identify here.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Novel nucleocytoplasmic protein O -fucosylation by SPINDLY regulates diverse developmental processes in plants

Here, in metazoans, protein O-fucosylation of Ser/Thr residues was only found in secreted or cell surface proteins, and this post-translational modification is catalyzed by ER-localized protein O-fucosyltransferases (POFUTs) in the GT65 family. Recently, a novel nucleocytoplasmic POFUT, SPINDLY (SPY), was identified in the reference plant Arabidopsis thaliana to modify nuclear transcription regulators DELLAs, revealing a new regulatory mechanism for gene expression. The paralog of AtSPY, SECRET AGENT (SEC), is an O-link-N-acetylglucosamine (GlcNAc) transferase (OGT), which O-GlcNAcylates Ser/Thr residues of target proteins. Both AtSPY and AtSEC are tetratricopeptide repeat-domain-containing glycosyltransferases in the GT41 family. The discovery that AtSPY is a POFUT clarified decades of miss-classification of AtSPY as an OGT. SPY and SEC play pleiotropic roles in plant development, and the interactions between SPY and SEC are complex. SPY-like genes are conserved in diverse organisms, except in fungi and metazoans, suggesting that O-fucosylation is a common mechanism in modulating intracellular protein functions.

59 BASIC BIOLOGICAL SCIENCES↗