Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)

De novo atomic protein structure modeling for cryoEM density maps using 3D transformer and HMM

Accurately building 3D atomic structures from cryo-EM density maps is a crucial step in cryo-EM-based protein structure determination. Converting density maps into 3D atomic structures for proteins lacking accurate homologous or predicted structures as templates remains a significant challenge. Here, we introduce Cryo2Struct, a fully automated de novo cryo-EM structure modeling method. Cryo2Struct utilizes a 3D transformer to identify atoms and amino acid types in cryo-EM density maps, followed by an innovative Hidden Markov Model (HMM) to connect predicted atoms and build protein backbone structures. Cryo2Struct produces substantially more accurate and complete protein structural models than the widely used ab initio method Phenix. Additionally, its performance in building atomic structural models is robust against changes in the resolution of density maps and the size of protein structures.

59 BASIC BIOLOGICAL SCIENCES

Artificial intelligence methods for protein structure and interaction prediction: Recent advances and challenges

Recent advances in artificial intelligence have introduced novel methods for high-accuracy prediction of protein tertiary structures, protein complex structures, and interactions between proteins and other biomolecules, such as small molecules and nucleic acids. Such advancements are accelerating biomedical research and the development of new protein design and bioengineering methods among many other important biotechnology applications. Here, in this review, we outline the recent advances in protein-centric biomolecular structure and interaction prediction, highlight some major challenges in the field, and discuss potential directions to address them.

Morehead, Alex [Lawrence Berkeley National Laborat

Signal sequences target enzymes and structural proteins to bacterial microcompartments and are critical for microcompartment formation

ABSTRACT Spatial organization of pathway enzymes has emerged as a promising tool to address several challenges in metabolic engineering, such as flux imbalances and off-target product formation. Bacterial microcompartments (MCPs) are a spatial organization strategy used natively by many bacteria to encapsulate metabolic pathways that produce toxic, volatile intermediates. Several recent studies have focused on engineering MCPs to encapsulate heterologous pathways of interest, but how this engineering affects MCP assembly and function is poorly understood. In this study, we investigated the role of signal sequences, short domains that target proteins to the MCP core, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterized two novel Pdu signal sequences on the structural proteins PduM and PduB, which constitute the first report of metabolosome signal sequences on structural proteins rather than enzymes. We then explored the role of enzymatic and structural Pdu signal sequences on MCP assembly by deleting their encoding sequences from the genome alone and in combination. Deleting enzymatic signal sequences decreased the MCP formation, but this defect could be recovered in some cases by overexpressing genes encoding the knocked-out signal sequence fused to a heterologous protein. By contrast, deleting structural signal sequences caused similar defects to knocking out the genes encoding the full-length PduM and PduB proteins. Our results contribute to a growing understanding of how MCPs form and function in bacteria and provide strategies to mitigate assembly disruption when encapsulating heterologous pathways in MCPs. IMPORTANCE Spatially organizing biosynthetic pathway enzymes is a promising strategy to increase pathway throughput and yield. Bacterial microcompartments (MCPs) are proteinaceous organelles that many bacteria natively use as a spatial organization strategy to encapsulate niche metabolic pathways, providing significant metabolic benefits. Encapsulating heterologous pathways of interest in MCPs could confer these benefits to industrially relevant pathways. Here, we investigate the role of signal sequences, short domains that target proteins for encapsulation in MCPs, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterize two novel signal sequences on structural proteins, constituting the first Pdu signal sequences found on structural proteins rather than enzymes, and perform knockout studies to compare the impacts of enzymatic and structural signal sequences on MCP assembly. Our results demonstrate that enzymatic and structural signal sequences play critical but distinct roles in Pdu MCP assembly and provide design rules for engineering MCPs while minimizing disruption to MCP assembly.

Johnson, Elizabeth R. (ORCID:0000000179236881)

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

Protein Structure Inspired Discovery of a Novel Inducer of Anoikis in Human Melanoma

Drug discovery historically starts with an established function, either that of compounds or proteins. This can hamper discovery of novel therapeutics. As structure determines function, we hypothesized that unique 3D protein structures constitute primary data that can inform novel discovery. Using a computationally intensive physics-based analytical platform operating at supercomputing speeds, we probed a high-resolution protein X-ray crystallographic library developed by us. For each of the eight identified novel 3D structures, we analyzed binding of sixty million compounds. Top-ranking compounds were acquired and screened for efficacy against breast, prostate, colon, or lung cancer, and for toxicity on normal human bone marrow stem cells, both using eight-day colony formation assays. Effective and non-toxic compounds segregated to two pockets. One compound, Dxr2-017, exhibited selective anti-melanoma activity in the NCI-60 cell line screen. In eight-day assays, Dxr2-017 had an IC50 of 12 nM against melanoma cells, while concentrations over 2100-fold higher had minimal stem cell toxicity. Dxr2-017 induced anoikis, a unique form of programmed cell death in need of targeted therapeutics. Our findings demonstrate proof-of-concept that protein structures represent high-value primary data to support the discovery of novel acting therapeutics. This approach is widely applicable.

Oncology

Electrochemically modulated single-molecule localization microscopy for in vitro imaging cytoskeletal protein structures

A new concept of electrochemically modulated single-molecule localization super-resolution imaging is developed. Applications of single-molecule localization super-resolution microscopy have been limited due to insufficient availability of qualified fluorophores with favorable low duty cycles. The key for the new concept is that the “On” state of a redox-active fluorophore with unfavorable high duty cycle could be driven to “Off” state by electrochemical potential modulation and thus become available for single-molecule localization imaging. The new concept was carried out using redox-active cresyl violet with unfavorable high duty cycle as a model fluorophore by synchronizing electrochemical potential scanning with a single-molecule localization microscope. The two cytoskeletal protein structures, the microtubules from porcine brain and the actins from rabbit muscle, were selected as the model target structures for the conceptual imaging in vitro. The super-resolution images of microtubules and actins were obtained from precise single-molecule localizations determined by modulating the On/Off states of single fluorophore molecules on the cytoskeletal proteins via electrochemical potential scanning. Importantly, this method could allow more fluorophores even with unfavorable photophysical properties to become available for a wider and more extensive application of single-molecule localization microscopy.

electrochemical modulation

Structural genomics of bacterial drug targets: Application of a high-throughput pipeline to solve 58 protein structures from pathogenic and related bacteria

Antibiotic resistance remains a leading cause of severe infections worldwide. Small changes in protein sequence can impact antibiotic efficacy. Here, we report deposition of 58 X-ray crystal structures of bacterial proteins that are known targets for antibiotics, which expands knowledge of structural variation to support future antibiotic discovery or modifications.

PDB

Deconvolution of dynamic heterogeneity in protein structure

Heterogeneity is intrinsic to the dynamic process of a chemical reaction. As reactants are converted to products via intermediates, the nature and extent of heterogeneity vary temporally throughout the duration of the reaction and spatially across the molecular ensemble. The goal of many biophysical techniques, including crystallography and spectroscopy, is to establish a reaction trajectory that follows an experimentally provoked dynamic process. It is essential to properly analyze and resolve heterogeneity inevitably embedded in experimental datasets. We have developed a deconvolution technique based on singular value decomposition (SVD), which we have rigorously practiced in diverse research projects. In this review, we recapitulate the motivation and challenges in addressing the heterogeneity problem and lay out the mathematical foundation of our methodology that enables isolation of chemically sensible structural signals. We also present a few case studies to demonstrate the concept and outcome of the SVD-based deconvolution. Finally, we highlight a few recent studies with mechanistic insights made possible by heterogeneity deconvolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato

Generating Protein Structures for Pathway Discovery Using Deep Learning

Resolving the intricate details of biological phenomena at the molecular level is fundamentally limited by both length- and time scales that can be probed experimentally. Molecular dynamics (MD) simulations at various scales are powerful tools frequently employed to offer valuable biological insights beyond experimental resolution. However, while it is relatively simple to observe long-lived, stable configurations of, for example, proteins, at the required spatial resolution, simulating the more interesting rare transitions between such states often takes orders of magnitude longer than what is feasible even on the largest supercomputers available today. One common aspect of this challenge is pathway discovery, where the start and end states of a scientific phenomenon are known or can be approximated, but the mechanistic details in between are unknown. Here, we propose a representation-learning-based solution that uses interpolation and extrapolation in an abstract representation space to synthesize potential transition states, which are automatically validated using MD simulations. The new simulations of the synthesized transition states are subsequently incorporated into the representation learning, leading to an iterative framework for targeted path sampling. Our approach is demonstrated by recovering the transition of a RAS-RAF protein domain (CRD) from membrane-free to interacting with the membrane using coarse-grain MD simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Serial-femtosecond crystallography reveals how a phytochrome variant couples chromophore and protein structural changes

The photoreaction and commensurate structural changes of a chromophore within biological photoreceptors elicit conformational transitions of the protein promoting the switch between deactivated and activated states. We investigated how this coupling is achieved in a bacterial phytochrome variant, Agp2-PAiRFP2. Contrary to classical protein crystallography, which only allows probing (cryo-trapped) stable states, we have used time-resolved serial femtosecond x-ray crystallography (tr-SFX) and pump-probe techniques with various illumination and delay times with respect to photoexcitation of the parent Pfr state. Thus, structural data for seven time frames were sorted into groups of molecular events along the reaction coordinate. They range from chromophore isomerization to the formation of Meta-F, the intermediate that precedes the functional relevant secondary structure transition of the tongue. Structural data for the early events were used to calculate the photoisomerization pathway to complement the experimental data. Late events allow identifying the molecular switch that is linked to the intramolecular proton transfer as a prerequisite for the following structural transitions.

59 BASIC BIOLOGICAL SCIENCES

N-terminal domain swapping: A new paradigm for spermidine/spermine N -acetyltransferase (SSAT) protein structures?

Enterococcus faecalis is a multi-drug-resistant human pathogen that is found in a variety of environments and is challenging to treat. Under stress conditions, some bacteria regulate intracellular polyamine concentrations via polyamine acetyltransferases to reduce their toxicity. The E. faecalis genome encodes two polyamine acetyltransferases: PmvE and BltD. Both of these proteins belong to the Gcn5-related N-acetyltransferase (GNAT) superfamily. It is unclear why there are two enzymes with similar substrate specificities in this organism. To better understand the structure/function relationship of the E. faecalis BltD enzyme, we determined its crystal structure and performed additional assays to explore its oligomeric state and enzymatic activity. The goal was to determine whether there were structural or catalytic differences between this enzyme and other polyamine acetyltransferases that could explain this redundancy and be exploited for future development of targeted inhibitors for this important human pathogen. We found the BltD enzyme was structurally unique due to its N-terminal domain swapped dimer. However, this enzyme adopts a catalytically active monomer rather than dimer in solution. This indicates the crystal structure we obtained may represent a state that forms at high protein and salt concentrations and at low pH used during crystallization. The BltD dimer found in the crystal may represent a unique view of how an inhibitory peptide or molecule could be designed to occupy its active site. Additionally, this structure shows the extensive flexibility of the N-terminal portion of the E. faecalis BltD enzyme.

59 BASIC BIOLOGICAL SCIENCES

AQuaRef: machine learning accelerated quantum refinement of protein structures

Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.

Zubatyuk, Roman [Carnegie Mellon University, Pitts

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH