Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

N-terminal domain swapping: A new paradigm for spermidine/spermine N -acetyltransferase (SSAT) protein structures?

Enterococcus faecalis is a multi-drug-resistant human pathogen that is found in a variety of environments and is challenging to treat. Under stress conditions, some bacteria regulate intracellular polyamine concentrations via polyamine acetyltransferases to reduce their toxicity. The E. faecalis genome encodes two polyamine acetyltransferases: PmvE and BltD. Both of these proteins belong to the Gcn5-related N-acetyltransferase (GNAT) superfamily. It is unclear why there are two enzymes with similar substrate specificities in this organism. To better understand the structure/function relationship of the E. faecalis BltD enzyme, we determined its crystal structure and performed additional assays to explore its oligomeric state and enzymatic activity. The goal was to determine whether there were structural or catalytic differences between this enzyme and other polyamine acetyltransferases that could explain this redundancy and be exploited for future development of targeted inhibitors for this important human pathogen. We found the BltD enzyme was structurally unique due to its N-terminal domain swapped dimer. However, this enzyme adopts a catalytically active monomer rather than dimer in solution. This indicates the crystal structure we obtained may represent a state that forms at high protein and salt concentrations and at low pH used during crystallization. The BltD dimer found in the crystal may represent a unique view of how an inhibitory peptide or molecule could be designed to occupy its active site. Additionally, this structure shows the extensive flexibility of the N-terminal portion of the E. faecalis BltD enzyme.

59 BASIC BIOLOGICAL SCIENCES

AQuaRef: machine learning accelerated quantum refinement of protein structures

Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.

Zubatyuk, Roman [Carnegie Mellon University, Pitts

Multistep modeling of protein structure: application towards refinement of tyr-tRNA synthetase

The scope of multistep modeling (MSM) is expanding by adding a least-squares minimization step in the procedure to fit backbone reconstruction consistent with a set of C-alpha coordinates. The analytical solution of Phi and Psi angles, that fits a C-alpha x-ray coordinate is used for tyr-tRNA synthetase. Phi and Psi angles for the region where the above mentioned method fails, are obtained by minimizing the difference in C-alpha distances between the computed model and the crystal structure in a least-squares sense. We present a stepwise application of this part of MSM to the determination of the complete backbone geometry of the 321 N terminal residues of tyrosine tRNA synthetase to a root mean square deviation of 0.47 angstroms from the crystallographic C-alpha coordinates.

NASA Discipline Exobiology

Phosphorylation of the budgerigar fledgling disease virus major capsid protein VP1

The structural proteins of the budgerigar fledgling disease virus, the first known nonmammalian polyomavirus, were analyzed by isoelectric focusing and sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). The major capsid protein VP1 was found to be composed of at least five distinct species having isoelectric points ranging from pH 6.45 to 5.85. By analogy with the murine polyomavirus, these species apparently result from different modifications of an initial translation product. Primary chicken embryo cells were infected in the presence of 32Pi to determine whether the virus structural proteins were modified by phosphorylation. SDS-PAGE of the purified virus structural proteins demonstrated that VP1 (along with both minor capsid proteins) was phosphorylated. Two-dimensional analysis of the radiolabeled virus showed phosphorylation of only the two most acidic isoelectric species of VP1, indicating that this posttranslational modification contributes to VP1 species heterogeneity. Phosphoamino acid analysis of 32P-labeled VP1 revealed that phosphoserine is the only phosphoamino acid present in the VP1 protein.

Non-NASA Center

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Early events in the folding of an amphipathic peptide: A multinanosecond molecular dynamics study

Folding of the capped LQQLLQQLLQL peptide is investigated at the water-hexane interface by molecular dynamics simulations for 161.5 ns. Initially placed in the aqueous phase as a beta-strand, the peptide rapidly adsorbs to the interface, where it adopts an amphipathic conformation. The marginal presence of nonamphipathic structures throughout the complete trajectory indicates that the corresponding conformations are strongly disfavored at the interface. It is further suggestive that folding in an interfacial environment proceeds through a pathway of successive amphipathic intermediates. The energetic and entropic penalties involved in the conformational changes along this pathway markedly increase the folding time scales of LQQLLQQLLQL, explaining why the alpha-helix, the hypothesized lowest free energy structure for a sequence with a hydrophobic periodicity of 3.6, has not been reached yet. The formation of a type I beta-turn at the end of the simulation confirms the importance of such motifs as initiation sites allowing the peptide to coalesce towards a secondary structure. Proteins 1999;36:383-399. Copyright 1999 Wiley-Liss, Inc.

NASA Center ARC

Composition of the carbohydrate granules of the cyanobacterium, Cyanothece sp. strain ATCC 51142

Cyanothece sp. strain ATCC 51142 is an aerobic, unicellular, diazotrophic cyanobacterium that temporally separates O2-sensitive N2 fixation from oxygenic photosynthesis. The energy and reducing power needed for N2 fixation appears to be generated by an active respiratory apparatus that utilizes the contents of large interthylakoidal carbohydrate granules. We report here on the carbohydrate and protein composition of the granules of Cyanothece sp. strain ATCC 51142. The carbohydrate component is a glucose homopolymer with branches every nine residues and is chemically identical to glycogen. Granule-associated protein fractions showed temporal changes in the number of proteins and their abundance during the metabolic oscillations observed under diazotrophic conditions. There also were temporal changes in the protein pattern of the granule-depleted supernatant fractions from diazotrophic cultures. None of the granule-associated proteins crossreacted with antisera directed against several glycogen-metabolizing enzymes or nitrogenase, although these proteins were tentatively identified in supernatant fractions. It is suggested that the granule-associated proteins are structural proteins required to maintain a complex granule architecture.

NASA Discipline Cell Biology

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES

Characterization of the DNA binding properties of polyomavirus capsid protein

The DNA binding properties of the polyomavirus structural proteins VP1, VP2, and VP3 were studied by Southwestern analysis. The major viral structural protein VP1 and host-contributed histone proteins of polyomavirus virions were shown to exhibit DNA binding activity, but the minor capsid proteins VP2 and VP3 failed to bind DNA. The N-terminal first five amino acids (Ala-1 to Lys-5) were identified as the VP1 DNA binding domain by genetic and biochemical approaches. Wild-type VP1 expressed in Escherichia coli (RK1448) exhibited DNA binding activity, but the N-terminal truncated VP1 mutants (lacking Ala-1 to Lys-5 and Ala-1 to Cys-11) failed to bind DNA. The synthetic peptide (Ala-1 to Cys-11) was also shown to have an affinity for DNA binding. Site-directed mutagenesis of the VP1 gene showed that the point mutations at Pro-2, Lys-3, and Arg-4 on the VP1 molecule did not affect DNA binding properties but that the point mutation at Lys-5 drastically reduced DNA binding affinity. The N-terminal (Ala-1 to Lys-5) region of VP1 was found to be essential and specific for DNA binding, while the DNA appears to be non-sequence specific. The DNA binding domain and the nuclear localization signal are located in the same N-terminal region.

NASA Discipline Cell Biology

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES