Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Unfolding of Proteins: Thermal and Mechanical Unfolding

We have employed a Hamiltonian model based on a self-consistent Gaussian appoximation to examine the unfolding process of proteins in external - both mechanical and thermal - force elds. The motivation was to investigate the unfolding pathways of proteins by including only the essence of the important interactions of the native-state topology. Furthermore, if such a model can indeed correctly predict the physics of protein unfolding, it can complement more computationally expensive simulations and theoretical work. The self-consistent Gaussian approximation by Micheletti et al. has been incorporated in our model to make the model mathematically tractable by signi cantly reducing the computational cost. All thermodynamic properties and pair contact probabilities are calculated by simply evaluating the values of a series of Incomplete Gamma functions in an iterative manner. We have compared our results to previous molecular dynamics simulation and experimental data for the mechanical unfolding of the giant muscle protein Titin (1TIT). Our model, especially in light of its simplicity and excellent agreement with experiment and simulation, demonstrates the basic physical elements necessary to capture the mechanism of protein unfolding in an external force field.

Hur, Joe S.

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

Kinetic Studies of Iron Deposition Catalyzed by Recombinant Human Liver Heavy, and Light Ferritins and Azotobacter Vinelandii Bacterioferritin Using O2 and H2O2 as Oxidants

The discrepancy between predicted and measured H2O2 formation during iron deposition with recombinant heavy human liver ferritin (rHF) was attributed to reaction with the iron protein complex [Biochemistry 40 (2001) 10832-10838]. This proposal was examined by stopped-flow kinetic studies and analysis for H2O2 production using (1) rHF, and Azotobacter vinelandii bacterial ferritin (AvBF), each containing 24 identical subunits with ferroxidase centers; (2) site-altered rHF mutants with functional and dysfunctional ferroxidase centers; and (3) rccombinant human liver light ferritin (rLF), containing 110 ferroxidase center. For rHF, nearly identical pseudo-first-order rate constants of 0.18 per second at pH 7.5 were measured for Fe(2+) oxidation by both O2 and H2O2, but for rLF, the rate with O2 was 200-fold slower than that for H2O2 (k-0.22 per second). A Fe(2+)/O2 stoichiometry near 2.4 was measured for rHF and its site altered forms, suggesting formation of H2O2. Direct measurements revealed no H2O2 free in solution 0.5-10 min after all Fe(2+) was oxidized at pH 6.5 or 7.5. These results are consistent with initial H2O2 formation, which rapidly reacts in a secondary reaction with unidentified solution components. Using measured rate constants for rHF, simulations showed that steady-state H2O2 concentrations peaked at 14 pM at approx. 600 ms and decreased to zero at 10-30 s. rLF did not produce measurable H2O2 but apparently conducted the secondary reaction with H2O2. Fe(2+)/O2 values of 4.0 were measured for AvBF. Stopped-flow measurements with AvBF showed that both H2O2 and O2 react at the same rate (k=0.34 per second), that is faster than the reactions with rHF. Simulations suggest that AvBF reduces O2 directly to H2O without intermediate H2O2 formation.

Bunker, Jared

Regulation of sarcomere formation and function in the healthy heart requires a titin intronic enhancer

Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.

Kim, Yuri

Computational Study on Full-length Human Ku70 with Double Stranded DNA: Dynamics, Interactions and Functional Implications

The Ku70/80 heterodimer is the first repair protein in the initial binding of double-strand break (DSB) ends following DNA damage, and is a component of nonhomologous end joining repair, the primary pathway for DSB repair in mammalian cells. In this study we constructed a full-length human Ku70 structure based on its crystal structure, and performed 20 ns conventional molecular dynamic (CMD) simulations on this protein and several other complexes with short DNA duplexes of different sequences. The trajectories of these simulations indicated that, without the topological support of Ku80, the residues in the bridge and C-terminal arm of Ku70 are more flexible than other experimentally identified domains. We studied the two missing loops in the crystal structure and predicted that they are also very flexible. Simulations revealed that they make an important contribution to the Ku70 interaction with DNA. Dislocation of the previously studied SAP domain was observed in several systems, implying its role in DNA binding. Targeted molecular dynamic (TMD) simulation was also performed for one system with a far-away 14bp DNA duplex. The TMD trajectory and energetic analysis disclosed detailed interactions of the DNA-binding residues during the DNA dislocation, and revealed a possible conformational transition for a DSB end when encountering Ku70 in solution. Compared to experimentally based analysis, this study identified more detailed interactions between DNA and Ku70. Free energy analysis indicated Ku70 alone is able to bind DNA with relatively high affinity, with consistent contributions from various domains of Ku70 in different systems. The functional implications of these domains in the processes of Ku heterodimerization and DNA damage recognition and repair can be characterized in detail based upon this analysis.

Hu, Shaowen

Enhancing chemical bioproduction with rational control of bacterial post-translational modifications

Efficient conversion of inexpensive feedstocks to valuable chemicals by microbes is critical for a robust bioeconomy, but the ability to rationally design bacteria is hampered by insufficient knowledge of how post translational modifications (PTMs) control bacterial protein function and thus bioproduction phenotypes. Our study will focus on the lysine acetylation, a ubiquitous bacterial PTM that can affect the function of enzymes in central metabolism that are often critical for bioproduction processes, disrupt transcriptional regulation, and reduce translation. However, most lysine acetylation data is observational, which means that we do not know when, how, and what specific acetylated residues affect protein function and bacterial physiology. For our model host, we will use a Pseudomonas putida strain that we previously engineered to convert lignocellulosic feedstocks into chemicals such as itaconic acid (ITA). With this strain, we use a dynamic two-stage bioproduction process in which ITA is produced during a non-growth associated production phase. Production is highest during growth stages when lysine acetylation is low in other organisms (early stationary phase) and stalls in conditions where acetylation is highest (late stationary phase). The switch from high to stalled ITA production is also correlated with an unexpected increase in acetate levels – the precursor to non-enzymatic lysine acetylation. As such, we predict that lysine acetylation plays a substantial role in regulating the metabolic pathways required for ITA production. We will develop a generalizable approach that combines high-throughput genetic screens and cutting-edge genome engineering with state-of-the-art proteomics, metabolomics, and genetic code expansion methods to identify and modulate lysine acetylation patterns in bacteria. Ultimately, these strategies aim to manipulate protein expression and acetylation patterns to enhance bioproduction phenotypes (e.g., sustained ITA production in late stationary phase).

60 APPLIED LIFE SCIENCES

Mechanosensitive Ion Channels in Bacteria: Functional Domains and Mechanisms of Gating

The past funding period was productive for the group. The progress in the mechanosensitive channel field was critically affected in the end of 1998 by the solution of the crystal structure of the mycobacterial homolog of MscL by our colleagues from Caltech. Having the structure of TbMscL in the closed state, we developed a detailed homology model of EcoMscL, and related the structural model with the wealth of functional phenomenology available for the E. coli version of the channel (EcoMscL). The biophysical properties of the open MscL helped to model the open conformation and infer the pathway for the entire gating transition. The following experiments provided strong support to the atomic model of the gating process, and allowed to make further predictions. The work has advanced our understanding of tension-driven conformational transitions in membrane-embedded mechanosensory proteins, determine major energetic contributions and set the stage for further exploration of the whole family of mechanosensitive channels. The results have been published in seven experimental and theoretical papers, with three other papers currently in press or in preparation.

Sukharev, Sergei

The interplay of DNA repair context with target sequence predictably biases Cas9-generated mutations

Abstract Repair of double-stranded breaks generated by CRISPR/Cas9 is highly dependent on the flanking DNA sequence. To learn about interactions between DNA repair and target sequence, we measure frequencies of over 236,000 distinct Cas9-generated mutational outcomes at over 2800 synthetic target sequences in 18 DNA repair deficient mouse embryonic stem cells lines. We classify the outcomes in an unbiased way, finding a specialised role forPrkdc(DNA-PKcs protein) andPolmin creating 1 bp insertions matching the nucleotide on the protospacer-adjacent motif side of the break, a variable involvement ofNbnandPolqin the creation of different deletion outcomes, and uni-directional deletions dependent on both end-protection and end-resection. Using our dataset, we build predictive models of the mutagenic outcomes of Cas9 scission that outperform the current standards. This work improves our understanding of DNA repair gene function, and provides avenues for more precise modulation of Cas9-generated mutations.

Science & Technology - Other Topics

Ceruloplasmin and cardiovascular disease

Transition metal ion-mediated oxidation is a commonly used model system for studies of the chemical, structural, and functional modifications of low-density lipoprotein (LDL). The physiological relevance of studies using free metal ions is unclear and has led to an exploration of free metal ion-independent mechanisms of oxidation. We and others have investigated the role of human ceruloplasmin (Cp) in oxidative processes because it the principal copper-containing protein in serum. There is an abundance of epidemiological data that suggests that serum Cp may be an important risk factor predicting myocardial infarction and cardiovascular disease. Biochemical studies have shown that Cp is a potent catalyst of LDL oxidation in vitro. The pro-oxidant activity of Cp requires an intact structure, and a single copper atom at the surface of the protein, near His(426), is required for LDL oxidation. Under conditions where inhibitory protein (such as albumin) is present, LDL oxidation by Cp is optimal in the presence of superoxide, which reduces the surface copper atom of Cp. Cultured vascular endothelial and smooth muscle cells also oxidize LDL in the presence of Cp. Superoxide release by these cells is a critical factor regulating the rate of oxidation. Cultured monocytic cells, when activated by zymosan, can oxidize LDL, but these cells are unique in their secretion of Cp. Inhibitor studies using Cp-specific antibodies and antisense oligonucleotides show that Cp is a major contributor to LDL oxidation by these cells. The role of Cp in lipoprotein oxidation and atherosclerotic lesion progression in vivo has not been directly assessed and is an important area for future studies.

NASA Discipline Cell Biology

Investigating Electron Conductivity Regimes in the Bacterial Cytochrome Wire OmcS

The anaerobic bacterium Geobacter sulfurreducens produces extracellular, electronically conductive cytochrome polymer wires that are conductive over micron length scales. Structure models from cryo-electron microscopy data show OmcS wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. However, the mechanism by which this heme arrangement supports long-range electronic conduction remains unknown. Structure models from cryo-electron microscopy data show these wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. Existing computational models using static heme redox potentials and coupling energies fail to explain experimental observations, predicting conductances 10,000 to 100,000 times lower than measured values. Here, we investigate how dynamic disorder affects site energies, interheme coupling, and long-range electronic conductivity within these cytochrome wires. We introduce an approach to extract charge carrier site information directly from Kohn–Sham density functional theory, without employing projector schemes, and show that site and coupling energies are highly sensitive to changes in interheme geometry and the surrounding electrostatic environment. Unlike models that incorporate dynamic disorder as a thermally averaged quantity, our quantum charge carrier model incorporates proxies for dynamic disorder through decoherence corrections, yielding predicted diffusion coefficient closer to what is expected from experiment and comparable with other organic-based electronic materials. Based on these simulations, we propose that the instantaneous fluctuations of the local electrostatic environment can transiently lift energy degeneracies and delocalize charge carriers. Furthermore, these studies reveal how incorporating dynamic fluctuations associated with the environment resolves the discrepancy between theory and experiment in microbial cytochrome wires and highlight design principles for bioinspired, heme-based conductive materials.

Bioinorganic chemistry

The Exoproteome and Surfaceome of Toxigenic Corynebacterium diphtheriae 1737 and Its Response to Iron Restriction and Growth on Human Hemoglobin

Toxin-producing Corynebacterium diphtheriae strains are the etiological agents of the severe upper respiratory disease, diphtheria. A global phylogenetic analysis revealed that biotype gravis is particularly lethal as it produces diphtheria toxin and a range of other virulence factors, particularly when it encounters low levels of iron at sites of infection. Here, to gain insight into how it colonizes its host, we have identified iron-dependent changes in the exoproteome and surfaceome of C. diphtheriae strain 1737 using a combination of whole-cell fractionation, intact cell surface proteolysis, and quantitative proteomics. In total, we identified 1414 of the predicted 2265 proteins (62%) encoded by its reference genome. For each protein, we quantified its degree of secretion and surface exposure, revealing that exoproteases and hydrolases predominate in the exoproteome, while the surfaceome is enriched with adhesins, particularly DIP2093. Our analysis provides insight into how components in the heme-acquisition system are positioned, showing pronounced surface exposure of the strain-specific ChtA/ChtC paralogues and high secretion of the species-conserved heme-binding HtaA protein, suggesting it functions as a hemophore. Profiling the response of the exoproteome and surfaceome after microbial exposure to human hemoglobin and iron limitation reveals potential virulence factors that may be expressed at sites of infection. Data are available via ProteomeXchange with identifier PXD051674.

cell envelope

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Carbon source–driven metabolic and regulatory remodeling defines phenomic states in Lipomyces starkeyi

Lipomyces is a genus of oleaginous yeasts with potential for contributing to reliable biomanufacturing supply chains. However, progress in advanced strain designs and engineering efforts are still constrained by a lack of understanding of the underlying molecular drivers of Lipomyces phenotypes. To address this gap, we collected a suite of multi-omic data to dissect how carbon source availability reshapes the metabolic network, lipid allocation, and regulatory architecture of Lipomyces starkeyi. We observed that glucose promotes biosynthetic and proliferative processes supported by abundant energy and carbon intermediates, xylose enhances redox-balancing mechanisms centered on the pentose phosphate pathway, and glycerol activates respiratory metabolism, ß-oxidation, and the glyoxylate cycle. Lipid species distributions remained consistent in both nitrogen replete and depleted conditions across the carbon sources, indicating robust production mechanisms. Regulatory protein identification and network analysis revealed glycerol-driven respiratory growth favors regulatory programs integrating stress tolerance, redox balance, and lipid-associated metabolism, whereas xylose growth activates compensatory transcriptional responses aimed at maintaining mitochondrial function. Nitrogen limitation modulates the strength of these responses but does not fundamentally alter their direction, reinforcing carbon source as the dominant driver of regulatory architecture. Taken together, this data enhances the understanding of Lipomyces molecular rearrangements and provides a foundation for further development of predictive phenotypic tools in this genus.

Biotechnology

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES

Coupling Microdroplet-Based Sample Preparation, Multiplexed Isobaric Labeling, and Nanoflow Peptide Fractionation for Deep Proteome Profiling of the Tissue Microenvironment

There is increasing interest in developing in-depth proteomic approaches for mapping tissue heterogeneity in a cell-type-specific manner to better understand and predict the function of complex biological systems such as human organs. Existing spatially resolved proteomics technologies cannot provide deep proteome coverage due to limited sensitivity and poor sample recovery. Herein, we seamlessly combined laser capture microdissection with a low-volume sample processing technology that includes a microfluidic device named microPOTS (microdroplet processing in one pot for trace samples), multiplexed isobaric labeling, and a nanoflow peptide fractionation approach. The integrated workflow allowed us to maximize proteome coverage of laser-isolated tissue samples containing nanogram levels of proteins. We demonstrated that the deep spatial proteomics platform can quantify more than 5000 unique proteins from a small-sized human pancreatic tissue pixel (∼60,000 μm2) and differentiate unique protein abundance patterns in pancreas. Furthermore, the use of the microPOTS chip eliminated the requirement for advanced microfabrication capabilities and specialized nanoliter liquid handling equipment, making it more accessible to proteomic laboratories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Modeling Clustered DNA Damage by Ionizing Radiation Using Multinomial Damage Probabilities and Energy Imparted Spectra

Simple and complex clustered DNA damage represent the critical initial damage caused by radiation. In this paper, a multinomial probability model of clustered damage is developed with probabilities dependent on the energy imparted to DNA and surrounding water molecules. The model consists of four probabilities: (A) direct damage of sugar-phosphate moieties leading to SSB, (B) OH− radical formation with subsequent SSB and BD formation, (C) direct damage to DNA bases, and (D) energy imparted to histone proteins and other molecules in a volume not leading to SSB or BD. These probabilities are augmented by introducing probabilities for the relative location of SSB using a ≤10 bp criteria for a double-strand break (DSB) and for the possible success of a radical attack that leads to SSB or BD. Model predictions for electrons, 4He, and 12C ions are compared to the experimental data and show good agreement. Thus, the developed model allows an accurate and rapid computational method to predict simple and complex clustered DNA damage as a function of radiation quality and to explore the resulting challenges to DNA repair.

Biochemistry & Molecular Biology