Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Extracellular matrix controls tubulin monomer levels in hepatocytes by regulating protein turnover

Cells have evolved an autoregulatory mechanism to dampen variations in the concentration of tubulin monomer that is available to polymerize into microtubules (MTs), a process that is known as tubulin autoregulation. However, thermodynamic analysis of MT polymerization predicts that the concentration of free tubulin monomer must vary if MTs are to remain stable under different mechanical loads that result from changes in cell adhesion to the extracellular matrix (ECM). To determine how these seemingly contradictory regulatory mechanisms coexist in cells, we measured changes in the masses of tubulin monomer and polymer that resulted from altering cell-ECM contacts. Primary rat hepatocytes were cultured in chemically defined medium on bacteriological petri dishes that were precoated with different densities of laminin (LM). Increasing the LM density from low to high (1-1000 ng/cm2), promoted cell spreading (average projected cell area increased from 1200 to 6000 microns2) and resulted in formation of a greatly extended MT network. Nevertheless, the steady-state mass of tubulin polymer was similar at 48 h, regardless of cell shape or ECM density. In contrast, round hepatocytes on low LM contained a threefold higher mass of tubulin monomer when compared with spread cells on high LM. Furthermore, similar results were obtained whether LM, fibronectin, or type I collagen were used for cell attachment. Tubulin autoregulation appeared to function normally in these cells because tubulin mRNA levels and protein synthetic rates were greatly depressed in round cells that contained the highest level of free tubulin monomer. However, the rate of tubulin protein degradation slowed, causing the tubulin half-life to increase from approximately 24 to 55 h as the LM density was lowered from high to low and cell rounding was promoted. These results indicate that the set-point for the tubulin monomer mass in hepatocytes can be regulated by altering the density of ECM contacts and changing cell shape. This finding is consistent with a mechanism of MT regulation in which the ECM stabilizes MTs by both accepting transfer of mechanical loads and altering tubulin degradation in cells that continue to autoregulate tubulin synthesis.

Non-NASA Center↗

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Insights into Mechanisms Underlying Mitochondrial and Bacterial Cytochrome c Synthases

Mitochondrial holocytochrome c synthase (HCCS) is an essential protein in assembling cytochrome c (cyt c) of the electron transport system. HCCS binds heme and covalently attaches the two vinyls of heme to two cysteine thiols of the cyt c CXXCH motif. Human HCCS recognizes both cyt c and cytochrome c1 of complex III (cytochrome bc1). HCCS is mutated in some human diseases and it has been investigated recombinantly by mutational, biochemical, and reconstitution studies in the past decade. Here, we employ structural prediction programs (e.g., AlphaFold 3) on HCCS and its two substrates, heme and cytochrome c. The results, when combined with spectroscopic and functional analyses of HCCS and variants, provide insights into the structural basis for heme binding, apocyt c binding, covalent attachment, and release of the holocyt c product. Results from in vitro reconstitution of purified human HCCS using cyt c and cyt c1 peptides as acceptors are consistent with the structural modeling of substrate binding. Reconstitution of HCCS and cyt c1 provides an approach to studying cyt c1 assembly, which has been refractile to recombinant in vivo reconstitution (unlike HCCS and cyt c). We propose a structural basis for release of the holocyt c product from HCCS based on in vitro studies and on cryoEM structures of the bacterial cyt c synthase (CcsBA) active site. We analyze the kinetoplastid mitochondrial synthase (KCCS), and hypothesize a molecular evolutionary path from mitochondrial endosymbiosis to the current HCCS.

Biochemistry & Molecular Biology↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗

Novel protein crystal growth technology: Proof of concept

A technology for crystal growth, which overcomes certain shortcomings of other techniques, is developed and its applicability to proteins is examined. There were several unknowns to be determined: the design of the apparatus for suspension of crystals of varying (growing) diameter, control of the temperature and supersaturation, the methods for seeding and/or controlling nucleation, the effect on protein solutions of the temperature oscillations arising from the circulation, and the effect of the fluid shear on the suspended crystals. Extensive effort was put forth to grow lysozyme crystals. Under conditions favorable to the growth of tetragonal lysozyme, spontaneous nucleation could be produced but the number of nuclei could not be controlled. Seed transfer techniques were developed and implemented. When conditions for the orthorhombic form were tried, a single crystal 1.5 x 0.5 x 0.2 mm was grown (after in situ nucleation) and successfully extracted. A mathematical model was developed to predict the flow velocity as a function of the geometry and the operating temperatures. The model can also be used to scaleup the apparatus for growing larger crystals of other materials such as water soluble non-linear optical materials. This crystal suspension technology also shows promise for high quality solution growth of optical materials such as TGS and KDP.

Nyce, Thomas A.↗

Knocking out the carboxyltransferase interactor 1 (CTI1) in Chlamydomonas boosted oil content by fivefold without affecting cell growth

Summary The first step in chloroplast de novo fatty acid synthesis is catalysed by acetyl‐CoA carboxylase (ACCase). As the rate‐limiting step for this pathway, ACCase is subject to both positive and negative regulation. In this study, we identify a Chlamydomonas homologue of the plant carboxyltransferase interactor 1 (CrCTI1) and show that this protein interacts with the Chlamydomonas α‐carboxyltransferase (Crα‐CT) subunit of the ACCase by yeast two‐hybrid protein–protein interaction assay. Three independent CRISPR‐Cas9 mediated knockout mutants for CrCTI1 each produced an ‘enhanced oil’ phenotype, accumulating 25% more total fatty acids and storing up to fivefold more triacylglycerols (TAGs) in lipid droplets. The TAG phenotype of the crcti1 mutants was not influenced by light but was affected by trophic growth conditions. By growing cells under heterotrophic conditions, we observed a crucial function of CrCTI1 in balancing lipid accumulation and cell growth. Mutating a previously mapped in vivo phosphorylation site (CrCTI1 Ser108 to either Ala or to Asp), did not affect the interaction with Crα‐CT. However, mutating all six predicted phosphorylation sites within Crα‐CT to create a phosphomimetic mutant reduced this pairwise interaction significantly. Comparative proteomic analyses of the crcti1 mutants and WT suggested a role for CrCTI1 in regulating carbon flux by coordinating carbon metabolism, antioxidant and fatty acid β‐oxidation pathways, to enable cells to adapt to carbon availability. Taken together, this study identifies CrCTI1 as a negative regulator of fatty acid synthesis in algae and provides a new molecular brick for the genetic engineering of microalgae for biotechnology purposes.

Li, Zhongze [Aix‐Marseille Université, CEA, CNRS, ↗

In-Flight Personalized Medication Management

Current medication selection for treatment of astronauts during spaceflight missions is primarily dictated by the task of efficiently treating the widest possible range of physiological conditions and illnesses with a limited set of medications. Dosage and recommendations on the combination of drugs are based on the assumption of genetically equal drug sensitivity and unchanged metabolism. To our knowledge, there was no pre-flight drug sensitivity testing on a genetic level for any of the previous manned NASA space missions. Although many of the common, binary drug-drug interactions are, most likely, already considered in the ISS Medical kit composition, multi-drug and multi-drug-gene factors are not incorporated in the medication selection or prescription. Furthermore, due to the physiological changes occurring in microgravity environments, astronauts might be susceptible to potential increased drug toxicity as a result of decreased clearance of numerous drugs. In particular, perturbation of CYP450 enzymes which contribute to the hepatic metabolism of the majority of drugs may have significant effects on therapeutic efficacy and increase treatment-related toxicity5. The genes encoding the CYP450 enzymes are highly variable in humans. Inheritable variations of CYP450 hepatic metabolizer enzymes and transport proteins play a crucial role in the inter-individual variability of drug efficiency and risks of adverse drug reactions5. Additionally, there are some reports that document changes in the levels of production of drug-metabolizing enzymes in microgravity. These data can be extrapolated to provide reasonable assumptions of decreased levels of expression for most CYP450 enzymes in human body during prolonged space travel. If the prescribed medication regiment is not fully effective or causes undesirable side effects, the ability of the astronauts to function and maintain peak performance levels during space flight could be seriously compromised. Therefore, technologies capable of predicting and managing medication side effects, interactions, and toxicity of drugs during spaceflight are needed. We propose to develop and customize for NASAs applications available on the market Personalized Prescribing System (PPS) that would provide a comprehensive, non-invasive solution for safer, targeted medication management for every crew member resulting in safer and more effective treatment and, consequently, better performance. PPS will function as both decision support and record-keeping tool for flight surgeons and astronauts in applying the recommended medications for situations arising in flight. The information on individual drug sensitivity will translate into personalized risk assessment for adverse drug reactions and treatment failures for each drug from the medication kit as well as predefined outcome of any combination of them. Dosage recommendations will also be made individually. The mobile app will facilitate ease of use by crew and medical professionals during training and flight missions.

Personalized Medication↗

Uncovering Sequence and Structural Characteristics of Fungal Expansin‐Related Proteins With Potential to Drive Substrate Targeting

Expansins loosen plant cell wall networks through disrupting non-covalent bonds between cellulose microfibrils and matrix polysaccharides. Whereas expansins were first discovered in plants, expansin-related proteins have since been identified in bacteria and fungi. The biological function of microbial expansins remains unclear; however, several studies have shown distinct binding preferences toward different structural polysaccharides. Earlier studies of bacterial expansin-related proteins uncovered sequence and structural features that correlate to substrate binding. Herein, 20 fungal expansin-related sequences were recombinantly produced in Komagataella phaffii, and the purified proteins were compared in terms of substrate binding to cellulosic and chitinous substrates. The impact of pH on the zeta potential of prioritized substrates was also measured, and Principal Component Analysis was performed to uncover correlations between protein characteristics (e.g., pI, hydrophobicity, surface charge distribution) and measured substrate binding preferences. Whereas acidic proteins with a predicted pI less than 5.0 preferentially bound to chitin, basic proteins with pI greater than 8.0 preferentially bound to xylan and xylan-containing fiber. Similar to many cellulases, binding to cellulose was correlated to relatively high aromatic amino acid content in the protein sequence and presence of a carbohydrate binding module (CBM), which in the case of expansins is a C-terminal CBM63. Whereas overall sequence characteristics could be correlated to substrate binding preference, the identity of amino acids occupying conserved positions that impact protein activity was better correlated with loosenin versus expansin classifications.

chitin↗

Integrating Large Scale Data Sets to Develop Predictive Hypotheses of Low-Dose Radiation-Induced Health Effects

Over one hundred years of radiation biology research has revealed much about the DNA damages induced by the deposition of energy from exposure to ionizing radiation and the subsequent cellular responses. However, there are still significant gaps in our understanding of how these might lead to detrimental health effects, particularly at low doses (100 mGy (milligray)). Recent advances in high throughput omics technologies enable interrogation of induced radiation effects at the genomic, proteomic and metabolomic levels. These include changes in gene expression, protein modifications, e.g., phosphorylation, acetylation, and methylation, and metabolic changes. We will discuss the integration of data obtained from multiple omics platforms to understand radiation dose, and dose rate effects in a complex human tissue model as a function of time. We will use as an example our results on the low dose responses in a 3D human skin model.

ionizing radiation↗

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES↗

Structural Heterogeneity and Hydrodynamics of an Intrinsically Disordered Protein Condensate

Biology demonstrates precise control over the free-energy landscape through the selective partitioning of biomacromolecules into membraneless organelles, enabling essential functions such as biochemical transformations, signaling cascades, and mechanical reinforcement. Although the function of these condensates depends on their underlying structure and hydrodynamics, molecular-scale information on these systems remains sparse. Here, in this study, neutron scattering is used to probe the organization and dynamics of the intrinsically disordered N-terminal domain of Galectin-3, an extracellular lectin responsible for facilitating liquid–liquid phase separation on the cellular surface, in both dilute and condensed phases. Dilute solutions contain isolated protein chains in equilibrium with mesoscopic clusters, whereas the condensed phase adopts a bicontinuous, microemulsion-like morphology. The dilute phase behavior is quantitatively described by coarse-grained polymer models from soft-matter physics, demonstrating their predictive power for complex biological proteins. At elevated concentrations, the proteins self-assemble akin to block copolymers, microphase separating through the aggregation of hydrophobic domains along the protein contour. The resulting condensate remains fluid-like despite a 25-fold increase in concentration; its internal hydrodynamics slow by only a factor of 3 relative to dilute protein chains. These results provide a molecular-level framework for how disordered proteins achieve both the structural complexity and dynamic fluidity of biomolecular condensates.

Carrick, Brian R. [Massachusetts Inst. of Technolo↗

Characterization of a widespread sugar phosphate-processing bacterial microcompartment

Many prokaryotes form Bacterial Microcompartments (BMCs) that encapsulate segments of specialized metabolic pathways to enhance catalysis. The various functions of metabolosomes, catabolic BMCs, are dictated by the signature enzyme that processes initial substrates of the confined pathway. The components and native functions of several metabolosomes have been experimentally characterized; however one of the most prevalent across all bacteria has yet to be studied. Sugar Phosphate Utilizing (SPU) BMC loci encode enzymes predicted to be involved in sugar phosphate metabolism. The SPU genetic loci are found in organisms occupying habitats ranging from soils to hot springs, highlighting the ubiquity of the SPU BMC. We bioinformatically characterized seven SPU subtypes, all which contain an enzyme unique to SPU BMCs, a deoxyribose 5-phosphate aldolase (DERA). Here, we define the fundamental characteristics of SPU BMCs and have expressed, purified, and characterized a set of SPU core enzymes. These include a protein-protein complex formed between a SPU BMC DERA and a predicted ribose 5-phosphate isomerase. Further, we show that the SPU BMC DERA is catalytically active and propose that it acts as the universal signature enzyme for the SPU BMC, with implications for fundamental understanding and biotechnological applications of SPU BMCs.

59 BASIC BIOLOGICAL SCIENCES↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Genome-wide transcriptional analysis of flagellar regeneration in Chlamydomonas reinhardtii identifies orthologs of ciliary disease genes

The important role that cilia and flagella play in human disease creates an urgent need to identify genes involved in ciliary assembly and function. The strong and specific induction of flagellar-coding genes during flagellar regeneration in Chlamydomonas reinhardtii suggests that transcriptional profiling of such cells would reveal new flagella-related genes. We have conducted a genome-wide analysis of RNA transcript levels during flagellar regeneration in Chlamydomonas by using maskless photolithography method-produced DNA oligonucleotide microarrays with unique probe sequences for all exons of the 19,803 predicted genes. This analysis represents previously uncharacterized whole-genome transcriptional activity profiling study in this important model organism. Analysis of strongly induced genes reveals a large set of known flagellar components and also identifies a number of important disease-related proteins as being involved with cilia and flagella, including the zebrafish polycystic kidney genes Qilin, Reptin, and Pontin, as well as the testis-expressed tubby-like protein TULP2.

Polycystic Kidney Diseases/genetics↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Large-scale prediction of outer-membrane multiheme cytochromes uncovers hidden diversity of electroactive bacteria and underlying pathways

Multi-heme cytochromes (MHCs), together with accessory proteins like porins and periplasmic cytochromes, enable microbes to transport electrons between the cytoplasmic membrane and extracellular substrates (e.g., minerals, electrodes, other cells). Extracellular electron transfer (EET) has been described in multiple systems; yet, the broad phylogenetic and mechanistic diversity of these pathways is less clear. One commonality in EET-capable systems is the involvement of MHCs, in the form of porin-cytochrome complexes, pili-like cytochrome polymers, and lipid-anchored extracellular cytochromes. Here, we put forth MHCscan—a software tool for identifying MHCs and identifying potential EET capability. Using MHCscan, we scanned ~60,000 bacterial and 2,000 archaeal assemblies, and identify a diversity of MHCs, many of which represent enzymes with no known function, and many found within organisms not previously known to be electroactive. In total, our scan identified ~1,400 unique enzymes, each encoding more than 10 heme-binding motifs. In our analysis, we also find evidence for modularity and flexibility in MHC-dependent EET pathways, and suggest that MHCs may be far more common than previously recognized, with many facets yet to be discovered. We present MHCscan as a lightweight and user-friendly software tool that is freely available: https://github.com/Arkadiy-Garber/MHCscan.

59 BASIC BIOLOGICAL SCIENCES↗

Origin of replication discovery for environmentally isolated Pantoea strain enables expression of heterologous proteins, pathways and products

Leveraging predicted origin sequences from a previously characterized groundwater plasmidome, we constructed a barcoded plasmid library to screen for previously unknown origins. Testing this library against a panel of representative bacterial strains led to the identification of 3 previously unknown origins that replicate in gram-negative bacteria not previously associated with these origin sequences. Experimental validation confirmed that a plasmid bearing origin 6911 as the sole origin could replicate with a copy number of 9 (±2) in Pantoea sp. MT58, a fast growing and metal tolerant, environmentally important bacterium. Plasmids based on this new origin were used to express the reporter protein GFP, and non-native metabolite pathways for the natural product indigoidine and the terpenoid compound isoprenol. Functional previously unknown origins of replication in such non-model organisms can expand the toolkit for genetic manipulations of both model and less-studied bacteria.

molecular biology↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗