Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Protein function prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Simple Math is Enough: Two Examples of Inferring Functional Associations from Genomic Data

Non-random features in the genomic data are usually biologically meaningful. The key is to choose the feature well. Having a p-value based score prioritizes the findings. If two proteins share a unusually large number of common interaction partners, they tend to be involved in the same biological process. We used this finding to predict the functions of 81 un-annotated proteins in yeast.

Liang, Shoudan

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Maize Rough Endosperm6 (rgh6) Encodes A Predicted Dead-Box RNA Helicase and Affects Mirna Processing in Endosperm Development

Maize rough endosperm (rgh) mutants have defective kernels with a rough, etched, or pitted endosperm surface. Molecular genetic analysis of this mutant class has identified multiple RNA processing proteins critical to endosperm development. Here, we report on the developmental and molecular function of the rgh6 locus. The rgh6 mutant was isolated from the UniformMu transposon tagging population. Mutant kernels have reduced endosperm size and defective embryos that develop in a more apical position than typical for defective embryos. TB translocation crosses revealed that rgh6 mutant endosperm inhibits normal embryo development. Positional cloning of the rgh6 locus found that it encodes a predicted DEAD-box RNA helicase. Consistent with a predicted function for RNA processing, transient expression of a RGH6-GFP fusion protein is localized to nucleolus and nuclear speckles in Nicotiana benthamiana leaves. Rgh6 transcripts are highly expressed in endosperm epidermal cell types such as the aleurone, basal endosperm cell layer, embryo surrounding region, and endosperm adjacent to scutellum. Markers of these cell types show increased levels in rgh6 mutant kernels. Mutant endosperm tissues have increased precursor microRNA (pre-miRNA) and decreased mature miRNA relative to normal sibling endosperm, indicating that rgh6 is required for miRNA processing. The transcript levels for most miRNA target genes accumulate to a higher level in rgh6 mutant tissue. These results suggest that miRNA processing and regulation of miRNA target genes are required for normal endosperm development.

Plant Sciences

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology

Cholesterol modulates membrane elasticity via unified biophysical laws

Cholesterol and lipid unsaturation underlie a balance of opposing forces that features prominently in adaptive cell responses to diet and environmental cues. These competing factors have resulted in contradictory observations of membrane elasticity across different measurement scales, requiring chemical specificity to explain incompatible structural and elastic effects. Here, we demonstrate that – unlike macroscopic observations – lipid membranes exhibit a unified elastic behavior in the mesoscopic regime between molecular and macroscopic dimensions. Using nuclear spin techniques and computational analysis, we find that mesoscopic bending moduli follow a universal dependence on the lipid packing density regardless of cholesterol content, lipid unsaturation, or temperature. Our observations reveal that compositional complexity can be explained by simple biophysical laws that directly map membrane elasticity to molecular packing associated with biological function, curvature transformations, and protein interactions. The obtained scaling laws closely align with theoretical predictions based on conformational chain entropy and elastic stress fields. These findings provide unique insights into the membrane design rules optimized by nature and unlock predictive capabilities for guiding the functional performance of lipid-based materials in synthetic biology and real-world applications.

Kumarage, Teshani [Virginia Polytechnic Inst. and

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost

Comparison of Two Bioinformatics Tools Used to Characterize the Microbial Diversity and Predictive Functional Attributes of Microbial Mats from Lake Obersee, Antarctica

In this study, using NextGen sequencing of the collective 16S rRNA genes obtained from two sets of samples collected from Lake Obersee, Antarctica, we compared and contrasted two bioinformatics tools, PICRUSt and Tax4Fun. We then developed an R script to assess the taxonomic and predictive functional profiles of the microbial communities within the samples. Taxa such as Pseudoxanthomonas, Planctomycetaceae, Cyanobacteria Subsection III, Nitrosomonadaceae, Leptothrix, and Rhodobacter were exclusively identified by Tax4Fun that uses SILVA database; whereas PICRUSt that uses Greengenes database uniquely identified Pirellulaceae, Gemmatimonadetes A1-B1, Pseudanabaena, Salinibacterium and Sinibacteraceae. Predictive functional profiling of the microbial communities using Tax4Fun and PICRUSt separately revealed common metabolic capabilities, while also showing specific functional IDs not shared between the two approaches. Combining these functional predictions using a customized R script revealed a more inclusive metabolic profile, such as hydrolases, oxidoreductases, transferases; enzymes involved in carbohydrate and amino acid metabolisms; and membrane transport proteins known for nutrient uptake from the surrounding environment. Our results present the first molecular-phylogenetic characterization and predictive functional profiles of the microbial mat communities in Lake Obersee, while demonstrating the efficacy of combining both the taxonomic assignment information and functional IDs using the R script created in this study for a more streamlined evaluation of predictive functional profiles of microbial communities.

Hyunmin Koo

A non-canonical fungal peroxisome PTS-1 signal, SYM, and its evolutionary aspects

Abstract Proteins localized to peroxisomes, particularly those expressed under specific conditions or in low abundance, are often undetected by routine proteomics methods due to detection sensitivity limits. In silico identification and experimental validation of peroxisomal targeting signals (PTSs) offer a reliable alternative. We demonstrate that SYM, a non-canonical plant PTS-1 signal, functions similarly inAspergillus nidulans, as GFP tagged with a SYM C-terminal tripeptide localizes to peroxisomes. One of two nativeA. nidulansproteins with C-terminal SYM tripeptide shows weak peroxisomal localization alongside cytoplasmic presence, indicating that only a subset of proteins with non-canonical signals access peroxisomes.In silicoanalysis of 1,010 fungal genomes identified diverse SYM-proteins with variable functions, suggesting that non-canonical PTS-1 signals may evolve spontaneously. Two-thirds of SYM-proteins are predicted to localize to specific intracellular compartments other than the peroxisome. We propose that despite their predicted localization, these proteins possessing SYM as a non-canonical peroxisomal signal might also have peroxisomal presence. Among SYM-proteins, pectinesterases, known plant pathogen virulence factors, were frequent. Notably, 25% of fungal pectinesterases harbor non-canonical PTS-1 signals, suggesting that partial peroxisomal localization of pectinesterases has evolved convergently. This suggests that partial peroxisomal localization may enhance protein functional flexibility, contributing to the organism’s adaptability.

Science & Technology - Other Topics

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES

Genetic Transfer in Action: Uncovering DNA Flow in an Extremophilic Microbial Community

ABSTRACT Horizontal genetic transfer (HGT) is a significant driver of genomic novelty in all domains of life. HGT has been investigated in many studies however, the focus has been on conspicuous protein‐coding DNA transfers that often prove to be adaptive in recipient organisms and are therefore fixed longer‐term in lineages. These results comprise a subclass of HGTs and do not represent exhaustive (coding and non‐coding) DNA transfer and its impact on ecology. Uncovering exhaustive HGT can provide key insights into the connectivity of genomes in communities and how these transfers may occur. In this study, we use the term frequency‐inverse document frequency (TF‐IDF) technique, that has been used successfully to mine DNA transfers within real and simulated high‐quality prokaryote genomes, to search for exhaustive HGTs within an extremophilic microbial community. We establish a pipeline for validating transfers identified using this approach. We find that most DNA transfers are within‐domain and involve non‐coding DNA. A relatively high proportion of the predicted protein‐coding HGTs appear to encode transposase activity, restriction‐modification system components, and biofilm formation functions. Our study demonstrates the utility of the TF‐IDF approach for HGT detection and provides insights into the mechanisms of recent DNA transfer.

Microbiology

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES

Co- and/or post-translational modifications are critical for TCH4 XET activity

TCH4 encodes a xyloglucan endotransglycosylase (XET) of Arabidopsis thaliana. XETs endolytically cleave and religate xyloglucan polymers; xyloglucan is one of the primary structural components of the plant cell wall. Therefore, XET function may affect cell shape and plant morphogenesis. To gain insight into the biochemical function of TCH4, we defined structural requirements for optimal XET activity. Recombinant baculoviruses were designed to produce distinct forms of TCH4. TCH4 protein engineered to be synthesized in the cytosol and thus lack normal co- and post-translational modifications is virtually inactive. TCH4 proteins, with and without a polyhistidine tag, that harbor an intact N-terminus are directed to the secretory pathway. Thus, as predicted, the N-terminal region of TCH4 functions as a signal peptide. TCH4 is shown to have at least one disulfide bond as monitored by a mobility shift in SDS-PAGE in the presence of dithiothreitol (DTT). This disulfide bond(s) is essential for full XET activity. TCH4 is glycosylated in vivo; glycosidases that remove N-linked glycosylation eliminated 98% of the XET activity. Thus, co- and/or post-translational modifications are critical for optimal TCH4 XET activity. Furthermore, using site-specific mutagenesis, we demonstrated that the first glutamate residue of the conserved DEIDFEFL motif (E97) is essential for activity. A change to glutamine at this position resulted in an inactive protein; a change to aspartic acid caused protein mislocalization. These data support the hypothesis that, in analogy to Bacillus beta-glucanases, this region may be the active site of XET enzymes.

NASA Discipline Cell Biology

Thermophilic Chassis-Enabled High-Throughput Selection of a Thermostable Fluorogenic Reporter

Thermostable proteins show increased shelf life and performance at elevated temperatures and under harsh conditions, resulting in lower costs for various industrial and biotechnological applications. However, due to a limited understanding of the relationship between stability and function, protein stabilization remains primarily a trial-and-error approach. Therefore, building a combinatorial library of mutations predicted to improve stability, followed by experimental testing, represents a markedly improved methodology. However, the lack of high-throughput approaches to screen even a moderately sized library presents a major bottleneck in the field. Here, in this study, we use a thermophile, Parageobacillus thermoglucosidasius (Ptherm) to rapidly screen combinatorial libraries consisting of rationally designed thermostabilizing mutations (∼10 3 –10 4 ) of a mesophilic fluorescent reporter, Y-FAST. On a Petri dish, microbial growth at an elevated temperature and exposure to fluorogen yielded several colonies of Ptherm that showed distinct fluorescence at 55 and 68 °C in our two sequentially generated libraries using Rosetta and ProteinMPNN, respectively. The Y-FAST variants isolated from fluorescent colonies were brighter than Y-FAST and showed higher resistance to thermal and chemical denaturation. AlphaFold-predicted structures and MD simulations revealed stability-enhancing salt bridges and hydrogen bond networks in the isolated FAST variants. The moderately thermostable FAST (tsFAST) and hyperstable FAST (hsFAST) were then demonstrated as translation reporters for protein expression and folding at elevated temperatures, such as 55 and 68 °C. Our approach of combinatorial library generation and high-throughput screening in a thermophilic chassis could, in principle, be extended to other proteins fused to these translation reporters. Furthermore, the hsFAST protein is small─half the size of the green fluorescent protein─and does not require oxygen for maturation, making it ideal for engineering extremophilic anaerobes for biosensing and bioconversion.

59 BASIC BIOLOGICAL SCIENCES

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric

PET-FBA: A lightweight enzyme allocation and thermodynamics-constrained flux analysis approach to explore Escherichia coli metabolic adaptation to intracellular acidification

Escherichia coli employs diverse strategies to adapt to acidic environments that disrupt enzyme activity and the thermodynamic feasibility of essential reactions. To understand the impact of pH stress on cell metabolism, we present the PET-FBA (pH-, Enzyme protein allocation-, and Thermodynamics-constrained Flux Balance Analysis) framework. PET-FBA extends genome-scale modeling by integrating enzyme protein costs and reaction Gibbs free energy changes. Additionally, by incorporating pH-dependent enzyme kinetics in response to intracellular acidification, this framework enables the simulation of E. coli's metabolic adjustments across varying external pH levels. The model's accuracy is validated by comparing in silico growth simulations with experimental measurements under both anaerobic and aerobic conditions, as well as in silico gene knockouts of essential genes. By explicitly incorporating pH effects, our model accurately replicates the metabolic shift towards lactate production as the primary fermentation product at low pH in anaerobic conditions. This shift is only predicted when enzyme kinetics are dynamically adjusted as a function of pH. Further analysis revealed that this shift can be attributed to the reduced protein efficiency of the acetyl-CoA branch compared to lactate dehydrogenase under acidic stress, which then becomes crucial for maintaining NAD regeneration and cell growth at low pH. Furthermore, we identified strategies for enhancing cell growth under acidic anaerobic conditions by improving the enzyme activity of lactate dehydrogenase and pyruvate formate lyase, which increases NAD production efficiency and reduces enzyme protein allocation costs. Designed as a lightweight yet versatile framework, PET-FBA enables efficient genome-scale metabolic analysis. Using E. coli as a model system, our framework provides a systematic approach to understanding metabolic responses to environmental stress, pinpointing key metabolic bottlenecks, and identifying potential targets for strain optimization.

42 ENGINEERING