Engineering Papers⌕ Search

Engineering topics

Palsson, Bernhard O.

Publications and source records attributed to Palsson, Bernhard O..

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

Integration of physiologically relevant photosynthetic energy flows into whole genome models of light‐driven metabolism

SUMMARY Characterizing photosynthetic productivity is necessary to understand the ecological contributions and biotechnology potential of plants, algae, and cyanobacteria. Light capture efficiency and photophysiology have long been characterized by measurements of chlorophyll fluorescence dynamics. However, these investigations typically do not consider the metabolic network downstream of light harvesting. By contrast, genome‐scale metabolic models capture species‐specific metabolic capabilities but have yet to incorporate the rapid regulation of the light harvesting apparatus. Here, we combine chlorophyll fluorescence parameters defining photosynthetic and non‐photosynthetic yield of absorbed light energy with a metabolic model of the pennate diatom Phaeodactylum tricornutum. This integration increases the model predictive accuracy regarding growth rate, intracellular oxygen production and consumption, and metabolic pathway usage. Through the quantification of excess electron transport, we uncover the sequential activation of non‐radiative energy dissipation processes, cross‐compartment electron shuttling, and non‐photochemical quenching as the rapid photoacclimation strategy in P. tricornutum. Interestingly, the photon absorption thresholds that trigger the transition between these mechanisms were consistent at low and high incident photon fluxes. We use this understanding to explore engineering strategies for rerouting cellular resources and excess light energy towards bioproducts in silico . Overall, we present a methodology for incorporating a common, informative data type into computational models of light‐driven metabolism and show its utilization within the design–build–test–learn cycle for engineering of photosynthetic organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-learning from Pseudomonas putida KT2440 transcriptomes reveals its transcriptional regulatory network

Bacterial gene expression is orchestrated by numerous transcription factors (TFs). Elucidating how gene expression is regulated is fundamental to understanding bacterial physiology and engineering it for practical use. In this study, a machine-learning approach was applied to uncover the genome-scale transcriptional regulatory network (TRN) in Pseudomonas putida KT2440, an important organism for bioproduction. We performed independent component analysis of a compendium of 321 high-quality gene expression profiles, which were previously published or newly generated in this study. We identified 84 groups of independently modulated genes (iModulons) that explain 75.7% of the total variance in the compendium. With these iModulons, we (i) expand our understanding of the regulatory functions of 39 iModulon associated TFs (e.g., HexR, Zur) by systematic comparison with 1993 previously reported TF-gene interactions; (ii) outline transcriptional changes after the transition from the exponential growth to stationary phases; (iii) capture group of genes required for utilizing diverse carbon sources and increased stationary response with slower growth rates; (iv) unveil multiple evolutionary strategies of transcriptome reallocation to achieve fast growth rates; and (v) define an osmotic stimulon, which includes the Type VI secretion system, as coordination of multiple iModulon activity changes. Taken together, this study provides the first quantitative genome-scale TRN for P. putida KT2440 and a basis for a comprehensive understanding of its complex transcriptome changes in a variety of physiological states.

09 BIOMASS FUELS↗

Machine Learning of All Mycobacterium tuberculosis H37Rv RNA-seq Data Reveals a Structured Interplay between Metabolism, Stress Response, and Infection

Mycobacterium tuberculosis is one of the most consequential human bacterial pathogens, posing a serious challenge to 21st century medicine. A key feature of its pathogenicity is its ability to adapt its transcriptional response to environmental stresses through its transcriptional regulatory network (TRN). While many studies have sought to characterize specific portions of the M. tuberculosis TRN, and some studies have performed system-level analysis, few have been able to provide a network-based model of the TRN that also provides the relative shifts in transcriptional regulator activity triggered by changing environments. Here, we compiled a compendium of nearly 650 publicly available, high quality M. tuberculosis RNA-sequencing data sets and applied an unsupervised machine learning method to obtain a quantitative, top-down TRN. It consists of 80 independently modulated gene sets known as “iModulons,” 41 of which correspond to known regulons. These iModulons explain 61% of the variance in the organism’s transcriptional response. We show that iModulons (i) reveal the function of poorly characterized regulons, (ii) describe the transcriptional shifts that occur during environmental changes such as shifting carbon sources, oxidative stress, and infection events, and (iii) identify intrinsic clusters of regulons that link several important metabolic systems, including lipid, cholesterol, and sulfur metabolism. This transcriptome-wide analysis of the M. tuberculosis TRN informs future research on effective ways to study and manipulate its transcriptional regulation and presents a knowledge-enhanced database of all published high-quality RNA-seq data for this organism to date.

59 BASIC BIOLOGICAL SCIENCES↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Generation of Pseudomonas putida KT2440 Strains with Efficient Utilization of Xylose and Galactose via Adaptive Laboratory Evolution

While Pseudomonas putida KT2440 has great potential for biomass-converting processes, its inability to utilize the biomass abundant sugars xylose and galactose has limited its applications. Here, in this study, we utilized Adaptive Laboratory Evolution (ALE) to optimize engineered KT2440 with heterologous expression of xylD encoding xylonate dehydratase from Caulobacter crescentus and galETKM encoding UDP-glucose 4-epimerase, galactose-1-phosphate uridylyltransferase, galactokinase, and galactose-1-epimerase from Escherichia coli K-12 MG1655. Poor starting strain growth (<0.1 h –1 or none) was evolutionarily optimized to rates of up to 0.25 h –1 on xylose and 0.52 h –1 on galactose. Whole-genome sequencing, transcriptomic analysis, and growth screens revealed significant roles of kguT encoding a 2-ketogluconate operon repressor and 2-ketogluconate transporter, and gtsABCD encoding an ATP-binding cassette (ABC) sugar transporting system in xylose and galactose growth conditions, respectively. Finally, we expressed the heterologous indigoidine production pathway in the evolved and unevolved engineered strains and successfully produced 3.2 g/L and 2.2 g/L from 10 g/L of either xylose or galactose in the evolved strains whereas the unevolved strains did not produce any detectable product. Thus, the generated KT2440 strains have the potential for broad application as optimized platform chassis to develop efficient microorganism-based biomass-utilizing bioprocesses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗