Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene expression networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Evolution of the regulatory subunits for the heteromeric acetyl-CoA carboxylase

The committed step for de novo fatty acid (FA) synthesis is the ATP-dependent carboxylation of acetyl-coenzyme A catalysed by acetyl-CoA carboxylase (ACCase). In most plants, ACCase is a multi-subunit complex orthologous to prokaryotes. However, unlike prokaryotes, the plant and algal orthologues are comprised both catalytic and additional dedicated regulatory subunits. Novel regulatory subunits, biotin lipoyl attachment domain-containing proteins (BADC) and carboxyltransferase interactors (CTI) (both three-gene families inArabidopsis) represent new effectors specific to plants and certain algal species. The evolutionary history of these genes in autotrophic eukaryotes remains elusive, making it an ongoing area of research. Analyses of potential protein–protein and co-occurrence interactions, informed by gene network patterns using the STRING database, inArabidopsis thalianaandChlamydomonas reinhardtiiunveil intricate gene associations with ACCase, suggesting a complex interplay between FA synthesis and other cellular processes. Among both species, a higher number of co-expressed genes was identified inArabidopsis, indicating a wider potential regulatory network of ACCase in plants. This review investigates the extent to which these genes arose in autotrophic eukaryotes and provides insights into their evolutionary trajectory. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics↗

Rewiring the specificity of extracytoplasmic function sigma factors

Significance Bacterial phenotypes require the concerted expression of multiple genes, usually coordinated by a transcriptional regulator. Although the functions of many genes in sequenced bacterial genomes can be inferred, the regulatory networks that coordinate their expression are only known in a few model systems. Using a bioinformatic and experimental approach, we solve the DNA-specificity code of extracytoplasmic function sigma factors (ECF σs), a major class of bacterial regulators. We develop and use a high-stringency pipeline to predict the genes regulated by 67% of ECF σs in >10,000 species, providing a comprehensive look at the role of a broadly distributed family of gene regulatory proteins. This conceptual and computational framework is potentially applicable to other bacterial regulators.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics analysis of drought response between obligate CAM and C 3 photosynthesis plants

Crassulacean acid metabolism (CAM) plants exhibit elevated drought and heat tolerance compared to C 3 and C 4 plants through an inverted pattern of day/night stomatal closure and opening for CO 2 assimilation. However, the molecular responses to water-deficit conditions remain unclear in obligate CAM species. In this study, we presented genome-wide transcription sequencing analysis using leaf samples of an obligate CAM species Kalanchoë fedtschenkoi under moderate and severe drought treatments at two-time points of dawn (2-h before the start of light period) and dusk (2-h before the dark period). Differentially expressed genes were identified in response to environmental drought stress and a whole genome wide co-expression network was created as well. We found that the expression of CAM-related genes was not regulated by drought stimuli in K. fedtschenkoi. Our comparative analysis revealed that CAM species (K. fedtschenkoi) and C 3 species (Arabidopsis thaliana, Populus deltoides ‘WV94’) share some common transcriptional changes in genes involved in multiple biological processes in response to drought stress, including ABA signaling and biosynthesis of secondary metabolites.

59 BASIC BIOLOGICAL SCIENCES↗

Improved recovery of cell-cycle gene expression in Saccharomyces cerevisiae from regulatory interactions in multiple omics data

Gene expression is regulated by DNA-binding transcription factors (TFs). Together with their target genes, these factors and their interactions collectively form a gene regulatory network (GRN), which is responsible for producing patterns of transcription, including cyclical processes such as genome replication and cell division. However, identifying how this network regulates the timing of these patterns, including important interactions and regulatory motifs, remains a challenging task. Results We employed four in vivo and in vitro regulatory data sets to investigate the regulatory basis of expression timing and phase-specific patterns cell-cycle expression in Saccharomyces cerevisiae . Specifically, we considered interactions based on direct binding between TF and target gene, indirect effects of TF deletion on gene expression, and computational inference. We found that the source of regulatory information significantly impacts the accuracy and completeness of recovering known cell-cycle expressed genes. The best approach involved combining TF-target and TF-TF interactions features from multiple datasets in a single model. In addition, TFs important to multiple phases of cell-cycle expression also have the greatest impact on individual phases. Important TFs regulating a cell-cycle phase also tend to form modules in the GRN, including two sub-modules composed entirely of unannotated cell-cycle regulators ( STE12-TEC1 and RAP1-HAP1-MSN4 ). Conclusion Our findings illustrate the importance of integrating both multiple omics data and regulatory motifs in order to understand the significance regulatory interactions involved in timing gene expression. This integrated approached allowed us to recover both known cell-cycles interactions and the overall pattern of phase-specific expression across the cell-cycle better than any single data set. Likewise, by looking at regulatory motifs in the form of TF-TF interactions, we identified sets of TFs whose co-regulation of target genes was important for cell-cycle expression, even when regulation by individual TFs was not. Overall, this demonstrates the power of integrating multiple data sets and models of interaction in order to understand the regulatory basis of established biological processes and their associated gene regulatory networks.

59 BASIC BIOLOGICAL SCIENCES↗

Expression quantitative trait loci mapping identified PtrXB38 as a key hub gene in adventitious root development in Populus

Summary Plant establishment requires the formation and development of an extensive root system with architecture modulated by complex genetic networks. Here, we report the identification of the PtrXB38 gene as an expression quantitative trait loci (eQTL) hotspot, mapped using 390 leaf and 444 xylem Populus trichocarpa transcriptomes. Among predicted targets of this trans ‐eQTL were genes involved in plant hormone responses and root development. Overexpression of PtrXB38 in Populus led to significant increases in callusing and formation of both stem‐born roots and base‐born adventitious roots. Omics studies revealed that genes and proteins controlling auxin transport and signaling were involved in PtrXB38‐mediated adventitious root formation. Protein–protein interaction assays indicated that PtrXB38 interacts with components of endosomal sorting complexes required for transport machinery, implying that PtrXB38‐regulated root development may be mediated by regulating endocytosis pathway. Taken together, this work identified a crucial root development regulator and sheds light on the discovery of other plant developmental regulators through combining eQTL mapping and omics approaches.

54 ENVIRONMENTAL SCIENCES↗

The transcriptional activator ClrB is crucial for the degradation of soybean hulls and guar gum in Aspergillus niger

Low-cost plant substrates, such as soybean hulls, are used for various industrial applications. Filamentous fungi are important producers of Carbohydrate Active enZymes (CAZymes) required for the degradation of these plant biomass substrates. CAZyme production is tightly regulated by several transcriptional activators and repressors. One such transcriptional activator is CLR-2/ClrB/ManR, which has been identified as a regulator of cellulase and mannanase production in several fungi. However, the regulatory network governing the expression of cellulase and mannanase encoding genes has been reported to differ between fungal species. Previous studies showed that Aspergillus niger ClrB is involved in the regulation of (hemi-)cellulose degradation, although its regulon has not yet been identified. To reveal its regulon, we cultivated an A. niger ΔclrB mutant and control strain on guar gum (a galactomannan-rich substrate) and soybean hulls (containing galactomannan, xylan, xyloglucan, pectin and cellulose) to identify the genes that are regulated by ClrB. Gene expression data and growth profiling showed that ClrB is indispensable for growth on cellulose and galactomannan and highly contributes to growth on xyloglucan in this fungus. Therefore, we show that A. niger ClrB is crucial for the utilization of guar gum and the agricultural substrate, soybean hulls. Moreover, we show that mannobiose is most likely the physiological inducer of ClrB in A. niger and not cellobiose, which is considered to be the inducer of N. crassa CLR-2 and A. nidulans ClrB.

59 BASIC BIOLOGICAL SCIENCES↗

Editorial: Functional microcircuits in the brain and in artificial intelligent systems

Fundamental principles underlying higher-order cognitive functions remain elusive, but recent breakthroughs in neurophysiology and deep learning offer new perspectives. First, experimental studies have uncovered neural circuit motifs consisting of various neuron types; see Brain Initiative Cell Census Network (https://www.nature.com/collections/cicghheddj). For example, inhibitory neuron types expressing exclusive genes have specific targets and distinct functions (Pfeffer et al., 2013). Furthermore, diverse neuron types in cortex and their connectomes were identified in cortical columns (Jiang et al., 2015); see also Barth et al. (2016) for a debate on neuron types. Second, artificial neural networks were originally inspired by structures of the brain (McCulloch and Pitts, 1943) and could be trained to perform complex functions similar to human perception/cognition by deep learning (DL) (Lecun et al., 2015).

59 BASIC BIOLOGICAL SCIENCES↗

Genomic diversifications of five Gossypium allopolyploid species and their impact on cotton improvement

Abstract Polyploidy is an evolutionary innovation for many animals and all flowering plants, but its impact on selection and domestication remains elusive. Here we analyze genome evolution and diversification for all five allopolyploid cotton species, including economically important Upland and Pima cottons. Although these polyploid genomes are conserved in gene content and synteny, they have diversified by subgenomic transposon exchanges that equilibrate genome size, evolutionary rate heterogeneities and positive selection between homoeologs within and among lineages. These differential evolutionary trajectories are accompanied by gene-family diversification and homoeolog expression divergence among polyploid lineages. Selection and domestication drive parallel gene expression similarities in fibers of two cultivated cottons, involving coexpression networks and N 6 -methyladenosine RNA modifications. Furthermore, polyploidy induces recombination suppression, which correlates with altered epigenetic landscapes and can be overcome by wild introgression. These genomic insights will empower efforts to manipulate genetic recombination and modify epigenetic landscapes and target genes for crop improvement.

59 BASIC BIOLOGICAL SCIENCES↗

Data augmentation and multimodal learning for predicting drug response in patient-derived xenografts from gene expressions and histology images

Patient-derived xenografts (PDXs) are an appealing platform for preclinical drug studies. A primary challenge in modeling drug response prediction (DRP) with PDXs and neural networks (NNs) is the limited number of drug response samples. We investigate multimodal neural network (MM-Net) and data augmentation for DRP in PDXs. The MM-Net learns to predict response using drug descriptors, gene expressions (GE), and histology whole-slide images (WSIs). We explore whether combining WSIs with GE improves predictions as compared with models that use GE alone. We propose two data augmentation methods which allow us training multimodal and unimodal NNs without changing architectures with a single larger dataset: 1) combine single-drug and drug-pair treatments by homogenizing drug representations, and 2) augment drug-pairs which doubles the sample size of all drug-pair samples. Unimodal NNs which use GE are compared to assess the contribution of data augmentation. The NN that uses the original and the augmented drug-pair treatments as well as single-drug treatments outperforms NNs that ignore either the augmented drug-pairs or the single-drug treatments. In assessing the multimodal learning based on the MCC metric, MM-Net outperforms all the baselines. Our results show that data augmentation and integration of histology images with GE can improve prediction performance of drug response in PDXs.

60 APPLIED LIFE SCIENCES↗

BdERECTA controls vasculature patterning and phloem-xylem organization in Brachypodium distachyon

Background: The vascular system of plants consists of two main tissue types, xylem and phloem. These tissues are organized into vascular bundles that are arranged into a complex network running through the plant that is essential for the viability of land plants. Despite their obvious importance, the genes involved in the organization of vascular tissues remain poorly understood in grasses. Results: We studied in detail the vascular network in stems from the model grass Brachypodium distachyon (Brachypodium) and identified a large set of genes differentially expressed in vascular bundles versus parenchyma tissues. To decipher the underlying molecular mechanisms of vascularization in grasses, we conducted a forward genetic screen for abnormal vasculature. We identified a mutation that severely affected the organization of vascular tissues. This mutant displayed defects in anastomosis of the vascular network and uncommon amphivasal vascular bundles. The causal mutation is a premature stop codon in ERECTA, a LRR receptor-like serine/threonine-protein kinase. Mutations in this gene are pleiotropic indicating that it serves multiple roles during plant development. This mutant also displayed changes in cell wall composition, gene expression and hormone homeostasis. Conclusion: In summary, ERECTA has a pleiotropic role in Brachypodium. We propose a major role of ERECTA in vasculature anastomosis and vascular tissue organization in Brachypodium.

59 BASIC BIOLOGICAL SCIENCES↗

Phenotypically anchored transcriptomics across diverse agrichemicals reveals conserved pathways and unique gene expression signatures in zebrafish

Agrichemicals such as herbicides, fungicides, insecticides, and biocides are widely used in agriculture, yet some are associated with adverse effects in humans and the environment. While many of these chemicals have been extensively studied in vitro and are included in the EPA’s ToxCast program, comprehensive in vivo comparisons using RNA sequencing across structurally diverse agrichemicals, in a single screening platform, are lacking. In this study, we examined structurally diverse agrichemicals found in the U.S. Environmental Protection Agency’s (EPA) Toxcast Phase I and II library by statically exposing early life stage zebrafish at 6 h post fertilization (hpf) until 120 hpf at concentrations ranging from 0.25 to 100 µM. Morphological outcomes were assessed at 120 hpf across 10 endpoints, including yolk sac edema, craniofacial malformations, and axis abnormalities. Chemicals that produced robust concentration-response relationships were selected for transcriptomic profiling. For transcriptomic analysis, zebrafish were statically exposed to each chemical and sampled at 48 hpf, prior to the onset of morphological effects observed at 120 hpf. Differential expression analysis identified between 0 and 4,538 differentially expressed genes (DEGs) per chemical, with no clear correlation to morphological severity. Both DEG and co-expression network analyses revealed chemical-specific expression patterns that converged on shared biological pathways, including neurodevelopment and cytoskeletal organization. Key regulatory genes such as mylpfa and krt4 were identified within co-expression modules, suggesting their potential role in conserved toxicity mechanisms. Semantic similarity analysis of enriched gene ontology (GO) terms, when compared to existing datasets, highlighted gaps in the annotation of neurodevelopmental processes, indicating that some in vivo effects may not be fully captured by current curated resources. The results provide new insights into the modes of action of diverse agrichemicals and establish a framework for understanding how agrichemical structure relates to biological function in a vertebrate model.

agrichemical↗

Disrupting autorepression circuitry generates “open-loop lethality” to yield escape-resistant antiviral agents

Across biological scales, gene-regulatory networks employ autorepression (negative feedback) to maintain homeostasis and minimize failure from aberrant expression. In this study, we present a proof of concept that disrupting transcriptional negative feedback dysregulates viral gene expression to therapeutically inhibit replication and confers a high evolutionary barrier to resistance. We find that nucleic-acid decoys mimicking cis-regulatory sites act as “feedback disruptors,” break homeostasis, and increase viral transcription factors to cytotoxic levels (termed “open-loop lethality”). Feedback disruptors against herpesviruses reduced viral replication >2-logs without activating innate immunity, showed sub-nM IC 50 , synergized with standard-of-care antivirals, and inhibited virus replication in mice. In contrast to approved antivirals where resistance rapidly emerged, no feedback-disruptor escape mutants evolved in long-term cultures. For SARS-CoV-2, disruption of a putative feedback circuit also generated open-loop lethality, reducing viral titers by >1-log. These results demonstrate that generating open-loop lethality, via negative-feedback disruption, may yield a class of antimicrobials with a high genetic barrier to resistance.

59 BASIC BIOLOGICAL SCIENCES↗

Logic-based analysis of gene expression data predicts association between TNF, TGFB1 and EGF pathways in basal-like breast cancer

For breast cancer, clinically important subtypes are well characterized at the molecular level in terms of gene expression profiles. In addition, signaling pathways in breast cancer have been extensively studied as therapeutic targets due to their roles in tumor growth and metastasis. However, it is challenging to put signaling pathways and gene expression profiles together to characterize biological mechanisms of breast cancer subtypes since many signaling events result from post-translational modifications, rather than gene expression differences. We designed a logic-based computational framework to explain the differences in gene expression profiles among breast cancer subtypes using Pathway Logic and transcriptional network information. Pathway Logic is a rewriting-logic-based formal system for modeling biological pathways including post-translational modifications. Our method demonstrated its utility by constructing subtype-specific path from key receptors (TNFR, TGFBR1 and EGFR) to key transcription factor (TF) regulators (RELA, ATF2, SMAD3 and ELK1) and identifying potential association between pathways via TFs in basal-specific paths, which could provide a novel insight on aggressive breast cancer subtypes. Lastly, codes and results are available at http://epigenomics.snu.ac.kr/PL/.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring Camelina sativa lipid metabolism regulation by combining gene co-expression and DNA affinity purification analyses

Camelina (Camelina sativa) is an annual oilseed plant that is gaining momentum as a biofuel cover crop. Understanding gene regulatory networks is essential to deciphering plant metabolic pathways, including lipid metabolism. Furthermore, we take advantage of a growing collection of gene expression datasets to predict transcription factors (TFs) associated with the control of Camelina lipid metabolism. We identified approximately 350 TFs highly co-expressed with lipid-related genes (LRGs). These TFs are highly represented in the MYB, AP2/ERF, bZIP, and bHLH families, including a significant number of homologs of well-known Arabidopsis lipid and seed developmental regulators. After prioritizing the top 22 TFs for further validation, we identified DNA-binding sites and predicted target genes for 16 out of the 22 TFs tested using DNA affinity purification followed by sequencing (DAP-seq). Enrichment analyses of targets supported the co-expression prediction for most TF candidates, and the comparison to Arabidopsis revealed some common themes, but also aspects unique to Camelina. Within the top potential lipid regulators, we identified CsaMYB1, CsaABI3AVP1-2, CsaHB1, CsaNAC2, CsaMYB3, and CsaNAC1 as likely involved in the control of seed fatty acid elongation and CsaABI3AVP1-2 and CsabZIP1 as potential regulators of the synthesis and degradation of triacylglycerols (TAGs), respectively. Altogether, the integration of co-expression data and DNA-binding assays permitted us to generate a high-confidence and short list of Camelina TFs involved in the control of lipid metabolism during seed development.

59 BASIC BIOLOGICAL SCIENCES↗

TULIP: An RNA-seq-based Primary Tumor Type Prediction Tool Using Convolutional Neural Networks

Background: With cancer as one of the leading causes of death worldwide, accurate primary tumor type prediction is critical in identifying genetic factors that can inhibit or slow tumor progression. There have been efforts to categorize primary tumor types with gene expression data using machine learning, and more recently with deep learning, in the last several years. Methods In this paper, we developed four 1-dimensional (1D) Convolutional Neural Network (CNN) models to classify RNA-seq count data as one of 17 highly represented primary tumor types or 32 primary tumor types regardless of imbalanced representation. Additionally, we adapted the models to take as input either all Ensembl genes (60,483) or protein coding genes only (19,758). Unlike previous work, we avoided selection bias by not filtering genes based on expression values. RNA-seq count data expressed as FPKM-UQ of 9,025 and 10,940 samples from The Cancer Genome Atlas (TCGA) were downloaded from the Genomic Data Commons (GDC) corresponding to 17 and 32 primary tumor types respectively for training and validating the models. Results: All 4 1D-CNN models had an overall accuracy of 94.7% to 97.6% on the test dataset. Further evaluation indicates that the models with protein coding genes only as features performed with better accuracy compared to the models with all Ensembl genes for both 17 and 32 primary tumor types. For all models, the accuracy by primary tumor type was above 80% for most primary tumor types. Conclusions: We packaged all 4 models as a Python-based deep learning classification tool called TULIP (TUmor CLassIfication Predictor) for performing quality control on primary tumor samples and characterizing cancer samples of unknown tumor type. Further optimization of the models is needed to improve the accuracy of certain primary tumor types.

Jones, Sara↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗