Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RNA-seq”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

A Multi-omics Longitudinal Study of the Murine Retinal Response to Chronic Low-dose Irradiation and/or Simulated Microgravity

The space environment includes unique hazards like radiation and microgravity which adversely affect physiology and behavior of humans and rodent models. To better characterize the retinal response to spaceflight, we assessed a multi-omics NASA GeneLab dataset where 6-month-old female mice were gamma irradiated and/or hindlimb unloaded for 21 days followed by whole transcriptome shotgun sequencing (RNA-Seq) and reduced representation bisulfite sequencing (RRBS) of retina samples collected at 7 days, 1 month or 4 months post-exposure. We compared time-matched epigenomic and transcriptomic retinal profiles revealing a total of 4,178 differentially methylated loci or regions, and 457 differentially expressed genes. Highest correlation in methylation differences was seen across different conditions at the same time point (e.g., between radiation exposure and hindlimb unloaded at 7 days). Biological processes related to nucleotide metabolism were enriched in all groups with activation at 1 month and suppression at 7 days and 4 months. Genes and processes related to Notch and Wnt signaling showed alterations 4 months post-exposure. Interestingly, Notch3 and Lrg1 showed differential patterns in the NASA Twins Study in-flight samples and in response to stressors in the murine retina in the current study. A total of 23 genes were both differentially methylated and expressed, including genes involved in retinal disease or cataract development (Crybb3, Fgfr1, Pitpnm3, Sipa1l3, Sox9) and inflammatory response (B4galt6, Ppm1a, Sphk1). To our knowledge, the current multi-omics analysis is the first multi-omics study to interrogate the epigenomic and transcriptomic impacts of radiation and hindlimb unloading on the retina in isolation and in combination. The results provide an insight into the retinal response to individual spaceflight hazard analogs and their interplay at different post-exposure stages and contributes towards a mechanistic understanding of spaceflight-induced vision impairment using ground-based models.

Prachi Kothiyal↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

The NASA Twins Study: The Effect of One Year in Space on Long-Chain Fatty Acid Desaturases and Elongases

Background: To date, there is no clear understanding of the effect of long-duration spaceflight on the major enzymes that govern the metabolism of omega-6 and omega-3 fatty acids. To address this gap in knowledge, we used data from the NASA Twins Study, which includes a multi-scale omic investigation of the changes that occurred during a year-long (340 days) human spaceflight. Embedded within the NASA Twins data are specific analytes associated with fatty acid metabolism. Objectives: To examine the long-chain fatty acid desaturases and elongases in a single human during one year in space. Method: One male twin was on board the International Space Station (ISS) for one year, while his monozygotic twin served as a genetically matched ground control. Longitudinal assessments included the genome, epigenome, transcriptome, proteome, metabolome, microbiome, and immunome during the mission, as well as six months before and after. The gene-specific fatty acid desaturase and elongase transcriptome data (FADS1, FADS2, ELOVL2 and ELOVL5) were extracted from untargeted RNA-seq measurements derived from white blood cell fractions. Results: Most data from the elongases and desaturases exhibited relatively similar expression profiles (R2>0.6) over time for the CD8, CD19, and LD cell fractions, indicating overall conservation of function within and between the subjects. Both cell-type and temporal specificity was observed in some cases, and some differences were also apparent between the poly-adenylated fraction (polyA) of processed RNAs vs. the ribo-depleted (ribo-) fraction. The flight subject showed a stronger enrichment of the Fatty Acid Metabolic processes pathway across almost all cell types (columns, CD4, CD8, CPT, LD), most especially in the ribodepleted fraction of RNA, but also with the polyA+ fraction of RNA. GSEA enrichment measures across three related Fatty Acid Metabolism pathways showed a differential between the ground and flight subject. Conclusions: There appears to be no persistent alteration of desaturase and elongase gene expression associated with one year in space. However, these data provide evidence that cellular lipid metabolism can be responsive and dynamic to spaceflight, even though it appears cell-type- and context-specific, most notably in terms of the fraction of RNA measured and the collection protocols. These results also provide new evidence of mid-flight spikes in expression of selected genes, which may indicate transient responses to specific insults during spaceflight.

Elongase↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics Processing Pipelines for Space Biology: An Open Source and Consensus-Driven Approach

Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.

Jonathan Oribello↗

Separation of life stages within anaerobic fungi (Neocallimastigomycota) highlights differences in global transcription and metabolism

Anaerobic gut fungi of the phylum Neocallimastigomycota are microbes proficient in valorizing low-cost but difficult-to-breakdown lignocellulosic plant biomass. Characterization of different fungal life stages and how they contribute to biomass breakdown are critical for biotechnological applications, yet we lack foundational knowledge about the transcriptional, metabolic, and enzyme secretion behavior of different life stages of anaerobic gut fungi: zoospores, germlings, immature thalli, and mature zoosporangia. A Miracloth-based technique was developed to enrich cell pellets with zoospores - the free-swimming, flagellated, young life stage of anaerobic gut fungi. By contrast, fungal mats contained relatively more vegetative, encysted, mature sporangia that form films. Global gene expression profiles were compared from two sample types (zoospore-enriched cell pellets vs. mature mats) harvested from the anaerobic gut fungal strain Neocallimastix californiae G1. Despite cultures being grown on glucose, the fungal zoospore-enriched samples were transcriptionally primed to encounter plant matter substrate, as evidenced by upregulation of catabolic carbohydrate-active enzymes and putative carbohydrate transporters. Furthermore, we report significant differential gene expression for gene annotation groups, including putative secondary metabolites and transcription factors. Understanding global gene expression differences between the fungal zoospore-enriched cells and mature fungi aid in characterizing fungal development, unmasking gene function, and guiding cultivation conditions and engineering targets to promote enzyme secretion.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes↗

Decoding crops one cell at a time: from cell atlases to single-cell genetics

Understanding the mechanisms underlying key agricultural traits remains a central challenge in crop research, but recent advances in technologies are providing powerful tools to address this issue. Among these, single-cell and spatial transcriptomics have revealed tissue heterogeneity and spatial organization, offering unique insights into cellular gene expression dynamics and the coordinated activity of multiple cell types. These approaches help uncover how specific cell types contribute to agricultural traits and refine candidate loci lists through integration with trait-associated loci. Additionally, single-cell and spatial transcriptomics have the potential to serve as cell-level readout platforms integrating cellular perturbations, enabling high-throughput discovery of causal relationships between genotype and gene expression at the cellular level in plants. Successful implementation will accelerate the identification of key genetic variants for crop improvement. Furthermore we review lessons learned from application of single-cell screening in mammalian cells, highlight major technical and biological barriers to its use in plants, and outline potential strategies to overcome these challenges. Together, the widespread application and integration of single-cell and spatial transcriptomics with other technologies enable not only the descriptive cataloging of cell states but also the causal interrogation of sequence functions and regulatory networks at cell type resolution, ultimately advancing gene function studies and accelerating crop improvement.

Cellular heterogeneity↗

Novosphingobium aromaticivorans LigR coordinates transcription of genes involved in metabolism of multiple types of aromatics

Aromatic compounds are a ubiquitous and diverse family of chemicals with functions as biomolecules, natural products, industrial chemicals, and pollutants. Novosphingobium aromaticivorans DSM 12444 uses multiple inducible pathways to catabolize H-, G-, and S-type aromatics that contain zero, one, or two methoxy groups, respectively. Here, we obtain a systems-level view of the transcriptional control of its aromatic metabolic pathways. Several in vitro analyses found that a N. aromaticivorans homolog of the Sphingobium lignivorans SYK-6 transcription factor LigR bound genomic DNA upstream of genes involved in metabolism of multiple aromatic types. We found that a ΔLigR mutant had growth defects on all three types of aromatics as sole carbon sources. Transcriptomic analysis revealed that LigR was required to increase expression of gene products that function in metabolism of all three aromatic types. We also found that, in media containing both glucose and an aromatic carbon source, the ΔLigR mutant directed intermediates through alternative aromatic metabolic pathways. Protein-DNA binding assays showed that N. aromaticivorans LigR binds immediately upstream of promoters of genes involved in aromatic metabolism. We found that N. aromaticivorans LigR coordinates the expression of enzymes that function in the catabolism of H-, G-, and S-type aromatics, and that there are differences in the role of LigR in N. aromaticivorans and S. lignivorans. A comparative genomic analysis predicted that LigR homologs and the aromatic-metabolizing genes that it directly regulates are often co-localized in the genomes of Sphingomonadales, but often not found in this arrangement in many other known aromatic metabolizing bacteria.

Aromatic Compound Degradation↗

Morpho-physiological and transcriptomic responses of field pennycress to waterlogging

Field pennycress (Thlaspi arvense) is a new biofuel winter annual crop with extreme cold hardiness and a short life cycle, enabling off-season integration into corn and soybean rotations across the U.S. Midwest. Pennycress fields are susceptible to winter snow melt and spring rainfall, leading to waterlogged soils. The objective of this research was to determine the extent to which waterlogging during the reproductive stage affected gene expression, morphology, physiology, recovery, and yield between two pennycress lines (SP32-10 and MN106). In a controlled environment, total pod number, shoot/root dry weight, and total seed count/weight were significantly reduced in SP32-10 in response to waterlogging, whereas primary branch number, shoot dry weight, and single seed weight were significantly reduced in MN106. This indicated waterlogging had a greater negative impact on seed yield in SP32-10 than MN106. We compared the transcriptomic response of SP32-10 and MN106 to determine the gene expression patterns underlying these different responses to seven days of waterlogging. The number of differentially expressed genes (DEGs) between waterlogged and control roots were doubled in MN106 (3,424) compared to SP32-10 (1,767). Functional enrichment analysis of upregulated DEGs revealed Gene Ontology (GO) terms associated with hypoxia and decreased oxygen, with genes in these categories encoding proteins involved in alcoholic fermentation and glycolysis. Additionally, downregulated DEGs revealed GO terms associated with cell wall biogenesis and suberin biosynthesis, indicating suppressed growth and energy conservation. Interestingly, MN106 waterlogged roots exhibited significant stronger regulation of these genes than SP32-10, displaying a more robust transcriptomic response overall. Together, these results reveal the reconfiguration of cellular and metabolic processes in response to the severe energy crisis invoked by waterlogging in pennycress.

ERF-VII↗

Zymomonas mobilis oxidative stress transcriptomics

Transcriptomic analysis of WT, a deletion of ZMO_0422, and a deletion of ZMO_1411 in Zymomonas mobilis ZM4 under aerobic and anaerobic growth conditions along with various oxidative stresses: Paraquate addition, No Iron, and hydrogen peroxide addition.

aerobic↗

Zymomonas mobilis oxidative stress transcriptomics

Zymomonas mobilis is an important bioenergy organism that has potential to produce biofuels, including ethanol, in high volumes. Here we examined the response of Zymomonas mobilis to various oxidative stresses using genome-scale transcriptomics data. We first examined the transcrpit abundance in WT aerobic growth compared to aerobic grown in paraquat, which forms superoxide. Under anaerobic growth conditions we compared WT Zymomonas mobilis with strains grown in media lacking iron as well as strains lacking iron that were treated with the iron chelator DIP before collection. Finally we examined transcript abundance in cells lacking ZMO_0422 (Rrf2 family transcription factor homolog) and ZMO_1411 (Fur homolog) grown under anaerobic conditions. Overall design: Transcriptomic analysis of WT, a deletion of ZMO_0422, and a deletion of ZMO_1411 in Zymomonas mobilis ZM4 under aerobic and anaerobic growth conditions along with various oxidative stresses: Paraquate addition, No Iron, and hydrogen peroxide addition.

aerobic↗

Transcriptomic analysis of ZMO_0422 in Zymomonas mobilis

Deletion of the IscR homolog ZMO_0422 was performed in Zymomonas mobilis to investigate the role of Fe-S cluster biogenesis in Zymomonas. Here we perform genome-wide transcirptomics study to examine transcript chagnes in delta-ZMO_0422 compared to WT Zymomonas mobilis under both aerobic and anaerobic growth conditions. Overall design: Transcriptomic analysis of WT and a deletion of ZMO_0422 of Zymomonas mobilis ZM4 under aerobic and anaerobic growth conditions.

aerobic↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Regulatory response to a hybrid ancestral nitrogenase in Azotobacter vinelandii

Biological nitrogen fixation, the microbial reduction of atmospheric nitrogen to bioavailable ammonia, represents both a major limitation on biological productivity and a highly desirable engineering target for synthetic biology. However, the engineering of nitrogen fixation requires an integrated understanding of how the gene regulatory dynamics of host diazotrophs respond across sequence-function space of its central catalytic metalloenzyme, nitrogenase. Here, we interrogate this relationship by analyzing the transcriptome of Azotobacter vinelandii engineered with a phylogenetically inferred ancestral nitrogenase protein variant. The engineered strain exhibits reduced cellular nitrogenase activity but recovers wild-type growth rates following an extended lag period. We find that expression of genes within the immediate nitrogen fixation network is resilient to the introduced nitrogenase sequence-level perturbations. Rather the sustained physiological compatibility with the ancestral nitrogenase variant is accompanied by reduced expression of genes that support trace metal and electron resource allocation to nitrogenase. Our results spotlight gene expression changes in cellular processes adjacent to nitrogen fixation as productive engineering considerations to improve compatibility between remodeled nitrogenase proteins and engineered host diazotrophs.

nitrogen fixation↗