Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RNAseq”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

28 records · Page 2

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Histological and Transcriptomic Analysis of Spaceflight-Induced Ocular Changes in the Mouse Retina

Anatomical changes have been observed in astronauts’ eyes after long duration spaceflight missions. These alterations can lead to visual impairment which in part constitutes the spaceflight-associated neuroocular syndrome (SANS), one of the top risk priorities for deep space missions. The HRP Systems Biology (SysBio) Translation Project will apply systems biology approaches utilizing current human physiological spaceflight data, molecular results from rodents, and future research with a multi-level, multi-system, and multi-species perspective to augment the existing research plan to resolve the SANS risk. Not much is known about SANS at the cellular and molecular level, but studies in mice and rats have recently begun to determine how spaceflight might affect the biology of the eye. Preliminary studies of mice that flew on the Space Shuttle, and more recently the International Space Station (ISS), have shown changes in retinal physiology as assessed by histology and gene expression analysis. The study presented here obtained samples from the CASIS sponsored Rodent Research 8 Experiment delivered to the ISS by SpaceX CRS-16 on 12/08/2018. Female BALB/cAnNTac mice flew on the ISS for 45 days, while ground controls were housed in a standard vivarium or animal enclosure module. Sacrifice and sample acquisition occurred once mice returned to Earth, possibly allowing for readaptation affecting retinal homeostasis. We applied standard transcriptomic (RNAseq) and histological approaches to characterize genes and pathways in the mouse retina affected by spaceflight or age. The differentially expressed gene (DEG) data was analyzed using Galaxy (GeneLab) and Ingenuity Pathway Analysis. Significant DEGs between flight and ground samples were relatively few but biologically meaningful. Pathways identified related to neuronal differentiation, cellular transport/movement, and wound healing. Age effects were detected between the young (10–12 weeks) and old (32 weeks) groups and between the baseline and end of experiment (~46 days). The biological relevance of specific DEGs were confirmed through immunohistochemical evaluation using fixed histological sections of the eye from four flight group mice and four habitat control mice. Staining was performed specific for synaptophysin, glial fibrillary acidic protein (GFAP), and neurofilament in the retinal periphery, equator, and peripapillary regions. For synaptophysin staining, the innerplexiform and outerplexiform layers were scored; for GFAP staining, Mueller cells and perivascular astrocytes were scored. Results show flight samples typically had more staining of GFAP and neurofilament while, conversely, the habitat control group had more staining of synaptophysin.

C. Perez↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Uncovering Unique Molecular Adaptations in the Arabidopsis Thaliana Cvi-0 Ecotype

This research proposal aims to investigate the unique molecular adaptations exhibited by Arabidopsis Thaliana, specifically focusing on the Cape Verde Islands (Cvi-0) ecotype, in response to microgravity conditions. The study examines data from NASA’s Open Science Data Repository and applies a multifaceted RNAseq analysis pipeline using tools in the UseGalaxy.org open platform. Through transcriptomic analysis, differential gene expression patterns were identified in Cvi-0, revealing an absence of heat shock protein (HSP) upregulation and an upregulation of Rubisco Activase (RCA) and chloroplast-related pathways. To test the hypothesis that these adaptations may contribute to Cvi-0’s increased adaptability in microgravity, a three-fold experimental design is proposed. Four experimental groups will be cultivated under simulated microgravity and ground control conditions, including Cvi-0, Col-0, and genetically modified Col-0 with silenced HSP genes, and genetically modified Col-0 with upregulated RCA gene. Growth parameters will be measured to assess plant resilience, and RNA sequencing will provide transcriptomic data for pathway analysis. Anticipated outcomes include improved markers of plant health (mass, growth, etc.) of Cvi-0 in simulated microgravity and enhanced resilience in genetically altered Col-0 variants, providing insights into potential mechanisms of adaptation. This research would bear significance for space agriculture, nutrition for extended space missions, and sustainable terrestrial crop enhancement. Moreover, the insights gained could reshape crop engineering on Earth, enhancing robustness to climate induced stresses and bolstering global food security. The proposal’s trajectory blends scientific curiosity with practical applicability, forging a path towards sustainable food production and improving human exploration beyond our planet.

GL4HS↗

Transcriptomic Response of Drosophila Melanogaster Pupae Developed in Hypergravity

The metamorphosis of Drosophila is evolutionarily adapted to Earth's gravity, and is a tightly regulated process. Deviation from 1g to microgravity or hypergravity can influence metamorphosis, and alter associated gene expression. Understanding the relationship between an altered gravity environment and developmental processes is important for NASA's space travel goals. In the present study, 20 female and 20 male synchronized (Canton S, 2 to 3day old) flies were allowed to lay eggs while being maintained in a hypergravity environment (3g). Centrifugation was briefly stopped to discard the parent flies after 24hrs of egg laying, and then immediately continued until the eggs developed into P6-staged pupae (25 - 43 hours after pupation initiation). Post hypergravity exposure, P6-staged pupae were collected, total RNA was extracted using Qiagen RNeasy mini kits. We used RNA-Seq and qRT-PCR techniques to profile global transcriptomic changes in early pupae exposed to chronic hypergravity. During the pupal stage, Drosophila relies upon gravitational cues for proper development. Assessing gene expression changes in the pupa under altered gravity conditions helps highlight gravity dependent genetic pathways. A robust transcriptional response was observed in hypergravity-exposed pupae compared to controls, with 1,513 genes showing a significant (q < 0.05) difference in gene expression. Five major biological processes were affected: ion transport, redox homeostasis, immune response, proteolysis, and cuticle development. This outlines the underlying molecular changes occurring in Drosophila pupae in response to hypergravity.

RNASeq↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗

Differential Gene Expression in A Cross-Feeding Two-Species Model Microbial Community Under Simulated Microgravity and Deep-Space Radiation

A long-term goal of space biology is to understand interspecies microbial interactions in space. Presently, little is known about the combined effect of microgravity and ionizing radiation on bacterial community response when species are interdependent through exchange of metabolites in fluid medium (cross-feeding). Microgravity is expected to slow interspecies mass transfer and growth in cross-feeding communities in the low-shear, diffusion-limited environment, while ionizing radiation may influence stress response to direct (DNA damage) and indirect damage (ROS). Using a well-understood, two-species (Escherichia coli and Salmonella enterica) microbial community engineered to be a model for studying cross-feeding, we simulated galactic cosmic rays (GCRsim) and microgravity to test the hypothesis: exposure to ionizing radiation causes cell damage or stress, altering transcriptomic community responses in metabolically interdependent cells, which is exacerbated by microgravity. We expect to see differential gene expression between cross-feeding and non-cross-feeding communities. We measured GCRsim effects on growth and gene expression in well-mixed versus simulated-microgravity conditions and in cross-feeding and non-cross-feeding medium. Microbial cultures were inoculated into liquid medium in rotating wall vessels (RWV) with different rotation rates: 5 RPM (simulated microgravity) and 50 RPM (well-mixed). The E. coli-S. enterica consortium, under simulated microgravity, were exposed to 500 mGy of Simplified 5-ion Galactic Cosmic Ray Simulation for 2 hours at Brookhaven National Lab. We harvested samples 40 minutes after irradiation for extraction and sequencing (NASA GeneLab). Here we present the differential gene expression analysis results, which reveal altered transcriptomic community responses, even where growth rate differences are not observed. Gene expression of these actively metabolizing microbial communities in GCRsim may illuminate molecular mechanisms of microbial interactions in space. Understanding how microbial community gene expression, metabolism, and other cellular processes are influenced by spaceflight stressors can inform the use of microbes in human life support for low Earth orbit missions and beyond.

microgravity↗

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab↗