Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Transcriptomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning

Histological and Transcriptomic Analysis of Spaceflight-Induced Ocular Changes in the Mouse Retina

Anatomical changes have been observed in astronauts’ eyes after long duration spaceflight missions. These alterations can lead to visual impairment which in part constitutes the spaceflight-associated neuroocular syndrome (SANS), one of the top risk priorities for deep space missions. The HRP Systems Biology (SysBio) Translation Project will apply systems biology approaches utilizing current human physiological spaceflight data, molecular results from rodents, and future research with a multi-level, multi-system, and multi-species perspective to augment the existing research plan to resolve the SANS risk. Not much is known about SANS at the cellular and molecular level, but studies in mice and rats have recently begun to determine how spaceflight might affect the biology of the eye. Preliminary studies of mice that flew on the Space Shuttle, and more recently the International Space Station (ISS), have shown changes in retinal physiology as assessed by histology and gene expression analysis. The study presented here obtained samples from the CASIS sponsored Rodent Research 8 Experiment delivered to the ISS by SpaceX CRS-16 on 12/08/2018. Female BALB/cAnNTac mice flew on the ISS for 45 days, while ground controls were housed in a standard vivarium or animal enclosure module. Sacrifice and sample acquisition occurred once mice returned to Earth, possibly allowing for readaptation affecting retinal homeostasis. We applied standard transcriptomic (RNAseq) and histological approaches to characterize genes and pathways in the mouse retina affected by spaceflight or age. The differentially expressed gene (DEG) data was analyzed using Galaxy (GeneLab) and Ingenuity Pathway Analysis. Significant DEGs between flight and ground samples were relatively few but biologically meaningful. Pathways identified related to neuronal differentiation, cellular transport/movement, and wound healing. Age effects were detected between the young (10–12 weeks) and old (32 weeks) groups and between the baseline and end of experiment (~46 days). The biological relevance of specific DEGs were confirmed through immunohistochemical evaluation using fixed histological sections of the eye from four flight group mice and four habitat control mice. Staining was performed specific for synaptophysin, glial fibrillary acidic protein (GFAP), and neurofilament in the retinal periphery, equator, and peripapillary regions. For synaptophysin staining, the innerplexiform and outerplexiform layers were scored; for GFAP staining, Mueller cells and perivascular astrocytes were scored. Results show flight samples typically had more staining of GFAP and neurofilament while, conversely, the habitat control group had more staining of synaptophysin.

C. Perez

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

NASA has employed high-throughput molecular assays to identify sub-cellular changes impacting human physiology during spaceflight. Machine learning (ML) methods hold the promise to improve our ability to identify important signals within highly dimensional molecular data. However, the inherent limitation of study subject numbers within a spaceflight mission minimizes the utility of ML approaches. To overcome the sample power limitations, data from multiple spaceflight missions must be aggregated while appropriately addressing intra- and inter-study variabilities. Here we describe an approach to log transform, scale and normalize data from six heterogeneous, mouse liver derived transcriptomics datasets (ntotal=137) which enabled ML-methods to perform well (AUC ≥ 0.87) in classifying spaceflown vs ground control animals rather than mission-of-origin. Concordance was found between liver-specific biological processes identified from harmonized ML-based analysis and study-by-study classical omics analysis. This work demonstrates the feasibility of applying ML methods on integrated, heterogeneous datasets of small sample size.

Machine Learning

Laminarin stimulates single cell rates of sulfate reduction whereas oxygen inhibits transcriptomic activity in coastal marine sediment

Abstract The chemical cycles carried out by bacteria and archaea living in coastal sediments are vital aspects of benthic ecology. These ecosystems are subject to physical disruption, which may allow for increased respiration and complex carbon consumption—impacting chemical cycling in this environment often thought to be a terminal place of deposition. We use the redox-enzyme sensitive probe RedoxSensor Green to measure rates of electron transfer physiology in individual sulfate reducer cells residing in anoxic sediment, subjected to transient exposure of oxygen and laminarin. We use index fluorescence activated cell sorting and single cell genomics sequencing to link those measurements to genomes of respiring cells. We measure per-cell sulfate reduction rates in marine sediments (0.01–4.7 fmol SO42− cell−1 h−1) and determine that cells within the Chloroflexota phylum are the most active in respiration. Chloroflexota respiration activity is also stimulated with the addition of laminarin, even in marine sediments already rich in organic matter. Evaluating metatranscriptomic data alongside this respiration-based technique, Chloroflexota genomes encode laminarinases indicating a likely ability to degrade laminarin. We also provide evidence that abundant Patescibacteria cells do not use electron transport pathways for energy, and instead likely carry out fermentation of polysaccharides. There is a decoupling of respiration-related activity rates from transcription, as respiration rates increase while transcription decreases with oxygen exposure. Overall, we reveal an active community of respiring Chloroflexota that cycles sulfate at potential rates of 23–40 nmol h−1 per cm3 sediment in incubation settings, and non-respiratory Patescibacteria that can cycle complex polysaccharides.

Lindsay, Melody R.

Transcriptomic and functional analyses uncover a conserved effector driving genotype-dependent virulence in the Sphaerulina musiva-Populus trichocarpa interaction

The introduction of invasive microbes compromises the structure, biodiversity, and function of naïve ecosystems. Sphaerulina musiva, a hemibiotrophic pathogen that causes leaf spot and stem cankers in Populus species, exemplifies an invasive fungal pathogen spread by human activities. However, the genetic mechanisms of pathogenicity and virulence are poorly understood, impeding mitigation strategies. We utilized RNA sequencing to identify fungal effectors linked to stem canker formation, informing the development of future strategies for effective disease management. Our analysis revealed 70 genes differentially expressed at 2 weeks and 110 genes at 3 weeks between inoculated trees and controls. Notably, the gene with the highest expression at 2 weeks and the second highest at 3 weeks was homologous to Extracellular protein 2 (Ecp2). Complementary genome-wide association studies linked sequence polymorphisms in this locus to phenotypic variation in disease severity. Infiltration of S. musiva Ecp2 into Populus trichocarpa leaves induced necrosis in susceptible genotypes. Gene disruption using a CRISPR-Cas9 RNP system resulted in a genotype-dependent reduction of stem canker and disease severity. Tracing the evolutionary history of this effector across the fungal kingdom, we uncovered clade-specific gene-family expansions and orthologs in new species. These findings raise questions about the function and adaptive significance of these gene families in fungal lifestyles. Our study provides the first tractable target for breeding resistant poplar genotypes, addressing the challenges of managing S. musiva and uncovering mechanisms that drive its virulence, and provides deeper insights into the evolutionary dynamics of a conserved small-secreted protein with a diversity of functions.

Sondreli, Kelsey L [Oregon State University]

Transcriptomic data sets examining several stress responses in Zymomonas mobilis strains

We used RNA-seq to compare gene expression from Zymomonas mobilis ZM4 grown under various conditions: aerobic ± paraquat, anaerobic ± hydrogen peroxide, or an iron chelator. We analyzed two mutant strains lacking predicted transcription factors (ZMO_0442 and ZMO_1411) grown under aerobic or anaerobic conditions. We report the RNA-seq data from these experiments.

Bacterial Stress Response

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive full-length, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

Loving, Rebekah K. (ORCID:0000000187250376)