Evaluation of sea salt aerosols in climate systems: global climate modeling and observation-based analyses*
Explore the source record for details and available documents.
Engineering topics
Publications and source records attributed to Yi-Chun Chen.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Transcriptome profiling by RNA sequencing (RNA-seq) is a powerful approach to identify gene expression changes in organisms exposed to unique environments such as spaceflight. One of the challenges of evaluating RNA-seq data both within and across different space-relevant studies is the ability to control for technical differences, including the use of different library preparation kits, sequencing platforms, RNA yield, and person-to-person variation. To help address this issue, the National Institute of Standards and Technology (NIST, nist.gov) initiated a consortium, at the request of industry and academia, to develop a set of controls for gene expression measurements. The result was a set of 92 unlabeled, polyadenylated transcripts that range from 250 – 2,000 nucleotides in length to mimic natural eukaryotic mRNAs. These External RNA Controls Consortium (ERCC) genes can be used in any RNA-seq experiment, by adding known concentrations of the ERCC genes to samples after RNA extraction, to offer a standard measurement for data comparison. At NASA GeneLab, we employ these controls as part of our standard operating procedures for every in-house RNA-seq study to assess the limit of detection, dynamic range, and power of differential expression analysis both within and across experiments. Here we will discuss the use, benefits, and limitations of ERCC genes and other types of controls, such as universal RNA references, to generate quality control information for RNA-seq studies conducted at GeneLab.
RNA sequencing (RNA-seq) data from space biology experiments promise to yield invaluable insights into the effects of spaceflight on terrestrial biology. However, sample numbers from each study are low due to limited crew availability, hardware, and space. To increase statistical power, spaceflight RNA-seq datasets from different missions are often aggregated together. However, this can introduce technical variation or "batch effects", often due to differences in sample handling, sample processing, and sequencing platforms. Several computational methods have been developed to correct for technical batch effects, thereby reducing their impact on true biological signals. In this study, we combined 7 mouse liver RNA-seq datasets from NASA GeneLab (part of the NASA Open Science Data Repository) to evaluate several common batch effect correction methods (ComBat and ComBat-seq from the sva R package, and Median Polish, Empirical Bayes, and ANOVA from the MBatch R package). We quantitatively evaluated the ability of these methods to correct for technical batch variables in space biology RNA-seq data using the following criteria: BatchQC, principal component analysis, dispersion separability criterion, log fold change correlation, and differential gene expression analysis. Each batch variable / correction method combination was then assessed using a custom scoring approach to identify the optimal correction method for the combined dataset, by geometrically probing the space of all allowable scoring functions to yield an aggregate volume-based scoring measure. Finally, we describe the way in which the GeneLab multi-study analysis and visualization portal will allow users to examine the presence or absence of batch effects using multiple metrics. If the user chooses to perform batch effect correction, the scoring approach described here can be implemented to identify the optimal correction method to use for their specific combined dataset prior to analysis.
Explore the source record for details and available documents.