Open Science for Life in Space
Explore the source record for details and available documents.
Engineering topics
Publications and source records attributed to Sylvain V Costes.
Explore the source record for details and available documents.
The NASA GeneLab Sample Processing Laboratory offers high-throughput sequencing services to NASA-funded space biology researchers. Space biology studies have specific challenges such as low sample numbers, introducing susceptibility to batch effects from sample handling. These issues are compounded by complex protocols such as single-nuclei isolation and sequencing, which has recently become an attractive methodology for assessing the cellular diversity within spaceflight samples. High quality single-nuclei sequencing requires reproducible protocols to dissociate tissue and generate clean suspension of intact single nuclei. Producing single-nuclei suspension from brain tissue is particularly challenging due to cell type heterogeneity and the myelin sheath that carries over into the nuclei suspension as debris. Current procedures tend to be time consuming and sometimes include steps that can alter gene expression and create cell-type bias. Commercially available nuclei isolation kits, such as the 10X Genomics nuclei isolation kit, offers a streamlined way to process samples for nuclei isolation, thereby minimizing batch effects and enabling reproducibility. In this study, we report on the performance of the 10X Genomics nuclei isolation kit and Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit to generate sequencing libraries from space-flown mouse brain samples. Single nuclei sequencing was performed on frozen mouse brain tissue from two spaceflight missions, Rodent Research-10 (RR-10) and RR Reference Mission-2 (RRRM-2). RR-10 mice were female B6129SF2/J, euthanized at 18-19 weeks whereas RRRM-2 mice were female C57BL/6NTac, euthanized at 20 or 37 weeks. Sequencing data was processed using standard GeneLab data processing pipelines. We report evaluation of the performance of the 10X Genomics nuclei isolation kit for spaceflight samples from mouse brain, and evaluation of reproducibility across different mouse strains and age groups. We also report preliminary scientific results including cell type inference, cell clustering, and differentially expressed genes and pathways between spaceflight and ground control samples.
NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.
NASA’s GeneLab project provides researchers open access to space-relevant multi-omics data via the Open Science Data Repository (OSDR) that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct (training students) and indirect (training educators) approaches. The GL4U pilot programs were conducted in June 2021 (direct training) and 2022 (indirect training). During the pilots, students and educators at Historically Black Colleges and Universities (HBCUs) and Minority Serving Institutions (MSIs) participated in a week-long (direct training) or two-week-long (indirect training) bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze space biology RNA sequencing data from OSDR. During the educator pilot, participants received materials, training, and the necessary compute resources to enable them to run the bootcamp at their home institutions, thereby extending the reach of this initiative. In July 2023, GeneLab is partnering with JPL to expand GL4U to include amplicon sequencing (Amp-Seq) analysis training. During the GL4U Amp-Seq bootcamp, student and educator participants will receive training on how to analyze and interpret Amp-Seq data using the NASA GeneLab data processing pipeline. All bootcamp material, including instructions for requesting compute resources, will be made publicly available on GitHub for educators to teach the GL4U content in subsequent semesters. GL4U provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. We present results from pre- and post-training surveys completed by all participants of the Amp-Seq bootcamp.
Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.
Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.
Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields in the last two decades, in part thanks to an increasing culture of open data sharing and reuse. Due to its capability for identifying complex relationships and patterns, AI/ML methodology is particularly well suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are many key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Even with the positive culture of Open Science and data sharing, inexperienced researchers working quickly without proper checks can produce models that perform poorly outside of the immediate training dataset. Lessons learned from biological AI/ML research indicate that Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.
The RadLab Project, initiated by the Open Science Data Repository (OSDR) for Space Biology at NASA Ames Research Center (osdr.nasa.gov), aims to be an ongoing compilation of radiation data relevant to human space flight. Sponsored by the NASA Human Research Program, RadLab serves as the latest database in this domain. RadLab is intended to serve the needs of both the space radiation detector and space radiation biology communities. The RadLab architecture is being designed to accommodate data both from low Earth orbit (LEO) and beyond low Earth orbit (BLEO). Our long-term vision entails the establishment of a sustainable database receiving continuous updates through APIs connecting to other databases, as well as individual investigator contributions via a RadLab submission portal, modeled on the OSDR (https://osdr.nasa.gov/bio/submission-sso-login.html). In addition to serving as a resource for space biologists and space radiation physicists, such a database would facilitate the deployment of AI algorithms to study the impact of location in space and within spacecraft on the ambient space radiation field. The ultimate goal is to develop predictive dosimetry algorithms for future BLEO missions.
Explore the source record for details and available documents.
NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.
Sensors on multiple spacecraft in and beyond low Earth orbit continuously monitor and collect space radiation data and transmit it back to Earth. These data are of vast importance to space biology research, as ionizing radiation affects living organisms—astronauts and non-human experiment subjects alike—placing them at higher risk of carcinogenesis, degenerative diseases, and radiation sickness. Therefore, knowledge of the biological effects of space radiation is essential for planning future crewed missions beyond low Earth orbit. The RadLab project, initiated by GeneLab and ALSDA (the Open Science Data Repository; OSDR) and sponsored by the NASA Human Research Program, is a new effort aimed at connecting dosimetry data from radiation detectors located on the International Space Station (ISS), as well as other spacecraft. To date, access to these data has been fragmented across space agencies and databases; to address this issue, we have developed an application programming interface (API) and an associated graphical user interface (GUI) designed to provide a single point of access to the data. As of now, OSDR has focused on the detectors located on the ISS, with the long-term goal to establish a self-sustained portal receiving continuous updates through APIs connecting to multiple radiation databases of varying scope, as well as individual investigator contributions. The RadLab API implements a request syntax enabling users to query data by craft, sensor type, timespan, etc, allowing for arbitrary combinations of original source data, thus providing programmatic access for use in computational pipelines, while the GUI facilitates data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.
Explore the source record for details and available documents.
The GeneLab data repository provides an unparalleled resource for exploring how spaceflight affects organisms with omics-level insights. However, two major interlinked challenges to capitalizing on the information within these data are their vast breadth and the often-specialized expertise that has been required in the past for their analysis. How do you compare responses within and between studies, especially if you are a non-bioinformatics specialist? This presentation will discuss how Space Biology data can be accessed using software to help provide these data resources to address research questions and generate new hypotheses. The presentation will cover a wide range of the available space life science tools but will focus on TOAST, MANGO, the MATRIX, RadBioApp and other interactive relational databases (https://genelab.nasa.gov/external-vis-apps). These exploration environments have been developed to search the GeneLab data repository for new insights that inform how model organisms respond to microgravity, radiation and other factors associated with spaceflight. The presentation will be interactive, and participants will have the opportunity to ask questions and learn more about the data viz and modeling tools that are available to them.
Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.
To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.
Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.
Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, to foster collaborations with equity. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within the Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate with community members and set new standards for space-relevant data and metadata. The AWGs welcome any ASGSR members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are encouraged to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).
The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.