Engineering PapersSearch

Engineering topics

Samrawit Gebre

Publications and source records attributed to Samrawit Gebre.

GeneLab: The NASA Systems Biology Platform for Space Omics Repository, Analysis and Visualization

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data, and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 220 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetery data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 120 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

Samrawit Gebre

GeneLab: The NASA System Biology Platform for Space Omics Repository, Analysis and Visualization

NASA’s GeneLab includes an open-access repository of some 250+ omics datasets generated by biological experiments relevant to spaceflight including simulated cosmic radiation and microgravity. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics background, GeneLab has become a knowledgebase platform converting raw genetic and proteomic signatures found in flight samples into biological and physiological meanings. A large community of more than 100 scientists has rallied behind GeneLab and organized into four Analysis Working Groups (AWGs: Animal, Plant, Microbe, and Multi-Omics). Together, the AWGs have gained scientific recognition worldwide by establishing a consortium in charge of adopting new complex standards for data analysis workflows and omics sample processing in a rapidly evolving field. We will demonstrate the usage of the repository with smart search capability, an online controlled-access toolshed "Galaxy" to process user data with vetted standard workflows, a workspace for data sharing and a data submission portal with ontology control for better metadata curation. The GeneLab visualization portal will also be demonstrated, showing how anyone without formal training in bioinformatics can now browse the space biology omics data to discover new biology and potential solutions to improve life in space.

GeneLab

NASA GeneLab: The NASA Systems Biology Platform for Spaceomics Repository, Analysis and Visualization

At NASA Ames Research Center, the GeneLab Open Science Project is on a mission to gather all large -omics datasets relevant to space biology research. These datasets come from various organisms flown in multiple space habitats such as the International Space Station or the Space Shuttle, in addition to mimicking space-like conditions on ground. Researchers and citizen scientists all around the world have used the data and the analytical tools put together by the GeneLab team to start deciphering new biological impact of microgravity, space ionizing radiation and other space stressors.

GeneLab

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 350 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab Sequencing Lab. The GLDS contains rich metadata about each experiment and has integrated radiation dosimetry data from experiments flown on the Space Shuttle, International Space Station, and Free Flying spacecrafts. With the increasing amount and complexity of omics data being generated, GeneLab utilizes community-defined, common models for metadata and terminology so that omics data and results are discoverable and reliably reproducible. GeneLab uses the ISA-Tab specification and semantic model for organizing and representing omics metadata. In addition to metadata standards, data files must be open-source file or common exchange formats to ensure accessibility and usability by all users. To ease data ingestion and transfer, the web-based submission tool allows PIs a user-friendly user interface to curate, organize, and publish their space relevant omics data. In the more recent years, data curation and submission portal has incorporated the FAIR principles making data findable, accessible, interoperable, and reusable. To increase reusability of data, GeneLab has implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 200 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. To train the next generation of scientists, NASA offers training programs such as GeneLab 4 High School (GL4HS) and GeneLab 4 Universities. NLM Curation at a Scale Workshop 2022 | NASA GeneLab (GL4U) to teach students bioinformatics and computational biology methods to analyze omics data. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

GeneLab

Batch Effect Correction Methods for NASA GeneLab Transcriptomic Datasets

RNA sequencing (RNA-seq) data from space biology experiments promise to yield invaluable insights into the effects of spaceflight on terrestrial biology. However, sample numbers from each study are low due to limited crew availability, hardware, and space. To increase statistical power, spaceflight RNA-seq datasets from different missions are often aggregated together. However, this can introduce technical variation or "batch effects", often due to differences in sample handling, sample processing, and sequencing platforms. Several computational methods have been developed to correct for technical batch effects, thereby reducing their impact on true biological signals. In this study, we combined 7 mouse liver RNA-seq datasets from NASA GeneLab (part of the NASA Open Science Data Repository) to evaluate several common batch effect correction methods (ComBat and ComBat-seq from the sva R package, and Median Polish, Empirical Bayes, and ANOVA from the MBatch R package). We quantitatively evaluated the ability of these methods to correct for technical batch variables in space biology RNA-seq data using the following criteria: BatchQC, principal component analysis, dispersion separability criterion, log fold change correlation, and differential gene expression analysis. Each batch variable / correction method combination was then assessed using a custom scoring approach to identify the optimal correction method for the combined dataset, by geometrically probing the space of all allowable scoring functions to yield an aggregate volume-based scoring measure. Finally, we describe the way in which the GeneLab multi-study analysis and visualization portal will allow users to examine the presence or absence of batch effects using multiple metrics. If the user chooses to perform batch effect correction, the scoring approach described here can be implemented to identify the optimal correction method to use for their specific combined dataset prior to analysis.

Lauren M. Sanders

Metadata Entry Optimization for NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Sample Repository

Metadata Entry Optimization For NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Biospecimen

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte

Transcriptomics Processing Pipelines for Space Biology: An Open Source and Consensus-Driven Approach

Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.

Jonathan Oribello

The Environmental Data Application for Analysis of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS and spacecrafts, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (OSDR) has developed a user interface for interrogation of telemetry data: the Environmental Data Application (EDA). The EDA provides the capability to visualize telemetry and radiation data collected on the International Space Station and corresponding ground platforms during the Rodent Research missions. Telemetry data includes temperature, relative humidity, and CO2 levels. Radiation data includes galactic cosmic rays, the contribution of the South Atlantic Anomaly, total radiation dose rate, and accumulated radiation dose. The application allows users to view single missions, compare multiple missions, and view and download summary or full data tables. In summary, the EDA provides GUIs for data visualization and exploration, as well as means for data export, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

telemetry

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado