Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

102 records · Page 6

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology↗

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

Does Collection Time Bias the Ecology of Cleanroom Air Samples?

Microbial monitoring of astromaterials collections has taken on increased importance with the return of biologically sensitive samples from the asteroids Ryugu and Bennu and the initiation of the Mars Sample Return Program. Terrestrial bacteria and fungi can alter the mineralogy and organic composition of our collections causing irreversible contamination of pristine samples and increasing the risk of false positives for life detection measurements. NASA has conducted routine microbial monitoring of its existing collections since 20181. Initial monitoring focused on surface samples collected with foam swabs. Although, airborne microbiology is often decoupled from surface microbiology in the built environment2 culture-based air sampling techniques like impactors were not compliant with existing contamination control requirements. Bringing organic rich media, gelatin or liquids into curation cleanrooms presents an unacceptable risk to pristine samples. In 2022 NASA purchased a materials complaint air sampler and began collecting air samples from the cleanrooms in addition to surface samples3. The new instrument uses an electret filter to collect samples that are suitable for cultivating organisms or for direct DNA sequencing. Preliminary DNA sequencing results appeared to indicate that longer sampling times biased the microbial community in favor of hearty, spore-forming bacteria3. We present the results of a study comparing overnight sampling (17 hours) to short (1 hour) sampling of unoccupied curation cleanrooms. The results will help us optimize our monitoring protocols and develop a more detailed inventory of the ecology of astromaterials curation cleanrooms. Methods: We analyzed 72 paired air samples from six different cleanrooms including the meteorite processing lab (ISO 7 equivalent, 16 samples), the lunar lab (ISO 6 equivalent, 10 samples), the stardust lab (ISO 5 equivalent 14 samples), the OSIRIS-REx lab (ISO 5 equivalent, 12 samples), the Hayabusa2 lab (ISO 5 equivalent, 14 samples), and the Genesis lab (ISO 4 equivalent, 6 samples). All the samples were collected with an InnovaPrep Bobcat air sampler operating at a sampling rate of 200 L/min. The sampler operates for 5 minutes out of every 20 minute period. Half of the samples were collected by filtering 3,000L (15 min. of active sampling) of air across an electret filter for one hour. The rest of the samples were collected by filtering approximately 51,000 L air across the filter overnight (~17 hours, 255 min. of active sampling). Cells were eluted from the filter using 6-7 ml of pressurized 0.15% tween 20 in PBS (phosphate buffered saline). This liquid was used to cultivate bacteria according to previously published methods1,4,5 and for DNA extraction and next generation sequencing. DNA was extracted with a Qiagen MagAttract PowerMicrobiome kit6. To identify bacteria and archaea, the 16S rRNA gene was amplified using Earth Microbiome primers for the V4 region 7. The amplified DNA was sequenced on an Illumina MiSeq using a V3 reagent kit. The resulting sequences were processed using DADA2 and QIIME2 as implemented on the EDGE bioinformatics platform8–10. Results: Only two of the 72 samples had no amplifiable DNA. Amplified DNA concentrations ranged from 2.67 – 0.272 ng/µl. The median concentration of amplified DNA for the 1 hour samples was 0.770 ± 0.368 ng/µl. The median concentration of amplified DNA for the overnight samples was 0.877 ± 0.434 ng/µl. On average the overnight samples had slightly more sequences (58,960 vs. 59,456) and ASV’s (amplicon sequence variants) (60 vs 64.5) than the one hour samples, but these differences are not statistically significant. The most abundant ASV in every sample mapped to the genus Cupravidus. ASV’s mapping to the genuses Bacillus, Schlegelella, Thermus, and Staphylococcus were also common. Discussion and Future Work: Alpha diversity statistics like Shannon Entropy and Faith Phylogenetic Diversity are used to describe the diversity of organisms in a single sample. If a longer sampling time was biasing the data, we would expect to see a change in these diversity statistics vs. sample time. However, we did not observe this in our data. The median Shannon entropy was slightly higher for the overnight samples (3.773 vs 3.611) as was the Faith Phylogenetic Diversity (4.042 vs 3.596), but both values were within a standard deviation of each other for the two sampling times (Fig. 1). It is unlikely, that the longer sampling time is introducing bias into our data. We do observe a significant decrease in diversity when comparing the air samples by lab. The Genesis lab (ISO 4 equivalent) has a lower median number of ASV’s (45.5) than the other labs (62). Median values for Shannon Entropy (3.717 vs. 3.430) and Faith Phylogenetic Diversity (3.796 vs. 3.548) are also lower for Genesis, but those values are with one standard deviation of each other for the different sampling times. This is consistent with previous culture-based results suggesting that the environment in cleanrooms tends to select for a core group of organisms capable of surviving under dry, low nutrient, conditions. The presence of the ASV’s mapping to Cupravidus and Thermus in our sequencing blanks and controls suggests that several of the most common organisms in our samples represent contaminants from the reagents used to perform the DNA extractions and sequencing. Further work is needed to identify these contaminants, remove them from our data and recalculate the diversity statistics. This is a systematic error. Therefore, we do not expect removing the sequencing contaminants to change our conclusions. Longer air sample collection times appear to result in slightly higher diversity and do not bias the results towards “hardy” bacteria like spore-formers. Based on these preliminary results we conclude that sampling at least 3,000 liters of air is sufficient to capture the microbial diversity of cleanrooms, and that air samples can also be collected overnight without negatively impacting diversity. These results allow us to be flexible when designing microbial monitoring plans so that they do not interfere with routine lab activity. References: 1. Regberg, A. B. et al. 49th Lunar and Planetary Science Conference (2018). 2. The United States Pharmacopeial Convention. USP General Chapter <1116> (2013). 3. Regberg, A. B., et al. 54th Lunar and Planetary Science Conference (2023). 4. Regberg, A. B. et al. 53rd Lunar and Planetary Science Conference ( 2022). 5. Davis, R. E.,et al. 50th Lunar and Planetary Science Conference (2019). 6. Qiagen. MagAttract® PowerMicrobiome® DNA/RNA EP Kit Handbook. (2018). 7. Walters, W. et al. mSystems 1, (2015). 8. Callahan, B. J. et al. Nat. Methods 13, 581–583 (2016). 9. Hall, M. & Beiko, R. G. Microbiome Analysis: Methods and Protocols113–129 (Springer, 2018). 10. Philipson, C. et al. Bio-Protoc. 7, e2622 (2017).

A. B. Regberg↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Benchmarking Computational Tools for Calling SNPs and Indels in Complex Microbial Populations

The NASA BioNutrients missions seek to understand the suitability of microorganisms for bioproduction during space flight. One topic of interest is the stability of microbial genomes during long-term ambient storage and subsequent rehydration and growth. To address these questions, samples from 8 species were flown to ISS for 5 years of desiccated storage at ambient temperature (Stasis Packs) and 2 species were packaged along with powdered media inside a bioreactor system to allow hydration and growth in microgravity (Production Packs). For both systems, Whole Genome Sequencing (WGS) of the DNA extracted from the returned samples and paired ground controls will be conducted to identify changes in genome stability due to time, storage conditions and growth in space. Across the technical replicates, ground controls, 10 timepoints, and multiple experimental conditions, ~300 samples have been selected for initial analysis with WGS sequencing to 100x coverage. A flexible and resource efficient mutation calling pipeline is needed to process this large dataset and allow for comparisons between species. Many bioinformatics tools for calling Indels and Single Nucleotide Variants (SNVs) are designed for use with pure isolates, where true variations from the reference genome are expected to dominate the reads aligning to the location of mutation. In contrast, DNA from the Stasis Pack (SP) samples was collected directly after recovery from desiccated storage and the Production Pack (PP) samples were collected after fermentation. In this context, reads with mutations are expected to be less frequent than reads that align with the reference genome, as each sample will include multiple lines of cells. Thus, BioNutrients samples are expected to be similar to samples from cancer cell or “pooled” sequencing approaches. In preparation for the analysis of the BioNutrients samples, we have tested three mutation calling tools (GATK for Microbes, BreSeq and DiscoSNP) designed for complex samples. A challenge of validating mutation identification pipelines is a lack of “Ground Truth” datasets, especially for complex samples. To compare these three tools, we sought to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO as part of the Space Algae-1 mission. Here we present a summary of these tools against the analysis originally conducted using the CRISP tool. Critical metrics are compared such as runtime, the number of SNPs, the number and size of Indels, and patterns of transversion and transitions identified by each tool are reported. By sharing these benchmarking results collected in support of the BioNutrients mission, we aim to guide others seeking to identify SNVs in similarly complex microbial samples.

Biology↗

Does the International Space Station Leak DNA? Preliminary Results from the ISS External Microorganisms Payload

Existing crewed spacecraft like the ISS (International Space Station) leak by design. The ISS routinely releases gas to maintain life support systems and when astronauts exit the station to perform space walks. The chemical component of this leakage is well characterized, but the biological components are not. The ISS is not subject to planetary protection requirements, but planned missions to Mars will use similar systems and will be subject to planetary protection requirements. If detectable microorganisms are escaping through vents and or airlocks we may need to redesign our crewed habitats to minimize this type of contamination. To test the hypothesis that microorganisms from inside ISS are detectable on exterior surfaces an astronaut used the ISS External Microorganisms sampling kit (Rucker et al. 2018) to sample exterior surfaces of the ISS during an EVA (Extra Vehicular Activity) in January of 2025. These samples were returned to Earth for DNA extraction and sequencing. We successfully, extracted and sequenced bacterial, fungal and viral DNA from these samples that was not present in the negative controls. These results should help NASA refine the planetary protection requirements for crewed missions. Methods: The samples were collected using sterile, DNA free, buccal swabs (23 mm. diameter) housed in custom canisters. Each canister uses a 0.2 μm Teflon filter to maintain sterility as the caddy, holding 8 swabs moves in and out of vacuum. The astronaut sampled the: 1) airlock vestibule, 2) airlock thermal cover, 3) a gap in the micrometeorite shielding near the airlock, 4) a handrail near the airlock, 5) the Carbon Dioxide Removal Assembly vent, and 6) the Vacuum Exhaust System vent. The seventh swab was exposed to vacuum during the EVA without touching it to a surface. The eighth swab, a negative control, was not opened until the caddy returned to Earth. DNA was extracted from the swabs using a QIamp UCP Pathogen kit and prepared for sequencing on an Aviti (Element Biosciences) sequencer (Arslan et al. 2024). The resulting sequences were analyzed using the EDGE Bioinformatics platform (Li et al. 2017). The sequences were analyzed individually using tools like BLAST, GOTTCHA2, Kraken2, and PanGIA. The data were also assembled into metagenome assembled genomes) using tools like CONCOCT, MaxBin2 and MetaBAT2. Results: We successfully extracted and sequenced bacterial, archaeal, fungal and viral DNA from all seven samples. The handrail swab had the lowest number of reads (768,651) and the airlock thermal cover had the highest number of reads (8,819,230). These samples contain DNA from human associated bacteria (e.g. Crynebacterium riegelii ), fungi (.e.g. Penicillium rubens ), and viruses (e.g Alphapapillomavirus ). Conclusion: Preliminary interpretation suggest that the airlock and the space suits themselves are the largest sources of contaminant DNA. Most if not all of the DNA is from organisms known to be present inside the ISS. Vents attached to life support systems may be a lesser source of biological contamination. Further analysis should help NASA address planetary protection knowledge gaps for crewed missions.

Aaron B Regberg↗

Genelab: Scientific Partnerships and an Open-Access Database to Maximize Usage of Omics Data from Space Biology Experiments

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. The GeneLab Data System (GLDS) is NASA's premier open-access omics data platform for biological experiments. GLDS houses standards-compliant, high-throughput sequencing and other omics data from spaceflight-relevant experiments. The GeneLab project at NASA-Ames Research Center is developing the database, and also partnering with spaceflight projects through sharing or augmentation of experiment samples to expand omics analyses on precious spaceflight samples. The partnerships ensure that the maximum amount of data is garnered from spaceflight experiments and made publically available as rapidly as possible via the GLDS. GLDS Version 1.0, went online in April 2015. Software updates and new data releases occur at least quarterly. As of October 2016, the GLDS contains 80 datasets and has search and download capabilities. Version 2.0 is slated for release in September of 2017 and will have expanded, integrated search capabilities leveraging other public omics databases (NCBI GEO, PRIDE, MG-RAST). Future versions in this multi-phase project will provide a collaborative platform for omics data analysis. Data from experiments that explore the biological effects of the spaceflight environment on a wide variety of model organisms are housed in the GLDS including data from rodents, invertebrates, plants and microbes. Human datasets are currently limited to those with anonymized data (e.g., from cultured cell lines). GeneLab ensures prompt release and open access to high-throughput genomics, transcriptomics, proteomics, and metabolomics data from spaceflight and ground-based simulations of microgravity, radiation or other space environment factors. The data are meticulously curated to assure that accurate experimental and sample processing metadata are included with each data set. GLDS download volumes indicate strong interest of the scientific community in these data. To date GeneLab has partnered with multiple experiments including two plant (Arabidopsis thaliana) experiments, two mice experiments, and several microbe experiments. GeneLab optimized protocols in the rodent partnerships for maximum yield of RNA, DNA and protein from tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected on the ground. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and as well as yield terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space environments.

bioinformatics↗

Increasing the Statistical Rigor of Cross-Species Differential Expression Analysis

Microgravity inflicts substantial, but undercharacterized, pressure on organisms that induces metabolic responses such as increased microbial virulence and antibiotic resistance, altered organ weights in developing rats, and loss of bone tissue in astronauts. Numerous studies have analyzed the effects of microgravity on specific organisms, tissues, or test conditions, but these projects are necessarily limited by the small sample size of space research. Increasing the sample size of spaceflight studies is non-trivial; however, pooling data from numerous studies can greatly increase the statistical rigor of comparative analyses. The GeneLab houses datasets from 73 spaceflight studies that performed transcription profiling assays. These data encompass a diverse array of organisms ranging from Escherichia coli to Mus musculus to Homo sapiens and comprise studies analyzing ionizing radiation, mammalian pregnancy, etc. Collectively, the GeneLab database contains a large quantity of transcription assays and RNA sequence data analyzing Differential Gene Expression (DGE) between microand normogravity. Xspecies, a cross-species analysis method for DGE developed by Kristiansson, et al. in 2012, identifies homologous genes between species that are universally up- or downregulated in response to test conditions. Previous work by an intern at GeneLab applied Xspecies to 19 datasets containing seven different species and identified 14 homologous groups differentially expressed under spaceflight conditions including several heat shock proteins and cytoskeletal components. Unfortunately, these results may be biased by the disproportionate number of studies on Arabidopsis thaliana (5) and Mus musculus (6) and the results are not normalized by evolutionary distances. Here, we present modifications to the Xspecies algorithm that permits incorporation of multi-omic data and normalizes data for effect size, directionality, and evolutionary distances. We then apply this algorithm to all currently available GeneLab studies

Xspecies↗

Graph Representation Learning for Dengue Forecasting

In 2017, the largest recorded dengue outbreak in Sri Lanka’s history occurred. Since then, dengue has continued to threaten national health across Sri Lanka. The development of an effective Early Warning System (EWS) for dengue outbreaks is essential for Sri Lanka’s Ministry of Health to take preventative measures. We propose the use of Graph Neural Networks as EWS. Using earth observational data from NASAs global satellites and dengue incidence data from Sri Lanka s Ministry of Health, we developed a series of traditional and graph representation EWS to forecast Dengue cases across Sri Lanka’s 25 districts between 2013 and 2022. We demonstrate empirically that Graph Neural Networks which incorporate spatiotemporal relations significantly outperform traditional EWS such as Autoregressive Integrated Moving Average (ARIMA), Random Forest, and Long Short-Term Memory (LSTM). Our source code is available on GitHub and will be provided in the final submission.

Graph Neural Networks↗