Engineering PapersSearch

SEARCH · Engineering Papers

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

GL4U: Bioinformatics training for students and educators using space omics data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler

GL4U: Bioinformatics Training for Students and Educators Using Space Omics Data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler

GL4U: Using Space Biology Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant multi-omics data via the Open Science Data Repository (OSDR) that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct (training students) and indirect (training educators) approaches. The GL4U pilot programs were conducted in June 2021 (direct training) and 2022 (indirect training). During the pilots, students and educators at Historically Black Colleges and Universities (HBCUs) and Minority Serving Institutions (MSIs) participated in a week-long (direct training) or two-week-long (indirect training) bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze space biology RNA sequencing data from OSDR. During the educator pilot, participants received materials, training, and the necessary compute resources to enable them to run the bootcamp at their home institutions, thereby extending the reach of this initiative. In July 2023, GeneLab is partnering with JPL to expand GL4U to include amplicon sequencing (Amp-Seq) analysis training. During the GL4U Amp-Seq bootcamp, student and educator participants will receive training on how to analyze and interpret Amp-Seq data using the NASA GeneLab data processing pipeline. All bootcamp material, including instructions for requesting compute resources, will be made publicly available on GitHub for educators to teach the GL4U content in subsequent semesters. GL4U provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. We present results from pre- and post-training surveys completed by all participants of the Amp-Seq bootcamp.

Amanda M Saravia-Butler

GeneLab: Omics Data System for Space Biology Research

During spaceflight, a complex set of detrimental factors impinge upon astronauts and other biological systems. To help understand this dynamic, NASA is developing GeneLab, an open access data system to encourage widespread analysis of spaceflight relevant omics data.

Omics

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab

GeneLab: Current and Future Omics Data Integration Between Space Biology and HRP

For the past five years, the Biological and Physical Sciences Division has pioneered Open Science in Space Biology by funding the NASA GeneLab project. Along with the Ames Life Sciences Data Archive, GeneLab has quickly become the world leader in archiving and scientifically curating spaceflight and spaceflight relevant multi-omics data. Specifically, the GeneLab Data System has become a full enterprise solution providing advanced mining capabilities, several application programming interfaces for data federation and machine learning approaches, and delivering to the world an analytical and visualization platform which has enabled collaboration within the scientific community. Over the past three years, large meta-analysis and modeling studies have been published by the GeneLab Analysis Working Groups (AWGs), which are comprised of ~200 volunteer scientists. One natural extension of GeneLab data reuse has recently turned towards linking animal data with human data, which is the next necessary step to further validate animal models for inferring biological risks to humans conducting LEO, lunar or Martian missions. As such, data from the Human Research Program are an essential component of GeneLab and ALSDA. At the moment, simulated space radiation experiments conducted at Brookhaven National Laboratory make the most of HRP GeneLab data, and the scientific community has been eager to also link their animal spaceflight results to actual Astronaut data and human analog data. We will discuss further the current status of knowledge and future approaches to accelerate our basics understanding of the impact of space stressors on humans using latest omics technology.

omics

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES

GeneLab: "Omics" Data Systems for Spaceflight and Simulated Spaceflight Environment

NASA's GeneLab Data System is a repository that hosts multi-omics datasets generated by biological experiments flown onboard the International Space Station. Strategies regarding how GeneLab envisions the involvement of the scientific community and the public at large will be discussed, and current and future capabilities of the system will be described. Information describing how scientists can participate in analyzing the current datasets on plants, microbes, invertebrates or mammals will be provided, and initial findings from the current datasets will be discussed during this presentation. Anyone interested in genomics, transcriptomics, epigenomics and proteomics, and systems biology, or who is curious to understand how space modifies living organisms should attend.

Omics

GeneLab: "Omics" Data System for Space Biology Research

Determining the biological impact of spaceflight through novel approaches is essential to reduce the health risks to astronauts for long-term space missions. The current established health risks due to spaceflight are only reflecting known symptomatic and physiologic responses and do not reflect early onset of other potential diseases. There are many unknown variables which still need to be identified to fully understand the health impacts due to the environmental factors in space. One method to uncover potential novel biological mechanisms responsible for health risks in astronauts is by utilizing NASA's GeneLab Data Systems (genelab.nasa.gov). GeneLab is public repository that hosts multiple omics datasets generated from space biology experiments that include experiments flown in space, simulated cosmic radiation experiments, and simulated microgravity experiments. This presentation will provide an overview of GeneLab and examples of analysis that are being done with GeneLab datasets. These example will include novel data and work being generated with various scientists around the world involved with GeneLab's Analysis Working Groups (AWG) that are assisting with the development of pipelines and advancing GeneLab to the next phase, a publication from GeneLab discovering novel Carbon Dioxide impact due to rodent habitats, another publication from GeneLab discovering a potential master regulator responsible for health risk associated due to spaceflight, and potential cardiovascular risk from space radiation.

microgravity

GeneLab: "Omics" Data Systems for Spaceflight and Simulated Spaceflight Environment

Determining the biological impact of spaceflight through novel approaches is essential to reduce the health risks to astronauts for long-term space missions. The current established health risks due to spaceflight are only reflecting known symptomatic and physiologic responses and do not reflect early onset of other potential diseases. There are many unknown variables which still need to be identified to fully understand the health impacts due to the environmental factors in space. One method to uncover potential novel biological mechanisms responsible for health risks in astronauts is by utilizing NASA's GeneLab Data Systems (genelab.nasa.gov). GeneLab is public repository that hosts multiple omics datasets generated from space biology experiments that include experiments flown in space, simulated cosmic radiation experiments, and simulated microgravity experiments. This presentation will provide an overview of GeneLab and examples of analysis that are being done with GeneLab datasets. These example will include novel data and work being generated with various scientists around the world involved with GeneLab's Analysis Working Groups (AWG) that are assisting with the development of pipelines and advancing GeneLab to the next phase, a publication from GeneLab discovering novel Carbon Dioxide impact due to rodent habitats, another publication from GeneLab discovering a potential master regulator responsible for health risk associated due to spaceflight, and potential cardiovascular risk from space radiation

spaceflight

Science Driven By Space Biology Omics Data Utilizing NASA's GeneLab Platform

Determining the biological impact of spaceflight through novel approaches is essential to reduce the health risks to astronauts for long-term space missions. The current established health risks due to spaceflight are only reflecting known symptomatic and physiologic responses and do not reflect early onset of other potential diseases. There are many unknown variables which still need to be identified to fully understand the health impacts due to the environmental factors in space. One method to uncover potential novel biological mechanisms responsible for health risks in astronauts is by utilizing NASA's GeneLab Data Systems (genelab.nasa.gov). GeneLab is public repository that hosts multiple omics datasets generated from space biology experiments that include experiments flown in space, simulated cosmic radiation.

Beheshti, Afshin

Multi-omics data resource: Data package 22 (Pck022)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines. The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and submitted for scRNA-seq. This study focused on understanding the heterogeneity of the cytokine-mediated response. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE156175 Publication: 10.26508/lsa.202000949

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 10 (Pck010)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 11 (Pck011)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab