Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Fast Plasma Instrument for MMS: Data Compression Simulation Results

Magnetospheric Multiscale (MMS) mission will study small-scale reconnection structures and their rapid motions from closely spaced platforms using instruments capable of high angular, energy, and time resolution measurements. To meet these requirements, the Fast Plasma Instrument (FPI) consists of eight (8) identical half top-hat electron sensors and eights (8) identical ion sensors and an Instrument Data Processing Unit (IDPU). The sensors (electron or ion) are grouped into pairs whose 6 deg x 180 deg fields-of-view (FOV) are set 90 deg apart. Each sensor is equipped with electrostatic aperture steering to allow the sensor to scan a 45 deg x 180 deg fan about its nominal viewing (0 deg deflection) direction. Each pair of sensors, known as the Dual Electron Spectrometer (DES) and the Dual Ion Spectrometer (DIS), occupies a quadrant on the MMS spacecraft and the combination of the eight electron/ion sensors, employing aperture steering, image the full-sky every 30-ms (electrons) and 150-ms (ions), respectively. To probe the results in the DES complement of a given spacecraft generating 6.5-Mbs(exp -1) of electron data while the DIS generates 1.1-Mbs(exp -1) of ion data yielding an FPI total data rate of 6.6-MBs(exp -1). The FPI electron/ion data is collected by the IDPU then transmitted to the Central Data Instrument Processor (CIDP) on the spacecraft for science interest ranking. Only data sequences that contain the greatest amount of temporal/spatial structure will be intelligently down-linked by the spacecraft. Currently, the FPI data rate allocation to the CIDP is 1.5-Mbs(exp -1). Consequently, the FPI-IDPU must employ data/image compression to meet this CIDP telemetry allocation. Here, we present simulations of the CCSDS 122.0-B-1 algorithm-based compression of the FPI-DES electron data. Compression analysis is based upon a seed of re-processed Cluster/PEACE electron measurements. Topics to be discussed include: review of compression algorithm; data quality; data formatting/organization; and, implications for data/matrix pruning. To conclude a presentation of the base-lined FPI data compression approach is provided.

Barrie, A.↗

Fast Plasma Investigation for MMS: Simulation of the Burst Triggering System

The Magnetospheric Multiscale (MMS) mission will study small-scale reconnection structures and their rapid motions from closely spaced platforms using instruments capable of high angular, energy, and time resolution measurements. To meet these requirements, the Fast Plasma Instrument (FPI) consists of eight (8) identical half top-hat electron sensors and eight (8) identical ion sensors and an Instrument Data Processing Unit (IDPU). The sensors (electron or ion) are grouped into pairs whose 6 degree x 180 degree fields-of-view (FOV) are set 90 degrees apart. Each sensor is equipped with electrostatic aperture steering to allow the sensor to scan a 45 degree x 180 degree fan about the its nominal viewing (0 deflection) direction. Each pair of sensors, known as the Dual Electron Spectrometer (DES) and the Dual Ion Spectrometer (DIS), occupies a quadrant on the MMS spacecraft and the combination of the eight electron/ion sensors, employing aperture steering, image the full-sky every 30-ms (electrons) and 150-ms (ions), respectively. To probe the diffusion regions of reconnection, the highest temporal/spatial resolution mode of FPI results in the DES complement of a given spacecraft generating 6.5-Mb (raised dot) per second of electron data while the DIS generates 1.1-Mb (raised dot) per second of ion data yielding an FPI total data rate of 6.6-Mb (raised dot) per second. The FPI electron/ion data is collected by the IDPU then transmitted to the Central Data Instrument Processor (CIDP) on the spacecraft for science interest ranking. Only data sequences that contain the greatest amount of temporal/spatial structure will be intelligently down-linked by the spacecraft. This requires a data ranking process known as the burst trigger system. The burst trigger system uses pseudo physical quantities to approximate the local plasma environments. As each pseudo quantity will have a different value, a set of two scaling factors is employed for each pseudo term. These pseudo quantities are then combined at the instrument, spacecraft, and observatory level leading to a final ranking of data based on expected scientific interest. Here, we present simulations of the fixed point burst trigger system for the FPI. A variety of data sets based on previous mission data as well as analytical formulations are tested. Comparisons of floating point calculations versus the fixed point hardware simulation are shown. Analysis of the potential sources of error from overflows, quantization, etc. are examined and mitigation methods are presented. Finally a series of calibration curves are presented, showing the expected error in pseudo quantities based solely on the scale parameters chosen and the expected data range. We conclude with a presentation of the current base-lined FPI burst trigger approach.

Barrie, A. C.↗

NASA Open Science Data Repository: Maximizing Spaceflight Bioscience Data

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.

space biology↗

Fast Plasma Instrument for MMS: Simulation Results

Magnetospheric Multiscale (MMS) mission will study small-scale reconnection structures and their rapid motions from closely spaced platforms using instruments capable of high angular, energy, and time resolution measurements. The Dual Electron Spectrometer (DES) of the Fast Plasma Instrument (FPI) for MMS meets these demanding requirements by acquiring the electron velocity distribution functions (VDFs) for the full sky with high-resolution angular measurements every 30 ms. This will provide unprecedented access to electron scale dynamics within the reconnection diffusion region. The DES consists of eight half-top-hat energy analyzers. Each analyzer has a 6 deg. x 11.25 deg. Full-sky coverage is achieved by electrostatically stepping the FOV of each of the eight sensors through four discrete deflection look directions. Data compression and burst memory management will provide approximately 30 minutes of high time resolution data during each orbit of the four MMS spacecraft. Each spacecraft will intelligently downlink the data sequences that contain the greatest amount of temporal structure. Here we present the results of a simulation of the DES analyzer measurements, data compression and decompression, as well as ground-based analysis using as a seed re-processed Cluster/PEACE electron measurements. The Cluster/PEACE electron measurements have been reprocessed through virtual DES analyzers with their proper geometrical, energy, and timing scale factors and re-mapped via interpolation to the DES angular and energy phase-space sampling measurements. The results of the simulated DES measurements are analyzed and the full moments of the simulated VDFs are compared with those obtained from the Cluster/PEACE spectrometer using a standard quadrature moment, a newly implemented spectral spherical harmonic method, and a singular value decomposition method. Our preliminary moment calculations show a remarkable agreement within the uncertainties of the measurements, with the results obtained by the Cluster/PEACE electron spectrometers. The data analyzed was selected because it represented a potential reconnection event as currently published.

Figueroa-Vinas, Adolfo↗

RECOVIR Software for Identifying Viruses

Most single-stranded RNA (ssRNA) viruses mutate rapidly to generate a large number of strains with highly divergent capsid sequences. Determining the capsid residues or nucleotides that uniquely characterize these strains is critical in understanding the strain diversity of these viruses. RECOVIR (an acronym for "recognize viruses") software predicts the strains of some ssRNA viruses from their limited sequence data. Novel phylogenetic-tree-based databases of protein or nucleic acid residues that uniquely characterize these virus strains are created. Strains of input virus sequences (partial or complete) are predicted through residue-wise comparisons with the databases. RECOVIR uses unique characterizing residues to identify automatically strains of partial or complete capsid sequences of picorna and caliciviruses, two of the most highly diverse ssRNA virus families. Partition-wise comparisons of the database residues with the corresponding residues of more than 300 complete and partial sequences of these viruses resulted in correct strain identification for all of these sequences. This study shows the feasibility of creating databases of hitherto unknown residues uniquely characterizing the capsid sequences of two of the most highly divergent ssRNA virus families. These databases enable automated strain identification from partial or complete capsid sequences of these human and animal pathogens.

Chakravarty, Sugoto↗

Steps Toward Improved Integration, Search, and Analysis of Heterogeneous Data in the Astrobiology Habitable Environments Database

The Astrobiology Habitable Environments Database (AHED) is a new data system being developed as a long-term, open-access repository for astrobiology data. AHED is intended to store user-contributed results from NASA or externally-funded research in astrobiology, and to encourage sharing and synergy within the astrobiology community. However, the interdisciplinary nature of astrobiology presents some specific challenges to data management, integration, and analysis within AHED. In some disciplines (e.g., genomics), open databases thrive because the contributed products are fairly uniform and standardized (e.g., sequence data). In astrobiology, each investigation produces a unique set of data products; this makes it difficult to search across different datasets to find similar data, or to combine results from separate investigations. With AHED, we are taking steps to ensure there is adequate metadata - both at the dataset and record levels - to facilitate search, integration, and analysis. At the dataset level, we are developing a new metadata standard for describing astrobiology datasets, with detailed information about content, funding source, and scientific relevance, along with a set of topical keywords for characterizing datasets. At the record level, we are encouraging users to provide more structured content and finer-grained metadata. In many user-contributed science data repositories, few restrictions are placed on the uploaded data format, and minimal or no record-level metadata is required; thus users are unburdened when it comes to data preparation. The tradeoff is that deep integration and search across datasets is almost impossible without standardized structures and metadata. Although AHED users are free to upload minimally-described datasets, they will be encouraged to use database authoring tools (supplied by the underlying platform - Open Data Repository's Data Publisher) plus a set of customizable astrobiology-specific templates to help structure their data and provide standardized metadata. In reward for their extra effort, AHED will be able to deliver enhanced search, discovery, and analysis capabilities.

astrobiology↗

Sequence characterization of 5S ribosomal RNA from eight gram positive procaryotes

Complete nucleotide sequences are presented for 5S rRNA from Bacillus subtilis, B. firmus, B. pasteurii, B. brevis, Lactobacillus brevis, and Streptococcus faecalis, and 5S rRNA oligonucleotide catalogs and partial sequence data are given for B. cereus and Sporosarcina ureae. These data demonstrate a striking consistency of 5S rRNA primary and secondary structure within a given bacterial grouping. An exception is B. brevis, in which the 5S rRNA sequence varies significantly from that of other bacilli in the tuned helix and the procaryotic loop. The localization of these variations suggests that B. brevis occupies an ecological niche that selects such changes. It is noted that this organism produces antibiotics which affect ribosome function.

Woese, C. R.↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗

LAPS Lidar Measurements at the ARM Alaska Northslope Site (Support to FIRE Project)

This report consists of data summaries of the results obtained during the May 1998 measurement period at Barrow Alaska. This report does not contain any data interpretation or analysis of the results which will follow this activity. This report is forwarded with a data set on magnetic media which contains the reduced data from the LAPS lidar in 15 minute intervals. The data was obtained during the period 15-30 May 1998. The measurement period overlapped with several aircraft flights conducted by NASA as part of the FIRE project. The report contains a summary list of the data obtained plus figures that have been prepared to help visualize the measurement periods. The order of the presentation is as follows: Section 1. A copy of the Statement of Work for the planned activity of the second measurement period at the ARM Northslope site is provided. Section 2. A list of the data collection periods shows the number of one minute data records stored during each hour of operation and the corresponding size (Mbytes) of the one hour data folders. The folder and file names are composed from the year, month, day, hour and minute. The date/time information is given in UTC for easier comparison with other data sets. Section 3. A set of 4 comparisons between the LAPS lidar results and the sondes released by the ARM scientists from a location nearby the lidar. The lidar results show the +/- 1 sigma statistical error on each of the independent 75 m altitude bins of the data. This set of 4 comparisons was used to set and validate the calibration value which was then used for the complete data set. Section 4. A set of false color figures with up to 10 hours of specific humidity measurements are shown in each graph. Two days of measurements are shown on each page. These plots are crude representations of the data and permit a survey which indicates when the clouds were very low or where interesting events may occur in the results. These plots are prepared using the real time sequence plot program which has no smoothing in either the altitude or time (except that you are allowed to pick the integration time and time step. All of these plots were prepared with 15 minute integration and 5 minute time step. Section 5. A set of time sequence data for all of the extended observation periods are shown with a smoothing algorithm from the Matlab plotting library. Most of these data are integrated for 5 minutes and stepped at I minute intervals but several plots are shown with both 15 minute integration and 5 minute steps. The upper level on these data was selected and converted to the white background where the error in the specific humidity reached 25%. Section 6. The set of one hour integrated plots shown with up to 4 hours per page are provided- from the real time analysis snapshot program. The only difference in these plots and the real time display is that the plots are stopped at an altitude where the error appears to be too large for the data to contain any meaningful information.

Philbrick, C. Russell↗

Simulator evaluation of display concepts for pilot monitoring and control of space shuttle approach and landing. Phase 2: Manual flight control

A study of the display requirements for final approach management of the space shuttle orbiter vehicle is presented. An experimental display concept, providing a more direct, pictorial representation of the vehicle's movement relative to the selected approach path and aiming points, was developed and assessed as an aid to manual flight path control. Both head-up, windshield projections and head-down, panel mounted presentations of the experimental display were evaluated in a series of simulated orbiter approach sequence. Data obtained indicate that the experimental display would enable orbiter pilots to exercise greater flexibility in implementing alternative final approach control strategies. Touchdown position and airspeed dispersion criteria were satisfied on 91 percent of the approach sequences, representing various profile and wind effect conditions. Flight path control and airspeed management satisfied operationally-relevant criteria for the two-segment, power-off orbiter approach and were consistently more accurate and less variable when the full set of experimental display elements was available to the pilot. Approach control tended to be more precise when the head-up display was used; however, the data also indicate that the head-down display would provide adequate support for the manual control task.

Gartner, W. B.↗

Integrated Cryogenic Satellite Communications Cross-Link Receiver Experiment

An experiment has been devised which will validate, in space, a miniature, high-performance receiver. The receiver blends three complementary technologies; high temperature superconductivity (HTS), pseudomorphic high electron mobility transistor (PHEMT) monolithic microwave integrated circuits (MMIC), and a miniature pulse tube cryogenic cooler. Specifically, an HTS band pass filter, InP MMIC low noise amplifier, HTS-sapphire resonator stabilized local oscillator (LO), and a miniature pulse tube cooler will be integrated into a complete 20 GHz receiver downconverter. This cooled downconverter will be interfaced with customized signal processing electronics and integrated onto the space shuttle's 'HitchHiker' carrier. A pseudorandom data sequence will be transmitted to the receiver, which is in low Earth orbit (LEO), via the Advanced Communication Technology Satellite (ACTS) on a 20 GHz carrier. The modulation format is QPSK and the data rate is 2.048 Mbps. The bit error rate (BER) will be measured in situ. The receiver is also equipped with a radiometer mode so that experiment success is not totally contingent upon the BER measurement. In this mode, the receiver uses the Earth and deep space as a hot and cold calibration source, respectively. The experiment closely simulates an actual cross-link scenario. Since the receiver performance depends on channel conditions, its true characteristics would be masked in a terrestrial measurement by atmospheric absorption and background radiation. Furthermore, the receiver's performance depends on its physical temperature, which is a sensitive function of platform environment, thermal design, and cryocooler performance. This empirical data is important for building confidence in the technology.

Romanofsky, R. R.↗

Word timing recovery in direct detection optical PPM communication systems with avalanche photodiodes using a phase lock loop

A technique for word timing recovery in a direct-detection optical PPM communication system is described. It tracks on back-to-back pulse pairs in the received random PPM data sequences with the use of a phase locked loop. The experimental system consisted of an 833-nm AlGaAs laser diode transmitter and a silicon avalanche photodiode photodetector, and it used Q = 4 PPM signaling at source data rate 25 Mb/s. The mathematical model developed to describe system performance is shown to be in good agreement with the experimental measurements. Use of this recovered PPM word clock with a slot clock recovery system caused no measurable penalty in receiver sensitivity. The completely self-synchronized receiver was capable of acquiring and maintaining both slot and word synchronizations for input optical signal levels as low as 20 average detected photons per information bit. The receiver achieved a bit error probability of 10 to the -6th at less than 60 average detected photons per information bit.

Sun, Xiaoli↗

Increasing the Statistical Rigor of Cross-Species Differential Expression Analysis

Microgravity inflicts substantial, but undercharacterized, pressure on organisms that induces metabolic responses such as increased microbial virulence and antibiotic resistance, altered organ weights in developing rats, and loss of bone tissue in astronauts. Numerous studies have analyzed the effects of microgravity on specific organisms, tissues, or test conditions, but these projects are necessarily limited by the small sample size of space research. Increasing the sample size of spaceflight studies is non-trivial; however, pooling data from numerous studies can greatly increase the statistical rigor of comparative analyses. The GeneLab houses datasets from 73 spaceflight studies that performed transcription profiling assays. These data encompass a diverse array of organisms ranging from Escherichia coli to Mus musculus to Homo sapiens and comprise studies analyzing ionizing radiation, mammalian pregnancy, etc. Collectively, the GeneLab database contains a large quantity of transcription assays and RNA sequence data analyzing Differential Gene Expression (DGE) between microand normogravity. Xspecies, a cross-species analysis method for DGE developed by Kristiansson, et al. in 2012, identifies homologous genes between species that are universally up- or downregulated in response to test conditions. Previous work by an intern at GeneLab applied Xspecies to 19 datasets containing seven different species and identified 14 homologous groups differentially expressed under spaceflight conditions including several heat shock proteins and cytoskeletal components. Unfortunately, these results may be biased by the disproportionate number of studies on Arabidopsis thaliana (5) and Mus musculus (6) and the results are not normalized by evolutionary distances. Here, we present modifications to the Xspecies algorithm that permits incorporation of multi-omic data and normalizes data for effect size, directionality, and evolutionary distances. We then apply this algorithm to all currently available GeneLab studies

Xspecies↗

Combustion Stability Characteristics of the Project Morpheus Liquid Oxygen / Liquid Methane Main Engine

The project Morpheus liquid oxygen (LOX) / liquid methane (LCH4) main engine is a Johnson Space Center (JSC) designed ~5,000 lbf-thrust, 4:1 throttling, pressure-fed cryogenic engine using an impinging element injector design. The engine met or exceeded all performance requirements without experiencing any in- ight failures, but the engine exhibited acoustic-coupled combustion instabilities during sea-level ground-based testing. First tangential (1T), rst radial (1R), 1T1R, and higher order modes were triggered by conditions during the Morpheus vehicle derived low chamber pressure startup sequence. The instability was never observed to initiate during mainstage, even at low power levels. Ground-interaction acoustics aggravated the instability in vehicle tests. Analysis of more than 200 hot re tests on the Morpheus vehicle and Stennis Space Center (SSC) test stand showed a relationship between ignition stability and injector/chamber pressure. The instability had the distinct characteristic of initiating at high relative injection pressure drop at low chamber pressure during the start sequence. Data analysis suggests that the two-phase density during engine start results in a high injection velocity, possibly triggering the instabilities predicted by the Hewitt stability curves. Engine ignition instability was successfully mitigated via a higher-chamber pressure start sequence (e.g., ~50% power level vs ~30%) and operational propellant start temperature limits that maintained \cold LOX" and \warm methane" at the engine inlet. The main engine successfully demonstrated 4:1 throttling without chugging during mainstage, but chug instabilities were observed during some engine shutdown sequences at low injector pressure drop, especially during vehicle landing.

Melcher, John C.↗

GL4U: Bioinformatics training for students and educators using space omics data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler↗

GL4U: Using Space Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. In June 2022, GL4U partnered with Jet Propulsion Laboratory’s (JPL) Planetary Protection Center of Excellence to conduct the indirect training pilot program by training educators at historically black colleges and universities (HBCUs) and minority serving institutions (MSIs). During the educator pilot, participants received materials, training, and will be provided the necessary compute resources to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U training program provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. Pre- and post-bootcamp surveys were completed by all participants and show the overwhelming success of the bootcamps.

Amanda M. Saravia-Butler↗

GL4U: Bioinformatics Training for Students and Educators Using Space Omics Data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler↗

GL4U: Using Space Biology Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant multi-omics data via the Open Science Data Repository (OSDR) that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct (training students) and indirect (training educators) approaches. The GL4U pilot programs were conducted in June 2021 (direct training) and 2022 (indirect training). During the pilots, students and educators at Historically Black Colleges and Universities (HBCUs) and Minority Serving Institutions (MSIs) participated in a week-long (direct training) or two-week-long (indirect training) bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze space biology RNA sequencing data from OSDR. During the educator pilot, participants received materials, training, and the necessary compute resources to enable them to run the bootcamp at their home institutions, thereby extending the reach of this initiative. In July 2023, GeneLab is partnering with JPL to expand GL4U to include amplicon sequencing (Amp-Seq) analysis training. During the GL4U Amp-Seq bootcamp, student and educator participants will receive training on how to analyze and interpret Amp-Seq data using the NASA GeneLab data processing pipeline. All bootcamp material, including instructions for requesting compute resources, will be made publicly available on GitHub for educators to teach the GL4U content in subsequent semesters. GL4U provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. We present results from pre- and post-training surveys completed by all participants of the Amp-Seq bootcamp.

Amanda M Saravia-Butler↗