Engineering PapersSearch

SEARCH · Engineering Papers

Results for “bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enabling Open and Interoperable Science: Multi-Omics Data Processing Platform with NASA GeneLab Standardized Bioinformatics Workflows for Space and Earth Research

Multi-omics biological data continues to be generated at an astounding pace. Genomics, transcriptomics, metabolomics, and proteomics, or collectively known as multi-omics data, are used to assess biological functions, and provide invaluable insights into human, animal, plant, and environmental health both on Earth and in Space. Despite the abundance of these valuable data, the need for bioinformatics expertise, particularly as it relates to the niche filed of space biology, and a lack of accessible resources for processing these data limit their usefulness in deriving biological insights. The NASA Open Science Data Repository (OSDR) provides access to omics data from various spaceflight and analog studies. To enhance the accessibility and reusability of these data, GeneLab (part of OSDR) designs and implements standardized, community-driven, open-source bioinformatics workflows to transform raw omics data into standardized processed data. Currently, GeneLab-processed data from hundreds of space studies have been reused for meta-analyses. This has led to new insights and scientific publications that extend beyond the initial research, thereby enriching our understanding of molecular-scale biological responses to the space environment. To make these bioinformatics workflows open and accessible, GeneLab teamed up with DOE-funded initiatives, including the National Microbiome Data Collaborative (NMDC), to create the NASA EDGE [Empowering the Development of Genomics Expertise] Bioinformatics web-based platform. NASA EDGE utilizes shared compute resources to run the GeneLab standardized bioinformatics workflows, which eliminates the need for researchers to have their own high performance computing cluster. The web-based platform makes complicated biological analyses incredibly easy to perform, thus expanding the reach of these analyses to bioinformatics novices, students, and even citizen scientists enabling them to contribute to scientific discoveries and progress. The authors will demonstrate how the NASA EDGE platform can be used to process microbial omics data hosted on OSDR as well as user-generated omics datasets using GeneLab’s standard workflows.

Amanda M. Saravia-Butler

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

Vascular Patterning Analysis by VESGEN 2D/3D with Bioinformatics: Updates for Rodent Tissues

Fractally branching vascular systems are a complex physiological requirement shared by humans with all higher terrestrial life forms, including other vertebrates, insects, and higher land plants. Vascular trees, networks, and tree-network composites are therefore mapped and quantified by the VESsel GENeration Analysis (VESGEN) software according to weighted physiological vascular rules that include vessel connectivity, tapering and bifurcational branching. According to fluid dynamics, successful vascular transport depends upon a complex distributed system of highly regulated laminar flow. VESGEN has elucidated changes in vascular patterning resulting from inflammatory, developmental and other signaling pathways within numerous tissues of major model organisms important for Space Biology, especially for rodents. Important early stage regenerative opportunities have been identified by VESGEN vascular analysis for visual impairments in the human retina, and is currently being used for research into astronaut visual and ocular disorders associated with long duration missions. The VESGEN 2D software is a mature, automated, widely published capability for which beta testing and public release by NASA is planned for the upcoming year. Early-stage capabilities for VESGEN 3D analysis are under development for the rodent retina and intestine as prototype tissues. A prototype VESGEN 2D Bioinformatics software capability has also been developed to associate phenotypic changes in molecular expression with vascular structure and function. By new VESGEN bioinformatic innovations, expression patterns of the genetic, transcriptional, protein and other markers for regulatory molecules such as vascular endothelial growth factor (VEGF) and their receptors, often indicators of tissue oxygenation status, are co-localized with alterations in vascular pattern. Biomarkers are therefore mapped and quantified as information dimensions directly correlated with the spatial dimensions of a vascular pattern. Further important technology innovations by NASA include substantial image segmentation advances for more automated binary extraction of the grayscale vascular patterns, together with informative associated image quality assessments. Vascular mapping and quantification capabilities for the rodent retina and intestine are illustrated for VESGEN 2D, along with technology status reports on VESGEN 3D and Bioinformatic capabilities. Research partially supported by Ames Center Innovation Awards.

Parsons-Wingerter, P.

AstroAmpSeq: Microbial Bioinformatics Education with NASA GeneLab’s Amplicon Pipeline

The prevalence and importance of large sequencing datasets in microbiology has led to a movement to share microbial ecology experimental data through open-access databases. This is particularly true of experiments that are difficult to replicate, such as those conducted in the spaceflight environment and shared via NASA GeneLab. It is now possible and indeed valuable for students to access and re-analyze these shared datasets for educational and research purposes. To provide students with experience utilizing microbial bioinformatics tools, GeneLab for Colleges and Universities (GL4U) has designed AstroAmpSeq, a week-long, virtually implemented project-based learning (PBL) minicourse to instruct undergraduate students on 16S amplicon sequencing. AstroAmpSeq was created to be accessible to students without prior bioinformatics or microbial ecology experience. During the minicourse students work in teams to process, analyze, and visualize a subsample of GeneLab dataset GLDS-280 using GeneLab’s standard amplicon processing pipeline, which is based in R. Students develop a hypothesis related to the dataset then generate and analyze figures to evaluate their hypothesis. Formative assessment of student learning is determined via pre- and post-evaluations, peer feedback, and self-reflection. Project and presentation rubrics serve as a summative assessment of student learning. GL4U AstroAmpSeq not only meets American Society for Microbiology Curriculum Guidelines, but also incites student interest in research by an inquiry-based approach and can be made part of a larger semester-long curriculum. GL4U AstroAmpSeq raises awareness of space microbiology and bioinformatics as a field and career path among undergraduates. Further, by using a GeneLab dataset and nesting microbiology techniques into the real-world application of space biology, AstroAmpSeq enforces deeper and longer-lasting student learning.

microbiology

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Bioinformatics for Exploration

For the purpose of this paper, bioinformatics is defined as the application of computer technology to the management of biological information. It can be thought of as the science of developing computer databases and algorithms to facilitate and expedite biological research. This is a crosscutting capability that supports nearly all human health areas ranging from computational modeling, to pharmacodynamics research projects, to decision support systems within autonomous medical care. Bioinformatics serves to increase the efficiency and effectiveness of the life sciences research program. It provides data, information, and knowledge capture which further supports management of the bioastronautics research roadmap - identifying gaps that still remain and enabling the determination of which risks have been addressed.

Johnson, Kathy A.

A Bioinformatics Facility for NASA

Building on an existing prototype, we have fielded a facility with bioinformatics technologies that will help NASA meet its unique requirements for biological research. This facility consists of a cluster of computers capable of performing computationally intensive tasks, software tools, databases and knowledge management systems. Novel computational technologies for analyzing and integrating new biological data and already existing knowledge have been developed. With continued development and support, the facility will fulfill strategic NASA s bioinformatics needs in astrobiology and space exploration. . As a demonstration of these capabilities, we will present a detailed analysis of how spaceflight factors impact gene expression in the liver and kidney for mice flown aboard shuttle flight STS-108. We have found that many genes involved in signal transduction, cell cycle, and development respond to changes in microgravity, but that most metabolic pathways appear unchanged.

Schweighofer, Karl

Robust Bioinformatics Recognition with VLSI Biochip Microsystem

A microsystem architecture for real-time, on-site, robust bioinformatic patterns recognition and analysis has been proposed. This system is compatible with on-chip DNA analysis means such as polymerase chain reaction (PCR)amplification. A corresponding novel artificial neural network (ANN) learning algorithm using new sigmoid-logarithmic transfer function based on error backpropagation (EBP) algorithm is invented. Our results show the trained new ANN can recognize low fluorescence patterns better than the conventional sigmoidal ANN does. A differential logarithmic imaging chip is designed for calculating logarithm of relative intensities of fluorescence signals. The single-rail logarithmic circuit and a prototype ANN chip are designed, fabricated and characterized.

bioinformatics

VLSI Microsystem for Rapid Bioinformatic Pattern Recognition

A system comprising very-large-scale integrated (VLSI) circuits is being developed as a means of bioinformatics-oriented analysis and recognition of patterns of fluorescence generated in a microarray in an advanced, highly miniaturized, portable genetic-expression-assay instrument. Such an instrument implements an on-chip combination of polymerase chain reactions and electrochemical transduction for amplification and detection of deoxyribonucleic acid (DNA).

Fang, Wai-Chi

Comparison of Two Bioinformatics Tools Used to Characterize the Microbial Diversity and Predictive Functional Attributes of Microbial Mats from Lake Obersee, Antarctica

In this study, using NextGen sequencing of the collective 16S rRNA genes obtained from two sets of samples collected from Lake Obersee, Antarctica, we compared and contrasted two bioinformatics tools, PICRUSt and Tax4Fun. We then developed an R script to assess the taxonomic and predictive functional profiles of the microbial communities within the samples. Taxa such as Pseudoxanthomonas, Planctomycetaceae, Cyanobacteria Subsection III, Nitrosomonadaceae, Leptothrix, and Rhodobacter were exclusively identified by Tax4Fun that uses SILVA database; whereas PICRUSt that uses Greengenes database uniquely identified Pirellulaceae, Gemmatimonadetes A1-B1, Pseudanabaena, Salinibacterium and Sinibacteraceae. Predictive functional profiling of the microbial communities using Tax4Fun and PICRUSt separately revealed common metabolic capabilities, while also showing specific functional IDs not shared between the two approaches. Combining these functional predictions using a customized R script revealed a more inclusive metabolic profile, such as hydrolases, oxidoreductases, transferases; enzymes involved in carbohydrate and amino acid metabolisms; and membrane transport proteins known for nutrient uptake from the surrounding environment. Our results present the first molecular-phylogenetic characterization and predictive functional profiles of the microbial mat communities in Lake Obersee, while demonstrating the efficacy of combining both the taxonomic assignment information and functional IDs using the R script created in this study for a more streamlined evaluation of predictive functional profiles of microbial communities.

Hyunmin Koo

GL4U: Bioinformatics training for students and educators using space omics data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler

GL4U: Using Space Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. In June 2022, GL4U partnered with Jet Propulsion Laboratory’s (JPL) Planetary Protection Center of Excellence to conduct the indirect training pilot program by training educators at historically black colleges and universities (HBCUs) and minority serving institutions (MSIs). During the educator pilot, participants received materials, training, and will be provided the necessary compute resources to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U training program provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. Pre- and post-bootcamp surveys were completed by all participants and show the overwhelming success of the bootcamps.

Amanda M. Saravia-Butler

GL4U: Bioinformatics Training for Students and Educators Using Space Omics Data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler

GL4U: Using Space Biology Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant multi-omics data via the Open Science Data Repository (OSDR) that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct (training students) and indirect (training educators) approaches. The GL4U pilot programs were conducted in June 2021 (direct training) and 2022 (indirect training). During the pilots, students and educators at Historically Black Colleges and Universities (HBCUs) and Minority Serving Institutions (MSIs) participated in a week-long (direct training) or two-week-long (indirect training) bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze space biology RNA sequencing data from OSDR. During the educator pilot, participants received materials, training, and the necessary compute resources to enable them to run the bootcamp at their home institutions, thereby extending the reach of this initiative. In July 2023, GeneLab is partnering with JPL to expand GL4U to include amplicon sequencing (Amp-Seq) analysis training. During the GL4U Amp-Seq bootcamp, student and educator participants will receive training on how to analyze and interpret Amp-Seq data using the NASA GeneLab data processing pipeline. All bootcamp material, including instructions for requesting compute resources, will be made publicly available on GitHub for educators to teach the GL4U content in subsequent semesters. GL4U provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. We present results from pre- and post-training surveys completed by all participants of the Amp-Seq bootcamp.

Amanda M Saravia-Butler

The GeneLab Buffet: A Bioinformatic MATRIX of MANGO and TOAST

The GeneLab data repository provides an unparalleled resource for exploring how spaceflight affects organisms with omics-level insights. However, two major interlinked challenges to capitalizing on the information within these data are their vast breadth and the often-specialized expertise that has been required in the past for their analysis. How do you compare responses within and between studies, especially if you are a non-bioinformatics specialist? This presentation will discuss how Space Biology data can be accessed using software to help provide these data resources to address research questions and generate new hypotheses. The presentation will cover a wide range of the available space life science tools but will focus on TOAST, MANGO, the MATRIX, RadBioApp and other interactive relational databases (https://genelab.nasa.gov/external-vis-apps). These exploration environments have been developed to search the GeneLab data repository for new insights that inform how model organisms respond to microgravity, radiation and other factors associated with spaceflight. The presentation will be interactive, and participants will have the opportunity to ask questions and learn more about the data viz and modeling tools that are available to them.

AstroBotany