Early Detection of Active Region Emergence in the Solar Interior Using Acoustic Power Maps and Machine Learning Data Analysis
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.
A short presentation highlighting using machine learning and topological data analysis to address the challenges of assuring and securing machine learning enabled systems.
We describe four pixel-based classifers that were developed to identify events in hyperspectral data onboard a spacecraft.
Explore the source record for details and available documents.
Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera, Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.
Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks [1]. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration [2]. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera [3], Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) [5] is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools [5, 7]. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool [8]. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers [5,6]. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.
"The escalating impact of climate change induced extreme weather events in urban, suburban, and rural environments demands a rethink of how we have been using the single event-based or use-case-based knowledge graph models. The lack of representation in interaction within environmental variables found in literature led to the development of a novel framework that reflects the true nature of the interconnectedness in our environment. We propose an Environmental Interaction Knowledge Graph (EIKG) framework. This general EIKG framework works as the basis for interconnected environmental events by knitting interrelated events such as hurricanes leading to storm surges, which lead to flood events that could cause mudslides, landslides, etc., The cascading nature of one event leading to another related event in the environment requires an adequate understanding of each event using contextual information before conducting any data-driven analytics. This vision paper showcases how the EIKG:floods, EIKG:wildfire EIKG:landslides, etc, can be derived from a base case framework of EIKG as those individual events are interconnected with some common denominator variables. As an example, the precipitation variable is used in the flood case study as well as in the wildfire case study, as excessive precipitation levels lead to floods, and lack of precipitation leads to droughts and wildfires. We identify the precipitation variable as a “common-denominator-variable” in extreme weather events that play a key role in modeling the environment leading to different extreme weather events based on the variability of that variable (varying values where low precipitation leads to drought, and high values lead to floods). We use the insights gained from EIKG to conduct classical and Quantum Machine Learning (QML) based data analysis on the research questions developed. Our preliminary study shows how the Variational Quantum Classifier (VQC) and Quantum Support Vector Classifier (QSVC) are used along with the classical machine learning models to compare the model accuracies. Our study elaborates on how a quantitative analysis uses state-of-the-art machine learning techniques that include implementing both classical and quantum machine learning models and developing the knowledge graph. The EIKG is used to organize heterogeneous datasets and integrate the relations to case-specific extreme weather events such as floods. The study uses datasets such as county-to-country residential mobility data, socioeconomic datasets from the US Census Bureau, climate and weather-related Earth Observational data from NASA, and critical infrastructure data from the Homeland Infrastructure datasets."
The goal of visual inference programming is to develop a software framework data analysis and to provide machine learning algorithms for inter-active data exploration and visualization. The topics include: 1) Intelligent Data Understanding (IDU) framework; 2) Challenge problems; 3) What's new here; 4) Framework features; 5) Wiring diagram; 6) Generated script; 7) Results of script; 8) Initial algorithms; 9) Independent Component Analysis for instrument diagnosis; 10) Output sensory mapping virtual joystick; 11) Output sensory mapping typing; 12) Closed-loop feedback mu-rhythm control; 13) Closed-loop training; 14) Data sources; and 15) Algorithms. This paper is in viewgraph form.
Solar sail deformation leads to disturbance torques from solar radiation pressure, driving performance requirements for momentum management systems. For the Solar Cruiser technology demonstrator mission, we have developed a model leveraging neural network-based machine learning to derive sail shape characteristics. The model uses torque and attitude telemetry simulated from a reduced-order tensor model of the deformed sail mesh over a characterization sequence. The machine learning model predicts sail boom deflection with comparable accuracy to that of an onboard context camera. This model can discover sail shape with no additional mass or data downlink requirements, allowing for validation of sail force modeling assumptions using in flight data. The results from the project hold promise for the further implementation of machine learning techniques in solar sail telemetry analysis and control.
Solar sail deformation leads to disturbance torques from solar radiation pressure, driving performance requirements for momentum management systems. For the Solar Cruiser technology demonstrator mission, we have developed a model leveraging neural network-based machine learning to derive sail shape characteristics. The model uses torque and attitude telemetry simulated from a reduced-order tensor model of the deformed sail mesh over a characterization sequence. The machine learning model predicts sail boom deflection with comparable accuracy to that of an onboard context camera. This model can discover sail shape with no additional mass or data downlink requirements, allowing for validation of sail force modeling assumptions using in flight data. The results from the project hold promise for the further implementation of machine learning techniques in solar sail telemetry analysis and control.
The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.
Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.
Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.
Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.
Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.
Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.
Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.