Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Quantifying Post-fire Vegetation Recovery on the Colorado Front Range

Forest composition and structure in the Colorado Front Range has been altered by changing wildfire regimes, particularly in response to increased fire severity. Subsequent reductions in post-fire tree regeneration can result in chronic impacts to ecological function and water quality. This project partnered with the U.S. Forest Service to estimate tree canopy recovery from 15 to 21 years after four Colorado Front Range fires that burned between 1996 and 2002— the Bobcat, Buffalo Creek, Hayman, and High Meadows fires—using Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper (ETM+), and Landsat 8 Operational Land Imager (OLI). First, we evaluated relationships between recovery metrics derived from spectral vegetation indices (i.e. post-fire vegetation, normalized burn ratio, percent unrecovered vegetation) and field-collected counts of post-fire tree seedling regeneration. Then, we used high-resolution imagery paired with Landsat timeseries spectral and geomorphometric data to map forest canopy cover 15-21 years post-fire. Finally, we assessed ecological drivers of recovery including topography, climate, soils, and fire severity using Random Forest. While we found that basic vegetation index metrics failed to distinguish field-measured tree regeneration, our data showed that both aspect and burn severity were important drivers of vegetation recovery. North-facing aspects tended to have higher pre-fire Normalized Difference Vegetation Index (NDVI) and thus higher burn severities. However, areas that burned at high severity experienced slower vegetation recovery, regardless of aspect suggesting that high severity fire can mute favorable site conditions on north-facing aspects (i.e. moister, more developed soils). While this work propels our understanding of variables that influence vegetative recovery, vegetation type, forest conversion, and watershed function, more work should focus on what vegetation indices identifying areas of recovery indicate across different forest types and burn severities.

NASA DEVELOP↗

Colorado Front Range Disasters: Understanding the Impact of Forest Management on the Cameron Peak and CalWood Fire

Along the Colorado Front Range, forest management has gained significant attention due to uncharacteristically large fires that burned late in 2020. The Cameron Peak Fire (largest in Colorado recorded history) and the CalWood Fire collectively burned an estimated 219,019 acres from August through December of 2020. Project partners at the Coalition for the Poudre River Watershed, Colorado State Forest Service, Ben Delatour Scout Ranch, The Nature Conservancy, Colorado Forest Restoration Institute, and Colorado State University were interested in understanding the effectiveness of previous forest treatments in reducing burn severity within the Cameron Peak Fire and the CalWood Fire. We first collated a forest treatment dataset from pre-existing datasets by reclassifying over 29,000 treatments, which occurred across the Northern Colorado Front Range between 1970-2020. Secondly, we mapped three burn severity indices using Landsat 8 OLI and Sentinel-2 MSI Earth observations and compared them to soil burn severity field data. Thirdly, a total of 35 topographic, disturbance, forest structure, and treatment predictor variables were generated across the fires. Finally, we assessed relationships between these predictor variables and burn severity using the random forest algorithm. Model results indicate that the primary drivers of burn severity were elevation and distance to treatment edge for the Cameron Peak Fire and fire area and forest canopy cover for the CalWood Fire. Further analysis of these variables paired with field data is necessary to understand the relationship between burn severity and treatments to guide future restoration efforts, improve forest resiliency, and mitigate fire risks.

Neal Swayze↗

Linear Subpixel Learning Algorithm for Land Cover Classification from WELD using High Performance Computing

In this work, we use a Fully Constrained Least Squares Subpixel Learning Algorithm to unmix global WELD (Web Enabled Landsat Data) to obtain fractions or abundances of substrate (S), vegetation (V) and dark objects (D) classes. Because of the sheer nature of data and compute needs, we leveraged the NASA Earth Exchange (NEX) high performance computing architecture to optimize and scale our algorithm for large-scale processing. Subsequently, the S-V-D abundance maps were characterized into 4 classes namely, forest, farmland, water and urban areas (with NPP-VIIRS-national polar orbiting partnership visible infrared imaging radiometer suite nighttime lights data) over California, USA using Random Forest classifier. Validation of these land cover maps with NLCD (National Land Cover Database) 2011 products and NAFD (North American Forest Dynamics) static forest cover maps showed that an overall classification accuracy of over 91 percent was achieved, which is a 6 percent improvement in unmixing based classification relative to per-pixel-based classification. As such, abundance maps continue to offer an useful alternative to high-spatial resolution data derived classification maps for forest inventory analysis, multi-class mapping for eco-climatic models and applications, fast multi-temporal trend analysis and for societal and policy-relevant applications needed at the watershed scale.

Subpixel↗

Modelling above-ground biomass stock over Norway using national forest inventory data with ArcticDEM and Sentinel-2 data

Boreal forests constitute a large portion of the global forest area, yet they are undersampled through field surveys, and only a few remotely sensed data sources provide structural information wall-to-wall throughout the boreal domain. ArcticDEM is a collection of high-resolution (2 m) space-borne stereogrammetric digital surface models (DSM) covering the entire land area north of 60° of latitude. The free-availability of ArcticDEM data offers new possibilities for aboveground biomass mapping (AGB) across boreal forests, and thus it is necessary to evaluate the potential for these data to map AGB over alternative open-data sources (i.e., Sentinel-2). This study was performed over the entire land area of Norway north of 60° of latitude, and the Norwegian national forest inventory (NFI) was used as a source of field data composed of accurately geolocated field plots (n=7710) systematically distributed across the study area. Separate random forest models were fitted using NFI data, and corresponding remotely sensed data consisting of either: i) a canopy height model (ArcticCHM) obtained by subtracting a high-quality digital terrain model (DTM) from the ArcticDEM DSM height values, ii) Sentinel-2 (S2), or iii) a combination of the two (ArcticCHM+S2). Furthermore, we assessed the effect of the forest- and terrain-specific factors on the models’ predictive accuracy. The best model (,i.e., ArcticCHM+S2) explained nearly 60% of the variance of the training set, which translated in the largest accuracy in terms of root mean square error (RMSE=41.4 t/ha). This result highlights the synergy between 3D and multispectral data in AGB modelling. Furthermore, this study showed that despite the importance of ArcticCHM variables, the S2 model performed slightly better than ArcticCHM model. This finding highlights some of the limitations of ArcticDEM, which, despite the unprecedented spatial resolution, is highly heterogeneous due to the blending of multiple acquisitions across different years and seasons. We found that both forest- and terrain-specific characteristics affected the uncertainty of the ArcticCHM+S2 model and concluded that the combined use of ArcticCHM and Sentinel-2 represents a viable solution for AGB mapping across boreal forests. The synergy between the two data sources allowed for a reduction of the saturation effects typical of multispectral data while ensuring the spatial consistency in the output predictions due to the removal of artifacts and data voids present in ArcticCHM data. While the main contribution of this study is to provide the first evidence of the best-case-scenario (i.e., availability of accurate terrain models) that ArcticDEM data can provide for large-scale AGB modelling, it remains critically important for other studies to investigate how ArcticDEM may be used in areas where no DTMs are available as is the case for large portions of the boreal zone.

space-borne imagery↗

Mapping Mangrove Extent and Change: A Globally Applicable Approach

This study demonstrates a globally applicable method for monitoring mangrove forest extent at high spatial resolution. A 2010 mangrove baseline was classified for 16 study areas using a combination of ALOS PALSAR and Landsat composite imagery within a random forests classifier. A novel map-to-image change method was used to detect annual and decadal changes in extent using ALOS PALSAR/JERS-1 imagery. The map-to-image method presented makes fewer assumptions of the data than existing methods, is less sensitive to variation between scenes due to environmental factors (e.g., tide or soil moisture) and is able to automatically identify a change threshold. Change maps were derived from the 2010 baseline to 1996 using JERS-1 SAR and to 2007, 2008 and 2009 using ALOS PALSAR. This study demonstrated results for 16 known hotspots of mangrove change distributed globally, with a total mangrove area of 2,529,760 ha. The method was demonstrated to have accuracies consistently in excess of 90% (overall accuracy: 92.2–93.3%, kappa: 0.86) for mapping baseline extent. The accuracies of the change maps were more variable and were dependent upon the time period between images and number of change features. Total change from 1996 to 2010 was 204,850 ha (127,990 ha gain, 76,860 ha loss), with the highest gains observed in French Guiana (15,570 ha) and the highest losses observed in East Kalimantan, Indonesia (23,003 ha). Changes in mangrove extent were the consequence of both natural and anthropogenic drivers, yielding net increases or decreases in extent dependent upon the study site. These updated maps are of importance to the mangrove research community, particularly as the continual updating of the baseline with currently available and anticipated spaceborne sensors. It is recommended that mangrove baselines are updated on at least a 5-year interval to suit the requirements of policy makers.

Thomas, Nathan↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Southern Wyoming Ecological Forecasting: Monitoring Cheatgrass in Southern Wyoming and Northern Colorado to Inform Management Efforts Post-Mullen Fire

Cheatgrass (Bromus tectorum) is a prominent invasive species in the Intermountain West that has the potential to out-compete native plant species, reduce biodiversity, and reduce the quality of habitat for ungulates. Furthermore, because cheatgrass readily establishes in disturbed landscapes, it can potentially increase fuel loads and exacerbate wildfire risk. In 2020, the Mullen Fire burned 176,878 acres in Carbon and Albany Counties, Wyoming and Jackson County, Colorado. Large fires such as this one raise concern for partners at the United States Forest Service and the United States Geological Survey Fort Collins Science Center, who are tasked with rapidly detecting and controlling invasive species in the post-fire environment. We developed a Random Forest model trained by in-situ field data and spectral indices such as Normalized Difference Vegetation Index (NDVI), Soil Adjusted Vegetation Index, and Enhanced Vegetation Index derived from Landsat 8 Operational Land Imager, Sentinel-2 MultiSpectral Instrument, and Shuttle Radar Topography Mission to detect and map cheatgrass presence during the 2021 growing season. The team successfully created a spectral cheatgrass detection map in the study area (RMSE = 13.71, R2 = 0.34). We also produced a NDVI time-series derived from Sentinel-2 MSI to analyze vegetation recovery patterns.

Dahlia Shahin↗

Forest Carbon Storage in the Western United States: Distribution, Drivers, and Trends

Abstract Forests are a large carbon sink and could serve as natural climate solutions that help moderate future warming. Thus, establishing forest carbon baselines is essential for tracking climate‐mitigation targets. Western US forests are natural climate solution hotspots but are profoundly threatened by drought and altered disturbance regimes. How these factors shape spatial patterns of carbon storage and carbon change over time is poorly resolved. Here, we estimate live and dead forest carbon density in 19 forested western US ecoregions with national inventory data (2005–2019) to determine: (a) current carbon distributions, (b) underpinning drivers, and (c) recent trends. Potential drivers of current carbon included harvest, wildfire, insect and disease, topography, and climate. Using random forests, we evaluated driver importance and relationships with current live and dead carbon within ecoregions. We assessed trends using linear models. Pacific Northwest (PNW) and Southwest (SW) ecoregions were most and least carbon dense, respectively. Climate was an important carbon driver in the SW and Lower Rockies. Fire reduced live and increased dead carbon, and was most important in the Upper Rockies and California. No ecoregion was unaffected by fire. Harvest and private ownership reduced carbon, particularly in the PNW. Since 2005, live carbon declined across much of the western US, likely from drought and fire. Carbon has increased in PNW ecoregions, likely recovering from past harvest, but recent record fire years may alter trajectories. Our results provide insight into western US forest carbon function and future vulnerabilities, which is vital for effective climate change mitigation strategies.

Environmental Sciences & Ecology↗

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment↗

Southern Colorado Disasters: Using NASA Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Project Summary↗

Southern Colorado Disasters: Using NASA Earth Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Technical Paper↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Colorado Ecological Forecasting: Monitoring Post-fire Cheatgrass (Bromus tectorum) Distribution to Inform Management Planning

Cheatgrass (Bromus tectorum) is a species of concern across the western United States as it has the potential to outcompete native plant species, reduce biodiversity, and diminish nutrient availability for ungulates. Furthermore, because cheatgrass can quickly dominate disturbed landscapes it has the potential to exacerbate wildfire risk by increasing fuel loads. In 2020, the Cameron Peak fire burned more than 200,000 acres on the Arapaho and Roosevelt National Forests in Colorado. These issues are of imminent concern for our partners at the Forest Service (USFS), as they are tasked with wildfire risk and invasive species mitigation. Disturbances such as wildfires can substantially increase the rate and extent of cheatgrass spread. Current cheatgrass mitigation methods rely on field crews to physically locate cheatgrass on the landscape, which takes time, money, and extensive manpower. Here, we developed two Random Forest models within the Software for Assisted Habitat Modeling (SAHM) using remote sensing predictors derived from Sentinel-2 MultiSpectral Instrument (MSI) and Shuttle Radar Topography Mission (SRTM). The first model identified suitable cheatgrass habitat while the other detected cheatgrass presence during the 2021 growing season. Topographic variables were found to be the most important in driving the habitat suitability model. Cheatgrass detection was also found to be possible within a short timespan with limited imagery surrounding a phenological shift of the plant. Maps produced from these models provide natural resource managers the ability to implement early detection and rapid response to prevent the spread of cheatgrass to new locations.

DEVELOP Tech Paper↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Maya Forest Water Resources I: Using NASA Earth Observations to Map Forested Inundation in the Maya Forest

As climate change increases the severity and frequency of extreme weather events in the tropics, it is vital for the safety of local communities and the health of ecosystems to monitor seasonal inundation. Forested inundation affects the ability of forested wetlands to provide ecosystem services, such as flood mitigation, water filtration, carbon storage, and erosion mitigation. While ground-based monitoring has traditionally been used to map inundation extent, those methods are costly and time-intensive. The NASA DEVELOP team focused on seasonal inundation throughout 2008 in the Maya Forest, when changes in inundation were drastic. To monitor seasonal inundation, our team used in situ field data and Earth observations from Landsat 7 Enhanced Thematic Mapper (ETM+), Advanced Land Observing Satellite (ALOS) Phased Array type L-band Synthetic Aperture Radar (PALSAR) 1, Shuttle Radar Topography Mission (SRTM), and products from the Ice, Cloud, and Land Elevation Satellite (ICESat). The team applied a Random Forest algorithm to Landsat 7 imagery, generating an object-level land cover classification with an overall accuracy of 72.1% and forest class with 100% recall and 78% precision. The team applied L-band backscatter thresholds from existing literature to forest-masked ALOS imagery and refined the thresholds in an iterative process using field data and hydrology models to delineate seasonal inundation extent. These publicly available data products help end users from Belize’s Land Information Center (LIC) and Forest Department, Guatemala’s Center for Monitoring and Evaluation (CEMEC), and Mexico’s El Colegio de la Frontera Sur (ECOSUR) to inform land management and protect community infrastructure.

Madelyn Savan↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗