Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Maya Forest Water Resources I: Using NASA Earth Observations to Map Forested Inundation in the Maya Forest

As climate change increases the severity and frequency of extreme weather events in the tropics, it is vital for the safety of local communities and the health of ecosystems to monitor seasonal inundation. Forested inundation affects the ability of forested wetlands to provide ecosystem services, such as flood mitigation, water filtration, carbon storage, and erosion mitigation. While ground-based monitoring has traditionally been used to map inundation extent, those methods are costly and time-intensive. The NASA DEVELOP team focused on seasonal inundation throughout 2008 in the Maya Forest, when changes in inundation were drastic. To monitor seasonal inundation, our team used in situ field data and Earth observations from Landsat 7 Enhanced Thematic Mapper (ETM+), Advanced Land Observing Satellite (ALOS) Phased Array type L-band Synthetic Aperture Radar (PALSAR) 1, Shuttle Radar Topography Mission (SRTM), and products from the Ice, Cloud, and Land Elevation Satellite (ICESat). The team applied a Random Forest algorithm to Landsat 7 imagery, generating an object-level land cover classification with an overall accuracy of 72.1% and forest class with 100% recall and 78% precision. The team applied L-band backscatter thresholds from existing literature to forest-masked ALOS imagery and refined the thresholds in an iterative process using field data and hydrology models to delineate seasonal inundation extent. These publicly available data products help end users from Belize’s Land Information Center (LIC) and Forest Department, Guatemala’s Center for Monitoring and Evaluation (CEMEC), and Mexico’s El Colegio de la Frontera Sur (ECOSUR) to inform land management and protect community infrastructure.

Madelyn Savan↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Gila Water Resources III - Modeling the Impacts of Post-fire Restoration Methods on Vegetation Recovery in the Gila National Forest

In recent years, wildfires in New Mexico’s Gila National Forest have become increasingly common and more severe. Wildfires can have powerful impacts on hydrology and soil stability, including erosion, flooding, and debris-flows that threaten lives and infrastructure downstream. Vegetation restoration treatments like seeding and mulching can mitigate these effects and facilitate ecosystem recovery. Understanding the effectiveness of various restoration methods is vital to planning a cost-effective and successful post-fire recovery strategy. The immediate response to a fire on US Forest Service land is coordinated by a Burned Area Emergency Response (BAER) team, a group responsible for mitigating immediate post-fire risks to human life, property, and critical natural and cultural resources. This study created a proof-of-concept methodology for a decision-support tool designed to help BAER teams identify the restoration treatments most likely to succeed in a given burned area. Leveraging random forest regression, Google Earth Engine, and Landsat 7 and 8 Earth observations, this study modeled vegetation recovery after the 2013 Silver Fire for seeded areas, seeded/mulched areas, and untreated areas. Treatment type and initial burn severity were the largest drivers of vegetation recovery across the landscape. Seeded/mulched areas showed higher recovery levels than untreated areas three months post-fire, but by four years post-fire, treated and untreated areas displayed similar recovery levels. To produce a robust predictive tool for the Gila National Forest, the model should be trained on many more fires and incorporate post-fire weather conditions into the process. Such a model will help partners ensure efficient resource use and plan effective post-fire restoration strategies.

DEVELOP Project Summary↗

Mapping Post-fire Conifer Regeneration using Snow-on Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled reforestation sites planted from 2001–2013. The field data and imagery were used to train a Random Forest (RF) model in Google Earth Engine. The RF model classified conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3% for 2021. Approximately 45.9% of the high-severity burn area had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration in a post-wildfire landscape.

Yeshey Seldon↗

Maya Forest Water Resources II: Mapping Inundation Below the Forest Canopy in the Maya Tri-National Forest

To monitor seasonal flooding within the tri-National Maya Forest the team completed the methodology started by the Summer 2021 term to analyze changes in inundation dynamic throughout 2017. The team analyzed inundation dynamics in Google Earth Engine (GEE) using Earth observation products from the Landsat 8 Operational Land Imager (OLI), Advanced Land Observing Satellite (ALOS) Phased Array type L-band Synthetic Aperture Radar (PALSAR) 2, and International Space Station (ISS) Global Ecosystem Dynamics Investigation LiDAR (GEDI). The team improved the landcover classification using the Random Forest algorithm in GEE by adding canopy height data derived from GEDI, elevation and slope data from Copernicus, and additional multi-spectral band ratios from Landsat 8. The pixel-based land cover classification produced an overall accuracy of 88%. Experiments measuring inundation extent using L-band SAR included comparing results with a priori knowledge, topography datasets, and auxiliary datasets. We iteratively tested and found threshold values for identifying forested inundation using the ratio for HH divided by HV. The resulting methodology and products helped end users from Belize’s Land Information Center (LIC) and Forest Department, Guatemala’s Center for Monitoring and Evaluation (CEMEC), and Mexico’s El Colegio de la Frontera Sur (ECOSUR) manage land and water resources and protect communities.

Stephanie Jiménez↗

Black Hills Wildfires: Mapping Post-fire Conifer Regeneration using Snow-on Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high-severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled additional reforestation sites planted in 2001–2013. In Google Earth Engine (GEE), the field data and imagery were used to train a Random Forest (RF) model. The RF model classified 2021 conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3%. Approximately 45.9% of the high-severity burn had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration.

Casey Menick​↗

Black Hills Wildfires: Mapping Post-Fire Conifer Regeneration using Snow-On Imagery

The 2000 Jasper Fire in the Black Hills of South Dakota was the largest wildfire to date in the region, burning over 83,000 acres of ponderosa pine forest. In collaboration with partners from the United States Forest Service (USFS) Black Hills Experimental Forest, USFS Rocky Mountain Research Station, and United States Geological Survey Geosciences and Environmental Change Science Center, we characterized post-fire forest regeneration within high-severity burn patches. We accomplished this by implementing novel conifer detection techniques using a snow index mask to create a winter, snow-on image composite from Landsat 8 Operational Land Imager (OLI) and Sentinel-2 Multispectral Instrument (MSI) data. We utilized 2015 USFS stem maps of field-observed regeneration plots and ocularly sampled additional reforestation sites planted in 2001–2013. In Google Earth Engine (GEE), the field data and imagery were used to train a Random Forest (RF) model. The RF model classified 2021 conifer regeneration density as low, medium, or high across the high-severity burn area with an overall accuracy of 81.3%. Approximately 45.9% of the high-severity burn had low or no regeneration (0-40 trees per acre) 20 years post-fire. Given our partners' desire to find easily accessible low conifer regeneration zones, we identified 4,079 acres of priority planting sites that were within 1,500 feet of roads, had not been planted previously, and were larger than 50 acres. This method supports the use of snow-on imagery as a successful technique to identify conifer regeneration.

Casey Menick↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States have experienced increasingly severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage production for their cattle and a need to adjust their business models, even considering abandoning their businesses altogether. The study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address stakeholder concerns of the efficacy of existing remotely sensed rangeland production estimation platforms and explore possible early warning climatic indicators of drought. The study identified two key rangeland platforms, the Rangeland Production Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP) and used in-situ data to statistically validate their efficacy. RAP outperformed RPMS in estimating in-situ biomass and was therefore used in our climate modeling. Our study performed a random forest analysis, sampling 1500 points across the study area, comparing monthly RAP biomass estimates to a variety of climatic variables, including mean precipitation, temperature, palmer drought severity index, snow water equivalent, wind speed and direction, and vapor pressure deficit. After analysis, our study determined that vapor pressure deficit is a key indicator in predicting forage production in MLRA-48. Our study recommends the use of RAP in estimating potential forage, with caution for its tendency to overestimate. Our climate analysis provided our partners with greater understanding of the influence of various climatic factors in determining forage production and allows them to assist landowners in planning for future drought.

Addie Gonzalez↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States experienced severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage for their cattle and a need to adjust their business models. This study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address concerns regarding the efficacy of remotely sensed rangeland production platforms and identify early warning climatic indicators of drought. The study identified two key platforms, The Rangeland Productivity Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP), which use NASA Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI, and Landsat 9 OLI-2 to estimate rangeland biomass. We regressed these with in situ biomass data to validate their efficacy and found that RAP was more effective than RPMS in estimating rangeland biomass, though it presents a tendency to overestimate. Our study performed a random forest analysis, comparing monthly RAP biomass estimates to a variety of climate variables, including mean precipitation, temperature, Palmer Drought Severity Index, snow water equivalent, snow persistence from Terra MODIS, wind speed and direction, and vapor pressure deficit. We determined that vapor pressure deficit and precipitation are key indicators in predicting forage production in MLRA-48. Our climate analysis provided our partners with greater understanding of the influence of various climate variables in determining rangeland production and allows them to assist land managers in drought mitigation.

remote sensing↗

California & Oregon Ecological Forecasting: Detecting and Forecasting Fog Occurrence, Frequency, and Change to Support Coast Redwood (Sequoia sempervirens) Habitat Assessments

Fog and low clouds play an important role in providing moisture to coastal ecosystems. Coast redwood (Sequoia sempervirens) forests are currently distributed along a narrow strip of coastline in California and Oregon and rely on the presence of marine fog for moisture availability during the dry season (June-October). Recent time series analyses presented an uncertain future of fog frequency; however, a decline in fog presence may have adverse effects on the coast redwood habitat. To support Save the Redwoods League, a non-profit organization dedicated to coast redwood forest management, the team analyzed hourly fog data from the Geospatial Operational Environmental Satellite 17 (GOES-17) Advanced Baseline Imager (ABI) and daily cloud cover data from the Moderate Resolution Imaging Spectroradiometer (MODIS) aboard the Terra satellite. To explore present day fog longevity, GOES-17 was utilized to map the number of fog hours per day for the 2019 and 2020 dry seasons. The MODIS cloud flag was used to map the presence or absence of daily fog, which was summarized to create a monthly fog frequency dataset and identify trends in fog presence between 2000-2020. Both datasets were used as inputs into the random forest machine learning algorithm to identify climatic drivers of fog presence and longevity over the landscape. The present-day models suggested that daily temperature difference is a driving force behind fog presence and longevity. Trends in fog presence from 2000-2020 indicated great interannual variability. Finally, fog presence was modeled under a 2080 climate projection to shed light on the future of fog presence under a projected warmer climate. Model results projected an overall decline in fog presence during the dry season in the 2080s. Decreased fog presence as a result of increased temperature difference under a warmer climate remains to be a topic of investigation as to the impact on future redwood habitat suitability.

DEVELOP Project Summary↗

California & Oregon Ecological Forecasting: Detecting and Forecasting Fog Occurrence, Frequency, and Change to Inform Coast Redwood (Sequoia sempervirens) Habitat Assessments

Fog and low clouds play an important role in providing moisture to coastal ecosystems. Coast redwood (Sequoia sempervirens) forests are currently distributed along a narrow strip of coastline in California and Oregon and rely on the presence of marine fog for moisture availability during the dry season (June-October). Recent time series analyses presented an uncertain future of fog frequency; however, a decline in fog presence may have adverse effects on the coast redwood habitat. To complement ongoing work by Save the Redwoods League, a non-profit organization dedicated to coast redwood forest management, the team analyzed hourly fog data from the Geospatial Operational Environmental Satellite 17 (GOES-17) Advanced Baseline Imager (ABI) and daily cloud cover data from the Moderate Resolution Imaging Spectroradiometer (MODIS) aboard the Terra satellite. To explore present day fog longevity, GOES-17 was utilized to map the number of fog hours per day for the 2019 and 2020 dry seasons. The MODIS cloud flag was used to map the presence or absence of daily fog, which was summarized to create a monthly fog frequency dataset and identify trends in fog presence between 2000-2020. Both datasets were used as inputs into the random forest machine learning algorithm to identify climatic drivers of fog presence and longevity over the landscape. The present-day models suggested that daily temperature difference is a driving force behind fog presence and longevity. Trends in fog presence from 2000-2020 indicated great interannual variability. Finally, fog presence was modeled under a 2080 climate projection to shed light on the future of fog presence under a projected warmer climate. Model results projected an overall decline in fog presence during the dry season in the 2080s. Decreased fog presence as a result of increased temperature difference under a warmer climate remains to be a topic of investigation as to the impact on future redwood habitat suitability.

DEVELOP Technical Paper↗

New York Ecological Forecasting: Utilizing NASA Earth Observations to Map Ash Distribution and Inform Emerald Ash Borer Control

Since their first sightings in the U.S. in 2002, emerald ash borer beetles (Agrilus planipennis; EAB) have killed millions of native ash (Fraxinus spp.) trees across 35 states. Infected ash stands frequently exhibit complete mortality, with the predicted result being the functional extinction of native ash in U.S. forests. In August of 2020, EAB was discovered in the 6.1-million-acre Adirondack Park. The team’s partners at the Adirondack Park Invasive Plant Program (APIPP) desired ash tree distribution and EAB susceptibility information to help improve EAB bio-control efficiency and apply the methodology to future invasive programs. To assist, the team mapped ash tree distribution using NASA Earth observations from Landsat 7 Enhanced Thematic Mapper Plus (ETM+) and Shuttle Radar Topography Mission (SRTM), along with hyperspectral imagery from the Airborne Visible/Infrared Imaging Spectrometer (AVIRIS). Field data from the Monitoring and Managing Ash (MaMA) project, iMapInvasives and iNaturalist databases, and the New York State Department of Environmental Conservation (NYSDEC) provided ground truthing for mapping and modeling. Results indicate that for ash detection, the team’s Spectral Angle Mapping (SAM) hyperspectral classification is slightly more sensitive but less accurate than multispectral Random Forest (RF) classification, though neither method was above a ~20% detection rate. End products include maps of ash extent derived from both imagery types, a model forecasting future spread scenarios based on current EAB presence, and outreach materials. These products inform APIPP’s management decisions and facilitate public awareness of EAB’s threat to communities within the region.

Liam Megraw↗

Northern Rockies Ecological Conservation: Leveraging Earth Observations to Monitor and Predict Populations of Federally Threatened Whitebark Pine (Pinus albicaulis) across the Intermountain West

Whitebark pine (WBP; Pinus albicaulis) is an ecologically important species in North America. As a federally listed threatened species, an understanding of WBP habitat, distribution, and health is important for the natural resource managers of the National Park Service, United States Forest Service, Bureau of Land Management, Fish and Wildlife Service, and non-profit organizations such as the Whitebark Pine Ecosystem Foundation. Previous attempts to develop models of WBP habitat suitability and distribution lack confidence in their validity and integrity for these organizations. The updated models of habitat suitability and distribution developed by this study would provide managers with a capability to be employed in the conservation and future research direction for WBP. Thus, we developed a habitat suitability model of WBP at a high spatial resolution (Landsat 9 Operational Land Image-2, National Land Cover Database, NASA Shuttle Radar Topography Mission; 30m pixels) using a generalized logistic regression with an area under the curve value of 0.754. We extracted spectral reflectance signatures from overlapped ground sample points and Sentinel-2 Multispectral Instrument. The spectral signature analysis indicates WBP is separable from other tree species. We also utilized a visual validation approach and random forest (RF) modeling to separate WBP from limber pine. Through visual validation the RF classifier successfully identified 8out of 10 WBP trees gathered through ground truth points. Additionally, we achieved an overall accuracy of 91%in our confusion matrix for the distribution model using a dependent validation approach. The derived products from this study allow project partners to assess current suitable habitat and apparent health status in areas of identified WBP occurrence, providing data to aid future research regarding WBP health.

Sentinel-2↗

Assessing Alaskan boreal forest landcover affected by climate-wildfire interactions from ground truth surveys and NASA airborne remote sensing

Alaska’s boreal forest is facing unprecedented challenges under rapid climate warming (increasingly severe fires, droughts, pest/disease outbreaks) that may destabilize its function as a global carbon sink. Forests near Fairbanks may be especially vulnerable, impacting air quality and ecosystem services. We combined GT (ground truthing) with Airborne Visible InfraRed Imaging Spectrometer (AVIRIS-NG) images collected by the NASA Arctic-Boreal Vulnerability Experiment (ABoVE) program (2017-2019) to assess landcover change at five recently burned sites (2001-2019) of different fire severities and moisture regimes within 30 miles of Fairbanks. GT included tree seedling counts, understory % cover and >50% leaf canopy color assessment. 36 circular plots (1/30 ha radius) including 6 moderate to severely burned plots were selected across sites. 31 additional sites including 12 burned sites were geotagged in photos. AVIRIS images were processed from 29 spectral bands selected to identify changes in chlorophyll and water content. Images were segmented into natural boundaries (polygons) using ENVI 5.5 software. A spectral library of 8 AVIRIS bands with high between-class/low within-class variation was used in two random forest models to predict vegetation classes (model 1: 12 classes, model 2: 14 classes) in each AVIRIS scene, using 20% of the data as training data. Model 2 classified 20% more polygons overall, but only 42% of GT/geotagged polygons were correctly classified by both models. More forest sites were correctly classified (63%) than open vegetation (32%) or post-fire sites (46%). 50% of aspen forest and post-fire polygons were misclassified as shrubland. GT revealed that post-fire plots supported 134,000 (± 48,000) tree seedlings and saplings ha-1 (0.2 - 4 m height, 64% deciduous) versus 2500 (± 2100) shrubs ha-1 (1-6 m height). > 50% canopy browning was observed in conifer forest (8 plots) with no signs of insect infestation. Canopy herbivory > 50% (leaf miner, leaf beetle) and moose herbivory of tree bark was seen across aspen sites. Our study suggests: 1) low canopy vegetation presents challenges for improved landcover classification, and 2) aspen forest should be differentiated in vegetation maps which would aid in tracking herbivory.

Alaska↗

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary↗

Evaluation of Algorithms for a Miles-in-Trail Decision Support Tool

Four machine learning algorithms were prototyped and evaluated for use in a proposed decision support tool that would assist air traffic managers as they set Miles-in-Trail restrictions. The tool would display probabilities that each possible Miles-in-Trail value should be used in a given situation. The algorithms were evaluated with an expected Miles-in-Trail cost that assumes traffic managers set restrictions based on the tool-suggested probabilities. Basic Support Vector Machine, random forest, and decision tree algorithms were evaluated, as was a softmax regression algorithm that was modified to explicitly reduce the expected Miles-in-Trail cost. The algorithms were evaluated with data from the summer of 2011 for air traffic flows bound to the Newark Liberty International Airport (EWR) over the ARD, PENNS, and SHAFF fixes. The algorithms were provided with 18 input features that describe the weather at EWR, the runway configuration at EWR, the scheduled traffic demand at EWR and the fixes, and other traffic management initiatives in place at EWR. Features describing other traffic management initiatives at EWR and the weather at EWR achieved relatively high information gain scores, indicating that they are the most useful for estimating Miles-in-Trail. In spite of a high variance or over-fitting problem, the decision tree algorithm achieved the lowest expected Miles-in-Trail costs when the algorithms were evaluated using 10-fold cross validation with the summer 2011 data for these air traffic flows.

Bloem, Michael↗