Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

47 records · Page 3

Advancements in Blowing Dust Detection at Night via Machine Learning

This presentation introduces operational users to a machine-learning based Dust Probability product developed by the NASA SPoRT program for the application of detecting and monitoring blowing dust plumes at night. Advances in earth observing satellites has improved monitoring and detection of dust both day and night through derived imagery such as the Dust RGB. However, limitations of the RGB at night result in less contrast between dust and land surface features, as seen by the user. A Machine Learning (ML) model has been developed and applied to GOES-16 ABI to overcome this limitation and improve nighttime dust detection. The ML capability is a subset of Artificial Intelligence methods. In this case the Dust ML model was developed using a simple Random Forest (RF) model, typically used to solve classification challenges (or to provide regression type output). The goal was to leverage the strengths of the RF model to learn how to identify blowing dust, and hence, overcome the limitation of a user trying to detect blowing dust within the satellite imagery by eye alone. A brief description of the ML model development will be provided. However, the focus of the presentation will be on the initial user feedback from the assessment of this tool for the 2022 blowing dust events of March through April. During this time several users across the U.S. Southwest collaborated to apply this Dust ML product at night as a complement to the existing Dust RGB in order to determine if it provided greater operational efficiency and value.

Machine Learning↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

Southern Rockies Western Slope Agriculture: Identifying Drivers of Rangeland Production for Drought Planning on the Western Slope of the Southern Rockies

Over the last decade, the southern Rocky Mountains of the United States experienced severe and variable drought. Local ranchers and landowners have reported strain on their operations, citing decreasing forage for their cattle and a need to adjust their business models. This study identified Major Land Resource Area-48 (MLRA-48) and northwestern Colorado as the key region for analysis. NASA DEVELOP partnered with the BLM Colorado River Field Office, Colorado State University Extension, USDA Forest Service, and the National Drought Mitigation Center to address concerns regarding the efficacy of remotely sensed rangeland production platforms and identify early warning climatic indicators of drought. The study identified two key platforms, The Rangeland Productivity Monitoring Service (RPMS) and Rangeland Analysis Platform (RAP), which use NASA Landsat 5 TM, Landsat 7 ETM+, Landsat 8 OLI, and Landsat 9 OLI-2 to estimate rangeland biomass. We regressed these with in situ biomass data to validate their efficacy and found that RAP was more effective than RPMS in estimating rangeland biomass, though it presents a tendency to overestimate. Our study performed a random forest analysis, comparing monthly RAP biomass estimates to a variety of climate variables, including mean precipitation, temperature, Palmer Drought Severity Index, snow water equivalent, snow persistence from Terra MODIS, wind speed and direction, and vapor pressure deficit. We determined that vapor pressure deficit and precipitation are key indicators in predicting forage production in MLRA-48. Our climate analysis provided our partners with greater understanding of the influence of various climate variables in determining rangeland production and allows them to assist land managers in drought mitigation.

remote sensing↗

Mapping tree canopy cover and canopy height with L-band SAR using LiDAR data and Random Forests

Light detection and ranging (LiDAR) data can provide direct measurements of vegetation structures but are limited by the sparse spatial coverage. Polarimetric synthetic aperture radar (SAR) can perform large-scale high-resolution mapping without weather constraints but the information about vegetation and ground subsurface are mixed in the backscatter data. In this paper, we adopted the Random Forests algorithm to train an upscaling function using tree canopy cover (TCC) and canopy height model (CHM) derived from Goddard’s LiDAR, Hyperspectral and Thermal Imager (G-LiHT) data. The regression model is then applied to the L-band Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) data acquired during the 2017 Arctic-Boreal Vulnerability Experiment (ABoVE) airborne campaign to map the TCC and CHM over the Delta Junction area in interior Alaska.

Moghaddam, Mahta↗

Northern Rockies Ecological Conservation: Leveraging Earth Observations to Monitor and Predict Populations of Federally Threatened Whitebark Pine (Pinus albicaulis) across the Intermountain West

Whitebark pine (WBP; Pinus albicaulis) is an ecologically important species in North America. As a federally listed threatened species, an understanding of WBP habitat, distribution, and health is important for the natural resource managers of the National Park Service, United States Forest Service, Bureau of Land Management, Fish and Wildlife Service, and non-profit organizations such as the Whitebark Pine Ecosystem Foundation. Previous attempts to develop models of WBP habitat suitability and distribution lack confidence in their validity and integrity for these organizations. The updated models of habitat suitability and distribution developed by this study would provide managers with a capability to be employed in the conservation and future research direction for WBP. Thus, we developed a habitat suitability model of WBP at a high spatial resolution (Landsat 9 Operational Land Image-2, National Land Cover Database, NASA Shuttle Radar Topography Mission; 30m pixels) using a generalized logistic regression with an area under the curve value of 0.754. We extracted spectral reflectance signatures from overlapped ground sample points and Sentinel-2 Multispectral Instrument. The spectral signature analysis indicates WBP is separable from other tree species. We also utilized a visual validation approach and random forest (RF) modeling to separate WBP from limber pine. Through visual validation the RF classifier successfully identified 8out of 10 WBP trees gathered through ground truth points. Additionally, we achieved an overall accuracy of 91%in our confusion matrix for the distribution model using a dependent validation approach. The derived products from this study allow project partners to assess current suitable habitat and apparent health status in areas of identified WBP occurrence, providing data to aid future research regarding WBP health.

Sentinel-2↗

Southern Colorado Disasters: Using NASA Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Project Summary↗

Southern Colorado Disasters: Using NASA Earth Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Technical Paper↗

Forest Biomass Mapping From Lidar and Radar Synergies

The use of lidar and radar instruments to measure forest structure attributes such as height and biomass at global scales is being considered for a future Earth Observation satellite mission, DESDynI (Deformation, Ecosystem Structure, and Dynamics of Ice). Large footprint lidar makes a direct measurement of the heights of scatterers in the illuminated footprint and can yield accurate information about the vertical profile of the canopy within lidar footprint samples. Synthetic Aperture Radar (SAR) is known to sense the canopy volume, especially at longer wavelengths and provides image data. Methods for biomass mapping by a combination of lidar sampling and radar mapping need to be developed. In this study, several issues in this respect were investigated using aircraft borne lidar and SAR data in Howland, Maine, USA. The stepwise regression selected the height indices rh50 and rh75 of the Laser Vegetation Imaging Sensor (LVIS) data for predicting field measured biomass with a R(exp 2) of 0.71 and RMSE of 31.33 Mg/ha. The above-ground biomass map generated from this regression model was considered to represent the true biomass of the area and used as a reference map since no better biomass map exists for the area. Random samples were taken from the biomass map and the correlation between the sampled biomass and co-located SAR signature was studied. The best models were used to extend the biomass from lidar samples into all forested areas in the study area, which mimics a procedure that could be used for the future DESDYnI Mission. It was found that depending on the data types used (quad-pol or dual-pol) the SAR data can predict the lidar biomass samples with R2 of 0.63-0.71, RMSE of 32.0-28.2 Mg/ha up to biomass levels of 200-250 Mg/ha. The mean biomass of the study area calculated from the biomass maps generated by lidar- SAR synergy 63 was within 10% of the reference biomass map derived from LVIS data. The results from this study are preliminary, but do show the potential of the combined use of lidar samples and radar imagery for forest biomass mapping. Various issues regarding lidar/radar data synergies for biomass mapping are discussed in the paper.

Sun, Guoqing↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

The use of space and high altitude aerial photography to classify forest land and to detect forest disturbances

In October 1969, an investigation was begun near Atlanta, Georgia, to explore the possibilities of developing predictors for forest land and stand condition classifications using space photography. It has been found that forest area can be predicted with reasonable accuracy on space photographs using ocular techniques. Infrared color film is the best single multiband sensor for this purpose. Using the Apollo 9 infrared color photographs taken in March 1969 photointerpreters were able to predict forest area for small units consistently within 5 to 10 percent of ground truth. Approximately 5,000 density data points were recorded for 14 scan lines selected at random from five study blocks. The mean densities and standard deviations were computed for 13 separate land use classes. The results indicate that forest area cannot be separated from other land uses with a high degree of accuracy using optical film density alone. If, however, densities derived by introducing red, green, and blue cutoff filters in the optical system of the microdensitometer are combined with their differences and their ratios in regression analysis techniques, there is a good possibility of discriminating forest from all other classes.

Aldrich, R. C.↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗