Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Remaining Useful Strength (RUS) Prediction of SiCf-SiCm Composite Materials Using Deep Learning and Acoustic Emission

Prognosis techniques for prediction of remaining useful life (RUL) are of crucial importance to the management of complex systems for they can lead to appropriate maintenance interventions and improvements in reliability. While various data-driven methods have been introduced to predict the remaining useful life (RUL) of machinery systems or batteries, no research has been reported on the remaining useful strength (RUS) prediction of silicon carbide fiber reinforced silicon carbide matrix (SiCf-SiCm) materials with pivotal role in its potential usage as a structural material in nuclear reactors and turbine engines. Knowledge of its degradation process is of the utmost importance to the manufacturers. For this purpose, two approaches based on the machine-learning techniques of random-forest (RF) and convolutional neural network (CNN) are proposed to predict the RUS of SiCf-SiCm using only acoustic emission (AE) signals generated during the material’s stress applying process. Experimental results show that the CNN models achieved better predictive performance than the RF models but the latter with expert-engineered features achieves better prediction for AE signals in the early stage of degradation. Additionally, our results demonstrate that both models can correctly predict the SiCf-SiCm RUS as evaluated by our robust testing method from which the best average root mean square error (RMSE) and Pearson correlation coefficient of 3.55 ksi units and 0.85 were obtained.

36 MATERIALS SCIENCE↗

Evaluation of cooling setpoint setback savings in commercial buildings using electricity and exterior temperature time series data

Commercial buildings account for a significant amount of total energy produced in the US, and the Heating Ventilation and Cooling (HVAC) systems are one of the most significant components of their overall consumption. In this study, we proposed a new data-driven approach to evaluate HVAC cooling systems in commercial buildings and identify savings opportunities. The focus is an investigation of the impact of thermostat setpoint setback but using only whole building, electricity data taken at 15-min intervals for the analysis. We conducted a comparative study of setpoint setback characteristics on 432 commercial buildings with 5 building usage types across the United States. To accomplish this, both piecewise and Random Forest regression algorithms were employed using electricity and exterior temperature datasets to identify operational characteristics and the effective setpoints in the building to determine the corresponding savings opportunities. Both occupied and unoccupied time periods were studied across cooling degree days (CDD), when air conditioning is typically operational. Here the results show that in commercial buildings, on average, cooling systems account for 9.5% of total consumption. When a one degree setback during the cooling season is applied, an average of approximately 1.1% of annual consumption is achieved; retail and office buildings demonstrate the highest potential for savings. Additionally, we identified that the number of cooling degree days and base to peak ratio (BPR) are the most important variables for predicting the magnitude of the consumption of cooling systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Understanding the Drivers of Mobility during the COVID-19 Pandemic in Florida, USA Using a Machine Learning Approach

As of March 2021, the State of Florida, U.S.A. had accounted for approximately 6.67% of total COVID-19 (SARS-CoV-2 coronavirus disease) cases in the U.S. The main objective of this research is to analyze mobility patterns during a three month period in summer 2020, when COVID-19 case numbers were very high for three Florida counties, Miami-Dade, Broward, and Palm Beach counties. To investigate patterns, as well as drivers, related to changes in mobility across the tri-county region, a random forest regression model was built using sociodemographic, travel, and built environment factors, as well as COVID-19 positive case data. Mobility patterns declined in each county when new COVID-19 infections began to rise, beginning in mid-June 2020. While the mean number of bar and restaurant visits was lower overall due to closures, analysis showed that these visits remained a top factor that impacted mobility for all three counties, even with a rise in cases. Our modeling results suggest that there were mobility pattern differences between counties with respect to factors relating, for example, to race and ethnicity (different population groups factored differently in each county), as well as social distancing or travel-related factors (e.g., staying at home behaviors) over the two time periods prior to and after the spike of COVID-19 cases.

60 APPLIED LIFE SCIENCES↗

Deep Learning Classification of Cheatgrass Invasion in the Western United States Using Biophysical and Remote Sensing Data

Cheatgrass (Bromus tectorum) invasion is driving an emerging cycle of increased fire frequency and irreversible loss of wildlife habitat in the western US. Yet, detailed spatial information about its occurrence is still lacking for much of its presumably invaded range. Deep learning (DL) has demonstrated success for remote sensing applications but is less tested on more challenging tasks like identifying biological invasions using sub-pixel phenomena. We compare two DL architectures and the more conventional Random Forest and Logistic Regression methods to improve upon a previous effort to map cheatgrass occurrence at >2% canopy cover. High-dimensional sets of biophysical, MODIS, and Landsat-7 ETM+ predictor variables are also compared to evaluate different multi-modal data strategies. All model configurations improved results relative to the case study and accuracy generally improved by combining data from both sensors with biophysical data. Cheatgrass occurrence is mapped at 30 m ground sample distance (GSD) with an estimated 78.1% accuracy, compared to 250-m GSD and 71% map accuracy in the case study. Furthermore, DL is shown to be competitive with well-established machine learning methods in a limited data regime, suggesting it can be an effective tool for mapping biological invasions and more broadly for multi-modal remote sensing applications.

54 ENVIRONMENTAL SCIENCES↗

Chatter detection in simulated machining data: a simple refined approach to vibration data

Vibration monitoring is a critical aspect of assessing the health and performance of machinery and industrial processes. This study explores the application of machine learning techniques, specifically the Random Forest (RF) classification model, to predict and classify chatter—a detrimental self-excited vibration phenomenon—during machining operations. While sophisticated methods have been employed to address chatter, this research investigates the efficacy of a novel approach to an RF model. The study leverages simulated vibration data, bypassing resource-intensive real-world data collection, to develop a versatile chatter detection model applicable across diverse machining configurations. The feature extraction process combines time-series features and Fast Fourier Transform (FFT) data features, streamlining the model while addressing challenges posed by feature selection. By focusing on the RF model’s simplicity and efficiency, this research advances chatter detection techniques, offering a practical tool with improved generalizability, computational efficiency, and ease of interpretation. The study demonstrates that innovation can reside in simplicity, opening avenues for wider applicability and accelerated progress in the machining industry.

42 ENGINEERING↗

Leveraging Machine Learning and Geo-Tagged Citizen Science Data to Disentangle the Factors of Avian Mortality Events at the Species Level

Abrupt environmental changes can affect the population structures of living species and cause habitat loss and fragmentations in the ecosystem. During August–October 2020, remarkably high mortality events of avian species were reported across the western and central United States, likely resulting from winter storms and wildfires. However, the differences of mortality events among various species responding to the abrupt environmental changes remain poorly understood. In this study, we focused on three species, Wilson’s Warbler, Barn Owl, and Common Murre, with the highest mortality events that had been recorded by citizen scientists. We leveraged the citizen science data and multiple remotely sensed earth observations and employed the ensemble random forest models to disentangle the species responses to winter storm and wildfire. We found that the mortality events of Wilson’s Warbler were primarily impacted by early winter storms, with more deaths identified in areas with a higher average daily snow cover. The Barn Owl’s mortalities were more identified in places with severe wildfire-induced air pollution. Both winter storms and wildfire had relatively mild effects on the mortality of Common Murre, which might be more related to anomalously warm water. Our findings highlight the species-specific responses to environmental changes, which can provide significant insights into the resilience of ecosystems to environmental change and avian conservations. Additionally, the study emphasized the efficiency and effectiveness of monitoring large-scale abrupt environmental changes and conservation using remotely sensed and citizen science data.

47 OTHER INSTRUMENTATION↗

Combining Machine Learning and Numerical Simulation for High-Resolution PM2.5 Concentration Forecast

Forecasting ambient PM2.5 concentrations with spatiotemporal coverage is key to alerting decision-makers of pollution episodes and preventing detrimental public exposure, especially in regions with limited ground air monitoring stations. The existing methods either rely on chemical transport models (CTMs) to forecast spatial distribution of PM2.5 with nontrivial uncertainty or statistical algorithms to forecast PM2.5 concentration time-series at air monitoring locations without continuous spatial coverage. In this study, we developed a PM2.5 forecast framework by combining the robust Random Forest algorithm with a publicly accessible global CTM forecast product – NASA’s Goddard Earth Observing System “Composition Forecasting” (GEOS-CF), providing spatiotemporally continuous PM2.5 concentration forecasts for the next five days at a 1-km spatial resolution. Our forecast experiment was conducted for a region in Central China including the populous and polluted Fenwei Plain. The forecast for the next two days had overall validation R2 of 0.76 and 0.64, respectively; the R2 was around 0.5 for the following three forecast days. Spatial cross-validation showed similar validation metrics. Our forecast model, with validation normalized mean bias close to zero, substantially reduced the large biases in GEOS-CF. The proposed framework requires minimal computational resources compared to running CTMs at urban scales, enabling near-real-time PM2.5 forecast in resource-restricted environments.

PM2.5↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Combined Effects of Stream Hydrology and Land Use on Basin‐Scale Hyporheic Zone Denitrification in the Columbia River Basin

Abstract Denitrification in the hyporheic zone (HZ) of river corridors is crucial to removing excess nitrogen in rivers from anthropogenic activities. However, previous modeling studies of the effectiveness of river corridors in removing excess nitrogen via denitrification were often limited to the reach‐scale and low‐order stream watersheds. We developed a basin‐scale river corridor model for the Columbia River Basin with random forest models to identify the dominant factors associated with the spatial variation of HZ denitrification. Our modeling results suggest that the combined effects of hydrologic variability in reaches and substrate availability influenced by land use are associated with the spatial variability of modeled HZ denitrification at the basin scale. Hyporheic exchange flux can explain most of spatial variation of denitrification amounts in reaches of different sizes, while among the reaches affected by different land uses, the combination of hyporheic exchange flux and stream dissolved organic carbon (DOC) concentration can explain the denitrification differences. Also, we can generalize that the most influential watershed and channel variables controlling denitrification variation are channel morphology parameters (median grain size (D50), stream slope), climate (annual precipitation and evapotranspiration), and stream DOC‐related parameters (percent of shrub area). The modeling framework in our study can serve as a valuable tool to identify the limiting factors in removing excess nitrogen pollution in large river basins where direct measurement is often infeasible.

54 ENVIRONMENTAL SCIENCES↗

What regulates decomposition in agroecosystems? Insights from reading the tea leaves

Litter decomposition is a critical Earth process, recycling nutrients and setting a portion of plant tissue on a path toward soil organic matter. Despite this importance, we still lack a good understanding of local factors that regulate decomposition, especially in agroecosystems where management plays an outsized role. Using a narrow range of climate and soils, we buried 1,308 pre-manufactured “litter bags” of differing residue quality (i.e., green and rooibos tea leaves) in 109 plots across several management practices to (1) explore the local controls on decomposition in agroecosystems and (2) test the robustness of the Tea Bag Index (TBI). We found that management practices intended to increase soil ecosystem services, that is, soil health, altered the decomposition of both teas. For example, adding nitrogen fertilizer and implementing perennial cropping decreased the extent of green tea decomposition (carbon-to-nitrogen ratio, or C:N = 12.8). No-tillage increased, but perennial cropping decreased, the rate of rooibos tea decomposition (C:N = 50.1). Cropped prairie accelerated green tea decomposition and increased the extent of red tea decomposition. A random forest regression model showed that soil temperature was the strongest predictor of green tea decomposition, but a soil health score also played a significant role in predicting the mass remaining. Soil texture and nutrient availability best predicted rooibos tea decomposition. Finer textured soils seemed to decelerate rooibos decomposition but increased the extent of decomposition. Furthermore, we demonstrated that the TBI metrics correlated somewhat well with empirically derived decomposition constants and were similarly sensitive to the effects of management. Still, the green tea stabilization factor had a substantial prediction bias. Our study increased our basic understanding of what regulates decomposition in agroecosystems. It also showed that the TBI can be a scientifically rigorous citizen science approach to monitoring changes in soil health.

60 APPLIED LIFE SCIENCES↗

Iona Ecological Conservation: Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan civil war in 2002, human and livestock populations have increased exponentially within Iona National Park. An ongoing drought since 2017 has brought these people and livestock into increasing competition with local wildlife for resources – highlighting a conservation challenge that will become more entrenched as the effects of anthropogenic climate change increase. In 2019, African Parks began co-managing Iona National Park in Angola with the Angolan government, hoping to enact scientifically grounded management strategies to meet this challenge. To accomplish this, African Parks needed contemporary and historic information on the spatial distribution of landcover types within Iona and adjacent areas. We constructed and applied a Random Forest classifier in Google Earth Engine to multispectral imagery gathered from Landsat 5, 7, 8 and Sentinel-1 and 2 to meet this need. Using the classifier, we generated a time-series of land cover maps between 1990–2023, from which landscape metrics and change detection analysis were calculated to show how certain habitats and formations had changed over time. The resulting maps have producer and user’s accuracies above 87% and show four broad landcover regions within the study area. Notably, we observed a decrease in the park’s diversity as per the Shannon Diversity Index – an index that considers the richness of classes, as well the evenness of their distribution. A lack of arid specific land cover indices and ground-truthed training data from earlier years limited the accuracy and resolution of our landcover maps. However, this project still demonstrates that Earth observations can be used to form the basis of conservation policy in arid environments, where ground-truth data may be difficult to obtain or non-existent.

remote sensing↗

Estimating Fine-Resolution Shortwave Broadband Albedo of Croplands from Harmonized Landsat and Sentinel-2 Data

Altered surface albedo due to land-cover conversions and management is a significant driver of global climate change. Albedo can be directly measured at ground stations, and remote sensing data can be used to scale-up albedo values to regional and global levels. Some previous studies have retrieved fine-resolution (10–30 m) instantaneous albedo and coarse-resolution (500–1000 m) daily mean albedo from remote sensing data, but they all required the input of Moderate Resolution Imaging Spectroradiometer (MODIS) albedo information at 500-m resolution, and none have assembled both instantaneous and daily albedo based exclusively on fine-resolution satellite data. Here, to address this issue, we compiled 387 instantaneous and 346 daily albedo records using field net radiometer measurements from the bioenergy croplands at the W. K. Kellogg Biological Station in southwest Michigan. We then connected these albedo records with a suite of variables derived from harmonized Landsat and Sentinel-2 data through two machine learning algorithms (random forest regression and extreme gradient boosting) to retrieve clear-sky instantaneous and daily shortwave broadband albedo. The performance statistics indicate reasonable accuracy of model results [root-mean-square error (RMSE)] around or below 0.03 except for snow-covered surfaces), suggesting that the retrieval of both instantaneous and daily albedo based exclusively on fine-resolution satellite data is promising. To facilitate the use of fine-resolution albedo products at the global level, future efforts need to include more albedo records of diverse surface cover types, as well as to accurately model daily albedo for cloudy days to address the “clear-sky bias.”

Harmonized Landsat and Sentinel-2↗

Empirical relationships between environmental factors and soil organic carbon produce comparable prediction accuracy as the Machine Learning

Accurate representation of environmental controllers of soil organic carbon (SOC) stocks in Earth System Model (ESM) land models could reduce uncertainties in future carbon-climate feedback projections. Using empirical relationships between environmental factors and SOC stocks to evaluate land models can help modelers understand prediction biases beyond what can be achieved with the observed SOC stocks alone. In this study, we used 31 observed environmental factors, field SOC observations (n = 6,213) from the continental US, and two Machine Learning approaches [Random Forest (RF) and Generalized Additive Modeling (GAM)] to (1) select important environmental predictors of SOC stocks, (2) derive empirical relationships between environmental factors and SOC stocks, and (3) use the derived relationships to predict SOC stocks and compare the prediction accuracy of simpler model developed with the machine learning predictions. Out of the 31 environmental factors we investigated, 12 were identified as important predictors of SOC stocks by the RF approach. In contrast, the GAM approach identified six (of those 12) environmental factors as important controllers of SOC stocks: potential evapotranspiration, normalized difference vegetation index, soil drainage condition, precipitation, elevation, and net primary productivity. The GAM approach showed minimal SOC predictive importance of the remaining six environmental factors identified by the RF approach. Our derived empirical relations produced comparable prediction accuracy as the GAM and RF approach using only a subset of environmental factors. The empirical relationships we derived using the GAM approach can serve as important benchmarks to evaluate environmental control representations of SOC stocks in ESMs, which could reduce uncertainty in predicting future carbon-climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Analysis of Hydrologic Exchange Flows and Transit Time Distributions in a Large Regulated River

Hydrologic exchange between river channels and adjacent subsurface environments is a key process that influences water quality and ecosystem function in river corridors. High-resolution numerical models were often used to resolve the spatial and temporal variations of exchange flows, which are computationally expensive. In this study, we adopt Random Forest (RF) and Extreme Gradient Boosting (XGB) approaches for deriving reduced order models of hydrologic exchange flows and associated transit time distributions, with integrated field observations (e.g., bathymetry) and hydrodynamic simulation data (e.g., river velocity, depth). The setup allows an improved understanding of the influences of various physical, spatial, and temporal factors on the hydrologic exchange flows and transit times. The predictors also contain those derived using hybrid clustering, leveraging our previous work on river corridor system hydromorphic classification. The machine learning-based predictive models are developed and validated along the Columbia River Corridor, and the results show that the top parameters are the thickness of the top geological formation layer, the flow regime, river velocity, and river depth; the RF and XGB models can achieve 70% to 80% accuracy and therefore are effective alternatives to the computational demanding numerical models of exchange flows and transit time distributions. Each machine learning model with its favorable configuration and setup have been evaluated. The transferability of the models to other river reaches and larger scales, which mostly depends on data availability, is also discussed.

97 MATHEMATICS AND COMPUTING↗

Machine learning-based inversion for acoustic impedance with large synthetic training data: Workflow and data characterization

Where wells are sparse or training data are difficult to label with high-quality wireline-derived impedance logs, machine learning (ML)-based inversion of acoustic impedance typically depends on small training data sets, leading to biased prediction. We have advanced a novel workflow that applies large synthetic seismic training data to reduce facies-related bias. Using a geologically realistic model as the truth model, we randomly select sparse seed wells to perform sequential Gaussian simulation (SGS) for impedance models of the same geometry and simulate facies variability. We implement random forest regression on 30 features extracted from the synthetic volume. We observe that more seed wells tend to reduce facies-induced bias by sampling more types of facies, resulting in a better prediction. We then focus on the responses of SGS models to facies changes, the number of seed wells necessary for a useful synthetic model, and how much a synthetic model can help ML-based inversion. Here, we observe that the SGS synthetic training model outperforms well-direct training in general. For modeled clastic shore-zone systems in Miocene Gulf of Mexico, two or more seed wells are necessary for a significant reduction of root-mean-square error and outliners, and improvement of facies imaging. In a field-data test, we apply a similar workflow to quantitatively predict acoustic impedance, which is then converted to a sand-volume map at a high-frequency sequence (10–100 m), revealing detailed facies and sandstone patterns. Such results are valuable in many geologic and engineering applications, such as hydrocarbon and CO 2 reservoir prospecting, reserve estimation, simulation, etc.

3D seismic↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

A Practical Comparison of Data-Driven Prognostics Methods for Energy Systems

This study explores data-driven prognostics for nuclear power plant (NPP) condensers, focusing on tube fouling. We utilized the Asherah nuclear power plant simulator (ANS) to compare four methods: Random Forest (RF), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory Neural Network (LSTM). By simulating various fouling scenarios in the ANS, we generated data with different degradation rates under transient operations. The models were trained and tested on these data, with performance evaluated visually and numerically including uncertainty assessment. The LSTM model excelled, exhibiting minimal prediction noise and the most accurate remaining useful life estimates across all degradation levels. Its ability to capture long-term dependencies and produce cleaner outputs makes it a strong candidate, although accurate training data across the entire component lifespan are crucial. The RF model emerged as a robust alternative, providing reliable predictions with high confidence. The FCNN and SVR models, while less effective overall, showed potential under specific conditions. FCNN offers a less complex alternative to LSTM and might benefit from larger datasets. SVR excels in precision when the quality of the training data is high. Furthermore, this study highlights the operational benefits of advanced prognostics in the energy sector and emphasizes the need for further research in NPP condenser health management through real-life experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Nominal 30-M Cropland Extent Map of Continental Africa by Integrating Pixel-Based and Object-Based Algorithms Using Sentinel-2 and Landsat-8 Data on Google Earth Engine

A satellite-derived cropland extent map at high spatial resolution (30-m or better) is a must for food and water security analysis. Precise and accurate global cropland extent maps, indicating cropland and non-cropland areas, is a starting point to develop high-level products such as crop watering methods (irrigated or rainfed), cropping intensities (e.g., single, double, or continuous cropping), crop types, cropland fallows, as well as assessment of cropland productivity (productivity per unit of land), and crop water productivity (productivity per unit of water). Uncertainties associated with the cropland extent map have cascading effects on all higher-level cropland products. However, precise and accurate cropland extent maps at high spatial resolution over large areas (e.g., continents or the globe) are challenging to produce due to the small-holder dominant agricultural systems like those found in most of Africa and Asia. Cloud-based Geospatial computing platforms and multi-date, multi-sensor satellite image inventories on Google Earth Engine offer opportunities for mapping croplands with precision and accuracy over large areas that satisfy the requirements of broad range of applications. Such maps are expected to provide highly significant improvements compared to existing products, which tend to be coarser in resolution, and often fail to capture fragmented small-holder farms especially in regions with high dynamic change within and across years. To overcome these limitations, in this research we present an approach for cropland extent mapping at high spatial resolution (30-m or better) using the 10-day, 10 to 20-m, Sentinel-2 data in combination with 16-day, 30-m, Landsat-8 data on Google Earth Engine (GEE). First, nominal 30-m resolution satellite imagery composites were created from 36,924 scenes of Sentinel-2 and Landsat-8 images for the entire African continent in 2015-2016. These composites were generated using a median-mosaic of five bands (blue, green, red, near-infrared, NDVI) during each of the two periods (period 1: January-June 2016 and period 2: July-December 2015) plus a 30-m slope layer derived from the Shuttle Radar Topographic Mission (SRTM) elevation dataset. Second, we selected Cropland/Non-cropland training samples (sample size 9791) from various sources in GEE to create pixel-based classifications. As supervised classification algorithm, Random Forest (RF) was used as the primary classifier because of its efficiency, and when over-fitting issues of RF happened due to the noise of input training data, Support Vector Machine (SVM) was applied to compensate for such defects in specific areas. Third, the Recursive Hierarchical Segmentation (RHSeg) algorithm was employed to generate an object-oriented segmentation layer based on spectral and spatial properties from the same input data. This layer was merged with the pixel-based classification to improve segmentation accuracy. Accuracies of the merged 30-m crop extent product were computed using an error matrix approach in which 1754 independent validation samples were used. In addition, a comparison was performed with other available cropland maps as well as with LULC maps to show spatial similarity. Finally, the cropland area results derived from the map were compared with UN FAO statistics. The independent accuracy assessment showed a weighted overall accuracy of 94, with a producers accuracy of 85.9 (or omission error of 14.1), and users accuracy of 68.5 (commission error of 31.5) for the cropland class. The total net cropland area (TNCA) of Africa was estimated as 313 Mha for the nominal year 2015.

Cropland mapping; cropland areas; 30-m; Landsat-8;↗