Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

240 records · Page 14

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

Model and remote-sensing-guided experimental design and hypothesis generation for monitoring snow-soil–plant interactions

In this study, we develop a machine-learning (ML)-enabled strategy for selecting hillslope-scale ecohydrological monitoring sites within snow-dominated mountainous watersheds, with a particular focus on snow-soil–plant interactions. Data layers rely on spatial data layers from both remote sensing and hydrological model simulations. Specifically, a Landsat-based foresummer drought sensitivity index is used to define the dependency of the annual peak plant productivity on the Palmer drought severity index in the early growing season. Hydrological simulations provide the spatiotemporal dynamics of near-surface soil moisture and snow depth. In this framework, a regression analysis identifies the key hydrological variables relevant to the spatial heterogeneity of drought sensitivity. We then apply unsupervised clustering to these key variables, using the Gaussian mixture model, to group hillslopes into several zones that have divergent relationships regarding soil moisture, snow dynamics, and drought sensitivity. Using the datasets collected in the East River Watershed (Crested Butte, Colorado, United States), results show that drought sensitivity is significantly correlated with model-derived soil moisture and snow-free timing over space and time. The relationship is, however, non-linear, such that the correlation decreases above a threshold elevation and in a heavy snow year due to large snowpacks, lateral flow, and soil storage limitations. Clustering is then able to define the zones that have high or low sensitivity to drought, as well as the mid-elevation regions where sensitivity is associated with the topographic aspect and net potential radiation. In addition, the algorithm identifies the most representative hillslopes with road/trail access within each zone for installing monitoring sites. Our method also aims to significantly increase the use of ML and model-simulation results to guide critical zone and watershed monitoring activities.

54 ENVIRONMENTAL SCIENCES↗

Integrating Reservoirs into the Dissolved Organic Matter Versus Primary Production Paradigm: How Does Chlorophyll-$a$ Change Across Dissolved Organic Carbon Concentrations in Reservoirs?

Primary production in freshwater ecosystems is largely a function of light and nutrient availability, both of which have been changing in many lakes and reservoirs in response to anthropogenic pressures. Recent studies focusing on natural lakes have found a hump-shaped response of primary production (sometimes measured as chlorophyll-$a$) to dissolved organic matter (DOM, measured as dissolved organic carbon, DOC), which has both light-absorbing chromophoric properties and DOM-bound nutrients. We used the United States National Lakes Assessment dataset to integrate reservoirs into this paradigm in comparison with natural lakes and assessed the relative differences in the predicted response’s model structure, regression parameter values, and drivers of the chlorophyll-$a$ residuals. We found that chlorophyll-$a$ in reservoirs exhibited a hump-shaped response to DOC, while natural lakes from this dataset were better fit with a linear response, differing from previous studies focused on boreal lakes. Despite this, reservoirs had a greater maximum chlorophyll-a response compared to natural lakes in this study (45.5 versus 33.8 μg L -1 ), which occurred at a lower DOC concentration threshold (18.3 versus 26.4 mg L -1 ) when compared using quadratic models. Reservoirs had lower median light:nutrient values compared to natural lakes, and greater median surface area and total phosphorus (TP), that can all influence the light environment and the peak chlorophyll-a responses. In both reservoirs and natural lakes, chlorophyll-$a$ residuals were most strongly influenced by TP, where TP < 25-30 µg L -1 suppressed chlorophyll-a residuals and higher TP amplified them. Light:nutrient values were somewhat important predictors, and patterns with chlorophyll-$a$ residuals supported previous work showing low light:nutrient values amplified chlorophyll-$a$ responses and higher values suppressed them. In conclusion, quantifying the shape of the response of primary production to DOM quantity and quality as well as the drivers of the residuals, namely TP for lakes and reservoirs in this dataset, will be important for understanding the effects that changes in water quality may have on primary production and freshwater ecosystem processes.

59 BASIC BIOLOGICAL SCIENCES↗