Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Southern Colorado Disasters: Using NASA Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Project Summary↗

Southern Colorado Disasters: Using NASA Earth Observations to Map Aspen Extent and Recovery Due to Wildfire

Quaking aspen (Populus tremuloides) is an important species for wildlife, watershed health, and ecosystem resilience across its range. Heavy ungulate browsing and factors influenced by a changing climate including seasonal temperature changes and moisture deficit have led to reduced post-fire aspen regeneration rates in southern Colorado. This project partnered with Trinchera Ranch and the Colorado State Forest Service to estimate aspen recovery after the Spring Creek Fire, which ignited in June of 2018. The Southern Colorado Disasters team utilized field measurements and satellite imagery from Landsat Operational Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and the Shuttle Radar Topography Mission (SRTM) to train and run several random forest models that detect pre- and post-fire aspen extent. Ocular sampling of over 500 points on high-resolution pre-fire and post-fire images identified percentage aspen cover in 30 x 30-meter grid cells. This process provided training data for regression models, which were able to detect aspen across the landscape for both time periods using multiple remote sensing vegetation health indices. In addition, landscape suitability for aspen regeneration was modeled to provide a guide for managers on where to monitor for aspen regeneration post-fire.

DEVELOP Technical Paper↗

Forest Biomass Mapping From Lidar and Radar Synergies

The use of lidar and radar instruments to measure forest structure attributes such as height and biomass at global scales is being considered for a future Earth Observation satellite mission, DESDynI (Deformation, Ecosystem Structure, and Dynamics of Ice). Large footprint lidar makes a direct measurement of the heights of scatterers in the illuminated footprint and can yield accurate information about the vertical profile of the canopy within lidar footprint samples. Synthetic Aperture Radar (SAR) is known to sense the canopy volume, especially at longer wavelengths and provides image data. Methods for biomass mapping by a combination of lidar sampling and radar mapping need to be developed. In this study, several issues in this respect were investigated using aircraft borne lidar and SAR data in Howland, Maine, USA. The stepwise regression selected the height indices rh50 and rh75 of the Laser Vegetation Imaging Sensor (LVIS) data for predicting field measured biomass with a R(exp 2) of 0.71 and RMSE of 31.33 Mg/ha. The above-ground biomass map generated from this regression model was considered to represent the true biomass of the area and used as a reference map since no better biomass map exists for the area. Random samples were taken from the biomass map and the correlation between the sampled biomass and co-located SAR signature was studied. The best models were used to extend the biomass from lidar samples into all forested areas in the study area, which mimics a procedure that could be used for the future DESDYnI Mission. It was found that depending on the data types used (quad-pol or dual-pol) the SAR data can predict the lidar biomass samples with R2 of 0.63-0.71, RMSE of 32.0-28.2 Mg/ha up to biomass levels of 200-250 Mg/ha. The mean biomass of the study area calculated from the biomass maps generated by lidar- SAR synergy 63 was within 10% of the reference biomass map derived from LVIS data. The results from this study are preliminary, but do show the potential of the combined use of lidar samples and radar imagery for forest biomass mapping. Various issues regarding lidar/radar data synergies for biomass mapping are discussed in the paper.

Sun, Guoqing↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Quantifying wildfire drivers and predictability in boreal peatlands using a two-step error-correcting machine learning framework in TeFire v1.0

Abstract. Wildfires are becoming an increasing challenge to the sustainability of boreal peatland (BP) ecosystems and can alter the stability of boreal carbon storage. However, predicting the occurrence of rare and extreme BP fires proves to be challenging, and gaining a quantitative understanding of the factors, both natural and anthropogenic, inducing BP fires remains elusive. Here, we quantified the predictability of BP fires and their primary controlling factors from 1997 to 2015 using a two-step correcting machine learning (ML) framework that combines multiple ML classifiers, regression models, and an error-correcting technique. We found that (1) the adopted oversampling algorithm effectively addressed the unbalanced data and improved the recall rate by 26.88 %–48.62 % when using multiple datasets, and the error-correcting technique tackled the overestimation of fire sizes during fire seasons; (2) nonparametric models outperformed parametric models in predicting fire occurrences, and the random forest machine learning model performed the best, with the area under the receiver operating characteristic curve ranging from 0.83 to 0.93 across multiple fire datasets; and (3) four sets of factor-control simulations consistently indicated the dominant role of temperature, air dryness, and climate extreme (i.e., frost) for boreal peatland fires, overriding the effects of precipitation, wind speed, and human activities. Our findings demonstrate the efficiency and accuracy of ML techniques in predicting rare and extreme fire events and disentangle the primary factors determining BP fires, which are critical for predicting future fire risks under climate change.

54 ENVIRONMENTAL SCIENCES↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The use of space and high altitude aerial photography to classify forest land and to detect forest disturbances

In October 1969, an investigation was begun near Atlanta, Georgia, to explore the possibilities of developing predictors for forest land and stand condition classifications using space photography. It has been found that forest area can be predicted with reasonable accuracy on space photographs using ocular techniques. Infrared color film is the best single multiband sensor for this purpose. Using the Apollo 9 infrared color photographs taken in March 1969 photointerpreters were able to predict forest area for small units consistently within 5 to 10 percent of ground truth. Approximately 5,000 density data points were recorded for 14 scan lines selected at random from five study blocks. The mean densities and standard deviations were computed for 13 separate land use classes. The results indicate that forest area cannot be separated from other land uses with a high degree of accuracy using optical film density alone. If, however, densities derived by introducing red, green, and blue cutoff filters in the optical system of the microdensitometer are combined with their differences and their ratios in regression analysis techniques, there is a good possibility of discriminating forest from all other classes.

Aldrich, R. C.↗

Rapid estimation of photosynthetic leaf traits of tropical plants in diverse environmental conditions using reflectance spectroscopy

Tropical forests are one of the main carbon sinks on Earth, but the magnitude of CO 2 absorbed by tropical vegetation remains uncertain. Terrestrial biosphere models (TBMs) are commonly used to estimate the CO 2 absorbed by forests, but their performance is highly sensitive to the parameterization of processes that control leaf-level CO 2 exchange. Direct measurements of leaf respiratory and photosynthetic traits that determine vegetation CO 2 fluxes are critical, but traditional approaches are time-consuming. Reflectance spectroscopy can be a viable alternative for the estimation of these traits and, because data collection is markedly quicker than traditional gas exchange, the approach can enable the rapid assembly of large datasets. However, the application of spectroscopy to estimate photosynthetic traits across a wide range of tropical species, leaf ages and light environments has not been extensively studied. Here, we used leaf reflectance spectroscopy together with partial least-squares regression (PLSR) modeling to estimate leaf respiration ( R dark25 ), the maximum rate of carboxylation by the enzyme Rubisco ( V cmax25 ), the maximum rate of electron transport ( J max25 ), and the triose phosphate utilization rate ( T p25 ), all normalized to 25°C. We collected data from three tropical forest sites and included leaves from fifty-three species sampled at different leaf phenological stages and different leaf light environments. Our resulting spectra-trait models validated on randomly sampled data showed good predictive performance for V cmax25 , J max25 , T p25 and R dark25 (RMSE of 13, 20, 1.5 and 0.3 μmol m -2 s -1 , and R 2 of 0.74, 0.73, 0.64 and 0.58, respectively). The models showed similar performance when applied to leaves of species not included in the training dataset, illustrating that the approach is robust for capturing the main axes of trait variation in tropical species. We discuss the utility of the spectra-trait and traditional gas exchange approaches for enhancing tropical plant trait studies and improving the parameterization of TBMs.

54 ENVIRONMENTAL SCIENCES↗

Fusion RF Modeling Machine Learning (FusionML_RF) v1.0

FusionML_RF consists of multiple codes and trained machine learning (ML) models that perform low-cost output modeling from the Genray-CQL3D. Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. For example, completing a single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time. On the other hand, these ML models achieve ~ms of inference time with high accuracy across the input parameter space. This software collection consists of multiple components. (1) codes that use ML methods and precomputed Genray-CQL3D simulation output to build regression models that enable approximate computations of Genray-CLQ3D outputs from arbitrary but physically meaningful input parameters (surrogate modeling); (2) three trained models created by the team, using a database of 16,000+ GENRAY/CQL3D simulations, to study the performance of ML models for surrogate modeling; (3) codes that load the trained models and simulation data, and then compute mean squared error between the models' predictions and the ground truth of simulation output data. This collection is being made available in conjunction with a scientific publication about the work to promote reusability and provide an artifact of the scientific work.

Bai, Zhe↗

Collective Risk Ranking of Highway Segments on the Basis of Severity-Weighted Crash Rates

This study is intended to focus on the major factors affecting traffic crash rates and severity levels, in addition to identifying crash-prone locations (i.e., black spots) based on the two indicators. The available crash data for different road segments used for the analysis were obtained from the Washington state database provided by the Highway Safety Information System (HSIS) for the years 2006 to 2011. A Random Forest (RF) classifier was used to predict the outcome level of crash severity, while crash rates were predicted by applying RF regressor. Certain features were selected for each model besides the abstraction of new features to check if there are unobserved correlations affecting the independent variables, such as accounting for the number and weight of crashes within 1 km2 area by implementing the Getis-Ord Gi∗ index. Moreover, to calculate the collective risk (CR) score, crash rates were adjusted to incorporate crash severity weights (cost per severity type) and regression-to-the-mean (RTM) bias via Empirical Bayes (EB) method. Finally, segments were ranked according to their CR score.

Li, Dawei↗

Predicting Intensive Care Unit Length of Stay and Mortality Using Patient Vital Signs: Machine Learning Model Development and Validation

Background: Patient monitoring is vital in all stages of care. In particular, intensive care unit (ICU) patient monitoring has the potential to reduce complications and morbidity, and to increase the quality of care by enabling hospitals to deliver higher-quality, cost-effective patient care, and improve the quality of medical services in the ICU. Objective: We here report the development and validation of ICU length of stay and mortality prediction models. The models will be used in an intelligent ICU patient monitoring module of an Intelligent Remote Patient Monitoring (IRPM) framework that monitors the health status of patients, and generates timely alerts, maneuver guidance, or reports when adverse medical conditions are predicted. Methods: We utilized the publicly available Medical Information Mart for Intensive Care (MIMIC) database to extract ICU stay data for adult patients to build two prediction models: one for mortality prediction and another for ICU length of stay. For the mortality model, we applied six commonly used machine learning (ML) binary classification algorithms for predicting the discharge status (survived or not). For the length of stay model, we applied the same six ML algorithms for binary classification using the median patient population ICU stay of 2.64 days. For the regression-based classification, we used two ML algorithms for predicting the number of days. We built two variations of each prediction model: one using 12 baseline demographic and vital sign features, and the other based on our proposed quantiles approach, in which we use 21 extra features engineered from the baseline vital sign features, including their modified means, standard deviations, and quantile percentages. Results: We could perform predictive modeling with minimal features while maintaining reasonable performance using the quantiles approach. The best accuracy achieved in the mortality model was approximately 89% using the random forest algorithm. The highest accuracy achieved in the length of stay model, based on the population median ICU stay (2.64 days), was approximately 65% using the random forest algorithm. Conclusions: The novelty in our approach is that we built models to predict ICU length of stay and mortality with reasonable accuracy based on a combination of ML and the quantiles approach that utilizes only vital signs available from the patient’s profile without the need to use any external features. This approach is based on feature engineering of the vital signs by including their modified means, standard deviations, and quantile percentages of the original features, which provided a richer dataset to achieve better predictive power in our models.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning using host/guest energy histograms to predict adsorption in metal–organic frameworks: Application to short alkanes and Xe/Kr mixtures

A machine learning (ML) methodology that uses a histogram of interaction energies has been applied to predict gas adsorption in metal–organic frameworks (MOFs) using results from atomistic grand canonical Monte Carlo (GCMC) simulations as training and test data. In this work, the method is first extended to binary mixtures of spherical species, in particular, Xe and Kr. In addition, it is shown that single-component adsorption of ethane and propane can be predicted in good agreement with GCMC simulation using a histogram of the adsorption energies felt by a methyl probe in conjunction with the random forest ML method. Here, the results for propane can be improved by including a small number of MOF textural properties as descriptors. We also discuss the most significant features, which provides physical insight into the most beneficial adsorption energy sites for a given application.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗