Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Airborne hyperspectral imaging of nitrogen deficiency on crop traits and yield of maize by machine learning and radiative transfer modeling

Nitrogen is an essential nutrient that directly affects plant photosynthesis, crop yield, and biomass production for bioenergy crops, but excessive application of nitrogen fertilizers can cause environmental degradation. To achieve sustainable nitrogen fertilizer management for precision agriculture, there is an urgent need for nondestructive and high spatial resolution monitoring of crop nitrogen and its allocation to photosynthetic proteins as that changes over time. Here, we used visible to shortwave infrared (400–2400 nm) airborne hyperspectral imaging with high spatial (0.5 m) and spectral (3–5 nm) resolutions to accurately estimate critical crop traits, i.e., nitrogen, chlorophyll, and photosynthetic capacity (CO 2 -saturated photosynthesis rate, V max,27 ), at leaf and canopy scales, and to assess nitrogen deficiency on crop yield. We conducted three airborne campaigns over a maize (Zea mays L.) field during the growing season of 2019. Physically based soil-canopy Radiative Transfer Modeling (RTM) and data-driven approaches i.e. Partial-Least Squares Regression (PLSR) were used to retrieve crop traits from hyperspectral reflectance, with ground truth of leaf nitrogen, chlorophyll, V max,27 , Leaf Area Index (LAI), and harvested grain yield. To improve computational efficiency of RTMs, Random Forest (RF) was used to mimic RTM simulations to generate machine learning surrogate models RTM-RF. The results show that prior knowledge of soil background and leaf angle distribution can significantly reduce the ill-posed RTM retrieval. RTM-RF achieved a high accuracy to predict leaf chlorophyll content (R 2 = 0.73) and LAI (R 2 = 0.75). Meanwhile, PLSR exhibited better accuracy to predict leaf chlorophyll content (R 2 = 0.79), nitrogen concentration (R 2 = 0.83), nitrogen content (R 2 = 0.77), and V max,27 (R 2 = 0.69) but required measured traits for model training. We also found that canopy structure signals can enhance the use of spectral data to predict nitrogen related photosynthetic traits, as combining RTM-RF LAI and PLSR leaf traits well predicted canopy-level traits (leaf traits × LAI) including canopy chlorophyll (R 2 = 0.80), nitrogen (R 2 = 0.85) and V max,27 (R 2 = 0.82). Compared to leaf traits, we further found that canopy-level photosynthetic traits, particularly canopy V max,27 , have higher correlation with maize grain yield. This study highlights the potential for synergistic use of process-based and data-driven approaches of hyperspectral imaging to quantify crop traits that facilitate precision agricultural management to secure food and bioenergy production.

54 ENVIRONMENTAL SCIENCES↗

Improving Solar and Solar+Storage Screening Techniques to Reduce Utility Interconnection Time and Costs (Final Technical Report)

Residential PV installations have increased rapidly over the last decade, and the increased application volume has caused permitting delays and lower overall adoption rates. In this project, we developed and evaluated whether data-driven secondary modeling and screening techniques can help utilities assess customer applications more accurately than traditional screening shortcuts. Secondary topologies are predicted using decision trees and commonly available information, such as service transformer, customer, and street locations. Conductors were predicted using a logistic regression method based on real world object (RWO) types, service transformer ratings, conductor length, and distance to transformer. After developing the combined primary and secondary distribution network model, hosting capacity results were used to train a random forest model to predict the pass/fail likelihood of a customer application. Powerflow based models with predicted secondaries and data-driven methods both increased the screening success rate, relative to common utility heuristics, by as much as 55 percentage points. Data-driven screening techniques were described by one utility as a "right-sized" approach for residential customers given the low-risk of small errors and the high-cost of accurate modeling.

14 SOLAR ENERGY↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Photometric redshift estimation of BASS DR3 quasars by machine learning

ABSTRACT Correlating Beijing–Arizona Sky Survey (BASS) data release 3 (DR3) catalogue with the ALLWISE data base, the data from optical and infrared information are obtained. The quasars from Sloan Digital Sky Survey are taken as training and test samples while those from LAMOST are considered as external test sample. We propose two schemes to construct the redshift estimation models with XGBoost, CatBoost, and Random Forest. One scheme (namely one-step model) is to predict photometric redshifts directly based on the optimal models created by these three algorithms; the other scheme (namely two-step model) is to first classify the data into low- and high-redshift data sets, and then predict photometric redshifts of these two data sets separately. For one-step model, the performance of these three algorithms on photometric redshift estimation is compared with different training samples, and CatBoost is superior to XGBoost and Random Forest. For two-step model, the performances of these three algorithms on the classification of low and high redshift subsamples are compared, and CatBoost still shows the best performance. Therefore, CatBoost is regarded as the core algorithm of classification and regression in two-step model. In contrast to one-step model, two-step model is optimal when predicting photometric redshift of quasars, especially for high-redshift quasars. Finally, the two models are applied to predict photometric redshifts of all quasar candidates of BASS DR3. The number of high-redshift quasar candidates is 3938 (redshift ≥3.5) and 121 (redshift ≥4.5) by two-step model. The predicted result will be helpful for quasar research and follow-up observation of high-redshift quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Quantifying wildfire drivers and predictability in boreal peatlands using a two-step error-correcting machine learning framework in TeFire v1.0

Abstract. Wildfires are becoming an increasing challenge to the sustainability of boreal peatland (BP) ecosystems and can alter the stability of boreal carbon storage. However, predicting the occurrence of rare and extreme BP fires proves to be challenging, and gaining a quantitative understanding of the factors, both natural and anthropogenic, inducing BP fires remains elusive. Here, we quantified the predictability of BP fires and their primary controlling factors from 1997 to 2015 using a two-step correcting machine learning (ML) framework that combines multiple ML classifiers, regression models, and an error-correcting technique. We found that (1) the adopted oversampling algorithm effectively addressed the unbalanced data and improved the recall rate by 26.88 %–48.62 % when using multiple datasets, and the error-correcting technique tackled the overestimation of fire sizes during fire seasons; (2) nonparametric models outperformed parametric models in predicting fire occurrences, and the random forest machine learning model performed the best, with the area under the receiver operating characteristic curve ranging from 0.83 to 0.93 across multiple fire datasets; and (3) four sets of factor-control simulations consistently indicated the dominant role of temperature, air dryness, and climate extreme (i.e., frost) for boreal peatland fires, overriding the effects of precipitation, wind speed, and human activities. Our findings demonstrate the efficiency and accuracy of ML techniques in predicting rare and extreme fire events and disentangle the primary factors determining BP fires, which are critical for predicting future fire risks under climate change.

54 ENVIRONMENTAL SCIENCES↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Rapid estimation of photosynthetic leaf traits of tropical plants in diverse environmental conditions using reflectance spectroscopy

Tropical forests are one of the main carbon sinks on Earth, but the magnitude of CO 2 absorbed by tropical vegetation remains uncertain. Terrestrial biosphere models (TBMs) are commonly used to estimate the CO 2 absorbed by forests, but their performance is highly sensitive to the parameterization of processes that control leaf-level CO 2 exchange. Direct measurements of leaf respiratory and photosynthetic traits that determine vegetation CO 2 fluxes are critical, but traditional approaches are time-consuming. Reflectance spectroscopy can be a viable alternative for the estimation of these traits and, because data collection is markedly quicker than traditional gas exchange, the approach can enable the rapid assembly of large datasets. However, the application of spectroscopy to estimate photosynthetic traits across a wide range of tropical species, leaf ages and light environments has not been extensively studied. Here, we used leaf reflectance spectroscopy together with partial least-squares regression (PLSR) modeling to estimate leaf respiration ( R dark25 ), the maximum rate of carboxylation by the enzyme Rubisco ( V cmax25 ), the maximum rate of electron transport ( J max25 ), and the triose phosphate utilization rate ( T p25 ), all normalized to 25°C. We collected data from three tropical forest sites and included leaves from fifty-three species sampled at different leaf phenological stages and different leaf light environments. Our resulting spectra-trait models validated on randomly sampled data showed good predictive performance for V cmax25 , J max25 , T p25 and R dark25 (RMSE of 13, 20, 1.5 and 0.3 μmol m -2 s -1 , and R 2 of 0.74, 0.73, 0.64 and 0.58, respectively). The models showed similar performance when applied to leaves of species not included in the training dataset, illustrating that the approach is robust for capturing the main axes of trait variation in tropical species. We discuss the utility of the spectra-trait and traditional gas exchange approaches for enhancing tropical plant trait studies and improving the parameterization of TBMs.

54 ENVIRONMENTAL SCIENCES↗

Fusion RF Modeling Machine Learning (FusionML_RF) v1.0

FusionML_RF consists of multiple codes and trained machine learning (ML) models that perform low-cost output modeling from the Genray-CQL3D. Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. For example, completing a single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time. On the other hand, these ML models achieve ~ms of inference time with high accuracy across the input parameter space. This software collection consists of multiple components. (1) codes that use ML methods and precomputed Genray-CQL3D simulation output to build regression models that enable approximate computations of Genray-CLQ3D outputs from arbitrary but physically meaningful input parameters (surrogate modeling); (2) three trained models created by the team, using a database of 16,000+ GENRAY/CQL3D simulations, to study the performance of ML models for surrogate modeling; (3) codes that load the trained models and simulation data, and then compute mean squared error between the models' predictions and the ground truth of simulation output data. This collection is being made available in conjunction with a scientific publication about the work to promote reusability and provide an artifact of the scientific work.

Bai, Zhe↗

Collective Risk Ranking of Highway Segments on the Basis of Severity-Weighted Crash Rates

This study is intended to focus on the major factors affecting traffic crash rates and severity levels, in addition to identifying crash-prone locations (i.e., black spots) based on the two indicators. The available crash data for different road segments used for the analysis were obtained from the Washington state database provided by the Highway Safety Information System (HSIS) for the years 2006 to 2011. A Random Forest (RF) classifier was used to predict the outcome level of crash severity, while crash rates were predicted by applying RF regressor. Certain features were selected for each model besides the abstraction of new features to check if there are unobserved correlations affecting the independent variables, such as accounting for the number and weight of crashes within 1 km2 area by implementing the Getis-Ord Gi∗ index. Moreover, to calculate the collective risk (CR) score, crash rates were adjusted to incorporate crash severity weights (cost per severity type) and regression-to-the-mean (RTM) bias via Empirical Bayes (EB) method. Finally, segments were ranked according to their CR score.

Li, Dawei↗

Predicting Intensive Care Unit Length of Stay and Mortality Using Patient Vital Signs: Machine Learning Model Development and Validation

Background: Patient monitoring is vital in all stages of care. In particular, intensive care unit (ICU) patient monitoring has the potential to reduce complications and morbidity, and to increase the quality of care by enabling hospitals to deliver higher-quality, cost-effective patient care, and improve the quality of medical services in the ICU. Objective: We here report the development and validation of ICU length of stay and mortality prediction models. The models will be used in an intelligent ICU patient monitoring module of an Intelligent Remote Patient Monitoring (IRPM) framework that monitors the health status of patients, and generates timely alerts, maneuver guidance, or reports when adverse medical conditions are predicted. Methods: We utilized the publicly available Medical Information Mart for Intensive Care (MIMIC) database to extract ICU stay data for adult patients to build two prediction models: one for mortality prediction and another for ICU length of stay. For the mortality model, we applied six commonly used machine learning (ML) binary classification algorithms for predicting the discharge status (survived or not). For the length of stay model, we applied the same six ML algorithms for binary classification using the median patient population ICU stay of 2.64 days. For the regression-based classification, we used two ML algorithms for predicting the number of days. We built two variations of each prediction model: one using 12 baseline demographic and vital sign features, and the other based on our proposed quantiles approach, in which we use 21 extra features engineered from the baseline vital sign features, including their modified means, standard deviations, and quantile percentages. Results: We could perform predictive modeling with minimal features while maintaining reasonable performance using the quantiles approach. The best accuracy achieved in the mortality model was approximately 89% using the random forest algorithm. The highest accuracy achieved in the length of stay model, based on the population median ICU stay (2.64 days), was approximately 65% using the random forest algorithm. Conclusions: The novelty in our approach is that we built models to predict ICU length of stay and mortality with reasonable accuracy based on a combination of ML and the quantiles approach that utilizes only vital signs available from the patient’s profile without the need to use any external features. This approach is based on feature engineering of the vital signs by including their modified means, standard deviations, and quantile percentages of the original features, which provided a richer dataset to achieve better predictive power in our models.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning using host/guest energy histograms to predict adsorption in metal–organic frameworks: Application to short alkanes and Xe/Kr mixtures

A machine learning (ML) methodology that uses a histogram of interaction energies has been applied to predict gas adsorption in metal–organic frameworks (MOFs) using results from atomistic grand canonical Monte Carlo (GCMC) simulations as training and test data. In this work, the method is first extended to binary mixtures of spherical species, in particular, Xe and Kr. In addition, it is shown that single-component adsorption of ethane and propane can be predicted in good agreement with GCMC simulation using a histogram of the adsorption energies felt by a methyl probe in conjunction with the random forest ML method. Here, the results for propane can be improved by including a small number of MOF textural properties as descriptors. We also discuss the most significant features, which provides physical insight into the most beneficial adsorption energy sites for a given application.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗