Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A novel random forest approach to revealing interactions and controls on chlorophyll concentration and bacterial communities during coastal phytoplankton blooms

Abstract Increasing occurrence of harmful algal blooms across the land–water interface poses significant risks to coastal ecosystem structure and human health. Defining significant drivers and their interactive impacts on blooms allows for more effective analysis and identification of specific conditions supporting phytoplankton growth. A novel iterative Random Forests (iRF) machine-learning model was developed and applied to two example cases along the California coast to identify key stable interactions: (1) phytoplankton abundance in response to various drivers due to coastal conditions and land-sea nutrient fluxes, (2) microbial community structure during algal blooms. In Example 1, watershed derived nutrients were identified as the least significant interacting variable associated with Monterey Bay phytoplankton abundance. In Example 2, through iRF analysis of field-based 16S OTU bacterial community and algae datasets, we independently found stable interactions of prokaryote abundance patterns associated with phytoplankton abundance that have been previously identified in laboratory-based studies. Our study represents the first iRF application to marine algal blooms that helps to identify ocean, microbial, and terrestrial conditions that are considered dominant causal factors on bloom dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Forecasting induced seismicity in Oklahoma using machine learning methods

Oklahoma earthquakes in the past decade have been mostly associated with wastewater injection. Here we use a machine learning technique—the Random Forest to forecast induced seismicity rate in Oklahoma based on injection-related parameters. We split the data into training (2011.01–2015.05) and test (2015.06–2020.12) periods. The model forecasts seismicity rate during the test period based on input features, including operational parameters (injection rate and pressure), geological information (depth to basement), and modeled pore pressure and poroelastic stress. The results show overall good match with observed seismicity rate (adjusted R 2 of 0.75). The model shows that pore pressure rate and poroelastic stressing rates are the two most important features in forecasting. The absolute values of pore pressure and poroelastic stress, and the injection rate itself, are less important than the stressing rates. These findings further emphasize that temporal changes of stressing rates would lead to significant changes in seismicity rates.

58 GEOSCIENCES↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Automating the interpretation of PM 2.5 time–resolved measurements using a data–driven approach

The rapid development of automated measurement equipment enables researchers to collect greater quantities of time-resolved data from indoor and outdoor environments. While significant, the interpretation of the resulting data can be a time-consuming effort. This paper introduces an automated process of interpreting PM 2.5 time-resolved data and differentiating PM 2.5 emissions resulting from indoor and outdoor sources. Here, we use Random Forest (RF), a machine learning approach, to study a dataset of 836 indoor emission events that occurred over a 2-week period in 18 apartments in California. In this paper, we show model development and evaluate its performance as the sample size and source vary. We discuss the characteristics of the dataset that tended to help the source identification and why. For example, we show that data from many events and from different apartments are essential for the model to be suitable for analyzing a new separate dataset. We also show that longitudinal data appear to be more helpful than the time frequency of measurements within a given apartment. We use the resulting RF model to analyze PM 2.5 data of an entirely separate dataset collected from 65 new homes in California. The RF model identifies 442 indoor emission events, with only a few misidentifications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fair Bagging Boosting Models [SWR-24-38]

Fair Bagging Boosting Models is a software implementation of a framework for building, measuring bias and correcting bias in 3 popular forest machine learning models: gradient boosted trees (GBT), random forest (RF), and XGBoost models, using the XGBoost library. The framework takes advantage of the flexibility in XGBoost library to represent gradient boosted tree and random forest models, as well as the ability to use custom loss function.

Ugirumurera, Juliette↗

Calibration and Rapid-Adoption Forecasting Techniques

CRAFT (Calibration and Rapid-Adoption Forecasting Techniques) CRAFT is a Python-based project for processing, analyzing, and modeling atmospheric or environmental data. It uses machine learning techniques, specifically Random Forest Regression, to create emulators for various environmental variables such as gross primary production and soil water content. It then uses these emulators to robustly test the parameter space of mechanistic models to provide posterior estimations of the free parameters.

Robins, Zachary↗

Using machine learning to derive cloud condensation nuclei number concentrations from commonly available measurements

Cloud condensation nuclei (CCN) number concentrations are an important aspect of aerosol–cloud interactions and the subsequent climate effects; however, their measurements are very limited. We use a machine learning tool, random decision forests, to develop a random forest regression model (RFRM) to derive CCN at 0.4 % supersaturation ([CCN0.4]) from commonly available measurements. The RFRM is trained on the long-term simulations in a global size-resolved particle microphysics model. Using atmospheric state and composition variables as predictors, through associations of their variabilities, the RFRM is able to learn the underlying dependence of [CCN0.4] on these predictors, which are as follows: eight fractions of PM 2.5 (NH 4 , SO 4 , NO 3, secondary organic aerosol (SOA), black carbon (BC), primary organic carbon (POC), dust, and salt), seven gaseous species (NO x , NH 3 , O 3 , SO 2 , OH, isoprene, and monoterpene), and four meteorological variables (temperature (T), relative humidity (RH), precipitation, and solar radiation). The RFRM is highly robust: it has a median mean fractional bias (MFB) of 4.4 % with ≈96.33 % of the derived [CCN0.4] within a good agreement range of -60% 2.5 speciation (NH 4 , SO 4 , NO 3 , and organic carbon (OC)), NO x , O 3 , SO 2 , T, and RH, as well as [CCN0.4] are available. We modify, optimize, and retrain the developed RFRM to make predictions from 19 to 9 of these available predictors. This retrained RFRM (RFRM-ShortVars) shows a reduction in performance due to the unavailability and sparsity of measurements (predictors); it captures the [CCN0.4] variability and magnitude at SGP with ≈67.02 % of the derived values in the good agreement range. This work shows the potential of using the more commonly available measurements of PM 2.5 speciation to alleviate the sparsity of CCN number concentrations' measurements.

54 ENVIRONMENTAL SCIENCES↗

Observational benchmarks inform representation of soil organic carbon dynamics in land surface models

Abstract. Representing soil organic carbon (SOC) dynamics in Earth system models (ESMs) is a key source of uncertainty in predicting carbon–climate feedbacks. Machine learning models can help identify dominant environmental controllers and establish their functional relationships with SOC stocks. The resulting knowledge can be integrated into ESMs to reduce uncertainty and improve predictions of SOC dynamics over space and time. In this study, we used a large number of SOC field observations (n=54 000), geospatial datasets of environmental factors (n=46), and two machine learning approaches (namely random forest, RF, and generalized additive modeling, GAM) to (1) identify dominant environmental controllers of global and biome-specific SOC stocks, (2) derive functional relationships between environmental controllers and SOC stocks, and (3) compare the identified environmental controllers and predictive relationships with those in models used in Phase 6 of the Coupled Model Intercomparison Project (CMIP6). Our results showed that the diurnal temperature, drought index, cation exchange capacity, and precipitation were important observed environmental predictors of global SOC stocks. While the RF model identified 14 environmental factors that describe climatic, vegetation, and edaphic conditions as important predictors of global SOC stocks (R2=0.61, RMSE = 0.46 kg m−2), current ESMs oversimplify the relationships between environmental factors and SOC, with precipitation, temperature, and net primary productivity explaining > 96 % of the variability in ESM-modeled SOC stocks. Further, our study revealed notable disparities among the functional relationships between environmental factors and SOC stocks simulated by ESMs compared with observed relationships. To improve SOC representations in ESMs, it is imperative to incorporate additional environmental controls, such as the cation exchange capacity, and refine the functional relationships to align more closely with observations.

54 ENVIRONMENTAL SCIENCES↗

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Power System Feature-Based Event Classification by Means of Multiple PMU Data

Abstract—Phasor Measurement Units (PMUs) provide time synchronized measurements across the power grid, enabling data driven event detection and classification for enhanced system monitoring and situational awareness. However, variations in event duration, spatial extent, and severity, along with coincident events, pose challenges for conventional classification models that require fixed-size inputs. This paper presents a feature-based framework that aggregates diverse attributes from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, and Multilayer Perceptron. A probabilistic post-processing scheme is further introduced to enable multi-label classification in the presence of overlapping events. Experiments using real-world PMU data demonstrate that the Random Forest model achieves 95% accuracy, while the proposed post-processing method yields an additional 3% improvement.

Nematirad, Reza↗

Network-Scale Ubiquitous Volume Estimation Using Tree-Based Ensemble Learning Methods

Currently ubiquitous volume data for roadway networks remains the key missing dimension in traffic operations. Most volume data are average annual daily traffic (AADT) measures derived from the Highway Performance Monitoring System (HPMS). Although methods to factor the AADT to hourly averages for typical day of week exist, actual volume data is limited to a sparse collection of locations in which volumes are continuously recorded. This paper/poster explores the use of state-of-art machine learning techniques to estimate accurate volume measures that span the highway network providing ubiquitous coverage in space, and point-in-time measures for a specific date and time. Three tree-based ensemble learning models, random forest (RF), gradient boost machine (GBM), and extreme gradient boost (XGBoost), were tested for volume estimation by learning from combined dataset of commercial probe data provided by TomTom, the FHWA's Travel Monitoring Analysis System (TMAS) data, and other infrastructure attributes such as number of lanes, speed limit, and weather. The methods were tested on major corridors and freeways in the metropolitan area of Denver. All three machine learning methods were able to provide hourly volume estimates 24 hours a day, 7 days a week, and 365 days a year with around 18% mean absolute error to true volume and about 5% of error with respect to roadway capacity. The low error measures allow the potential application by transportation agencies.

33 ADVANCED PROPULSION SYSTEMS↗

iRF v2.0

A predictive, stable, and interpretable machine learning tool: the iterative random forest algorithm (iRF). iRF discovers high-order interactions among variables with the same order of computational cost as random forests (RF). We have demonstrated the utility of iRF in several applications in the biological and environmental sciences. It is a general purpose machine learning framework for building "explainable" predictive engines.

Brown, JamesB.↗

Modeling Spatial Distribution of Snow Water Equivalent by Combining Meteorological and Satellite Data with Lidar Maps

Abstract An accurate characterization of the water content of snowpack, or snow water equivalent (SWE), is necessary to quantify water availability and constrain hydrologic and land surface models. Recently, airborne observations (e.g., lidar) have emerged as a promising method to accurately quantify SWE at high resolutions (scales of ∼100 m and finer). However, the frequency of these observations is very low, typically once or twice per season in the Rocky Mountains of Colorado. Here, we present a machine learning framework that is based on random forests to model temporally sparse lidar-derived SWE, enabling estimation of SWE at unmapped time points. We approximated the physical processes governing snow accumulation and melt as well as snow characteristics by obtaining 15 different variables from gridded estimates of precipitation, temperature, surface reflectance, elevation, and canopy. Results showed that, in the Rocky Mountains of Colorado, our framework is capable of modeling SWE with a higher accuracy when compared with estimates generated by the Snow Data Assimilation System (SNODAS). The mean value of the coefficient of determination R 2 using our approach was 0.57, and the root-mean-square error (RMSE) was 13 cm, which was a significant improvement over SNODAS (mean R 2 = 0.13; RMSE = 20 cm). We explored the relative importance of the input variables and observed that, at the spatial resolution of 800 m, meteorological variables are more important drivers of predictive accuracy than surface variables that characterize the properties of snow on the ground. This research provides a framework to expand the applicability of lidar-derived SWE to unmapped time points. Significance Statement Snowpack is the main source of freshwater for close to 2 billion people globally and needs to be estimated accurately. Mountainous snowpack is highly variable and is challenging to quantify. Recently, lidar technology has been employed to observe snow in great detail, but it is costly and can only be used sparingly. To counter that, we use machine learning to estimate snowpack when lidar data are not available. We approximate the processes that govern snowpack by incorporating meteorological and satellite data. We found that variables associated with precipitation and temperature have more predictive power than variables that characterize snowpack properties. Our work helps to improve snowpack estimation, which is critical for sustainable management of water resources.

54 ENVIRONMENTAL SCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Probabilistic Forecasting of Generators Startups and Shutdowns in the MISO System Based on Random Forest

Solving security constrained unit commitment (SCUC) problems to plan an economical generation schedule for day-head electricity market has been an important research topic in recent years. Mixed integer programming method (MIP), the-state-of-art approach for solving SCUC problem, is known computationally hard when the number of binary status variables is large. In this paper, a machine learning-based algorithm - random forest (RF), was applied to forecast the startups (SU) and shutdowns (SD) hours of generators, based on historical hourly system condition observations in the Midcontinent Independent System Operator (MISO) system. The main purpose is to reduce the number of binary status variables, by fixing the SU/SD hours to a narrow range of high confidence. This would significantly reduce the size of the decision space, and therefore speed up SCUC solutions with reduced uncertainty.

Lin, Xinming↗

On the estimation of boundary layer heights: a machine learning approach

Abstract. The planetary boundary layer height (zi) is a key parameter used in atmospheric models for estimating the exchange of heat, momentum, and moisture between the surface and the free troposphere. Near-surface atmospheric and subsurface properties (such as soil temperature, relative humidity, etc.) are known to have an impact on zi. Nevertheless, precise relationships between these surface properties and zi are less well known and not easily discernible from the multi-year dataset. Machine learning approaches, such as random forest (RF), which use a multi-regression framework, help to decipher some of the physical processes linking surface-based characteristics to zi. In this study, a 4-year dataset from 2016 to 2019 at the Southern Great Plains site is used to develop and test a machine learning framework for estimating zi. Parameters derived from Doppler lidars are used in combination with over 20 different surface meteorological measurements as inputs to a RF model. The model is trained using radiosonde-derived zi values spanning the period from 2016 through 2018 and then evaluated using data from 2019. Results from 2019 showed significantly better agreement with the radiosonde compared to estimates derived from a thresholding technique using Doppler lidars only. Noteworthy improvements in daytime zi estimates were observed using the RF model, with a 50 % improvement in mean absolute error and an R2 of greater than 85 % compared to the Tucker method zi. We also explore the effect of zi uncertainty on convective velocity scaling and present preliminary comparisons between the RF model and zi estimates derived from atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

Towards fast and accurate predictions of radio frequency power deposition and current profile via data-driven modelling: applications to lower hybrid current drive

Three machine learning techniques (multilayer perceptron, random forest and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modelling and real-time control applications. The machine learning models use a database of more than 16 000 GENRAY/CQL3D simulations for training, validation and testing. Latin hypercube sampling methods ensure that the database covers the range of nine input parameters ( $n_{e0}$ , $T_{e0}$ , $I_p$ , $B_t$ , $R_0$ , $n_{\|}$ , $Z_{{\rm eff}}$ , $V_{{\rm loop}}$ and $P_{{\rm LHCD}}$ ) with sufficient density in all regions of parameter space. The surrogate models reduce the inference time from minutes to $\sim$ ms with high accuracy across the input parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards Fast and Accurate Predictions of Radio Frequency Power Deposition and Current Profile via Data-driven Modeling

Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modeling and real-time control applications. The machine learning models use a database of 16,000+ GENRAY/CQL3D simulations for training, validation, and testing. Latin hypercube sampling methods ensure that the database covers the range of 9 input parameters ($n_{e0}$, $T_{e0}$, $I_p$, $B_t$, $R_0$, $n_{||}$, $Z_{eff}$, $V_{loop}$, $P_{LHCD}$) with sufficient density in all regions of parameter space. The surrogate models reduce the inference time from minutes to ~ms with high accuracy across the input parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗