Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Climatology of Linear Mesoscale Convective System Morphology in the United States based on Random Forests Method

This study uses machine learning methods, specifically the random forest (RF), on a radar-based mesoscale convective system (MCS) tracking dataset to classify the five types of linear MCS morphology in the contiguous United States during the period 2004-2016. The algorithm is trained using radar- and satellite-derived spatial and morphological parameters, and reanalysis environmental information from 5-26yr manually identified nonlinear and five linear MCS modes. The algorithm is then used to automate the classification of linear MCSs over 8 years with high accuracy, providing a systematic, long-term climatology of linear MCSs. Results reveal that nearly 40% of MCSs are classified as linear MCSs, in which half of the linear events belong to the type of system having a leading convective line. The occurrence of linear MCSs shows large annual and seasonal variations. On average, 113 linear MCSs occur annually during the warm season (through March to October), with most of these events clustered from May through August in the central eastern Great Plains. MCS characteristics, including duration, propagation speed, orientation, and system cloud size, have large variability among the different linear modes. The systems having a trailing convective line and the systems having a back-building area of convection typically move more slowly and have higher precipitation rate, and thus have higher potential in producing extreme rainfall and flash flooding. Analysis of the environmental conditions associated with linear MCSs show that the storm-relative flow is of most importance in determining the organization mode of linear MCSs.

Cui, Wenjun↗

Analysis of Random Forest Modeling Strategies for Multi-Step Wind Speed Forecasting

Although the random forest (RF) model is a powerful machine learning tool that has been utilized in many wind speed/power forecasting studies, there has been no consensus on optimal RF modeling strategies. This study investigates three basic questions which aim to assist in the discernment and quantification of the effects of individual model properties, namely: (1) using a standalone RF model versus using RF as a correction mechanism for the persistence approach, (2) utilizing a recursive versus direct multi-step forecasting strategy, and (3) training data availability on model forecasting accuracy from one to six hours ahead. These questions are investigated utilizing data from the FINO1 offshore platform and Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) C1 site, and testing results are compared to the persistence method. At FINO1, due to the presence of multiple wind farms and high inter-annual variability, RF is more effective as an error-correction mechanism for the persistence approach. The direct forecasting strategy is seen to slightly outperform the recursive strategy, specifically for forecasts three or more steps ahead. Finally, increased data availability (up to ~8 equivalent years of hourly training data) appears to continually improve forecasting accuracy, although changing environmental flow patterns have the potential to negate such improvement. We hope that the findings of this study will assist future researchers and industry professionals to construct accurate, reliable RF models for wind speed forecasting.

54 ENVIRONMENTAL SCIENCES↗

A Machine Learning Initializer for Newton-Raphson AC Power Flow Convergence

Power flow computations are fundamental to many power system studies. Obtaining a converged power flow case is not a trivial task especially in large power grids due to the non-linear nature of the power flow equations. One key challenge is that the widely used Newton based power flow methods are sensitive to the initial voltage magnitude and angle estimates, and a bad initial estimate would lead to non-convergence. This paper addresses this challenge by developing a random-forest (RF) machine learning model to provide better initial voltage magnitude and angle estimates towards achieving power flow convergence. This method was implemented on a real ERCOT 6102 bus system under various operating conditions. By providing better Newton-Raphson initialization, the RF model precipitated the solution of 2,106 cases out of 3,899 non-converging dispatches. These cases could not be solved from flat start or by initialization with the voltage solution of a reference case. Finally, results obtained from the RF initializer performed better when compared with DC power flow initialization, Linear regression, and Decision Trees.

random forest↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

The importance of round-robin validation when assessing machine-learning-based vertical extrapolation of wind speeds

The extrapolation of wind speeds measured at a meteorological mast to wind turbine rotor heights is a key component in a bankable wind farm energy assessment and a significant source of uncertainty. Industry-standard methods for extrapolation include the power-law and logarithmic profiles. The emergence of machine-learning applications in wind energy has led to several studies demonstrating substantial improvements in vertical extrapolation accuracy in machine-learning methods over these conventional power-law and logarithmic profile methods. In all cases, these studies assess relative model performance at a measurement site where, critically, the machine-learning algorithm requires knowledge of the rotor-height wind speeds in order to train the model. This prior knowledge provides fundamental advantages to the site-specific machine-learning model over the power-law and log profiles, which, by contrast, are not highly tuned to rotor-height measurements but rather can generalize to any site. Furthermore, there is no practical benefit in applying a machine-learning model at a site where winds at the heights relevant for wind energy production are known; rather, its performance at nearby locations (i.e., across a wind farm site) without rotor-height measurements is of most practical interest. To more fairly and practically compare machine-learning-based extrapolation to standard approaches, we implemented a round-robin extrapolation model comparison, in which a random-forest machine-learning model is trained and evaluated at different sites and then compared against the power-law and logarithmic profiles. We consider 20 months of lidar and sonic anemometer data collected at four sites between 50 and 100 km apart in the central United States. We find that the random forest outperforms the standard extrapolation approaches, especially when incorporating surface measurements as inputs to include the influence of atmospheric stability. When compared at a single site (the traditional comparison approach), the machine-learning improvement in mean absolute error was 28 % and 23 % over the power-law and logarithmic profiles, respectively. Using the round-robin approach proposed here, this improvement drops to 20 % and 14 %, respectively. These latter values better represent practical model performance, and we conclude that round-robin validation should be the standard for machine-learning-based wind speed extrapolation methods.

17 WIND ENERGY↗

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Landsat 8 monitoring of multi-depth suspended sediment concentrations in Lake Erie’s Maumee River using machine learning

Satellite remote sensing has been widely used to map suspended sediment concentration (SSC) in waterbodies. However, due to the complexity of sediment-water interactions, it has been difficult to derive linear and non-linear regression equations to reliably predict SSC, especially when trying to estimate depth of integrated sediment. Herein, this study uses Landsat 8 OLI (Operational Land Imager) sensor to map SSC within the Maumee River in Ohio, USA, at multiple depth intervals (15, 61, 91, and 182 cm). Simple linear least squares regression (LLSR), and three common machine learning models: random forest (RF), support vector regression (SVR), and model averaged neural network (MANN) were used to estimate SSC at the depth intervals. All machine learning models significantly outperformed LLSR while RF performed the best. In both RF and MANN, R2 (coefficient of determination) increases with depth with a maximum R2 of 0.89 and 0.83, respectively, at a depth of 0–182 cm. The results show that machine learning models can implement nonlinear relationships that produce better predictions than traditional linear regression methods in estimating depth integrated SSC, especially when samples are limited.

47 OTHER INSTRUMENTATION↗

Observations and Machine-Learned Models of Near-Surface Permafrost along the Koyukuk River, Alaska, USA

This dataset contains GeoTIFs (raster) and GeoPackages (vector) that map observations of near-surface permafrost and not-permafrost from a field campaign conducted near the village of Huslia, AK along the Koyukuk River and its floodplain in July 2018. These data were collected as part of a campaign to understand if and how permafrost impacts riverbank erosion. This problem cannot be assessed without knowing where permafrost exists. Permafrost was observed via frost probing (to a maximum depth of one meter), coring (to a maximum depth of two meters) and bank/bar excavations. An additional boat survey was performed wherein expert (Joel Rowland) judgment assessed the presence or absence of distinctive permafrost features (e.g., overhanging tundra mats, thermoerosional niching, ice wedges, active drainage of ice melt from soils). This dataset also contains the input features and results of two machine learning models (random forest and convolutional neural network) that extrapolate the observations to the full floodplain that may be useful for building, testing, or validating other machine-learned permafrost models. Permafrost data are provided as georasters of the same shape and geovectors (polylines/polygons) and are all projected into EPSG:32605. All data can be visualized with a GIS (QGIS, ArcGIS, etc.).

54 ENVIRONMENTAL SCIENCES↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Empirical relationships between environmental factors and soil organic carbon produce comparable prediction accuracy as the Machine Learning

Accurate representation of environmental controllers of soil organic carbon (SOC) stocks in Earth System Model (ESM) land models could reduce uncertainties in future carbon-climate feedback projections. Using empirical relationships between environmental factors and SOC stocks to evaluate land models can help modelers understand prediction biases beyond what can be achieved with the observed SOC stocks alone. In this study, we used 31 observed environmental factors, field SOC observations (n = 6,213) from the continental US, and two Machine Learning approaches [Random Forest (RF) and Generalized Additive Modeling (GAM)] to (1) select important environmental predictors of SOC stocks, (2) derive empirical relationships between environmental factors and SOC stocks, and (3) use the derived relationships to predict SOC stocks and compare the prediction accuracy of simpler model developed with the machine learning predictions. Out of the 31 environmental factors we investigated, 12 were identified as important predictors of SOC stocks by the RF approach. In contrast, the GAM approach identified six (of those 12) environmental factors as important controllers of SOC stocks: potential evapotranspiration, normalized difference vegetation index, soil drainage condition, precipitation, elevation, and net primary productivity. The GAM approach showed minimal SOC predictive importance of the remaining six environmental factors identified by the RF approach. Our derived empirical relations produced comparable prediction accuracy as the GAM and RF approach using only a subset of environmental factors. The empirical relationships we derived using the GAM approach can serve as important benchmarks to evaluate environmental control representations of SOC stocks in ESMs, which could reduce uncertainty in predicting future carbon-climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

PlasmidHostFinder: Prediction of Plasmid Hosts Using Random Forest

Plasmids play a major role facilitating the spread of antimicrobial resistance between bacteria. Understanding the host range and dissemination trajectories of plasmids is critical for surveillance and prevention of antimicrobial resistance. Identification of plasmid host ranges could be improved using automated pattern detection methods compared to homology-based methods due to the diversity and genetic plasticity of plasmids. In this study, we developed a method for predicting the host range of plasmids using machine learning—specifically, random forests. We trained the models with 8,519 plasmids from 359 different bacterial species per taxonomic level; the models achieved Matthews correlation coefficients of 0.662 and 0.867 at the species and order levels, respectively. Our results suggest that despite the diverse nature and genetic plasticity of plasmids, our random forest model can accurately distinguish between plasmid hosts. This tool is available online through the Center for Genomic Epidemiology (https://cge.cbs.dtu.dk/services/PlasmidHostFinder/).

59 BASIC BIOLOGICAL SCIENCES↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

Development of a Random-Forest Cloud-Regime Classification Model Based on Surface Radiation and Cloud Products

Various methods have been developed to characterize cloud type, otherwise referred to as cloud regime. These include manual sky observations, combining radiative and cloud vertical properties observed from satellite, surface-based remote sensing, and digital processing of sky imagers. While each method has inherent advantages and disadvantages, none of these cloud-typing methods actually includes measurements of surface shortwave or longwave radiative fluxes. Here, a method that relies upon detailed, surface-based radiation and cloud measurements and derived data products to train a random-forest machine-learning cloud classification model is introduced. Measurements from five years of data from the ARM Southern Great Plains site were compiled to train and independently evaluate the model classification performance. A cloud-type accuracy of approximately 80% using the random-forest classifier reveals that the model is well suited to predict climatological cloud properties. Furthermore, an analysis of the cloud-type misclassifications is performed. While physical cloud types may be misreported, the shortwave radiative signatures are similar between misclassified cloud types. From this, we assert that the cloud-regime model has the capacity to successfully differentiate clouds with comparable cloud–radiative interactions. Therefore, we conclude that the model can provide useful cloud-property information for fundamental cloud studies, inform renewable energy studies, and be a tool for numerical model evaluation and parameterization improvement, among many other applications.

54 ENVIRONMENTAL SCIENCES↗

Estimating Fine-Resolution Shortwave Broadband Albedo of Croplands from Harmonized Landsat and Sentinel-2 Data

Altered surface albedo due to land-cover conversions and management is a significant driver of global climate change. Albedo can be directly measured at ground stations, and remote sensing data can be used to scale-up albedo values to regional and global levels. Some previous studies have retrieved fine-resolution (10–30 m) instantaneous albedo and coarse-resolution (500–1000 m) daily mean albedo from remote sensing data, but they all required the input of Moderate Resolution Imaging Spectroradiometer (MODIS) albedo information at 500-m resolution, and none have assembled both instantaneous and daily albedo based exclusively on fine-resolution satellite data. Here, to address this issue, we compiled 387 instantaneous and 346 daily albedo records using field net radiometer measurements from the bioenergy croplands at the W. K. Kellogg Biological Station in southwest Michigan. We then connected these albedo records with a suite of variables derived from harmonized Landsat and Sentinel-2 data through two machine learning algorithms (random forest regression and extreme gradient boosting) to retrieve clear-sky instantaneous and daily shortwave broadband albedo. The performance statistics indicate reasonable accuracy of model results [root-mean-square error (RMSE)] around or below 0.03 except for snow-covered surfaces), suggesting that the retrieval of both instantaneous and daily albedo based exclusively on fine-resolution satellite data is promising. To facilitate the use of fine-resolution albedo products at the global level, future efforts need to include more albedo records of diverse surface cover types, as well as to accurately model daily albedo for cloudy days to address the “clear-sky bias.”

Harmonized Landsat and Sentinel-2↗

Application of Machine Learning Algorithms to Identify Problematic Nuclear Data

In this work we aim to show that Machine learning algorithms are promising tools for the identification of nuclear data that contribute to increased errors in transport simulations. We demonstrate this through an application of a machine learning algorithm (Random Forest) to the Whisper/MCNP6 criticality validation library to identify nuclear data that are associated with an increase of the bias (simulated - experimental $k_{eff}$) in the calculations. Specifically, the $k_{eff}$ sensitivity profiles (w.r.t. nuclear data) of 233 U solution benchmarks are used to predict the bias and Shapley Additive Explanations (SHAP) are used to explain how the sensitivities are related to the predicted bias. The SHAP values can be interpreted as sensitivity coefficients of the machine learning model to the $k_{eff}$ sensitivities which are used to make predictions of bias. Using the SHAP values we can identify specific subsets of nuclear data which have the highest probability of influencing bias. We demonstrate the utility of this method by showing how SHAP values were used to identify an inconsistency in the 19 F inelastic scattering nuclear data. The methodology presented here is not limited to transport problems and can be applied to other simulations if there are experimental measurements to compare against, simulations of those experimental measurements, and the ability to calculate sensitivities of the model output with respect to the data inputs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Machine learning-based inversion for acoustic impedance with large synthetic training data: Workflow and data characterization

Where wells are sparse or training data are difficult to label with high-quality wireline-derived impedance logs, machine learning (ML)-based inversion of acoustic impedance typically depends on small training data sets, leading to biased prediction. We have advanced a novel workflow that applies large synthetic seismic training data to reduce facies-related bias. Using a geologically realistic model as the truth model, we randomly select sparse seed wells to perform sequential Gaussian simulation (SGS) for impedance models of the same geometry and simulate facies variability. We implement random forest regression on 30 features extracted from the synthetic volume. We observe that more seed wells tend to reduce facies-induced bias by sampling more types of facies, resulting in a better prediction. We then focus on the responses of SGS models to facies changes, the number of seed wells necessary for a useful synthetic model, and how much a synthetic model can help ML-based inversion. Here, we observe that the SGS synthetic training model outperforms well-direct training in general. For modeled clastic shore-zone systems in Miocene Gulf of Mexico, two or more seed wells are necessary for a significant reduction of root-mean-square error and outliners, and improvement of facies imaging. In a field-data test, we apply a similar workflow to quantitatively predict acoustic impedance, which is then converted to a sand-volume map at a high-frequency sequence (10–100 m), revealing detailed facies and sandstone patterns. Such results are valuable in many geologic and engineering applications, such as hydrocarbon and CO 2 reservoir prospecting, reserve estimation, simulation, etc.

3D seismic↗