Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust PCA-Deep Belief Network Surrogate Model for Distribution System Topology Identification with DERs

With the expansion of distribution networks and increased penetration of distributed energy resources (DERs), it is becoming increasingly important to obtain accurate distribution network topology in real-time. In this paper, a robust principal component analysis coupled deep belief network (PCA-DBN) surrogate model is proposed for distribution system topology identification. It integrates the benefits of robust feature extraction from PCA to deal with data quality issues and filter out noise, and the strength of DBN in capturing the nonlinear relationship between voltage amplitudes and the binary states of switchable connections. This also significantly reduces the DBN training complexity without loss of accuracy. It is shown that the widely used standard deviation of voltage drop and the voltage covariance matrix features yield less accuracy as compared to that of the voltage amplitudes in presence of high penetration of DERs and ZIP loads. Comparison results with other alternatives, such as the random forest (RF), multi-output regression (MOR) and the traditional DBN methods demonstrate that the proposed method can achieve a much higher topology identification accuracy while maintaining robustness to missing data and measurement noise under various penetration levels of DERs.

deep belief network↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Improving E3SM Land Model Photosynthesis Parameterization via Satellite SIF, Machine Learning, and Surrogate Modeling

The parameterization of key photosynthesis parameters is one of the key uncertain sources in modeling ecosystem gross primary productivity (GPP). Solar-induced chlorophyll fluorescence (SIF) offers a good proxy for GPP since it marks the actual process of photosynthesis; while machine learning (ML) provides a robust approach to model the GPP-SIF relationship. Here, we trained the boosted regressing tree (BRT) and the Random Forest ML models with Greenhouse Gases Observing Satellite SIF data and in situ GPP observations from 49 eddy covariance towers. These trained ML GPP-SIF models were fed into the Energy Exascale Earth System Model (E3SM) Land Model (ELM) to generate ELM-simulated global SIF estimates, which were then benchmarked against satellite SIF observations with a surrogate modeling approach. Our results indicated good modeling performance of the ML-based GPP-SIF relationship. The ELM model when fed with the ML GPP-SIF models also can well predict the spatial-temporal variations in SIF. We also found high model accuracy for the surrogate modeling. Model parameter sensitivity analysis suggested that the fraction of leaf nitrogen in RuBisCO (flnr) is the most sensitive parameter to the SIF; other sensitive parameters include the Ball-Berry stomatal conductance slope (mbbopt) and the vcmax entropy (vcmaxse). The posterior uncertainty in simulated GPP was greatly reduced after benchmarking, and the model produced improved spatial patterns of mean GPP relative to FLUXCOM GPP. Our integrated approach provides a new avenue for improving land models and using remote-sensing SIF, which can be further improved in the future with more ground- and satellite-based observations.

54 ENVIRONMENTAL SCIENCES↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

Integration of LIBS with Machine Learning for Real-Time Monitoring of Feedstock in H 2 Gasification Applications

This project, funded by the U.S. Department of Energy (DOE) – Office of Fossil Energy under Award Number DE-FE0032177, aimed to assess the feasibility of an integrated Laser-Induced Breakdown Spectroscopy (LIBS) system with advanced machine learning (ML) models for real-time characterization and potential control of hydrogen gasifiers running on waste materials as feedstocks. This was a multidisciplinary effort that encompassed the acquisition and standardized analysis of individual and blended feedstocks—comprising biomass, coal waste, and plastic waste, followed by the development of a dynamic LIBS bench system for material sample analysis and development of predictive ML models. Comprehensive laboratory testing enabled the creation of a robust elemental dataset that served as the foundation for ML model training. Techniques such as Random Forest, Gradient Boosting, Support Vector Regression, and Neural Networks were employed to predict key feedstock properties, including higher heating value (HHV), moisture content, thermal conductivity, and ash composition with high accuracy. The results were validated against experimental data and demonstrated strong potential for real-time application in gasifier control systems. The project concluded with a study on the integration of the LIBS+ML approach for gasifier control and a techno-economic analysis of the implementation of the approach into hydrogen (H 2 ) gasification systems. Dissemination of results was carried out at a DOE meeting. This work establishes a scalable framework for automated, in-line feedstock quality assessment, offering significant implications for process optimization and emissions reduction in hydrogen production.

01 COAL, LIGNITE, AND PEAT↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Stand validation of lidar forest inventory modeling for a managed southern pine forest

We evaluated area-based approaches (ABAs) to light detection and ranging (lidar) predictions of plot- and stand-level forest attributes (tree count, height, basal area, volume, aboveground biomass, broadleaf/conifer, and diameter at breast height — “diameter”). ABA methods included post-stratification (PS), ordinary least squares (OLSs) regression, k nearest neighbors ( kNN), and random forest (RF). This study was conducted on the Savannah River Site in South Carolina, USA. Plot- and stand-level predictions were validated against fixed-radius 0.04 ha (0.1 acre) plots in 49 ≈2.0 ha (5 acre) stands. Our findings demonstrate that lidar can be incorporated operationally into forest inventory systems to provide stand-level inferences for a wide range of forest attributes. Volume predictions for specific diameter classes, however, often fared poorly (root mean squared error (RMSE) > 100%) for the methods we explored, especially for larger (less common) diameter trees. Stand-level results were consistently better than pixel-level results (10–200+ percentage points). kNN and RF performed similarly and better than OLS and PS, but RF was the most robust to model configurations, while kNN has practical advantages such as simultaneous predictions of many attributes.

Forestry↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Applying NIR and MIR spectroscopy for C and soil property prediction in northern cold-region ecosystems. Which approach works better?

Here, developing reliable predictions of soil attributes is necessary to understand northern cold-region climate-soil feedback. Calibration models using near-infrared (NIR) and mid-infrared (MIR) spectroscopy were developed to predict eight commonly measured soil properties for 119 soil samples representing a range of vegetation types, parent materials, and soil types spanning >23° of latitude from southeast Alaska to the Canadian high Arctic. In order to obtain a more accurate prediction, this study compared the performance of linear and non-linear calibration techniques, including lasso regression (Lasso), support vector machine (SVM), random forest (RF) and classic partial least squares (PLS) to predict different soil properties of these soils. Comparing the four models, we noticed that their performance was quite similar for MIR overall, while NIR achieved better results with a PLS model for our dataset. PLS coupled with MIR showed a better performance for soil parameters, such as total organic carbon (TOC), total nitrogen (TN), cation exchange capacity (CEC) and clay (R-squared of 0.9, 0.81, 0.80, and 0.84) when compared with NIR (R-squared of 0.85, 0.72, 0.81 and 0.68). However, using either MIR or NIR spectroscopy, PLS predictions for bulk density (BD) and sand content were not accurate. The variable importance analysis based on the PLS model successfully estimated the relative contribution of wavelengths influencing soil property predictions most. Overall, TOC, TN, CEC and clay mineral predictions are closely related to the occurrence of specific spectral bands in the MIR region. For example, wavelengths at 2978 and 1761 cm -1 for TOC and TN, as well as at 3064 cm -1 for CEC, were selected as the most influential predictor variables. We demonstrated that MIR spectroscopy is a powerful tool for more extensive monitoring in soils of the northern cold climate region; however, NIR could be utilized for rapid estimates when the highest accuracy is not essential.

54 ENVIRONMENTAL SCIENCES↗

The Global LAnd Surface Satellite (GLASS) evapotranspiration product Version 5.0: Algorithm development and preliminary validation

An accurate estimation of spatially and temporally continuous global terrestrial evapotranspiration (ET) is essential in the assessment of surface energy, water and carbon cycles. The Global LAnd Surface Satellite (GLASS) ET product Version 4.0 (v4.0) based on the Bayesian model averaging (BMA) method was generated to estimate global terrestrial ET. However, certain uncertainty for the GLASS ET product v4.0 limits its application. In this study, we introduced the deep neural networks (DNN) merging framework to improve terrestrial ET estimation for GLASS ET product Version 5.0 (v5.0) generation by integrating five satellite-derived ET products [Moderate Resolution Imaging Spectroradiometer (MODIS) ET product (MOD16), Shuttleworth–Wallace dual-source ET product (SW), Priestley–Taylor-based ET product (PT-JPL), modified satellite-based Priestley–Taylor ET product (MS-PT) and simple hybrid ET product (SIM)]. We compared the performance of DNN method against other merging methods, including GLASS ET algorithm v4.0 (BMA), the gradient boosting regression tree (GBRT) method and the random forest (RF) method, based on 195 global eddy covariance (EC) flux towers covering observations from 2000 through 2015. Validations indicated that the DNN had the highest accuracy among four merging methods across different land cover types, yielding the highest average determination coefficients (R 2 , 0.62), root-mean-squared-error (RMSE, 24.1 W/m 2 ) and Kling–Gupta efficiency (KGE, 0.77) with a of 99% confidence interval. Compared with GLASS ET algorithm v4.0, the DNN improved on the R 2 by approximately 7% (p < 0.01) and the KGE by 10%. Based on the DNN, we then generated 8-day GLASS ET product v5.0 globally with a 1 km spatial resolution from 2001 to 2015 driven by GLASS vegetation and surface net radiation (R n ) datasets and Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA2) datasets. Finally, this global terrestrial ET product provides a valuable dataset for monitoring regional and global water resources and environmental changes.

54 ENVIRONMENTAL SCIENCES↗

Use of Satellite, Surface Observations and Numerical Weather Prediction Model Data to Improve Cloud Base Height and Cloud Base Vertical Velocity Estimation

Cloud base height (CBH) and cloud base vertical velocity (CBVV) are important variables that impact the overall climate in a region as they influence the formulation, longevity, and evolution of clouds. Retrieval of both parameters have long used ground instrumentation (e.g., Doppler lidar (DL), ground base radar); however, retrieving CBH from satellites is particularly challenging given that space-based instruments only observe cloud tops. In this manuscript, CBH is retrieved using a multi-linear regression equation, while CBVV used a random forests model. Both retrievals combine satellite and numerical weather prediction data. The satellite data used are the Visible Infrared Imaging Radiometer Suite imagery, while measurements of CBH and CBVV include DL and radiosonde data at the Southern Great Plains (SGP) Atmospheric Radiation Measurement observatory. Data from 83 summer days (May-August) in 2018–2021 featuring cumulus clouds forced by solar heating were examined and used to train the models, with years 2022–2023 used for validation. Various spatial domains were defined with one large (2.4° longitude by 2.0° latitude) SGP domain being split into smaller sections (smallest being 0.99° and 0.61° longitude and latitude respectably). CBH and CBVV values obtained from the DL as compared to the models show root mean square errors between 150 and 200 m, with CBVV values between 0.45 and 1 ms -1 . Finally, it was found that the CBH formulation performs well over all domains, while the CBVV retrievals become less accurate due to more turbulence being introduced into the observations as the number of DL stations decreases in the smaller domains.

54 ENVIRONMENTAL SCIENCES↗

Machine learning prediction of incidence of Alzheimer’s disease using large-scale administrative health data

Nationwide population-based cohort provides a new opportunity to build an automated risk prediction model based on individuals’ history of health and healthcare beyond existing risk prediction models. We tested the possibility of machine learning models to predict future incidence of Alzheimer’s disease (AD) using large-scale administrative health data. From the Korean National Health Insurance Service database between 2002 and 2010, we obtained de-identified health data in elders above 65 years (N = 40,736) containing 4,894 unique clinical features including ICD-10 codes, medication codes, laboratory values, history of personal and family illness and socio-demographics. To define incident AD we considered two operational definitions: “definite AD” with diagnostic codes and dementia medication (n = 614) and “probable AD” with only diagnosis (n = 2026). We trained and validated random forest, support vector machine and logistic regression to predict incident AD in 1, 2, 3, and 4 subsequent years. For predicting future incidence of AD in balanced samples (bootstrapping), the machine learning models showed reasonable performance in 1-year prediction with AUC of 0.775 and 0.759, based on “definite AD” and “probable AD” outcomes, respectively; in 2-year, 0.730 and 0.693; in 3-year, 0.677 and 0.644; in 4-year, 0.725 and 0.683. The results were similar when the entire (unbalanced) samples were used. Important clinical features selected in logistic regression included hemoglobin level, age and urine protein level. This study may shed a light on the utility of the data-driven machine learning model based on large-scale administrative health data in AD risk prediction, which may enable better selection of individuals at risk for AD in clinical trials or early detection in clinical settings.

97 MATHEMATICS AND COMPUTING↗

A Comparison of Machine Learning Methods to Forecast Tropospheric Ozone Levels in Delhi

Ground-level ozone is a pollutant that is harmful to urban populations, particularly in developing countries where it is present in significant quantities. It greatly increases the risk of heart and lung diseases and harms agricultural crops. This study hypothesized that, as a secondary pollutant, ground-level ozone is amenable to 24 h forecasting based on measurements of weather conditions and primary pollutants such as nitrogen oxides and volatile organic compounds. We developed software to analyze hourly records of 12 air pollutants and 5 weather variables over the course of one year in Delhi, India. To determine the best predictive model, eight machine learning algorithms were tuned, trained, tested, and compared using cross-validation with hourly data for a full year. The algorithms, ranked by R2 values, were XGBoost (0.61), Random Forest (0.61), K-Nearest Neighbor Regression (0.55), Support Vector Regression (0.48), Decision Trees (0.43), AdaBoost (0.39), and linear regression (0.39). When trained by separate seasons across five years, the predictive capabilities of all models increased, with a maximum R 2 of 0.75 during winter. Bidirectional Long Short-Term Memory was the least accurate model for annual training, but had some of the best predictions for seasonal training. Out of five air quality index categories, the XGBoost model was able to predict the correct category 24 h in advance 90% of the time when trained with full-year data. Separated by season, winter is considerably more predictable (97.3%), followed by post-monsoon (92.8%), monsoon (90.3%), and summer (88.9%). These results show the importance of training machine learning methods with season-specific data sets and comparing a large number of methods for specific applications.

54 ENVIRONMENTAL SCIENCES↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

Predicting measures of soil health using the microbiome and supervised machine learning

Soil health encompasses a range of biological, chemical, and physical soil properties that sustain the commercial and ecological value of agroecosystems. Monitoring soil health requires a comprehensive set of diagnostics that can be cost-prohibitive for routine analyses. The soil microbiome provides a rich source of information about soil properties, which can be assayed in a high-throughput, cost-effective way. We evaluated the accuracy of random forest (RF) and support vector machine (SVM) regression and classification models in predicting 12 measures of soil health, tillage status, and soil texture from 16S rRNA gene amplicon data with an operationally relevant sample set. We validated the efficacy of the best performing models against independent datasets and also tested best practices for processing microbiome data for use in machine learning. Soil health metrics could be predicted from microbiome data with the best models achieving a Kappa value of ~0.65, for categorical assessments, and a R2 value of ~0.8, for numerical scores. Biological health ratings were better predicted than chemical or physical ratings. Validation with independent datasets revealed that models had general predictive value for soil properties, including yield. The ecological profiles of several taxa important for model accuracy matched the observed relationships with soil health, including Pyrinomonadaceae, Nitrososphaeraceae, and Candidatus Udeaobacter. Models trained at the highest taxonomic resolution proved most accurate, with losses in accuracy resulting from rarefying, sparsity filtering, and aggregating at higher taxonomic ranks. Furthermore, our study provides the groundwork for developing scalable technology to use microbiome-based diagnostics for the assessment of soil health.

16S rRNA gene↗

Machine learning prediction of self-diffusion in Lennard-Jones fluids

In this work, different machine learning (ML) methods were explored for the prediction of self-diffusion in Lennard-Jones (LJ) fluids. Using a database of diffusion constants obtained from the molecular dynamics simulation literature, multiple Random Forest (RF) and Artificial Neural Net (ANN) regression models were developed and characterized. The role and improved performance of feature engineering coupled to the RF model development was also addressed. The performance of these different ML models was evaluated by comparing the prediction error to an existing empirical relationship used to describe LJ fluid diffusion. It was found that the ANN regression models provided superior prediction of diffusion in comparison to the existing empirical relationships.

74 ATOMIC AND MOLECULAR PHYSICS↗