Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Predictive models of long COVID

Background: The cause and symptoms of long COVID are poorly understood. It is challenging to predict whether a given COVID-19 patient will develop long COVID in the future. Methods: We used electronic health record (EHR) data from the National COVID Cohort Collaborative to predict the incidence of long COVID. We trained two machine learning (ML) models — logistic regression (LR) and random forest (RF). Features used to train predictors included symptoms and drugs ordered during acute infection, measures of COVID-19 treatment, pre-COVID comorbidities, and demographic information. We assigned the ‘long COVID’ label to patients diagnosed with the U09.9 ICD10-CM code. The cohorts included patients with (a) EHRs reported from data partners using U09.9 ICD10-CM code and (b) at least one EHR in each feature category. We analysed three cohorts: all patients (n = 2,190,579; diagnosed with long COVID = 17,036), inpatients (149,319; 3,295), and outpatients (2,041,260; 13,741). Findings: LR and RF models yielded median AUROC of 0.76 and 0.75, respectively. Ablation study revealed that drugs had the highest influence on the prediction task. The SHAP method identified age, gender, cough, fatigue, albuterol, obesity, diabetes, and chronic lung disease as explanatory features. Models trained on data from one N3C partner and tested on data from the other partners had average AUROC of 0.75. Interpretation: ML-based classification using EHR information from the acute infection period is effective in predicting long COVID. SHAP methods identified important features for prediction. Cross-site analysis demonstrated the generalizability of the proposed methodology.

60 APPLIED LIFE SCIENCES↗

Non-mercury methylating microbial taxa are integral to understanding links between mercury methylation and elemental cycles in marine and freshwater sediments

The goal of this study was to explore the role of non-mercury (Hg) methylating taxa in mercury methylation and to identify potential links between elemental cycles and Hg methylation. Statistical approaches were utilized to investigate the microbial community and biochemical functions in relation to methylmercury (MeHg) concentrations in marine and freshwater sediments. Sediments were collected from the methylation zone (top 15 cm) in four Hg-contaminated sites. Both abiotic (e.g., sulfate, sulfide, iron, salinity, total organic matter, etc.) and biotic factors (e.g., hgcA, abundances of methylating and non-methylating taxa) were quantified. Random forest and stepwise regression were performed to assess whether non-methylating taxa were significantly associated with MeHg concentration. Co-occurrence and functional network analyses were constructed to explore associations between taxa by examining microbial community structure, composition, and biochemical functions across sites. Regression analysis showed that approximately 80% of the variability in sediment MeHg concentration was predicted by total mercury concentration, the abundances of Hg methylating taxa, and the abundances of the non-Hg methylating taxa. The co-occurrence networks identified Paludibacteraceae and Syntrophorhabdaceae as keystone non Hg methylating taxa in multiple sites, indicating the potential for syntrophic interactions with Hg methylators. Strong associations were also observed between methanogens and sulfate-reducing bacteria, which were likely symbiotic associations. The functional network results suggested that non-Hg methylating taxa play important roles in sulfur respiration, nitrogen respiration, and the carbon metabolism-related functions methylotrophy, methanotrophy, and chemoheterotrophy. Interestingly, keystone functions varied by site and did not involve carbon- and sulfur-related functions only. In conclusion, our findings highlight associations between methylating and non-methylating taxa and sulfur, carbon, and nitrogen cycles in sediment methylation zones, with implications for predicting and understanding the impact of climate and land/sea use changes on Hg methylation.

54 ENVIRONMENTAL SCIENCES↗

Yield strength prediction of high-entropy alloys using machine learning

Yield strength at high temperature is an important parameter in the design and application of high entropy alloys (HEAs). However, the experimental measurement of yield strength at high temperature is quite costly, complicated, and time-consuming. Therefore, it is essential to identify and apply a robust method for the accurate prediction of yield strength at high temperature from the available experimental and simulation data. In this study, for the first time, a machine learning (ML) method based on the regression technique of random forest (RF) regressor is used to predict the yield strength of HEAs at the desired temperature. Further, the yield strengths of MoNbTaTiW and HfMoNbTaTiZr at 800 °C and 1200 °C, are predicted using the RF regressor model. We find that the results are consistent with the experimental reports, showing that the RF regressor model predicts the yield strength of HEAs at the desired temperatures with high accuracy.

36 MATERIALS SCIENCE↗

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Peak Rain Rate Sensitivity to Observed Cloud Condensation Nuclei and Turbulence in Continental Warm Shallow Clouds During CACTI

Abstract Warm clouds strongly affect Earth's energy budget but remain imperfectly represented in climate models, partly due to the complexity and covariability of relevant processes influencing warm rain. This work presents a detailed analysis of different factors affecting rain rate peak intensity (RR) in continental warm clouds. Clouds were identified with vertically pointing radar and lidar observations and categorized via a temperature‐based cloud type classification algorithm from which warm clouds were isolated. Observations and retrievals of liquid water path (LWP), cloud condensation nuclei concentration (N CCN ), cloud depth, and cloud duration of more than 3,000 separate warm clouds sampled during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign are analyzed in this work. Multiple linear regression (MLR) and random forest (RF) models are applied to assess the relative impact of these variables on RR. Overall, RR tends to increase as cloud depth, LWP, and cloud duration increase, or N CCN decreases. Cloud depth affects RR the most while N CCN impacts it the least. When considering over 170 warm clouds observed at least 1 hr in which in‐cloud turbulence is retrieved, the effect of N CCN on RR remains most likely suppressive, but it is not significant at a 75% level for MLR and is highly uncertain for RF. The impact of in‐cloud turbulence depends on the moment and location it is sampled. Cloud base turbulence around the time of RR suppresses RR, while cloud top turbulence effects are inconclusive. Possible difficulties in isolating robust CCN and turbulence effects on RR are discussed.

54 ENVIRONMENTAL SCIENCES↗

Predicting solid state material platforms for quantum technologies

Semiconductor materials provide a compelling platform for quantum technologies (QT). However, identifying promising material hosts among the plethora of candidates is a major challenge. Therefore, we have developed a framework for the automated discovery of semiconductor platforms for QT using material informatics and machine learning methods. Different approaches were implemented to label data for training the supervised machine learning (ML) algorithms logistic regression, decision trees, random forests and gradient boosting. We find that an empirical approach relying exclusively on findings from the literature yields a clear separation between predicted suitable and unsuitable candidates. In contrast to expectations from the literature focusing on band gap and ionic character as important properties for QT compatibility, the ML methods highlight features related to symmetry and crystal structure, including bond length, orientation and radial distribution, as influential when predicting a material as suitable for QT.

36 MATERIALS SCIENCE↗

Learning electric vehicle driver range anxiety with an initial state of charge-oriented gradient boosting approach

This manuscript focuses on the modeling of electric vehicle (EV) driver’s range anxiety, a fear that a vehicle does not have sufficient range, or state of charge (SOC) of the battery pack, to reach its destination and would strand its occupants. Despite numerous research studies on the modeling of charging behaviors, modeling efforts to understand at what battery percentages do EV drivers charge their vehicles, and what are the associated contributing factors, are rather limited. To this end, an ensemble learning model based on gradient boosting is developed. The model sequentially fits new predictors to new residuals of the previous prediction and, then, minimizes the loss when adding the latest prediction. A total of 18 features are defined and extracted from the multisource data, which cover information on driver, vehicles, stations, traffic conditions, as well as spatial-temporal context information of the charging events. The analyzed dataset includes 4.5-year’s charging event log data from 3,096 users and 468 public charging stations in Kansas City Missouri, and the macroscopic travel demand model maintained by the metropolitan planning organization. Here, the result shows the proposed model achieved a satisfactory result with a R square value of 0.54 and root mean square error of 0.14, both better than multiple linear regression model and random forest model. To reduce range anxiety, it is suggested that the priorities of deploying new charging facilities should be given to the areas with higher daily traffic prediction, with more conservative EV users or that are further from residential areas.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

A Machine Learning-Based Vulnerability Analysis for Cascading Failures of Integrated Power-Gas Systems

This article proposes a cascading failure simulation (CFS) method and a hybrid machine learning method for vulnerability analysis of integrated power-gas systems (IPGSs). The CFS method is designed to study the propagating process of cascading failures between the two systems, generating data for machine learning with initial states randomly sampled. The proposed method considers generator and gas well ramping, transmission line and gas pipeline tripping, island issue handling and load shedding strategies. Then, a hybrid machine learning model with a combined random forest (RF) classification and regression algorithms is proposed to investigate the impact of random initial states on the vulnerability metrics of IPGSs. Extensive case studies are carried out on three test IPGSs to verify the proposed models and algorithms. Simulation results show that the proposed models and algorithms can achieve high accuracy for the vulnerability analysis of IPGSs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Surface temperatures reveal the patterns of vegetation water stress and their environmental drivers across the tropical Americas

Vegetation is a key component in the global carbon cycle as it stores ~450 GtC as biomass, and removes about a third of anthropogenic CO2 emissions. However, in some regions, the rate of plant carbon uptake is beginning to slow, largely because of water stress. Here, we develop a new observation-based methodology to diagnose vegetation water stress and link it to environmental drivers. We used the ratio of remotely sensed land surface to near surface atmospheric temperatures (LST/Tair) to represent vegetation water stress, and built regression tree models (random forests) to assess the relationship between LST/Tair and the main environmental drivers of surface energy fluxes in the tropical Americas. We further determined ecosystem traits associated with water stress and surface energy partitioning, pinpointed critical thresholds for water stress, and quantified changes in ecosystem carbon uptake associated with crossing these critical thresholds. Additionally, we found that the top drivers of LST/Tair, explaining over a quarter of its local variability in the study region, are (1) radiation, in 58% of the study region; (2) water supply from precipitation, in 30% of the study region; and (3) atmospheric water demand from vapor pressure deficits (VPD), in 22% of the study region. Regions in which LST/Tair variation is driven by radiation are located in regions of high aboveground biomass or at high elevations, while regions in which LST/Tair is driven by water supply from precipitation or atmospheric demand tend to have low species richness. Carbon uptake by photosynthesis can be reduced by up to 80% in water-limited regions when critical thresholds for precipitation and air dryness are exceeded simultaneously, that is, as compound events. Our results demonstrate that vegetation structure and diversity can be important for regulating surface energy and carbon fluxes over tropical regions.

59 BASIC BIOLOGICAL SCIENCES↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse↗

Online LIBS–ML Framework for Dynamic Characterization of Heterogeneous Waste-Derived Gasification Feedstocks

LIBS−ML framework for real time feedstock characterization during continuous conveyor transport Heterogeneous waste derived feedstocks (e.g., waste coal, biomass and blends) introduce rapid variability in heating value and ash chemistry that affect gasifier operation, yet conventional laboratory characterization techniques are too slow to support proactive control. To address this gap, this study reports on an online, in situ, dynamic characterization framework that couple’s laser-induced breakdown spectroscopy (LIBS) with leakage safe machine learning (ML) regression to deliver real time, decision quality predictions of gasifier relevant properties. A controlled sample matrix spanning two different waste coals, two different biomasses, and engineered blends under two particle size conditions were constructed and benchmarked using standardized laboratory analyses for proximate/ultimate properties and ash composition. LIBS spectra were acquired dynamically as material flowed on a conveyor belt, using high energy 1064 nm laser ablation and shot averaging to improve repeatability and precision. Supervised regression models (multi layer perceptron (MLP) /artificial neural network (ANN), random forest (RF), and support vector regression (SVR)) and an optimized weighted ensemble were trained on emission line feature sets using nested cross validation with Bayesian hyperparameter tuning and validated against an independent hold out set. The proposed LIBS−ML workflow achieves near laboratory predictive fidelity across parametric targets (including higher heating value (HHV), ash content, fixed carbon, sulfur, major ash forming oxides, and initial deformation temperature (IDT)), with the weighted ensemble providing a robust default predictor under dynamic measurement conditions. These results demonstrate a practical pathway for real time feedstock characterization that can enable feedforward adjustments and more resilient gasifier operation for variable quality waste derived fuels.

Biomass↗

Machine learning predictions of high-Curie-temperature materials

Technologies that function at room temperature often require magnets with a high Curie temperature, $T$ C , and can be improved with better materials. Discovering magnetic materials with a substantial $T$ C is challenging because of the large number of candidates and the cost of fabricating and testing them. Using the two largest known datasets of experimental Curie temperatures, we develop machine-learning models to make rapid $T$ C predictions solely based on the chemical composition of a material. We train a random-forest model and a k -NN one and predict on an initial dataset of over 2500 materials and then validate the model on a new dataset containing over 3000 entries. The accuracy is compared for multiple compounds' representations (“descriptors”) and regression approaches. A random-forest model provides the most accurate predictions and is not improved by dimensionality reduction or by using more complex descriptors based on atomic properties. Further, a random-forest model trained on a combination of both datasets shows that cobalt-rich and iron-rich materials have the highest Curie temperatures for all binary and ternary compounds. An analysis of the model reveals systematic error that causes the model to over-predict low-$T$ C materials and under-predict high-$T$ C materials. For exhaustive searches to find new high-$T$ C materials, analysis of the learning rate suggests either that much more data is needed or that more efficient descriptors are necessary.

36 MATERIALS SCIENCE↗

Deep Learning Estimation of Daily Ground–Level NO 2 Concentrations from Remote Sensing Data

The limited number of nitrogen dioxide (NO 2 ) surface measurements calls for the development of highly accurate approaches to estimating surface NO 2 concentrations. In this study, we leverage a new satellite instrument, the TROPOspheric Monitoring Instrument (TROPOMI), along with other predictor variables, to estimate daily surface NO 2 concentrations over Texas in 2019. We use the deep convolutional neural network (Deep-CNN), an advanced deep learning algorithm, to obtain estimates and achieve a correlation coefficient (R) of 0.91, an index of agreement (IOA) of 0.95, and a mean absolute bias (MAB) of 1.75 ppb in surface NO 2 estimation. Additionally, we leverage a novel approach, SHapley Additive exPlanations (SHAP), to describe how Deep-CNN understands each predictor variable. The SHAP results show that the Deep-CNN model has an advanced understanding of the dataset, revealing that TROPOMI closely captures levels of NO 2 . In addition, we show the superiority of our Deep-CNN model at estimating surface NO 2 over other well-known machine learning and regression models in the field, including the support vector machines (SVM), random forest (RF), and multiple linear regression (MLR). Although SVM and RF show strong capabilities at estimating surface NO 2 concentrations, their accuracy is inferior to that of the Deep-CNN model, ranking second and third in model accuracy in this study. The MLR, however, shows a poor ability at NO 2 estimation and ranks last among all models. Furthermore, testing the impact of sample size on model performance, we also show that, compared to other models, Deep-CNN needs more samples to trigger its strength at surface NO 2 estimation.

54 ENVIRONMENTAL SCIENCES↗