Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Estimating Compressional Velocity and Bulk Density Logs in Marine Gas Hydrates Using Machine Learning

Compressional velocity (Vp) and bulk density (ρb) logs are essential for characterizing gas hydrates and near-seafloor sediments; however, it is sometimes difficult to acquire these logs due to poor borehole conditions, safety concerns, or cost-related issues. We present a machine learning approach to predict either compressional Vp or ρb logs with high accuracy and low error in near-seafloor sediments within water-saturated intervals, in intervals where hydrate fills fractures, and intervals where hydrate occupies the primary pore space. We use scientific-quality logging-while-drilling well logs, gamma ray, ρb, Vp, and resistivity to train the machine learning model to predict Vp or ρb logs. Of the six machine learning algorithms tested (multilinear regression, polynomial regression, polynomial regression with ridge regularization, K nearest neighbors, random forest, and multilayer perceptron), we find that the random forest and K nearest neighbors algorithms are best suited to predicting Vp and ρb logs based on coefficients of determination (R2) greater than 70% and mean absolute percentage errors less than 4%. Given the high accuracy and low error results for Vp and ρb prediction in both hydrate and water-saturated sediments, we argue that our model can be applied in most LWD wells to predict Vp or ρb logs in near-seafloor siliciclastic sediments on continental slopes irrespective of the presence or absence of gas hydrate.

Naim, Fawz↗

On the estimation of boundary layer heights: a machine learning approach

Abstract. The planetary boundary layer height (zi) is a key parameter used in atmospheric models for estimating the exchange of heat, momentum, and moisture between the surface and the free troposphere. Near-surface atmospheric and subsurface properties (such as soil temperature, relative humidity, etc.) are known to have an impact on zi. Nevertheless, precise relationships between these surface properties and zi are less well known and not easily discernible from the multi-year dataset. Machine learning approaches, such as random forest (RF), which use a multi-regression framework, help to decipher some of the physical processes linking surface-based characteristics to zi. In this study, a 4-year dataset from 2016 to 2019 at the Southern Great Plains site is used to develop and test a machine learning framework for estimating zi. Parameters derived from Doppler lidars are used in combination with over 20 different surface meteorological measurements as inputs to a RF model. The model is trained using radiosonde-derived zi values spanning the period from 2016 through 2018 and then evaluated using data from 2019. Results from 2019 showed significantly better agreement with the radiosonde compared to estimates derived from a thresholding technique using Doppler lidars only. Noteworthy improvements in daytime zi estimates were observed using the RF model, with a 50 % improvement in mean absolute error and an R2 of greater than 85 % compared to the Tucker method zi. We also explore the effect of zi uncertainty on convective velocity scaling and present preliminary comparisons between the RF model and zi estimates derived from atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automatic Waveform Quality Control for Surface Waves Using Machine Learning

Surface-wave seismograms are widely used by researchers to study Earth’s interior and earthquakes. To extract information reliably and robustly from a suite of surface waveforms, the signals require quality control screening to reduce artifacts from signal complexity and noise. This process has usually been completed by human experts labeling each waveform visually, which is time consuming and tedious for large data sets. We explore automated approaches to improve the efficiency of waveform quality control processing by investigating logistic regression, support vector machines, K-nearest neighbors, random forests (RF), and artificial neural networks (ANN) algorithms. To speed up signal quality assessment, we trained these five machine learning (ML) methods using nearly 400,000 human-labeled waveforms. The ANN and RF models outperformed other algorithms and achieved a test accuracy of 92%. We evaluated these two best-performing models using seismic events from geographic regions not used for training. The results show that the two trained models agree with labels from human analysts but required only 0.4% of the time. Although the original (human) quality assignments assessed general waveform signal-to-noise, the ANN or RF labels can help facilitate detailed waveform analysis. Our investigations demonstrate the capability of the automated processing using these two ML models to reduce outliers in surface-wave-related measurements without human quality control screening.

58 GEOSCIENCES↗

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Analysis and prediction of intersection traffic violations using automated enforcement system data

We report that the automated enforcement system (AES) is an effective way of supplementing traditional traffic enforcement, and the traffic violation data from AES can also be effectively used for safety research. In this study, traffic violation data were used to analyze the influencing factors associated with traffic violations and to predict the probability of violations at intersections. The potential factors influencing violations include 24 independent factors related to time, space, traffic and weather. Results from a logistic model showed that the midday period, weekends, residential districts, collector roads, congested traffic conditions, high traffic flow, lower wind speed and low temperature would increase the probability of traffic violations. The probability of violations was predicted by the random forest algorithm, which was proven to be the best traffic violation prediction model among logistic regression, Gaussian naive Bayes, and support vector machine. Moreover, the proximity weighted synthetic oversampling technique (ProWSyn) method was applied to reduce the impact of the imbalance ratio (IR) and improve the model’s prediction performance. The receiver operating characteristics (ROC) curves and Precision-Recall (PR) curves illustrated that the random forest algorithm using oversampling data had the best classifier prediction performance than undersampling data. The area under curve (AUC) and out-of-bag (OOB) error with IR = 1 reached 0.914 and 0.0787, which showed the better performance of the random forest algorithm using ProWSyn in dealing with imbalanced traffic violation data.

42 ENGINEERING↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Field-scale dynamics of planting dates in the US Corn Belt from 2000 to 2020

Crop planting dates are a dynamic feature of agricultural systems that respond to short- and long-term climate signals, crop and cultivar selection, and technology changes. Planting date records are essential for yield gap analyses, accurate crop modeling, and tracking farmer adaptations to weather and climate change. Although planting dates have high variation at local scales due to heterogeneity in farm resources and decision-making, available long-term data on planting dates is largely restricted to aggregated regional statistics or, at best, satellite-derived datasets with limited spatiotemporal extent and at resolutions unable to distinguish individual fields (> 250 m). Here, we generated retrospective annual field-scale (30 m) planting date maps for both maize and soybeans spanning 2000-2020 across a 12 state region in the United States Corn Belt based on Landsat satellite data and a large ground sample of over 28,000 maize and soybean fields. Using training data from 2015-2020 for model selection, we found that planting date predictions improved with harmonic regression of Landsat data and additional annual weather covariates. The preferred random forests model approximately doubled performance compared to a null model based on state median planting dates, capturing 47% of field-level variation for maize (mean absolute error, MAE = 7.4 days) and 44% for soybeans (MAE = 7.5 days) against held-out ground truth test data for 2008-2014. We also evaluated the full 2000-2020 dataset with state agricultural statistics, finding strong agreement with median planting dates for maize (R 2 = 0.76, MAE = 4.4 days) and slightly lower agreement for soybeans (R 2 = 0.65, MAE = 5.4 days) when aggregated to the state level. We then used this new dataset to analyze environmental determinants of planting dates at a finer-scale than previously possible, controlling for unobserved variation at the sub-state district level. We found that during 2000-2020, each standard deviation increase in rainfall delayed planting by ~ 2.5 days, and fields with higher soil productivity ratings tended to be planted earlier. We did not find meaningful trends over the last two decades in planting dates for maize or soybeans, in contrast to trends towards earlier planting dates late last century and predicted for this period in climate adaptation studies. We hypothesize increases in early season rainfall may have inhibited these shifts towards earlier planting. Remotely sensed planting dates will be a useful tool for yield gap analyses, crop simulation modeling, and ongoing assessment of climate adaptation.

54 ENVIRONMENTAL SCIENCES↗

Angular clustering properties of the DESI QSO target selection using DR9 Legacy Imaging Surveys

ABSTRACT The quasar target selection for the upcoming survey of the Dark Energy Spectroscopic Instrument (DESI) will be fixed for the next 5 yr. The aim of this work is to validate the quasar selection by studying the impact of imaging systematics as well as stellar and galactic contaminants, and to develop a procedure to mitigate them. Density fluctuations of quasar targets are found to be related to photometric properties such as seeing and depth of the Data Release 9 of the DESI Legacy Imaging Surveys. To model this complex relation, we explore machine learning algorithms (random forest and multilayer perceptron) as an alternative to the standard linear regression. Splitting the footprint of the Legacy Imaging Surveys into three regions according to photometric properties, we perform an independent analysis in each region, validating our method using extended Baryon Oscillation Spectroscopic Survey (eBOSS) EZ-mocks. The mitigation procedure is tested by comparing the angular correlation of the corrected target selection on each photometric region to the angular correlation function obtained using quasars from the Sloan Digital Sky Survey (SDSS) Data Release 16. With our procedure, we recover a similar level of correlation between DESI quasar targets and SDSS quasars in two-thirds of the total footprint and we show that the excess of correlation in the remaining area is due to a stellar contamination that should be removed with DESI spectroscopic data. We derive the Limber parameters in our three imaging regions and compare them to previous measurements from SDSS and the 2dF QSO Redshift Survey.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating Random Forest Classification on GPU and FPGA

Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification. In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that for high accuracy targets, our GPU implementation yields 5-9x speedup over CSR, and up to a 2x speedup over cuML.

FPGA, Xilinx FPGA, GPU, Random Forest classificati↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests

Female reproductive disorders (FRDs) are common health conditions that may present with significant symptoms. Diet and environment are potential areas for FRD interventions. We utilized a knowledge graph (KG) method to predict factors associated with common FRDs (for example, endometriosis, ovarian cyst, and uterine fibroids). We harmonized survey data from the Personalized Environment and Genes Study (PEGS) on internal and external environmental exposures and health conditions with biomedical ontology content. We merged the harmonized data and ontologies with supplemental nutrient and agricultural chemical data to create a KG. We analyzed the KG by embedding edges and applying a random forest for edge prediction to identify variables potentially associated with FRDs. We also conducted logistic regression analysis for comparison. Across 9765 PEGS respondents, the KG analysis resulted in 8535 significant or suggestive predicted links between FRDs and chemicals, phenotypes, and diseases. Amongst these links, 32 were exact matches when compared with the logistic regression results, including comorbidities, medications, foods, and occupational exposures. Mechanistic underpinnings of predicted links documented in the literature may support some of our findings. Our KG methods are useful for predicting possible associations in large, survey-based datasets with added information on directionality and magnitude of effect from logistic regression. These results should not be construed as causal but can support hypothesis generation. This investigation enabled the generation of hypotheses on a variety of potential links between FRDs and exposures. Future investigations should prospectively evaluate the variables hypothesized to impact FRDs.

60 APPLIED LIFE SCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

A Machine Learning Initializer for Newton-Raphson AC Power Flow Convergence

Power flow computations are fundamental to many power system studies. Obtaining a converged power flow case is not a trivial task especially in large power grids due to the non-linear nature of the power flow equations. One key challenge is that the widely used Newton based power flow methods are sensitive to the initial voltage magnitude and angle estimates, and a bad initial estimate would lead to non-convergence. This paper addresses this challenge by developing a random-forest (RF) machine learning model to provide better initial voltage magnitude and angle estimates towards achieving power flow convergence. This method was implemented on a real ERCOT 6102 bus system under various operating conditions. By providing better Newton-Raphson initialization, the RF model precipitated the solution of 2,106 cases out of 3,899 non-converging dispatches. These cases could not be solved from flat start or by initialization with the voltage solution of a reference case. Finally, results obtained from the RF initializer performed better when compared with DC power flow initialization, Linear regression, and Decision Trees.

random forest↗