Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Code Description for "Brief Communication: Monitoring snow depth using small, cheap, and easy-to-deploy ground surface temperature sensors"

Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We train a random forest machine learning model to predict snow depth from variability in ground surface temperature. To our knowledge, this is the first time that small ground surface temperature sensors have been used to estimate snow depth. The model performs well at sites where the model was trained and at pan-arctic evaluation sites (RMSE <= 0.15 m). Small temperature sensors are cheap and easy-to-deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring to an extent previously infeasible. The model is flexible and can be applied to datasets retroactively to retrieve snow depth estimates at additional sites. This code package includes a *.joblib file of the trained random forest model and a *.ipynb file showing how to clean input data, train the random forest model, and apply the model.

Bachand, Claire↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

Utilizing physics-based input features within a machine learning model to predict wind speed forecasting error

Machine learning is quickly becoming a commonly used technique for wind speed and power forecasting. Many machine learning methods utilize exogenous variables as input features, but there remains the question of which atmospheric variables are most beneficial for forecasting, especially in handling non-linearities that lead to forecasting error. This question is addressed via creation of a hybrid model that utilizes an autoregressive integrated moving-average (ARIMA) model to make an initial wind speed forecast followed by a random forest model that attempts to predict the ARIMA forecasting error using knowledge of exogenous atmospheric variables. Variables conveying information about atmospheric stability and turbulence as well as inertial forcing are found to be useful in dealing with non-linear error prediction. Streamwise wind speed, time of day, turbulence intensity, turbulent heat flux, vertical velocity, and wind direction are found to be particularly useful when used in unison for hourly and 3 h timescales. The prediction accuracy of the developed ARIMA–random forest hybrid model is compared to that of the persistence and bias-corrected ARIMA models. The ARIMA–random forest model is shown to improve upon the latter commonly employed modeling methods, reducing hourly forecasting error by up to 5 % below that of the bias-corrected ARIMA model and achieving an R 2 value of 0.84 with true wind speed.

17 WIND ENERGY↗

Fair Bagging Boosting Models [SWR-24-38]

Fair Bagging Boosting Models is a software implementation of a framework for building, measuring bias and correcting bias in 3 popular forest machine learning models: gradient boosted trees (GBT), random forest (RF), and XGBoost models, using the XGBoost library. The framework takes advantage of the flexibility in XGBoost library to represent gradient boosted tree and random forest models, as well as the ability to use custom loss function.

Ugirumurera, Juliette↗

Selection of high-redshift Lyman-Break Galaxies from broadband and wide photometric surveys

In this paper, we investigate the possibility of selecting high-redshift Lyman-Break Galaxies (LBG) using current and future broadband wide photometric surveys, such as the Ultraviolet Near Infrared Optical Northern Survey (UNIONS) or the Vera C. Rubin Legacy Survey of Space and Time (LSST), using a Random Forest algorithm. This work is conducted in the context of future large-scale structure spectroscopic surveys like DESI-II, the next phase of the Dark Energy Spectroscopic Instrument (DESI), which will start around 2029.We use deep imaging data from the Hyper Suprime Camera (HSC) and the Canada-France-Hawaii Telescope Large Area U-band Deep Survey (CLAUDS) on the COSMOS and XMM-LSS fields. To predict the selection performance of LBGs with image quality similar to UNIONS, we degrade the u,g,r,i and z bands to UNIONS depth.The Random Forest algorithm is trained with the u,g,r,i and z bands to classify LBGs in the 2.5 < z < 3.5 range.We find that fixing a target density budget of 1,100 deg$^{-2}$, the Random Forest approach gives a density of z > 2 targets of 873 deg$^{-2}$, and a density of 493 deg$^{-2}$ of confirmed LBGs after spectroscopic confirmation with DESI. This UNIONS-like selection was tested in a dedicated spectroscopic observation campaign of 1,000 targets with DESI on the COSMOS field, providing a safe spectroscopic sample with a mean redshift of 3. This sample is used to derive forecasts for DESI-II, assuming a sky coverage of 5,000 deg$^{2}$. We predict uncertainties on Alcock-Paczynski parameters α$_{⊥}$ and α$_{∥}$ to be 0.7% and 1% for 2.6 < z < 3.2, resulting in a potential 2% measurement of the dark energy fraction at high redshift. Additionally, we estimate the uncertainty in local non-Gaussianity and predict σ$_{fNL}$ ≈ 7, which would be comparable to the current best precision achieved by Planck. The latter forecast suggests that achieving the precision required to place stringent constraints on inflationary models (σ$_{fNL}$ ≈ 1) using spectroscopic galaxy surveys necessitates the development of a next-generation (Stage V) spectroscopic survey.

79 ASTRONOMY AND ASTROPHYSICS↗

Selection of high-redshift Lyman-Break Galaxies from broadband and wide photometric surveys

Here, in this paper, we investigate the possibility of selecting high-redshift Lyman-Break Galaxies (LBG) using current and future broadband wide photometric surveys, such as the Ultraviolet Near Infrared Optical Northern Survey (UNIONS) or the Vera C. Rubin Legacy Survey of Space and Time (LSST), using a Random Forest algorithm. This work is conducted in the context of future large-scale structure spectroscopic surveys like DESI-II, the next phase of the Dark Energy Spectroscopic Instrument (DESI), which will start around 2029. We use deep imaging data from the Hyper Suprime Camera (HSC) and the Canada-France-Hawaii Telescope Large Area U-band Deep Survey (CLAUDS) on the COSMOS and XMM-LSS fields. To predict the selection performance of LBGs with image quality similar to UNIONS, we degrade the u,g,r,i and z bands to UNIONS depth. The Random Forest algorithm is trained with the u,g,r,i and z bands to classify LBGs in the 2.5 < z < 3.5 range. We find that fixing a target density budget of 1,100 deg -2 , the Random Forest approach gives a density of z > 2 targets of 873 deg -2 , and a density of 493 deg -2 of confirmed LBGs after spectroscopic confirmation with DESI. This UNIONS-like selection was tested in a dedicated spectroscopic observation campaign of 1,000 targets with DESI on the COSMOS field, providing a safe spectroscopic sample with a mean redshift of 3. This sample is used to derive forecasts for DESI-II, assuming a sky coverage of 5,000 deg 2 . We predict uncertainties on Alcock-Paczynski parameters α ⊥ and α ∥ to be 0.7% and 1% for 2.6 < z < 3.2, resulting in a potential 2% measurement of the dark energy fraction at high redshift. Additionally, we estimate the uncertainty in local non-Gaussianity and predict σ fNL ≈ 7, which would be comparable to the current best precision achieved by Planck. The latter forecast suggests that achieving the precision required to place stringent constraints on inflationary models (σ fNL ≈ 1) using spectroscopic galaxy surveys necessitates the development of a next-generation (Stage V) spectroscopic survey.

cosmological parameters from LSS↗

Estimating Compressional Velocity and Bulk Density Logs in Marine Gas Hydrates Using Machine Learning

Compressional velocity (Vp) and bulk density (ρb) logs are essential for characterizing gas hydrates and near-seafloor sediments; however, it is sometimes difficult to acquire these logs due to poor borehole conditions, safety concerns, or cost-related issues. We present a machine learning approach to predict either compressional Vp or ρb logs with high accuracy and low error in near-seafloor sediments within water-saturated intervals, in intervals where hydrate fills fractures, and intervals where hydrate occupies the primary pore space. We use scientific-quality logging-while-drilling well logs, gamma ray, ρb, Vp, and resistivity to train the machine learning model to predict Vp or ρb logs. Of the six machine learning algorithms tested (multilinear regression, polynomial regression, polynomial regression with ridge regularization, K nearest neighbors, random forest, and multilayer perceptron), we find that the random forest and K nearest neighbors algorithms are best suited to predicting Vp and ρb logs based on coefficients of determination (R2) greater than 70% and mean absolute percentage errors less than 4%. Given the high accuracy and low error results for Vp and ρb prediction in both hydrate and water-saturated sediments, we argue that our model can be applied in most LWD wells to predict Vp or ρb logs in near-seafloor siliciclastic sediments on continental slopes irrespective of the presence or absence of gas hydrate.

Naim, Fawz↗

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of the Planetary Boundary Layer Height From ERA5 Reanalysis With MOSAiC Observations Over the Arctic Ocean

The planetary boundary layer height (PBLH) is a crucial indicator reflecting the region of the atmosphere characterized by continuous turbulence. Here, we use radiosonde and surface meteorological observations (4–7 times per day, year-round measurements) during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition to derive the PBLH (PBLH MOSAiC ), and further evaluate the PBLH from the ERA5 reanalysis (PBLH ERA5 ). Comparisons between PBLH MOSAiC and PBLH ERA5 from different perspectives reveal that: (a) The overestimation of PBLH ERA5 when the sea ice concentration is >90% is significant with the centered root mean squared error reaching up to 201 m; (b) The difference between the two products is notably pronounced in cold seasons, while it is comparatively diminished in warm seasons; (c) In neutral boundary layers, differences in PBLH ERA5 are larger compared with stable and convective boundary layers. In addition, the analysis of error sources indicates that the bias of PBLH ERA5 is sensitive to the bias of vertical thermal structure and wind speed profiles in ERA5 data sets in all conditions. Finally, we find a Random Forest model effectively reduces the bias of PBLH ERA5 with the index of agreement reaching up to 0.71 in the test data set, while a multiple linear regression demonstrates comparable performance to the Random Forest model.

54 ENVIRONMENTAL SCIENCES↗

Effects of forest structural and compositional change on forest microclimates across a gradient of disturbance severity

Forest structural diversity and community composition are key in regulating forest microclimates. When disturbance affects structural diversity or composition, forest microclimates may be altered due to changes in soil temperature, soil water content, and light availability. It is unclear however which structural or compositional components, when changed or to what extent, result in microclimatic change. To address this question, we used data from a large scale, manipulative stem-girdling experiment in northern, lower Michigan—the Forest Resilience and Threshold Experiment (FoRTE). FoRTE follows a factorial design with multiple levels of disturbance severity (0, 45, 65, 85%) based on targeted reductions in gross leaf area index via stem-girdling induced mortality. These disturbance severity treatments are applied in two ways: either as top-down (largest trees are killed) or bottom-up (small to medium trees killed) treatments. We examined how multiple components of structural diversity and community composition changed as a product of disturbance severity and type, and then tested for resulting effects on forest microclimates (light availability, soil temperature, and soil water), using a multivariate, Random Forest framework. We found that measures of community composition (species richness, species evenness, and Shannon-Wiener Diversity Index) and stand structure (basal area, standard deviation of DBH, tree size diversity) declined more following disturbance than did measures of canopy cover, heterogeneity, arrangement, or height. However, when changes in each variable from pre- to post-disturbance, measured as log change, were employed in a multivariate, Random Forest regression framework, structural diversity measures of heterogeneity (rugosity, top rugosity), cover (canopy cover), and arrangement (porosity) were the most influential variables, but with differences among bottom-up and top-down treatments We found that the death of large trees from disturbance impacts soil temperature, water, and light environments more substantially and uniformly across disturbance gradients than does the death of smaller trees. Furthermore, our results have implications for both statistical and process-based modeling of forest disturbance.

54 ENVIRONMENTAL SCIENCES↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

“Thought I’d Share First” and Other Conspiracy Theory Tweets from the COVID-19 Infodemic: Exploratory Study

Background: The COVID-19 outbreak has left many people isolated within their homes; these people are turning to social media for news and social connection, which leaves them vulnerable to believing and sharing misinformation. Health-related misinformation threatens adherence to public health messaging, and monitoring its spread on social media is critical to understanding the evolution of ideas that have potentially negative public health impacts. Objective: The aim of this study is to use Twitter data to explore methods to characterize and classify four COVID-19 conspiracy theories and to provide context for each of these conspiracy theories through the first 5 months of the pandemic. Methods: We began with a corpus of COVID-19 tweets (approximately 120 million) spanning late January to early May 2020. We first filtered tweets using regular expressions (n=1.8 million) and used random forest classification models to identify tweets related to four conspiracy theories. Our classified data sets were then used in downstream sentiment analysis and dynamic topic modeling to characterize the linguistic features of COVID-19 conspiracy theories as they evolve over time. Results: Analysis using model-labeled data was beneficial for increasing the proportion of data matching misinformation indicators. Random forest classifier metrics varied across the four conspiracy theories considered (F1 scores between 0.347 and 0.857); this performance increased as the given conspiracy theory was more narrowly defined. We showed that misinformation tweets demonstrate more negative sentiment when compared to non-misinformation tweets and that theories evolve over time, incorporating details from unrelated conspiracy theories as well as real-world events. Conclusions: Although we focus here on health-related misinformation, this combination of approaches is not specific to public health and is valuable for characterizing misinformation in general, which is an important first step in creating targeted messaging to counteract its spread. Initial messaging should aim to preempt generalized misinformation before it becomes widespread, while later messaging will

5g↗

Application of Machine Learning Techniques to an Agent-Based Model of Pantoea

Agent-based modeling (ABM) is a powerful simulation technique which describes a complex dynamic system based on its interacting constituent entities. While the flexibility of ABM enables broad application, the complexity of real-world models demands intensive computing resources and computational time; however, a metamodel may be constructed to gain insight at less computational expense. Here, we developed a model in NetLogo to describe the growth of a microbial population consisting of Pantoea . We applied 13 parameters that defined the model and actively changed seven of the parameters to modulate the evolution of the population curve in response to these changes. We efficiently performed more than 3,000 simulations using a Python wrapper, NL4Py . Upon evaluation of the correlation between the active parameters and outputs by random forest regression, we found that the parameters which define the depth of medium and glucose concentration affect the population curves significantly. Subsequently, we constructed a metamodel, a dense neural network, to predict the simulation outputs from the active parameters and found that it achieves high prediction accuracy, reaching an R 2 coefficient of determination value up to 0.92. Our approach of using a combination of ABM with random forest regression and neural network reduces the number of required ABM simulations. The simplified and refined metamodels may provide insights into the complex dynamic system before their transition to more sophisticated models that run on high-performance computing systems. The ultimate goal is to build a bridge between simulation and experiment, allowing model validation by comparing the simulated data to experimental data in microbiology.

59 BASIC BIOLOGICAL SCIENCES↗

On the discovery of stars, quasars, and galaxies in the Southern Hemisphere with S-PLUS DR2

ABSTRACT This paper provides a catalogue of stars, quasars, and galaxies for the Southern Photometric Local Universe Survey Data Release 2 (S-PLUS DR2) in the Stripe 82 region. We show that a 12-band filter system (5 Sloan-like and 7 narrow bands) allows better performance for object classification than the usual analysis based solely on broad bands (regardless of infrared information). Moreover, we show that our classification is robust against missing values. Using spectroscopically confirmed sources retrieved from the Sloan Digital Sky Survey DR16 and DR14Q, we train a random forest classifier with the 12 S-PLUS magnitudes + 4 morphological features. A second random forest classifier is trained with the addition of the W1 (3.4 $\mu\mathrm{m} $) and W2 (4.6 $\mu\mathrm{m} $) magnitudes from the Wide-field Infrared Survey Explorer (WISE). Forty-four per cent of our catalogue have WISE counterparts and are provided with classification from both models. We achieve 95.76 per cent (52.47 per cent) of quasar purity, 95.88 per cent (92.24 per cent) of quasar completeness, 99.44 per cent (98.17 per cent) of star purity, 98.22 per cent (78.56 per cent) of star completeness, 98.04 per cent (81.39 per cent) of galaxy purity, and 98.8 per cent (85.37 per cent) of galaxy completeness for the first (second) classifier, for which the metrics were calculated on objects with (without) WISE counterpart. A total of 2926 787 objects that are not in our spectroscopic sample were labelled, obtaining 335 956 quasars, 1347 340 stars, and 1243 391 galaxies. From those, 7.4 per cent, 76.0 per cent, and 58.4 per cent were classified with probabilities above 80 per cent. The catalogue with classification and probabilities for Stripe 82 S-PLUS DR2 is available for download.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗