Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Fair Bagging Boosting Models [SWR-24-38]

Fair Bagging Boosting Models is a software implementation of a framework for building, measuring bias and correcting bias in 3 popular forest machine learning models: gradient boosted trees (GBT), random forest (RF), and XGBoost models, using the XGBoost library. The framework takes advantage of the flexibility in XGBoost library to represent gradient boosted tree and random forest models, as well as the ability to use custom loss function.

Ugirumurera, Juliette↗

Selection of high-redshift Lyman-Break Galaxies from broadband and wide photometric surveys

In this paper, we investigate the possibility of selecting high-redshift Lyman-Break Galaxies (LBG) using current and future broadband wide photometric surveys, such as the Ultraviolet Near Infrared Optical Northern Survey (UNIONS) or the Vera C. Rubin Legacy Survey of Space and Time (LSST), using a Random Forest algorithm. This work is conducted in the context of future large-scale structure spectroscopic surveys like DESI-II, the next phase of the Dark Energy Spectroscopic Instrument (DESI), which will start around 2029.We use deep imaging data from the Hyper Suprime Camera (HSC) and the Canada-France-Hawaii Telescope Large Area U-band Deep Survey (CLAUDS) on the COSMOS and XMM-LSS fields. To predict the selection performance of LBGs with image quality similar to UNIONS, we degrade the u,g,r,i and z bands to UNIONS depth.The Random Forest algorithm is trained with the u,g,r,i and z bands to classify LBGs in the 2.5 < z < 3.5 range.We find that fixing a target density budget of 1,100 deg$^{-2}$, the Random Forest approach gives a density of z > 2 targets of 873 deg$^{-2}$, and a density of 493 deg$^{-2}$ of confirmed LBGs after spectroscopic confirmation with DESI. This UNIONS-like selection was tested in a dedicated spectroscopic observation campaign of 1,000 targets with DESI on the COSMOS field, providing a safe spectroscopic sample with a mean redshift of 3. This sample is used to derive forecasts for DESI-II, assuming a sky coverage of 5,000 deg$^{2}$. We predict uncertainties on Alcock-Paczynski parameters α$_{⊥}$ and α$_{∥}$ to be 0.7% and 1% for 2.6 < z < 3.2, resulting in a potential 2% measurement of the dark energy fraction at high redshift. Additionally, we estimate the uncertainty in local non-Gaussianity and predict σ$_{fNL}$ ≈ 7, which would be comparable to the current best precision achieved by Planck. The latter forecast suggests that achieving the precision required to place stringent constraints on inflationary models (σ$_{fNL}$ ≈ 1) using spectroscopic galaxy surveys necessitates the development of a next-generation (Stage V) spectroscopic survey.

79 ASTRONOMY AND ASTROPHYSICS↗

Selection of high-redshift Lyman-Break Galaxies from broadband and wide photometric surveys

Here, in this paper, we investigate the possibility of selecting high-redshift Lyman-Break Galaxies (LBG) using current and future broadband wide photometric surveys, such as the Ultraviolet Near Infrared Optical Northern Survey (UNIONS) or the Vera C. Rubin Legacy Survey of Space and Time (LSST), using a Random Forest algorithm. This work is conducted in the context of future large-scale structure spectroscopic surveys like DESI-II, the next phase of the Dark Energy Spectroscopic Instrument (DESI), which will start around 2029. We use deep imaging data from the Hyper Suprime Camera (HSC) and the Canada-France-Hawaii Telescope Large Area U-band Deep Survey (CLAUDS) on the COSMOS and XMM-LSS fields. To predict the selection performance of LBGs with image quality similar to UNIONS, we degrade the u,g,r,i and z bands to UNIONS depth. The Random Forest algorithm is trained with the u,g,r,i and z bands to classify LBGs in the 2.5 < z < 3.5 range. We find that fixing a target density budget of 1,100 deg -2 , the Random Forest approach gives a density of z > 2 targets of 873 deg -2 , and a density of 493 deg -2 of confirmed LBGs after spectroscopic confirmation with DESI. This UNIONS-like selection was tested in a dedicated spectroscopic observation campaign of 1,000 targets with DESI on the COSMOS field, providing a safe spectroscopic sample with a mean redshift of 3. This sample is used to derive forecasts for DESI-II, assuming a sky coverage of 5,000 deg 2 . We predict uncertainties on Alcock-Paczynski parameters α ⊥ and α ∥ to be 0.7% and 1% for 2.6 < z < 3.2, resulting in a potential 2% measurement of the dark energy fraction at high redshift. Additionally, we estimate the uncertainty in local non-Gaussianity and predict σ fNL ≈ 7, which would be comparable to the current best precision achieved by Planck. The latter forecast suggests that achieving the precision required to place stringent constraints on inflationary models (σ fNL ≈ 1) using spectroscopic galaxy surveys necessitates the development of a next-generation (Stage V) spectroscopic survey.

cosmological parameters from LSS↗

Estimating Compressional Velocity and Bulk Density Logs in Marine Gas Hydrates Using Machine Learning

Compressional velocity (Vp) and bulk density (ρb) logs are essential for characterizing gas hydrates and near-seafloor sediments; however, it is sometimes difficult to acquire these logs due to poor borehole conditions, safety concerns, or cost-related issues. We present a machine learning approach to predict either compressional Vp or ρb logs with high accuracy and low error in near-seafloor sediments within water-saturated intervals, in intervals where hydrate fills fractures, and intervals where hydrate occupies the primary pore space. We use scientific-quality logging-while-drilling well logs, gamma ray, ρb, Vp, and resistivity to train the machine learning model to predict Vp or ρb logs. Of the six machine learning algorithms tested (multilinear regression, polynomial regression, polynomial regression with ridge regularization, K nearest neighbors, random forest, and multilayer perceptron), we find that the random forest and K nearest neighbors algorithms are best suited to predicting Vp and ρb logs based on coefficients of determination (R2) greater than 70% and mean absolute percentage errors less than 4%. Given the high accuracy and low error results for Vp and ρb prediction in both hydrate and water-saturated sediments, we argue that our model can be applied in most LWD wells to predict Vp or ρb logs in near-seafloor siliciclastic sediments on continental slopes irrespective of the presence or absence of gas hydrate.

Naim, Fawz↗

Classifying Forest Type in the National Forest Inventory Context with Airborne Hyperspectral and Lidar Data

Forest structure and composition regulate a range of ecosystem services, including biodiversity, water and nutrient cycling, and wood volume for resource extraction. Forest type is an important metric measured in the US Forest Service Forest Inventory and Analysis (FIA) program, the national forest inventory of the USA. Forest type information can be used to quantify carbon and other forest resources within specific domains to support ecological analysis and forest management decisions, such as managing for disease and pests. In this study, we developed a methodology that uses a combination of airborne hyperspectral and lidar data to map FIA-defined forest type between sparsely sampled FIA plot data collected in interior Alaska. To determine the best classification algorithm and remote sensing data for this task, five classification algorithms were tested with six different combinations of raw hyperspectral data, hyperspectral vegetation indices, and lidar-derived canopy and topography metrics. Models were trained using forest type information from 632 FIA subplots collected in interior Alaska. Of the thirty model and input combinations tested, the random forest classification algorithm with hyperspectral vegetation indices and lidar-derived topography and canopy height metrics had the highest accuracy (78% overall accuracy). This study supports random forest as a powerful classifier for natural resource data. It also demonstrates the benefits from combining both structural (lidar) and spectral (imagery) data for forest type classification.

random forest↗

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of the Planetary Boundary Layer Height From ERA5 Reanalysis With MOSAiC Observations Over the Arctic Ocean

The planetary boundary layer height (PBLH) is a crucial indicator reflecting the region of the atmosphere characterized by continuous turbulence. Here, we use radiosonde and surface meteorological observations (4–7 times per day, year-round measurements) during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition to derive the PBLH (PBLH MOSAiC ), and further evaluate the PBLH from the ERA5 reanalysis (PBLH ERA5 ). Comparisons between PBLH MOSAiC and PBLH ERA5 from different perspectives reveal that: (a) The overestimation of PBLH ERA5 when the sea ice concentration is >90% is significant with the centered root mean squared error reaching up to 201 m; (b) The difference between the two products is notably pronounced in cold seasons, while it is comparatively diminished in warm seasons; (c) In neutral boundary layers, differences in PBLH ERA5 are larger compared with stable and convective boundary layers. In addition, the analysis of error sources indicates that the bias of PBLH ERA5 is sensitive to the bias of vertical thermal structure and wind speed profiles in ERA5 data sets in all conditions. Finally, we find a Random Forest model effectively reduces the bias of PBLH ERA5 with the index of agreement reaching up to 0.71 in the test data set, while a multiple linear regression demonstrates comparable performance to the Random Forest model.

54 ENVIRONMENTAL SCIENCES↗

Effects of forest structural and compositional change on forest microclimates across a gradient of disturbance severity

Forest structural diversity and community composition are key in regulating forest microclimates. When disturbance affects structural diversity or composition, forest microclimates may be altered due to changes in soil temperature, soil water content, and light availability. It is unclear however which structural or compositional components, when changed or to what extent, result in microclimatic change. To address this question, we used data from a large scale, manipulative stem-girdling experiment in northern, lower Michigan—the Forest Resilience and Threshold Experiment (FoRTE). FoRTE follows a factorial design with multiple levels of disturbance severity (0, 45, 65, 85%) based on targeted reductions in gross leaf area index via stem-girdling induced mortality. These disturbance severity treatments are applied in two ways: either as top-down (largest trees are killed) or bottom-up (small to medium trees killed) treatments. We examined how multiple components of structural diversity and community composition changed as a product of disturbance severity and type, and then tested for resulting effects on forest microclimates (light availability, soil temperature, and soil water), using a multivariate, Random Forest framework. We found that measures of community composition (species richness, species evenness, and Shannon-Wiener Diversity Index) and stand structure (basal area, standard deviation of DBH, tree size diversity) declined more following disturbance than did measures of canopy cover, heterogeneity, arrangement, or height. However, when changes in each variable from pre- to post-disturbance, measured as log change, were employed in a multivariate, Random Forest regression framework, structural diversity measures of heterogeneity (rugosity, top rugosity), cover (canopy cover), and arrangement (porosity) were the most influential variables, but with differences among bottom-up and top-down treatments We found that the death of large trees from disturbance impacts soil temperature, water, and light environments more substantially and uniformly across disturbance gradients than does the death of smaller trees. Furthermore, our results have implications for both statistical and process-based modeling of forest disturbance.

54 ENVIRONMENTAL SCIENCES↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

“Thought I’d Share First” and Other Conspiracy Theory Tweets from the COVID-19 Infodemic: Exploratory Study

Background: The COVID-19 outbreak has left many people isolated within their homes; these people are turning to social media for news and social connection, which leaves them vulnerable to believing and sharing misinformation. Health-related misinformation threatens adherence to public health messaging, and monitoring its spread on social media is critical to understanding the evolution of ideas that have potentially negative public health impacts. Objective: The aim of this study is to use Twitter data to explore methods to characterize and classify four COVID-19 conspiracy theories and to provide context for each of these conspiracy theories through the first 5 months of the pandemic. Methods: We began with a corpus of COVID-19 tweets (approximately 120 million) spanning late January to early May 2020. We first filtered tweets using regular expressions (n=1.8 million) and used random forest classification models to identify tweets related to four conspiracy theories. Our classified data sets were then used in downstream sentiment analysis and dynamic topic modeling to characterize the linguistic features of COVID-19 conspiracy theories as they evolve over time. Results: Analysis using model-labeled data was beneficial for increasing the proportion of data matching misinformation indicators. Random forest classifier metrics varied across the four conspiracy theories considered (F1 scores between 0.347 and 0.857); this performance increased as the given conspiracy theory was more narrowly defined. We showed that misinformation tweets demonstrate more negative sentiment when compared to non-misinformation tweets and that theories evolve over time, incorporating details from unrelated conspiracy theories as well as real-world events. Conclusions: Although we focus here on health-related misinformation, this combination of approaches is not specific to public health and is valuable for characterizing misinformation in general, which is an important first step in creating targeted messaging to counteract its spread. Initial messaging should aim to preempt generalized misinformation before it becomes widespread, while later messaging will

5g↗

Application of Machine Learning Techniques to an Agent-Based Model of Pantoea

Agent-based modeling (ABM) is a powerful simulation technique which describes a complex dynamic system based on its interacting constituent entities. While the flexibility of ABM enables broad application, the complexity of real-world models demands intensive computing resources and computational time; however, a metamodel may be constructed to gain insight at less computational expense. Here, we developed a model in NetLogo to describe the growth of a microbial population consisting of Pantoea . We applied 13 parameters that defined the model and actively changed seven of the parameters to modulate the evolution of the population curve in response to these changes. We efficiently performed more than 3,000 simulations using a Python wrapper, NL4Py . Upon evaluation of the correlation between the active parameters and outputs by random forest regression, we found that the parameters which define the depth of medium and glucose concentration affect the population curves significantly. Subsequently, we constructed a metamodel, a dense neural network, to predict the simulation outputs from the active parameters and found that it achieves high prediction accuracy, reaching an R 2 coefficient of determination value up to 0.92. Our approach of using a combination of ABM with random forest regression and neural network reduces the number of required ABM simulations. The simplified and refined metamodels may provide insights into the complex dynamic system before their transition to more sophisticated models that run on high-performance computing systems. The ultimate goal is to build a bridge between simulation and experiment, allowing model validation by comparing the simulated data to experimental data in microbiology.

59 BASIC BIOLOGICAL SCIENCES↗

On the discovery of stars, quasars, and galaxies in the Southern Hemisphere with S-PLUS DR2

ABSTRACT This paper provides a catalogue of stars, quasars, and galaxies for the Southern Photometric Local Universe Survey Data Release 2 (S-PLUS DR2) in the Stripe 82 region. We show that a 12-band filter system (5 Sloan-like and 7 narrow bands) allows better performance for object classification than the usual analysis based solely on broad bands (regardless of infrared information). Moreover, we show that our classification is robust against missing values. Using spectroscopically confirmed sources retrieved from the Sloan Digital Sky Survey DR16 and DR14Q, we train a random forest classifier with the 12 S-PLUS magnitudes + 4 morphological features. A second random forest classifier is trained with the addition of the W1 (3.4 $\mu\mathrm{m} $) and W2 (4.6 $\mu\mathrm{m} $) magnitudes from the Wide-field Infrared Survey Explorer (WISE). Forty-four per cent of our catalogue have WISE counterparts and are provided with classification from both models. We achieve 95.76 per cent (52.47 per cent) of quasar purity, 95.88 per cent (92.24 per cent) of quasar completeness, 99.44 per cent (98.17 per cent) of star purity, 98.22 per cent (78.56 per cent) of star completeness, 98.04 per cent (81.39 per cent) of galaxy purity, and 98.8 per cent (85.37 per cent) of galaxy completeness for the first (second) classifier, for which the metrics were calculated on objects with (without) WISE counterpart. A total of 2926 787 objects that are not in our spectroscopic sample were labelled, obtaining 335 956 quasars, 1347 340 stars, and 1243 391 galaxies. From those, 7.4 per cent, 76.0 per cent, and 58.4 per cent were classified with probabilities above 80 per cent. The catalogue with classification and probabilities for Stripe 82 S-PLUS DR2 is available for download.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Frost prediction using machine learning and deep neural network models

This study describes accurate, computationally efficient models that can be implemented for practical use in predicting frost events for point-scale agricultural applications. Frost damage in agriculture is a costly burden to farmers and global food security alike. Timely prediction of frost events is important to reduce the cost of agricultural frost damage and traditional numerical weather forecasts are often inaccurate at the field-scale in complex terrain. In this paper, we developed machine learning (ML) algorithms for the prediction of such frost events near Alcalde, NM at the point-scale. ML algorithms investigated include deep neural network, convolution neural networks, and random forest models at lead-times of 6–48 h. Our results show promising accuracy (6-h prediction RMSE = 1.53–1.72°C) for use in frost and minimum temperature prediction applications. Seasonal differences in model predictions resulted in a slight negative bias during Spring and Summer months and a positive bias in Fall and Winter months. Additionally, we tested the model transferability by continuing training and testing using data from sensors at a nearby farm. We calculated the feature importance of the random forest models and were able to determine which parameters provided the models with the most useful information for predictions. We determined that soil temperature is a key parameter in longer term predictions (>24 h), while other temperature related parameters provide the majority of information for shorter term predictions. The model error compared favorable to previous ML based frost studies and outperformed the physically based High Resolution Rapid Refresh forecasting system making our ML-models attractive for deployment toward real-time monitoring of frost events and damage at commercial farming operations.

97 MATHEMATICS AND COMPUTING↗