Engineering PapersSearch

SEARCH · Engineering Papers

Results for “random forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Selection of high-redshift Lyman-Break Galaxies from broadband and wide photometric surveys

Here, in this paper, we investigate the possibility of selecting high-redshift Lyman-Break Galaxies (LBG) using current and future broadband wide photometric surveys, such as the Ultraviolet Near Infrared Optical Northern Survey (UNIONS) or the Vera C. Rubin Legacy Survey of Space and Time (LSST), using a Random Forest algorithm. This work is conducted in the context of future large-scale structure spectroscopic surveys like DESI-II, the next phase of the Dark Energy Spectroscopic Instrument (DESI), which will start around 2029. We use deep imaging data from the Hyper Suprime Camera (HSC) and the Canada-France-Hawaii Telescope Large Area U-band Deep Survey (CLAUDS) on the COSMOS and XMM-LSS fields. To predict the selection performance of LBGs with image quality similar to UNIONS, we degrade the u,g,r,i and z bands to UNIONS depth. The Random Forest algorithm is trained with the u,g,r,i and z bands to classify LBGs in the 2.5 < z < 3.5 range. We find that fixing a target density budget of 1,100 deg -2 , the Random Forest approach gives a density of z > 2 targets of 873 deg -2 , and a density of 493 deg -2 of confirmed LBGs after spectroscopic confirmation with DESI. This UNIONS-like selection was tested in a dedicated spectroscopic observation campaign of 1,000 targets with DESI on the COSMOS field, providing a safe spectroscopic sample with a mean redshift of 3. This sample is used to derive forecasts for DESI-II, assuming a sky coverage of 5,000 deg 2 . We predict uncertainties on Alcock-Paczynski parameters α ⊥ and α ∥ to be 0.7% and 1% for 2.6 < z < 3.2, resulting in a potential 2% measurement of the dark energy fraction at high redshift. Additionally, we estimate the uncertainty in local non-Gaussianity and predict σ fNL ≈ 7, which would be comparable to the current best precision achieved by Planck. The latter forecast suggests that achieving the precision required to place stringent constraints on inflationary models (σ fNL ≈ 1) using spectroscopic galaxy surveys necessitates the development of a next-generation (Stage V) spectroscopic survey.

cosmological parameters from LSS

Background subtraction in inelastic scattering measurements using machine learning

Identifying, isolating, and subtracting background from the signal of interest is vital for nuclear physics experiments. These backgrounds introduce unwanted uncertainties that must be accounted for properly to extract accurate results from the signals. In nuclear reaction measurements, the typical contaminants are carbon and oxygen, contributing to background signals, and complicating the measurement of the light ejectiles. For instance, in the inelastic scattering measurement of a 20.9-MeV proton beam on 96 Mo, the 96 Mo target was contaminated with carbon and oxygen. Here, we used random forest, a machine learning algorithm commonly used for classification and regression tasks, to separate the inelastic scattering on the carbon and oxygen contaminants from the data of interest resulting from 96 Mo(p, p').

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State

Power System Feature-Based Event Classification by Means of Multiple PMU Data

Abstract—Phasor Measurement Units (PMUs) provide time synchronized measurements across the power grid, enabling data driven event detection and classification for enhanced system monitoring and situational awareness. However, variations in event duration, spatial extent, and severity, along with coincident events, pose challenges for conventional classification models that require fixed-size inputs. This paper presents a feature-based framework that aggregates diverse attributes from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, and Multilayer Perceptron. A probabilistic post-processing scheme is further introduced to enable multi-label classification in the presence of overlapping events. Experiments using real-world PMU data demonstrate that the Random Forest model achieves 95% accuracy, while the proposed post-processing method yields an additional 3% improvement.

Nematirad, Reza

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]

Automated ICRF heating surrogate modeling via machine learning

This work introduces automated machine learning workflows that address critical bottlenecks in surrogate model development for Ion Cyclotron Range of Frequencies (ICRF) heating applications. The automated framework includes data analysis tools that transform raw datasets into actionable insights in seconds, replacing weeks of manual exploratory effort and ensuring consistent, reproducible dataset characterization. By integrating advanced hyperparameter optimization (HPO) methods including Bayesian optimization via BoTorch and Tree-structured Parzen Estimators (TPE), the framework significantly reduces model development time from weeks to hours, decreasing computational cost and required expertise, while enabling high-accuracy surrogate models. Compared to traditional hyperparameter scanning (HPS) techniques such as methodical, randomized, and grid searches, HPO methods achieve superior convergence and predictive performance, even when compared to already well-tuned reference models. On NSTX High Harmonic Fast Wave (HHFW) heating datasets, both Random Forest Regressor (RFR) and neural network surrogates demonstrate improved accuracy, achieving R 2 values beyond 0.97 and 0.98, respectively. The results show that while HPO gains are modest for robust architectures like RFR, they become essential for more sensitive models such as neural networks, highlighting the trade-offs across optimization strategies. Through automated workflows that eliminate manual hyperparameter tuning and require minimal ML expertise, this work enables widespread adoption of high-fidelity surrogate models across the fusion community for real-time plasma control, uncertainty quantification, rapid experimental scenario development, and integrated system optimization.

Sanchez-Villar, Alvaro [Princeton Plasma Physics L

Digital twin framework for PIP-II linac: AI-driven multi-scale modeling from ion source to 800 MeV

The PIP-II superconducting linac at Fermilab is designed to deliver multi-megawatt proton beams for neutrino physics and other high-intensity applications. To expedite commissioning and enhance operational reliability, we have developed an EPICS-based data flow framework that seamlessly integrates digital twins (DT) with physical twins (PT). These digital twins comprise high-fidelity beam dynamics models or data-driven surrogate models connected to their physical counterparts through real-time diagnostics and advanced machine-learning algorithms.Central to this framework is Linac_Gen, an accelerated simulation tool that incorporates convolutional neural networks, random forests, and genetic algorithms to provide up to a tenfold speedup in optimizing the accelerator geometry model. An EPICS translator layer ensures interoperability by efficiently mapping lattice parameters across diverse simulation platforms.Our EPICS-based framework supports multiple operational modes—monitoring, passive learning, closed-loop control, and online learning—covering the entire machine lifecycle. By leveraging HPC resources and multi-objective optimization techniques, the digital twin enables adaptive trajectory correction, real-time fault detection, and predictive modeling of beam stability. This comprehensive approach paves the way for robust, high-intensity operation and data-driven accelerator R&D at Fermilab.

Pathak, Abhishek [Fermilab]

Environmental Factors Associated With Fall Phytoplankton Blooms in the Northern Bering and Chukchi Seas

This study investigates environmental drivers of fall phytoplankton blooms in the Arctic, focusing on the northern Bering and Chukchi seas. Random Forests models were used to analyze covariates of fall phytoplankton blooms from 2013 to 2018, incorporating shipboard, remote sensing, and modeled environmental properties. Four regional models and one comprehensive all-station model considered fall as well as midsummer conditions. Midsummer properties included suspended particulate matter, chlorophyll-a, and the proportion of degraded pheophytin to chlorophyll-a used as a proxy for bloom stage. Open water duration was one of the highest ranked factors in predicting fall blooms. Open water duration also influences the stage of midsummer (July) blooms as indicated by pheophytin proportions, which in turn were the highest-ranked factor for predicting fall bloom events in the Chirikov Basin (northern Bering Sea between St. Lawrence Island and the Bering Strait) and the Chukchi Sea. Wind direction, specifically easterly winds, was an important predictor in the northern Bering Sea. Maximum wind speed ranked highly at stations located within the nutrient-poor Alaska Coastal Current in the Chukchi Sea. However, stormy days, average and maximum wind speeds generally ranked low in importance as a predictor of fall bloom events. Other parameters, including photosynthetic active radiation, modeled nutrient concentrations, mixed layer depth, and time since sea ice breakup date showed strong but regionally varying relationships with fall blooms. Altogether, results from these Random Forests models suggest that high wind events and storms in the absence of sea ice provide an incomplete narrative for initiating fall bloom events.

Gaffey, C. B. [Clark University, Worcester, MA (Un

Anomaly Identification of Synchronized Voltage Waveform for Situational Awareness of Low Inertia Systems

Inverter-based resources (IBRs) such as photovoltaics (PVs), wind turbines, and battery energy storage systems (BESSs) are widely deployed in low-carbon power systems. However, these resources typically do not provide the inertia needed for grid stability, resulting in a low-inertia power system. IBRs and lack of inertia have been known to cause anomalies such as waveform distortions and wideband oscillations in power systems due to the limited inertia level, leading to increased generation trips and load shedding. Here, to achieve effective anomaly identification, this paper proposes a synchro-waveform-based algorithm utilizing real-time synchronized voltage waveform measurements from waveform measurement units (WMUs). In the proposed method, different physical characteristics, as well as statistical features, are extracted from synchronized voltage waveform measurements to filter anomalies. Then, the anomaly identification approach based on the random forest is developed and deployed into the FNET/GridEye system considering trade-offs among accuracy, computational burden, and deployment cost. Moreover, four WMUs are specially designed and deployed on Kauai Island to receive instantaneous synchronized voltage waveform measurements. To verify the performance of the proposed algorithm, different experiments are carried out with collected field test data. The result demonstrates that the performance of the proposed synchro-waveform-based anomaly categorization algorithm can accurately identify anomalies 95.35% of the time, which has comparable performance among benchmarking algorithms.

Situational awareness