Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Model and remote-sensing-guided experimental design and hypothesis generation for monitoring snow-soil–plant interactions

In this study, we develop a machine-learning (ML)-enabled strategy for selecting hillslope-scale ecohydrological monitoring sites within snow-dominated mountainous watersheds, with a particular focus on snow-soil–plant interactions. Data layers rely on spatial data layers from both remote sensing and hydrological model simulations. Specifically, a Landsat-based foresummer drought sensitivity index is used to define the dependency of the annual peak plant productivity on the Palmer drought severity index in the early growing season. Hydrological simulations provide the spatiotemporal dynamics of near-surface soil moisture and snow depth. In this framework, a regression analysis identifies the key hydrological variables relevant to the spatial heterogeneity of drought sensitivity. We then apply unsupervised clustering to these key variables, using the Gaussian mixture model, to group hillslopes into several zones that have divergent relationships regarding soil moisture, snow dynamics, and drought sensitivity. Using the datasets collected in the East River Watershed (Crested Butte, Colorado, United States), results show that drought sensitivity is significantly correlated with model-derived soil moisture and snow-free timing over space and time. The relationship is, however, non-linear, such that the correlation decreases above a threshold elevation and in a heavy snow year due to large snowpacks, lateral flow, and soil storage limitations. Clustering is then able to define the zones that have high or low sensitivity to drought, as well as the mid-elevation regions where sensitivity is associated with the topographic aspect and net potential radiation. In addition, the algorithm identifies the most representative hillslopes with road/trail access within each zone for installing monitoring sites. Our method also aims to significantly increase the use of ML and model-simulation results to guide critical zone and watershed monitoring activities.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Downscaling of SoilMERGE in the United States Southern Great Plains

SoilMERGE (SMERGE) is a root-zone soil moisture (RZSM) product that covers the entire continental United States and spans 1978 to 2019. Machine learning techniques, Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Gradient Boost (GBoost) downscaled SMERGE to spatial resolutions straddling the field scale domain (100 to 3000 m). Study area was northern Oklahoma and southern Kansas. The coarse resolution of SMERGE (0.125 degree) limits this product’s utility. To validate downscaled results in situ data from four sources were used that included: United States Department of Energy Atmospheric Radiation Measurement (ARM) observatory, United States Climate Reference Network (USCRN), Soil Climate Analysis Network (SCAN), and Soil moisture Sensing Controller and oPtimal Estimator (SoilSCAPE). In addition, RZSM retrievals from NASA’s Airborne Microwave Observatory of Subcanopy and Surface (AirMOSS) campaign provided a nearly spatially continuous comparison. Three periods were examined: era 1 (2016 to 2019), era 2 (2012 to 2015), and era 3 (2003 to 2007). During eras 1 and 2, RF outperformed XGBoost and GBoost, whereas during era 3 no model dominated. Performance was better during eras 1 and 2 as opposed to the pre-L band era 3. Improvements across all eras, regions, and models realized from downscaling included an increase in correlation from 0.03 to 0.42 and a decrease in ub RMSE from -0.0005 to -0.0118 m 3 /m 3 . This study demonstrates the feasibility of SMERGE downscaling opening the prospect for the development of a long-term RZSM dataset at a more desirable field-scale resolution with the potential to support diverse hydrometeorological and agricultural applications.

54 ENVIRONMENTAL SCIENCES↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Chatter detection in simulated machining data: a simple refined approach to vibration data

Vibration monitoring is a critical aspect of assessing the health and performance of machinery and industrial processes. This study explores the application of machine learning techniques, specifically the Random Forest (RF) classification model, to predict and classify chatter—a detrimental self-excited vibration phenomenon—during machining operations. While sophisticated methods have been employed to address chatter, this research investigates the efficacy of a novel approach to an RF model. The study leverages simulated vibration data, bypassing resource-intensive real-world data collection, to develop a versatile chatter detection model applicable across diverse machining configurations. The feature extraction process combines time-series features and Fast Fourier Transform (FFT) data features, streamlining the model while addressing challenges posed by feature selection. By focusing on the RF model’s simplicity and efficiency, this research advances chatter detection techniques, offering a practical tool with improved generalizability, computational efficiency, and ease of interpretation. The study demonstrates that innovation can reside in simplicity, opening avenues for wider applicability and accelerated progress in the machining industry.

42 ENGINEERING↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Deep Learning for In-Situ Layer Quality Monitoring during Laser-Based Directed Energy Deposition (LB-DED) Additive Manufacturing Process

Defects are a leading issue for the rejection of parts manufactured through the Directed Energy Deposition (DED) Additive Manufacturing (AM) process. In an attempt to illuminate and advance in situ quality monitoring and control of workpieces, we present an innovative data-driven method that synchronously collects sensing data and AM process parameters with a low sampling rate during the DED process. The proposed data-driven technique determines the important influences that individual printing parameters and sensing features have on prediction at the inter-layer qualification to perform feature selection. Three Machine Learning (ML) algorithms including Random Forest (RF), Support Vector Machine (SVM), and Convolutional Neural Network (CNN) are used. During post-production, a threshold is applied to detect low-density occurrences such as porosity sizes and quantities from CT scans that render individual layers acceptable or unacceptable. This information is fed to the ML models for training. Training/testing are completed offline on samples deemed “high-quality” and “low-quality”, utilizing only features recorded from the build process. CNN results show that the classification of acceptable/unacceptable layers can reach between 90% accuracy while training/testing on a “high-quality” sample and dip to 65% accuracy when trained/tested on “low-quality”/“high-quality” (respectively), indicating over-fitting but showing CNN as a promising inter-layer classifier.

36 MATERIALS SCIENCE↗

A learning-augmented approach for AC optimal power flow

Because of the high nonlinearity of AC optimal power flow (OPF), numerous efforts have been made in recent decades to find efficient methods. Machine learning (ML) has proven to significantly reduce the computational costs in many real-world problems. Thus, this paper develops a learning-augmented method for solving AC OPF, which integrates both power network equations and ML to yield near-optimal solutions. More specifically, ML models are developed to first predict bus voltage magnitudes and angles. Then, physics-based network equations are employed to calculate the power injection at different buses. Three ML algorithms, i.e., random forest, multi-target decision tree, and extreme learning machine, are explored and compared. To evaluate the efficiency of the proposed learning-augmented AC OPF solver, the MATPOWER Interior Point Solver is adopted as a baseline. Case studies on both 500-bus and 4918-bus test networks show that the proposed learning-augmented method has reduced the computational time by 15–100 times depending on the network size with a minimal loss in optimality.

42 ENGINEERING↗

Interpreting High-resolution Spectroscopy of Exoplanets using Cross-correlations and Supervised Machine Learning

We present a new method for performing atmospheric retrieval on ground-based, high-resolution data of exoplanets. Our method combines cross-correlation functions with a random forest, a supervised machine-learning technique, to overcome challenges associated with high-resolution data. A series of cross-correlation functions are concatenated to give a “CCF-sequence” for each model atmosphere, which reduces the dimensionality by a factor of ∼100. The random forest, trained on our grid of ∼65,000 models, provides a likelihood-free method of retrieval. The precomputed grid spans 31 values of both temperature and metallicity, and incorporates a realistic noise model. We apply our method to HARPS-N observations of the ultra-hot Jupiter KELT-9b and obtain a metallicity consistent with solar (logM = − 0.2 ± 0.2). Our retrieved transit chord temperature (T=6000{sub −200}{sup +0}K) is unreliable as strong ion lines lie outside of the extent of the training set, which we interpret as being indicative of missing physics in our atmospheric model. We compare our method to traditional nested sampling, as well as other machine-learning techniques, such as Bayesian neural networks. We demonstrate that the likelihood-free aspect of the random forest makes it more robust than nested sampling to different error distributions, and that the Bayesian neural network we tested is unable to reproduce complex posteriors. We also address the claim in Cobb et al. 2019 that our random forest retrieval technique can be overconfident but incorrect. We show that this is an artifact of the training set, rather than of the machine-learning method, and that the posteriors agree with those obtained using nested sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]↗

Design Space Exploration of Emerging Memory Technologies for Machine Learning Applications

Memory design space exploration methods study memory systems’ performances and limitations before implementation. The computer memory design space has grown exponentially because of the enormous growth of memory types, memory controllers, and application software. Computer simulators are commonly used for memory design space exploration. However, complex memory simulations take an enormous amount of time. Hence, in this paper, we proposed a machine learning-based design space exploration method for dynamic random-access memory and non-volatile memory systems. We applied our method to the CosmoGAN and LeNet applications to predict the following six memory response parameters: (i) bandwidth, (ii) power, (iii) average latency, (iv) average total latency, (v) memory reads, and (vi) memory writes. Our experimental results show that machine learning models can predict memory response parameter values faster than simulations. We used support vector machine, random forest, and gradient boosting machine learning models. We observed that the support vector machine provides better performance for bandwidth, average latency, and average total latency. The random forest model works better for memory reads and writes. The gradient boosting model provides superior prediction performance for power. We provide a detailed discussion on learning curve characteristics, error analysis, and memory type recommendation.

Hasan, S M Shamimul↗

Forest Soil Carbon Efflux Evaluation Across China: A New Estimate With Machine Learning

Forest soil respiration (Rs) plays an important role in the carbon balance of terrestrial ecosystems. China's forest occupies a large part of the world's forest. However, due to the lack of integrated observation data and appropriate upscaling methodologies, substantial uncertainties exist in the Rs evaluation, which limits our understanding of the carbon balance. Here, we re-evaluated the total soil carbon effluxes across China by combining field observations from 634 published annual Rs with a machine learning technique (i.e., Random Forest (RF)). Our results revealed that the combination of systematic measurements with the RF model allowed a definite estimate. The average annual Rs was 773.6 g C m -2 yr -1 , ranging from 404.6 to 1,772.3 g C m -2 yr -1 . Total forest soil carbon effluxes amounted to 1.17 Pg C yr -1 in China. Geographically, annual Rs showed a clear spatially increasing trend from northeast to southwest. Forest type is an important factor in determining the soil respiration rate. Bamboo and Evergreen broadleaf forests were higher than other types of forests. In conclusion, these results provide a unique insight into the magnitudes and mechanisms of soil CO 2 emissions in China's forest ecosystems.

58 GEOSCIENCES↗

Machine learning feature analysis illuminates disparity between E3SM climate models and observed climate change

In September of 2020, Arctic sea ice extent was the second-lowest on record. State of the art climate prediction uses Earth system models (ESMs), driven by systems of differential equations representing the laws of physics. Previously, these models have tended to underestimate Arctic sea ice loss. The issue is grave because accurate modeling is critical for economic, ecological, and geopolitical planning. We use machine learning techniques, including random forest regression and Gini importance, to show that the Energy Exascale Earth System Model (E3SM) relies too heavily on just one of the ten chosen climatological quantities to predict September sea ice averages. Furthermore, E3SM gives too much importance to six of those quantities when compared to observed data. Finally, identifying the features that climate models incorrectly rely on should allow climatologists to improve prediction accuracy.

54 ENVIRONMENTAL SCIENCES↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

Correcting for filter-based aerosol light absorption biases at the Atmospheric Radiation Measurement program's Southern Great Plains site using photoacoustic measurements and machine learning

Abstract. Measurement of light absorption of solar radiation by aerosols is vital for assessing direct aerosol radiative forcing, which affects local and global climate. Low-cost and easy-to-operate filter-based instruments, such as the Particle Soot Absorption Photometer (PSAP), that collect aerosols on a filter and measure light attenuation through the filter are widely used to infer aerosol light absorption. However, filter-based absorption measurements are subject to artifacts that are difficult to quantify. These artifacts are associated with the presence of the filter medium and the complex interactions between the filter fibers and accumulated aerosols. Various correction algorithms have been introduced to correct for the filter-based absorption coefficient measurements toward predicting the particle-phase absorption coefficient (Babs). However, the inability of these algorithms to incorporate into their formulations the complex matrix of influencing parameters such as particle asymmetry parameter, particle size, and particle penetration depth results in prediction of particle-phase absorption coefficients with relatively low accuracy. The analytical forms of corrections also suffer from a lack of universal applicability: different corrections are required for rural and urban sites across the world. In this study, we analyzed and compared 3 months of high-time-resolution ambient aerosol absorption data collected synchronously using a three-wavelength photoacoustic absorption spectrometer (PASS) and PSAP. Both instruments were operated on the same sampling inlet at the Department of Energy's Atmospheric Radiation Measurement program's Southern Great Plains (SGP) user facility in Oklahoma. We implemented the two most commonly used analytical correction algorithms, namely, Virkkula (2010) and the average of Virkkula (2010) and Ogren (2010)–Bond et al. (1999) as well as a random forest regression (RFR) machine learning algorithm to predict Babs values from the PSAP's filter-based measurements. The predicted Babs was compared against the reference Babs measured by the PASS. The RFR algorithm performed the best by yielding the lowest root mean square error of prediction. The algorithm was trained using input datasets from the PSAP (transmission and uncorrected absorption coefficient), a co-located nephelometer (scattering coefficients), and the Aerosol Chemical Speciation Monitor (mass concentration of non-refractory aerosol particles). A revised form of the Virkkula (2010) algorithm suitable for the SGP site has been proposed; however, its performance yields approximately 2-fold errors when compared to the RFR algorithm. To generalize the accuracy and applicability of our proposed RFR algorithm, we trained and tested it on a dataset of laboratory measurements of combustion aerosols. Input variables to the algorithm included the aerosol number size distribution from the Scanning Mobility Particle Sizer, absorption coefficients from the filter-based Tricolor Absorption Photometer, and scattering coefficients from a multiwavelength nephelometer. The RFR algorithm predicted Babs values within 5 % of the reference Babs measured by the multiwavelength PASS during the laboratory experiments. Thus, we show that machine learning approaches offer a promising path to correct for biases in long-term filter-based absorption datasets and accurately quantify their variability and trends needed for robust radiative forcing determination.

54 ENVIRONMENTAL SCIENCES↗

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗