Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme gradient boost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Decoy selection for protein structure prediction via extreme gradient boosting and ranking

Background: Identifying one or more biologically-active/native decoys from millions of non-native decoys is one of the major challenges in computational structural biology. The extreme lack of balance in positive and negative samples (native and non-native decoys) in a decoy set makes the problem even more complicated. Consensus methods show varied success in handling the challenge of decoy selection despite some issues associated with clustering large decoy sets and decoy sets that do not show much structural similarity. Recent investigations into energy landscape-based decoy selection approaches show promises. However, lack of generalization over varied test cases remains a bottleneck for these methods. Results: We propose a novel decoy selection method, ML-Select, a machine learning framework that exploits the energy landscape associated with the structure space probed through a template-free decoy generation. The proposed method outperforms both clustering and energy ranking-based methods, all the while consistently offering better performance on varied test-cases. Moreover, ML-Select shows promising results even for the decoy sets consisting of mostly low-quality decoys. Conclusions: ML-Select is a useful method for decoy selection. This work suggests further research in finding more effective ways to adopt machine learning frameworks in achieving robust performance for decoy selection in template-free protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

A novel improved model for building energy consumption prediction based on model integration

Building energy consumption prediction plays an irreplaceable role in energy planning, management, and conservation. Constantly improving the performance of prediction models is the key to ensuring the efficient operation of energy systems. Moreover, accuracy is no longer the only factor in revealing model performance, it is more important to evaluate the model from multiple perspectives, considering the characteristics of engineering applications. Based on the idea of model integration, this paper proposes a novel improved integration model (stacking model) that can be used to forecast building energy consumption. The stacking model combines advantages of various base prediction algorithms and forms them into “meta-features” to ensure that the final model can observe datasets from different spatial and structural angles. Two cases are used to demonstrate practical engineering applications of the stacking model. A comparative analysis is performed to evaluate the prediction performance of the stacking model in contrast with existing well-known prediction models including Random Forest, Gradient Boosted Decision Tree, Extreme Gradient Boosting, Support Vector Machine, and K-Nearest Neighbor. The results indicate that the stacking method achieves better performance than other models, regarding accuracy (improvement of 9.5%–31.6% for Case A and 16.2%–49.4% for Case B), generalization (improvement of 6.7%–29.5% for Case A and 7.1%-34.6% for Case B), and robustness (improvement of 1.5%–34.1% for Case A and 1.8%–19.3% for Case B). The proposed model enriches the diversity of algorithm libraries of empirical models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Using Machine Learning to Predict Cloud Turbulent Entrainment–Mixing Processes

Different turbulent entrainment–mixing mechanisms between clouds and environment are essential to cloud–related processes; however, accurate representation of entrainment–mixing in weather/climate models still poses a challenge. This study exploits the use of machine learning (ML) to address this challenge. Four ML (Light Gradient Boosting Machine [LGB], eXtreme Gradient Boosting, Random Forest, and Support Vector Regression) are examined and compared. It is found that LGB performs best, and thus is selected to understand the impact of entrainment–mixing on microphysics using simulation data from Explicit Mixing Parcel Model. Compared with traditional parameterizations, the trained LGB provides more accurate microphysical properties (number concentration and cloud droplet spectral dispersion). The partial dependences of predicted microphysics on features exhibit a strong alignment with physical mechanisms and expectations, as determined by the interpreting method, thus overcoming the limitations of the “black box” scheme. The underlying mechanisms are that the smaller number concentration and larger spectral dispersion correspond to more inhomogeneous entrainment–mixing. Specifically, number concentration after entrainment–mixing is positively correlated with adiabatic number concentration and liquid water content affected by entrainment–mixing, and inversely correlated with adiabatic volume mean radius. Spectral dispersion after entrainment–mixing is negatively correlated with liquid water content affected by entrainment–mixing, turbulent dissipation rate and relative humidity of entrained air. Sensitivity analysis further suggests that number concentration is mainly determined by cloud microphysical properties whereas spectral dispersion is influenced by both cloud microphysical properties and environmental variables. The results indicate that the LGB scheme has the potential to enhance the representation of entrainment–mixing in weather/climate models.

54 ENVIRONMENTAL SCIENCES↗

A Predictive Prescription Framework for Stochastic Unit Commitment Using Boosting Ensemble Learning Algorithms

To take unit commitment (UC) decisions under uncertain load, most existing stochastic optimization (SO) frameworks adopt a generic representation of uncertainty. While load levels that materialize on a particular day are influenced by various covariates (such as the day of the week or temperature), SO frameworks typically disregard such side observations, wasting actionable information that could significantly enhance decision quality. Here, this article proposes a contextual SO (CSO) framework for UC under uncertain load, which can effectively exploit covariate observations in conjunction with a class of machine learning (ML) algorithms to improve the out-of-sample performance of UC decisions. It shows how three ML algorithms, adaptive boosting, gradient boosted trees, and extreme gradient boosting, can be used to this end, constituting the first application of these algorithms in any CSO framework. Using real-world data harvested from the New York ISO grid, we measure the out-of-sample performance of the framework in terms of total operation cost, shed load values, locational marginal prices, and total payments by the loads, against several benchmark methods proposed in the literature. The article has an online companion (Yurdakul et al.), wherein we present additional results and lay out further mathematical formulations used in this work.

42 ENGINEERING↗

Network-Scale Ubiquitous Volume Estimation Using Tree-Based Ensemble Learning Methods

Currently ubiquitous volume data for roadway networks remains the key missing dimension in traffic operations. Most volume data are average annual daily traffic (AADT) measures derived from the Highway Performance Monitoring System (HPMS). Although methods to factor the AADT to hourly averages for typical day of week exist, actual volume data is limited to a sparse collection of locations in which volumes are continuously recorded. This paper/poster explores the use of state-of-art machine learning techniques to estimate accurate volume measures that span the highway network providing ubiquitous coverage in space, and point-in-time measures for a specific date and time. Three tree-based ensemble learning models, random forest (RF), gradient boost machine (GBM), and extreme gradient boost (XGBoost), were tested for volume estimation by learning from combined dataset of commercial probe data provided by TomTom, the FHWA's Travel Monitoring Analysis System (TMAS) data, and other infrastructure attributes such as number of lanes, speed limit, and weather. The methods were tested on major corridors and freeways in the metropolitan area of Denver. All three machine learning methods were able to provide hourly volume estimates 24 hours a day, 7 days a week, and 365 days a year with around 18% mean absolute error to true volume and about 5% of error with respect to roadway capacity. The low error measures allow the potential application by transportation agencies.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning for fundamental spectroscopic and thermodynamic data of actinides and lanthanides

Accurately modeling optical spectra with absolute radiometric intensities is vital for nuclear forensics applications that depend on characterizing optical emissions from energetic nuclear phenomena. This requires precise knowledge of the individual atomic transition probabilities, known as Einstein A-coefficients, for each emission line. Obtaining these values theoretically or experimentally is often impractical due to the complex electronic structures and the number of transitions involved in atoms relevant to nuclear applications. In this study, we explore the use of machine learning to predict the Einstein A coefficients for atomic transitions. Seven models were evaluated that ranged from deep learning to decision tree algorithms, and found that gradient boosting performed best, specifically the Extreme Gradient Boosting (XGB) architecture, achieving a precision of 86% across transitions of 36 elements. Furthermore, the model was cross-validated using published transition probabilities reported in the literature and applied to estimate Pu plasma temperatures from a previous experiment conducted at Savannah River National Laboratory.

Atomic spectroscopy↗

Delegated Regressor, A Robust Approach for Automated Anomaly Detection in the Soil Radon Time Series Data

We propose a new method based on the idea of delegating regressors for predicting the soil radon gas concentration (SRGC) and anomalies in radon or any other time series data. The proposed method is compared to different traditional boosting e.g., Extreme Gradient Boosting (EGB) and simple regression methods e.g., support vector regressors with linear kernel and radial kernel in terms of accurate predictions. R language has been used for the statistical analysis of radon time series (RTS) data. The results obtained show that the proposed methodology predicts SRGC more accurately when compared to different traditional boosting and regression methods. The best correlation is found between the actual and predicted radon concentration for window size of 2 i.e., two days before and after the start of seismic activities. RTS data was collected from 05 February 2017 to 16 February 2018, including 7 seismic events recorded during the study period. Findings of study show that the proposed methodology predicts the SRGC with more precision, for all the window sizes, by overlapping predicted with the actual radon time series concentrations.

54 ENVIRONMENTAL SCIENCES↗

Automatic Search of Cataclysmic Variables Based on LightGBM in LAMOST-DR7

The search for special and rare celestial objects has always played an important role in astronomy. Cataclysmic Variables (CVs) are special and rare binary systems with accretion disks. Most CVs are in the quiescent period, and their spectra have the emission lines of Balmer series, HeI, and HeII. A few CVs in the outburst period have the absorption lines of Balmer series. Owing to the scarcity of numbers, expanding the spectral data of CVs is of positive significance for studying the formation of accretion disks and the evolution of binary star system models. At present, the research for astronomical spectra has entered the era of Big Data. The Large Sky Area Multi-Object Fiber Spectroscopy Telescope (LAMOST) has produced more than tens of millions of spectral data. the latest released LAMOST-DR7 includes 10.6 million low-resolution spectral data in 4926 sky regions, providing ideal data support for searching CV candidates. To process and analyze the massive amounts of spectral data, this study employed the Light Gradient Boosting Machine (LightGBM) algorithm, which is based on the ensemble tree model to automatically conduct the search in LAMOST-DR7. Finally, 225 CV candidates were found and four new CV candidates were verified by SIMBAD and published catalogs. This study also built the Gradient Boosting Decision Tree (GBDT), Adaptive Boosting (AdaBoost), and eXtreme Gradient Boosting (XGBoost) models and used Accuracy, Precision, Recall, the F1-score, and the ROC curve to compare the four models comprehensively. Experimental results showed that LightGBM is more efficient. The search for CVs based on LightGBM not only enriches the existing CV spectral library, but also provides a reference for the data mining of other rare celestial objects in massive spectral data.

79 ASTRONOMY AND ASTROPHYSICS↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Fault location in High Voltage Multi-terminal dc Networks Using Ensemble Learning

Precise location of faults for large distance power transmission networks is essential for faster repair and restoration process. High Voltage direct current (HVdc) networks using modular multi-level converter (MMC) technology has found its prominence for interconnected multi-terminal networks. This allows for large distance bulk power transmission at lower costs. However, they cope with the challenge of dc faults. Fast and efficient methods to isolate the network under dc faults have been widely studied and investigated. After successful isolation, it is essential to precisely locate the fault. The post-fault voltage and current signatures are a function of multiple factors and thus accurately locating faults on a multi-terminal network is challenging. In this paper, we discuss a novel data-driven ensemble learning based approach for accurate fault location. Here we utilize the eXtreme Gradient Boosting (XGB) method for accurate fault location. The sensitivity of the proposed algorithm to measurement noise, fault location, resistance and current limiting inductance are performed on a radial three-terminal MTdc network designed in Power System Computer Aided Design (PSCAD)/Electromagnetic Transients including dc (EMTdc).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning Analysis of Hydrologic Exchange Flows and Transit Time Distributions in a Large Regulated River

Hydrologic exchange between river channels and adjacent subsurface environments is a key process that influences water quality and ecosystem function in river corridors. High-resolution numerical models were often used to resolve the spatial and temporal variations of exchange flows, which are computationally expensive. In this study, we adopt Random Forest (RF) and Extreme Gradient Boosting (XGB) approaches for deriving reduced order models of hydrologic exchange flows and associated transit time distributions, with integrated field observations (e.g., bathymetry) and hydrodynamic simulation data (e.g., river velocity, depth). The setup allows an improved understanding of the influences of various physical, spatial, and temporal factors on the hydrologic exchange flows and transit times. The predictors also contain those derived using hybrid clustering, leveraging our previous work on river corridor system hydromorphic classification. The machine learning-based predictive models are developed and validated along the Columbia River Corridor, and the results show that the top parameters are the thickness of the top geological formation layer, the flow regime, river velocity, and river depth; the RF and XGB models can achieve 70% to 80% accuracy and therefore are effective alternatives to the computational demanding numerical models of exchange flows and transit time distributions. Each machine learning model with its favorable configuration and setup have been evaluated. The transferability of the models to other river reaches and larger scales, which mostly depends on data availability, is also discussed.

97 MATHEMATICS AND COMPUTING↗

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass↗

Estimation of the Surface Fluxes for Heat and Momentum in Unstable Conditions with Machine Learning and Similarity Approaches for the LAFE Data Set

Abstract Measurements of three flux towers operated during the land atmosphere feedback experiment (LAFE) are used to investigate relationships between surface fluxes and variables of the land–atmosphere system. We study these relations by means of two machine learning (ML) techniques: multilayer perceptrons (MLP) and extreme gradient boosting (XGB). We compare their flux derivation performance with Monin–Obukhov similarity theory (MOST) and a similarity relationship using the bulk Richardson number (BRN). The ML approaches outperform MOST and BRN. Best agreement with the observations is achieved for the friction velocity. For the sensible heat flux and even more so for the latent heat flux, MOST and BRN deviate from the observations while MLP and XGB yield more accurate predictions. Using MOST and BRN for latent heat flux, the root mean square errors (RMSE) are 107 Wm $$^{-2}$$ - 2 and 121 Wm $$^{-2}$$ - 2 , respectively, as well as the intercepts of the regression lines are $$\approx 110$$ ≈ 110 Wm $$^{-2}$$ - 2 . For the ML methods, the RMSEs reduce to 31 Wm $$^{-2}$$ - 2 for MLP and 33 Wm $$^{-2}$$ - 2 for XGB as well as the intercepts to just 4 Wm $$^{-2}$$ - 2 for MLP and $$-1$$ - 1 Wm $$^{-2}$$ - 2 for XGB with slopes of the regression lines close to 1, respectively. These results indicate significant deficiencies of MOST and BRN, particularly for the derivation of the latent heat flux. In fact, in contrast to the established theories, feature importance weighting demonstrates that the ML methods base their improved derivations on net radiation, the incoming and outgoing shortwave radiations, the air temperature gradient, and the available water contents, but not on the water vapor gradient. The results imply that further studies of surface fluxes and other turbulent variables with ML techniques provide great promise for deriving advanced flux parameterizations and their implementation in land–atmosphere system models.

54 ENVIRONMENTAL SCIENCES↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗