Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “xgboost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Characterizing Seasonal Variation of the Atmospheric Mixing Layer Height Using Machine Learning Approaches

As machine learning becomes more integrated into atmospheric science, XGBoost has gained popularity for its ability to assess the relative contributions of influencing factors in the atmospheric boundary layer height. To examine how these factors vary across seasons, a seasonal analysis is necessary. However, dividing data by season reduces the sample size, which can affect result reliability and complicate factor comparisons. To address these challenges, this study replaces default parameters with grid search optimization and incorporates cross-validation to mitigate dataset limitations. Using XGBoost with four years of data from the atmospheric radiation measurement (ARM) (Southern Great Plains (SGP) C1 site, cross-validation stabilizes correlation coefficient fluctuations from 0.3 to within 0.1. With optimized parameters, the R value can reach 0.81. Analysis of the C1 site reveals that the relative importance of different factors changes across seasons. Lower tropospheric stability (LTS, ~0.53) is the dominant factor at C1 throughout the year. However, during DJF, latent heat flux (LHF, 0.44) surpasses LTS (0.22). In SON, LTS (0.58) becomes more influential than LHF (0.18). Further comparisons among the four long-term SGP sites (C1, E32, E37, and E39) show seasonal variations in relative importance. Notably, during JJA, the differences in the relative importance of the three factors across all sites are lower than in other seasons. This suggests that boundary layer development in the summer is not dominated by a single factor, reflecting a more intricate process likely influenced by seasonal conditions such as enhanced convective activity, higher temperatures, and humidity, which collectively contribute to a balanced distribution of parameter impacts. Furthermore, the relative importance of LTS gradually increases from morning to noon, indicating that LTS becomes more significant as the boundary layer approaches its maximum height. Consequently, the LTS in the early morning in autumn exhibits greater relative importance compared to other seasons. This reflects a faster development of the mixing layer height (MLH) in autumn, suggesting that it is easier to retrieve the MLH from the previous day during this period. The findings enhance understanding of boundary layer evolution and contribute to improved boundary layer parameterization.

54 ENVIRONMENTAL SCIENCES↗

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling↗

Contributions of vegetation heterogeneity within tower footprint to CO 2 flux estimations through graph neural network modeling

Net ecosystem exchange of CO 2 (Fc) measured directly by eddy covariance towers is based on various assumptions, including large, flat and homogenous land cover type. In reality, often a tower site is not large enough for flux measurements, and landscapes consist of patches of different land cover types within the flux footprint. In addition, some portions of fluxes are contributed by different cover types when a footprint exceeds the size of the target ecosystem. The contributions of non-dominant patches to Fc are often ignored. Here, in this study, we propose a novel integrated modeling framework that combines random forest (RF) and XGBoost with a residual correction module based on a deep graph convolutional network (DeeperGCN) to simulate Fc for seven flux measurement sites in southwest Michigan. High-resolution remote sensing vegetation indices, soil properties, meteorological variables, and footprint-weighted spatial features were used as model inputs at three spatial resolutions (10 m, 20 m, 30 m), and their importance in predicting Fc with DeeperGCN was assessed. We found that residual correction using DeeperGCN significantly improved prediction accuracy, with the R 2 increasing from 0.9098 to 0.9479 for RF and from 0.9235 to 0.9433 for XGBoost. At site level, the maximum improvement in R 2 reached 0.1617. Paired t-tests confirmed that these improvements were statistically significant (p < 0.05). Among all predictors, leaf area index and incoming shortwave radiation emerged as the dominant drivers of spatial residual variation, followed by precipitation, relative humidity, and selected vegetation indices. The 20 m resolution yielded the best balance between model performance and computational efficiency. In conclusion, our modeling framework effectively captures both spatial heterogeneity and nonlinear interactions, offering a robust solution for spatially explicit flux modeling in structurally diverse ecosystems beyond the study sites.

footprint model↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

AutoSourceID-Classifier: Star-galaxy classification using a convolutional neural network with spatial information

Aims.Traditional star-galaxy classification techniques often rely on feature estimation from catalogs, a process susceptible to introducing inaccuracies, thereby potentially jeopardizing the classification’s reliability. Certain galaxies, especially those not manifesting as extended sources, can be misclassified when their shape parameters and flux solely drive the inference. We aim to create a robust and accurate classification network for identifying stars and galaxies directly from astronomical images. Methods.The AutoSourceID-Classifier (ASID-C) algorithm developed for this work uses 32x32 pixel single filter band source cutouts generated by the previously developed AutoSourceID-Light (ASID-L) code. By leveraging convolutional neural networks (CNN) and additional information about the source position within the full-field image, ASID-C aims to accurately classify all stars and galaxies within a survey. Subsequently, we employed a modified Platt scaling calibration for the output of the CNN, ensuring that the derived probabilities were effectively calibrated, delivering precise and reliable results. Results.We show that ASID-C, trained on MeerLICHT telescope images and using the Dark Energy Camera Legacy Survey (DECaLS) morphological classification, is a robust classifier and outperforms similar codes such as SourceExtractor. To facilitate a rigorous comparison, we also trained an eXtreme Gradient Boosting (XGBoost) model on tabular features extracted by SourceExtractor. While this XGBoost model approaches ASID-C in performance metrics, it does not offer the computational efficiency and reduced error propagation inherent in ASID-C’s direct image-based classification approach. ASID-C excels in low signal-to-noise ratio and crowded scenarios, potentially aiding in transient host identification and advancing deep-sky astronomy.

Astronomy & Astrophysics↗

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

The Role of Nuclear Data Sensitivities in Prompt α-Eigenvalue Predictions of Delayed Critical Benchmarks

Alpha (α) eigenvalues, which describe the logarithmic time derivative of the neutron population in a multiplying system, are integral to time-dependent behavior and diagnostic applications. However, uncertainties in the evaluated nuclear data can significantly impact the accuracy of transport simulations for such quantities. This work explores the use of machine learning models to predict two key outputs, α-eigenvalues and keff bias, using input features derived from α-eigenvalue sensitivities to nuclear data. The criticality safety benchmark models used in this study come from the International Handbook of Evaluated Criticality Safety Benchmark Experiments. Three models, random forest, XGBoost, and NGBoost, are trained on both energy-resolved and energy-summed α sensitivities. For the α-eigenvalue bias prediction, NGBoost achieved the highest R 2 (0.9476) using energy-resolved features, while XGBoost performed best using summed sensitivities. In contrast, when predicting the keff bias, all the models showed moderate predictive capability (best R 2 ≈ 0.72), as the mapping from the static α-sensitivities to the static keff bias was less direct. SHAP (SHapley Additive exPlanations) analysis was used to interpret the model predictions. Across both prediction tasks, the features associated with neutron capture [H-1 (n, γ)], uranium scattering reactions (such as 235 U elastic/inelastic), and actinide capture/fission reactions (such as 239 Pu and 234 U) were consistently identified as the most impactful. This highlights the key role of specific nuclear reactions and energy ranges in shaping both time-dependent and steady-state criticality behavior. These results demonstrated that α-sensitivities, despite being computed for time-dependent metrics, can provide valuable insights for predicting both α-eigenvalues and the keff bias. Moreover, machine learning models offer a promising pathway for uncovering important nuclear data dependencies and guiding future data evaluation efforts.

Nuclear data↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

Photometric redshift estimation of BASS DR3 quasars by machine learning

ABSTRACT Correlating Beijing–Arizona Sky Survey (BASS) data release 3 (DR3) catalogue with the ALLWISE data base, the data from optical and infrared information are obtained. The quasars from Sloan Digital Sky Survey are taken as training and test samples while those from LAMOST are considered as external test sample. We propose two schemes to construct the redshift estimation models with XGBoost, CatBoost, and Random Forest. One scheme (namely one-step model) is to predict photometric redshifts directly based on the optimal models created by these three algorithms; the other scheme (namely two-step model) is to first classify the data into low- and high-redshift data sets, and then predict photometric redshifts of these two data sets separately. For one-step model, the performance of these three algorithms on photometric redshift estimation is compared with different training samples, and CatBoost is superior to XGBoost and Random Forest. For two-step model, the performances of these three algorithms on the classification of low and high redshift subsamples are compared, and CatBoost still shows the best performance. Therefore, CatBoost is regarded as the core algorithm of classification and regression in two-step model. In contrast to one-step model, two-step model is optimal when predicting photometric redshift of quasars, especially for high-redshift quasars. Finally, the two models are applied to predict photometric redshifts of all quasar candidates of BASS DR3. The number of high-redshift quasar candidates is 3938 (redshift ≥3.5) and 121 (redshift ≥4.5) by two-step model. The predicted result will be helpful for quasar research and follow-up observation of high-redshift quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

Demystifying Cyberattacks: Potential for Securing Energy Systems With Explainable AI : Preprint

Modernization of energy systems has led to in- creased interactions among multiple critical infrastructures and diverse stakeholders making the challenge of operational decision making more complex and at times beyond cognitive capabilities of human operators. The state-of-the-art machine learning and deep learning approaches show promise of supporting users with complex decision-making challenges, such as those occurring in our rapidly transforming cyber-physical energy systems. However, successful adoption of data-driven decision support technology for critical infrastructure will be dependent on the ability of these technologies to be trustworthy and contextually interpretable. In this paper, we investigate the feasibility of implementing XAI for interpretable detection of cyberattacks in the energy system. Leveraging a proof-of-concept simulation use case of detection of a data falsification attack on a photovoltaic system using XGBoost algorithm, we demonstrate how Local Interpretable Model-Agnostic Explanations (LIME), a flavor XAI approach, can help provide contextual and actionable interpretation of cyberattack detection.

artificial intelligence↗

Cross-Market Price Difference Forecast Using Deep Learning for Electricity Markets

Price forecasting is in the center of decision making in electricity markets. Many researches have been done in forecasting energy prices while little research has been reported on forecasting price difference between day-ahead and realtime markets due to its high volatility, which however plays a critical role in virtual trading. To this end, this paper takes the first attempt to employ novel deep learning architecture with Bidirectional Long-Short Term Memory (LSTM) units to forecast the price difference between day-ahead and real-time markets for the same node. The raw data is collected from PJM market, processed and fed into the proposed network. The Root Mean Squared Error (RMSE) and customized performance metric are used to evaluate the performance of the proposed method. Case studies show that it outperforms the traditional statistical models like ARIMA, and machine learning models like XGBoost and SVR methods in both RMSE and the capability of forecasting the sign of price difference. Additionally, to cross-market price difference forecast, the proposed approach has the potential to be applied to solve other forecasting problems such as price spread forecast in DA market for Financial Transmission Right (FTR) trading purpose.

DA/RT price difference↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different datasets, i.e., 1) the unit generation dataset and 2) the system total inertia and headroom dataset, for the prediction of the frequency nadir. The training and testing datasets are generated through extensive generation scheduling simulations using Multi-timescale Integrated Dynamic and Scheduling (MI-DAS) toolbox on the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. Numerical results show that all five machine learning methods perform well in predicting the nadir frequency of the system. Among them, the gradient boosting and the XGBoost are clear winners yielding the best prediction accuracy in terms of four evaluation metrics.

data driven↗

Six Machine-Learning Methods for Predicting Hospital-Stay Duration for Patients with Sepsis: A Comparative Study

Sepsis is a life-threatening medical condition that, if not treated promptly, can result in tissue damage, organ failure, and death. According to the Centers for Disease Control, about 270,000 individuals die of sepsis in the US each year. Further, sepsis expenditures accounted for 13% of total US hospital costs in 2013, totaling more than $24 billion. Our project objectives were to determine if Machine Learning algorithms could reliably predict hospital stay duration for patients with sepsis. The data set we used has been de-identified and is freely available through the BupaR package. The data includes 1050 cases, 15214 events, and 16 types of actions related to sepsis patient care. First, we used process mining to determine how long each patient was in the hospital. Using BupaR’s functions, we created several process model graphs. These process models depict the movement of patients at a hospital and provide duration data for each patent case. Second, we identified outlier data and created two dataset versions: one with and one without outliers. We then applied the following analysis methods: Linear Regression, Random Forest, K-Nearest Neighbors, Neural Networks, XGBoost, and lightGBM. We compared the model validations for the six machine learning models using the same data-splitting method. We found that the XGBoost model had the best prediction accuracy of 73.9 percent for cases with outliers, and 79 percent for cases without outliers. We also found that the lightGBM model had the lowest mean absolute error between prediction and actual duration in days with 3.66 days for the case with outliers, and 2.4 days for the case without outliers. These two models outperformed the other four models. This work will be enhanced in the future by exploring new prediction algorithms and comparing them with the results of this study.

Chen, Lingtao↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating Experiments and Well Logs to Predict Caney Shale Static Mechanical Properties During Production with Supervised Machine Learning

Caney shale is one of the emerging oil reservoirs in Oklahoma. Understanding the impact of effective stress on its mechanical properties is critical for predicting hydraulic fracture geometry and overall hydrocarbon production. The objective of our study is to evaluate the impact of effective stress on the dynamic Young’s modulus using ultrasonic velocity measurements for Caney shale samples. A triaxial cell was utilized to measure ultrasonic (P-wave and S-wave) velocities for ten downhole Caney shale samples under various effective stresses to indirectly assess the impact of pore pressure change. The dynamic Young’s moduli estimated from these measurements were integrated with available conventional well logs (excluding sonic logs) and triaxial test results from Benge et al. (2021) to predict the static Young’s modulus using Random Forest (RF) and Extreme Gradient Boosting (XGBoost) models. The results showed that the estimated dynamic Young’s moduli from ultrasonic measurements were higher than the corresponding static Young’s modulus of cores from the same vertical well at similar depths. With increasing effective stress, the dynamic Young’s modulus increased for all samples. The estimated dynamic-to-static correction factor tended to be higher in zones of high neutron porosity (PHIN) and low density compared to other zones. Finally, SHapley Additive exPlanations (SHAP) for RF and XGBoost models identified depth, gamma ray (GR), and PHIN as key features for predicting the static Young’s modulus. This study enhances our understanding of the dynamic and static Young’s moduli for the Caney shale interval, as a function of effective stress and conventional well logs. The findings from this study can improve predictions of production throughout the well's lifespan by offering insights into the mechanical property degradation resulting from pore pressure depletion.

Kholy, Sherif M.↗