Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “xgboost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Improving Low‐Cloud Fraction Prediction Through Machine Learning

Abstract In this study, we evaluated the performance of machine learning (ML) models (XGBoost) in predicting low‐cloud fraction (LCF), compared to two generations of the community atmospheric model (CAM5 and CAM6) and ERA5 reanalysis data, each having a different cloud scheme. ML models show a substantial enhancement in predicting LCF regarding root mean squared errors and correlation coefficients. The good performance is consistent across the full spectrums of atmospheric stability and large‐scale vertical velocity. Employing an explainable ML approach, we revealed the importance of including the amount of available moisture in ML models for representing spatiotemporal variations in LCF in the midlatitudes. Also, ML models demonstrated marked improvement in capturing the LCF variations during the stratocumulus‐to‐cumulus transition (SCT). This study suggests ML models' great potential to address the longstanding issues of “too few” low clouds and “too rapid” SCT in global climate models.

Geology↗

Aerosol Influences on Cloud Water: Insights From ARM EPCAPE Observations With Explainable Machine Learning

This study employs an explainable machine learning (ML) framework (XGBoost-SHapley Additive exPlanations analysis) to investigate controlling factors on cloud liquid water path (LWP) using EPCAPE observations near the California coast. Aerosols are found to be the dominant factor explaining LWP variability, surpassing meteorological factors (MFs). By isolating aerosol effects from meteorological influences, the ML reveals a negative linear relationship between LWP and cloud droplet number concentration (Nd) in log space, likely driven by entrainment drying via evaporation-entrainment feedback. This aligns with the negative regime of the inverted-V relationship reported in previous studies, while no positive LWP responses are found due to a limited number of precipitating cases in EPCAPE. Furthermore, the sensitivity of LWP to Nd shows a non-linear dependence on MFs like moisture contrast between surface and free troposphere and lower-tropospheric stability. This occurs due to the interplay between the MFs' direct effects on entrainment drying and indirect effects through LWP adjustments.

EPCAPE observations↗

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

Characterize traction–separation relation and interfacial imperfections by data-driven machine learning models

Abstract Interfacial mechanical properties are important in composite materials and their applications, including vehicle structures, soft robotics, and aerospace. Determination of traction–separation (T–S) relations at interfaces in composites can lead to evaluations of structural reliability, mechanical robustness, and failures criteria. Accurate measurements on T–S relations remain challenging, since the interface interaction generally happens at microscale. With the emergence of machine learning (ML), data-driven model becomes an efficient method to predict the interfacial behaviors of composite materials and establish their mechanical models. Here, we combine ML, finite element analysis (FEA), and empirical experiments to develop data-driven models that characterize interfacial mechanical properties precisely. Specifically, eXtreme Gradient Boosting (XGBoost) multi-output regressions and classifier models are harnessed to investigate T–S relations and identify the imperfection locations at interface, respectively. The ML models are trained by macroscale force–displacement curves, which can be obtained from FEA and standard mechanical tests. The results show accurate predictions of T–S relations ( R 2 = 0.988) and identification of imperfection locations with 81% accuracy. Our models are experimentally validated by 3D printed double cantilever beam specimens from different materials. Furthermore, we provide a code package containing trained ML models, allowing other researchers to establish T–S relations for different material interfaces.

97 MATHEMATICS AND COMPUTING↗

Data-driven analysis and prediction of stable phases for high-entropy alloy design

High-entropy alloys (HEAs) represent a promising class of materials with exceptional structural and functional properties. However, their design and optimization pose challenges due to the large composition-phase space coupled with the complex and diverse nature of the phase formation dynamics. In this study, a data-driven approach that utilizes machine learning (ML) techniques to predict HEA phases and their composition-dependent phases is proposed. By employing a comprehensive dataset comprising 5692 experimental records encompassing 50 elements and 11 phase categories, we compare the performance of various ML models. Our analysis identifies the most influential features for accurate phase prediction. Furthermore, the class imbalance is addressed by employing data augmentation methods, raising the number of records to 1500 in each category, and ensuring a balanced representation of phase categories. The results show that XGBoost and Random Forest consistently outperform the other models, achieving 86% accuracy in predicting all phases. Additionally, this work provides an extensive analysis of HEA phase formers, showing the contributions of elements and features to the presence of specific phases. We also examine the impact of including different phases on ML model accuracy and feature significance. Notably, the findings underscore the need for ML model selection based on specific applications and desired predictions, as feature importance varies across models and phases. This study significantly advances the understanding of HEA phase formation, enabling targeted alloy design and fostering progress in the field of materials science.

36 MATERIALS SCIENCE↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Machine learning-enabled discovery of ionic liquid–solvent electrolytes exhibiting high ionic conductivity

Ionic liquids (ILs), which are a class of materials with versatile nature and growing popularity, are facing impediments toward widespread usage as electrolytes due to various factors such as low ionic conductivity, high viscosity, high market price etc. One of the ways these limitations can be addressed is by mixing ILs with a molecular solvent. In a combinatorial sense, there exists an immense number of specific IL–solvent combinations. An exhaustive experimental or even simulation-based investigation of the chemical space spanned by such combinations can be extremely time-consuming, expensive, and nearly impossible. An alternative approach is to employ machine learning-based models developed from available databases. Although there exists prior literature that integrates machine learning to investigate mixtures of specific solvents with ILs, these models lack generalization necessitating development of a large number of ML models to handle various solvents. To remedy this shortcoming, as a part of designing green electrolytes with high ionic conductivity that can have potential applications in next-generation batteries and solar cells, this work aims to develop a unified machine learning model to predict ionic conductivity of any IL–solvent mixture system. In this regard, three models, namely, Random Forest, extreme gradient boosting (XGBoost), and artificial neural network (ANN) were formulated using the NIST ILThermo database. The dataset contained 549 unique ionic liquids from 16 cation families and 81 unique solvents, representing a total of 23 712 datapoints. SHAPLEY additive explanation (SHAP) method was used to assess the impact of various features on model prediction and their significance was compared with literature to gain physical insight about the model behavior. Finally, using the developed models, approximately 2.5 million IL–solvent mixtures at five different compositions were screened at room temperature. The high-throughput screening yielded nearly 19 000 IL–solvent mixtures for which ionic conductivity was found to exceed the ionic conductivity of conventional Li-ion battery electrolyte.

25 ENERGY STORAGE↗

Photometric redshift-aided classification using ensemble learning

We present SHEEP, a new machine learning approach to the classic problem of astronomical source classification, which combines the outputs from the XGBoost, LightGBM, and CatBoost learning algorithms to create stronger classifiers. A novel step in our pipeline is that prior to performing the classification, SHEEP first estimates photometric redshifts, which are then placed into the data set as an additional feature for classification model training; this results in significant improvements in the subsequent classification performance. SHEEP contains two distinct classification methodologies: (i) Multi-class and (ii) one versus all with correction by a meta-learner. We demonstrate the performance of SHEEP for the classification of stars, galaxies, and quasars using a data set composed of SDSS and WISE photometry of 3.5 million astronomical sources. The resulting F1 -scores are as follows: 0.992 for galaxies; 0.967 for quasars; and 0.985 for stars. In terms of the F1-scores for the three classes, SHEEP is found to outperform a recent RandomForest-based classification approach using an essentially identical data set. Our methodology also facilitates model and data set explainability via feature importances; it also allows the selection of sources whose uncertain classifications may make them interesting sources for follow-up observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploring the Effects of Population and Employment Characteristics on Truck Flows: An Analysis of NextGen NHTS Origin-Destination Data

Truck transportation remains the dominant mode of US freight transportation because of its advantages, such as the flexibility of accessing pickup and drop-off points and faster delivery. Because of the massive freight volume transported by trucks, understanding the effects of population and employment characteristics on truck flows is critical for better transportation planning and investment decisions. The US Federal Highway Administration published a truck travel origin-destination data set as part of the Next Generation National Household Travel Survey program. This data set contains the total number of truck trips in 2020 within and between 583 predefined zones encompassing metropolitan and nonmetropolitan statistical areas within each state and Washington, DC. In this study, origin-destination-level truck trip flow data was augmented to include zone-level population and employment characteristics from the US Census Bureau. Census population and County Business Patterns data were included. The final data set was used to train a machine learning algorithm-based model, Extreme Gradient Boosting (XGBoost), where the target variable is the number of total truck trips. Shapley Additive ExPlanation (SHAP) was adopted to explain the model results. Results showed that the distance between the zones was the most important variable and had a nonlinear relationship with truck flows.

Uddin, Majbah↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Accurate prediction of global-density-dependent range-separation parameters based on machine learning

In this work, we develop an accurate and efficient XGBoost machine learning model for predicting the global-density-dependent range-separation parameter, ωGDD, for long-range corrected functional (LRC)-ωPBE. This ωGDDML model has been built using a wide range of systems (11 466 complexes, ten different elements, and up to 139 heavy atoms) with fingerprints for the local atomic environment and histograms of distances for the long-range atomic correlation for mapping the quantum mechanical range-separation values. The promising performance on the testing set with 7046 complexes shows a mean absolute error of 0.001 117 a0−1 and only five systems (0.07%) with an absolute error larger than 0.01 a0−1, which indicates the good transferability of our ωGDDML model. In addition, the only required input to obtain ωGDDML is the Cartesian coordinates without electronic structure calculations, thereby enabling rapid predictions. LRC-ωPBE(ωGDDML) is used to predict polarizabilities for a series of oligomers, where polarizabilities are sensitive to the asymptotic density decay and are crucial in a variety of applications, including the calculations of dispersion corrections and refractive index, and surpasses the performance of all other popular density functionals except for the non-tuned LRC-ωPBE. Finally, LRC-ωPBE (ωGDDML) combined with (extended) symmetry-adapted perturbation theory is used in calculating noncovalent interactions to further show that the traditional ab initio system-specific tuning procedure can be bypassed. The present study not only provides an accurate and efficient way to determine the range-separation parameter for LRC-ωPBE but also shows the synergistic benefits of fusing the power of physically inspired density functional LRC-ωPBE and the data-driven ωGDDML model.

Chemistry↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

APSO-enhanced algebraic derivative estimation approach for real-time traffic flow prediction on critical road sections during wildfire evacuation

In rapid-onset disaster scenarios such as wildfires, evacuation traffic often significantly deviates from historical patterns, rendering conventional data-driven forecasting methods less effective. To address this challenge, we propose an improved algebraic derivative estimation (ADE) incorporating particle swarm optimization (PSO) for real-time traffic flow prediction. Our approach dynamically adjusts the ADE prediction time window at each step by minimizing a cost function based on the mean and variance of accumulated forecasting errors within the window, thereby balancing bias and variability. We evaluate the method using traffic data from the January 2025 California wildfires, focusing on key road segments critical for large-scale evacuations. The results demonstrate that our approach surpasses established machine learning and deep learning models—XGBoost, LSTM, and GRU—in predictive accuracy and maintains high computational efficiency. Notably, the proposed method eliminates the need for offline model training. Moreover, rapid PSO-based tuning enables real-time deployment, which provides a crucial advantage in scenarios where evacuation timings and road closures change dynamically. In conclusion, these findings highlight the benefits of the PSO-enhanced ADE framework for emergency traffic management, where rapid, data-sparse forecasts are essential for effective evacuation planning.

Algebraic derivative estimation↗

Carbon-enhanced metal-poor star candidates from BP/RP spectra in Gaia DR3

ABSTRACT Carbon-enhanced metal-poor (CEMP) stars comprise almost a third of stars with [Fe/H] < −2, although their origins are still poorly understood. It is highly likely that one sub-class (CEMP-s stars) is tied to mass-transfer events in binary stars, while another sub-class (CEMP-no stars) are enriched by the nucleosynthetic yields of the first generations of stars. Previous studies of CEMP stars have primarily concentrated on the Galactic halo, but more recently they have also been detected in the thick disc and bulge components of the Milky Way. Gaia DR3 has provided an unprecedented sample of over 200 million low-resolution (R ≈ 50) spectra from the BP and RP photometers. Training on the CEMP catalogue from the SDSS/SEGUE database, we use XGBoost to identify the largest all-sky sample of CEMP candidate stars to date. In total, we find 58 872 CEMP star candidates, with an estimated contamination rate of 12 per cent. When comparing to literature high-resolution catalogues, we positively identify 60–68 per cent of the CEMP stars in the data, validating our results and indicating a high completeness rate. Our final catalogue of CEMP candidates spans from the inner to outer Milky Way, with distances as close as r ∼ 0.8 kpc from the Galactic centre, and as far as r > 30 kpc. Future higher resolution spectroscopic follow-up of these candidates will provide validations of their classification and enable investigations of the frequency of CEMP-s and CEMP-no stars throughout the Galaxy, to further constrain the nature of their progenitors.

79 ASTRONOMY AND ASTROPHYSICS↗

Towards an astronomical foundation model for stars with a transformer-based model

ABSTRACT Rapid strides are currently being made in the field of artificial intelligence using transformer-based models like Large Language Models (LLMs). The potential of these methods for creating a single, large, versatile model in astronomy has not yet been explored. In this work, we propose a framework for data-driven astronomy that uses the same core techniques and architecture as used by LLMs. Using a variety of observations and labels of stars as an example, we build a transformer-based model and train it in a self-supervised manner with cross-survey data sets to perform a variety of inference tasks. In particular, we demonstrate that a single model can perform both discriminative and generative tasks even if the model was not trained or fine-tuned to do any specific task. For example, on the discriminative task of deriving stellar parameters from Gaia XP spectra, we achieve an accuracy of 47 K in Teff, 0.11 dex in log g, and 0.07 dex in [M/H], outperforming an expert XGBoost model in the same setting. But the same model can also generate XP spectra from stellar parameters, inpaint unobserved spectral regions, extract empirical stellar loci, and even determine the interstellar extinction curve. Our framework demonstrates that building and training a single foundation model without fine-tuning using data and parameters from multiple surveys to predict unmeasured observations and parameters is well within reach. Such ‘Large Astronomy Models’ trained on large quantities of observational data will play a large role in the analysis of current and future large surveys.

Leung, Henry W. (ORCID:0000000200362752)↗

Machine learning predicts new anti-CRISPR proteins

The increasing use of CRISPR–Cas9 in medicine, agriculture, and synthetic biology has accelerated the drive to discover new CRISPR–Cas inhibitors as potential mechanisms of control for gene editing applications. Many anti-CRISPRs have been found that inhibit the CRISPR–Cas adaptive immune system. However, comparing all currently known anti-CRISPRs does not reveal a shared set of properties for facile bioinformatic identification of new anti-CRISPR families. Here, we describe AcRanker, a machine learning based method to aid direct identification of new potential anti-CRISPRs using only protein sequence information. Using a training set of known anti-CRISPRs, we built a model based on XGBoost ranking. We then applied AcRanker to predict candidate anti-CRISPRs from predicted prophage regions within self-targeting bacterial genomes and discovered two previously unknown anti-CRISPRs: AcrllA20 (ML1) and AcrIIA21 (ML8). We show that AcrIIA20 strongly inhibits Streptococcus iniae Cas9 (SinCas9) and weakly inhibits Streptococcus pyogenes Cas9 (SpyCas9). We also show that AcrIIA21 inhibits SpyCas9, Streptococcus aureus Cas9 (SauCas9) and SinCas9 with low potency. The addition of AcRanker to the anti-CRISPR discovery toolkit allows researchers to directly rank potential anti-CRISPR candidate genes for increased speed in testing and validation of new anti-CRISPRs. A web server implementation for AcRanker is available online at http://acranker.pythonanywhere.com/.

59 BASIC BIOLOGICAL SCIENCES↗

A Machine Learning based Approach of Estimating Equivalent Circuit Model Parameters at Different SoCs of Li-ion Batteries from Voltage Relaxation

Abstract: In this study, an approach of estimating the equivalent circuit model (ECM) parameters for Li-ion batteries (LIBs) is proposed based on the voltage value at different intervals while relaxing the LIB after discharge. The typical approach for estimating ECM parameters of a LIB is to conduct electrochemical impedance spectroscopy (EIS) measurements at different frequencies and fit them to a predefined circuit model, which requires additional measuring arrangements and specialized devices. The proposed methodology utilizes four different voltages at 0s, 60s, 360s, and 1800s alongside the specific state of charge (SoC) value for a specific constant discharge current value of ~1C until the relaxation stage to train and evaluate three regression-based machine learning models— Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Gaussian Process Regression (GPR)—for estimating the ECM parameters of the selected model. Bayesian optimization is employed for hyperparameter tuning to achieve optimal performance for all the regressor models, among which, the GPR provided the best performance with the root-mean-squared error (RMSE) of less than 4x10-4 on average for the resistive components and less than 0.27 for capacitive components with excellent R2 scores. The simplicity of the approach enables it to eliminate the need for sophisticated measuring equipment and computation power.

Sagar, Md. Samiul [The University of Alabama (UA)]↗

Performance Comparison of Clipping Detection Techniques in AC Power Time Series

In this research, a variety of methods were developed to detect clipping periods in AC power time series. AC power data streams associated with 36 unique systems across the United States were collected, and data points representing clipping periods were manually labeled by experts. Using this data set for training and validation, novel logic-based and machine learning (ML) approaches were developed to classify time series values as clipping or non-clipping. These approaches were compared to the RdTools method for detecting clipping periods. The logic-based and ML XGBoost approaches achieved F-scores of 85.0 and 77.6, respectively, when cross-validated against the manually labeled data, as compared to the current RdTools approach (F-score of 56.4), indicating a significant improvement at detecting clipping periods. Additionally, the effects of each clipping filter when evaluating system degradation rates were assessed, using 31 unique systems across the United States. Results indicate that estimated system degradation rate can vary based on the type of clipping filter used, by up to 0.6% degradation rate for some cases.

clipping↗