Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boost machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

Predicting solid state material platforms for quantum technologies

Semiconductor materials provide a compelling platform for quantum technologies (QT). However, identifying promising material hosts among the plethora of candidates is a major challenge. Therefore, we have developed a framework for the automated discovery of semiconductor platforms for QT using material informatics and machine learning methods. Different approaches were implemented to label data for training the supervised machine learning (ML) algorithms logistic regression, decision trees, random forests and gradient boosting. We find that an empirical approach relying exclusively on findings from the literature yields a clear separation between predicted suitable and unsuitable candidates. In contrast to expectations from the literature focusing on band gap and ionic character as important properties for QT compatibility, the ML methods highlight features related to symmetry and crystal structure, including bond length, orientation and radial distribution, as influential when predicting a material as suitable for QT.

36 MATERIALS SCIENCE↗

Revealing EDL-driven reduction mechanisms in binary, ternary, and quaternary fluorinated electrolytes via an integrated MD–DFT–ML framework

Accurately predicting solid electrolyte interphase (SEI) formation requires explicitly resolving the electric double layer (EDL) structure, which deviates significantly from that of the bulk electrolyte. Although an established molecular dynamics (MD) and Density Functional Theory (DFT) framework can model SEI formation by evaluating reduction reactions of local clusters in the EDL, it suffers from a combinatorial computational bottleneck. To overcome this limitation, we introduce a machine-learning-accelerated simulation workflow (MD–DFT–ML), integrating a gradient-boosted regression model trained on EDL composition data to efficiently predict reduction potentials. We apply this framework to seven fluorinated electrolytes comprising fluorinated anions, a fluorinated ester solvent, two types of diluent (ion-solvating ester vs. non-solvating ether), and an FEC additive. The analysis shows that the EDL selectively accumulates cation-binding species; consequently, the non–cation-binding ether diluent rarely enters the EDL and makes minimal contributions to SEI formation. DFT calculations on statistically representative EDL clusters provide reduction potentials and fluorine-release pathways, while the ML model, which substantially reduces the DFT workload, predicts cluster reduction energies with a mean absolute error of 0.1 eV. The combined MD–DFT–ML approach also quantifies contributions from different sources to LiF formation in the SEI. This methodology establishes a generalizable route for multiscale modeling electrolyte and interphase design for next-generation electrochemical energy-storage systems.

DFT-MD-ML workflow↗

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis↗

Unraveling the Correlation between Raman and Photoluminescence in Monolayer MoS 2 through Machine‐Learning Models

Abstract 2D transition metal dichalcogenides (TMDCs) with intense and tunable photoluminescence (PL) have opened up new opportunities for optoelectronic and photonic applications such as light‐emitting diodes, photodetectors, and single‐photon emitters. Among the standard characterization tools for 2D materials, Raman spectroscopy stands out as a fast and non‐destructive technique capable of probing material's crystallinity and perturbations such as doping and strain. However, a comprehensive understanding of the correlation between photoluminescence and Raman spectra in monolayer MoS 2 remains elusive due to its highly nonlinear nature. Here, the connections between PL signatures and Raman modes are systematically explored, providing comprehensive insights into the physical mechanisms correlating PL and Raman features. This study's analysis further disentangles the strain and doping contributions from the Raman spectra through machine‐learning models. First, a dense convolutional network (DenseNet) to predict PL maps by spatial Raman maps is deployed. Moreover, a gradient boosted trees model (XGBoost) with Shapley additive explanation (SHAP) to bridge the impact of individual Raman features in PL features is applied. Last, a support vector machine (SVM) to project PL features on Raman frequencies is adopted. This work may serve as a methodology for applying machine learning to characterizations of 2D materials.

Lu, Ang‐Yu↗

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

ANN↗

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif↗

Machine Learning for the Prediction of Local Asteroid Damages

Risk assessment studies of local asteroid hazards traditionally simulate the physics of meteors with engineering models tailored to analyze tens-of-millions of scenarios. However, these simplified approaches still need to solve time-dependent ODEs to model the entry process and the resulting ground damage. With a computational cost of O(0.01 CPU.s) per scenario, simulating these large numbers of potential entry conditions in risk assessment studies can take several days on local computers. To improve computational efficiency, we propose in this paper an orthogonal approach based on machine learning models to predict the size of damaged areas given a list of entry parameters. We train 5 machine learning methods and compare the predictions to the outputs of the PAIR model, first only with primitive entry condition variables, and then with more advanced features. We find that complex models like neural networks are well-suited to estimate blast hazards, while simpler linear models can accurately assess thermal damage. For both types of hazards, the radii of damaged areas can be predicted with around 10% average errors and a coefficient of determination (R2) of 0.99. The CPU time is decreased by a factor O(10 3 ) compared to the PAIR model, which enables the simulation of millions of scenarios in minutes, on a local computer. We then use the same machine learning approaches for a classification task where the models are trained to predict if an asteroid will produce a given level of damage. Results show that complex models like the gradient boosting classifier and the neural network can perform this task with 98% accuracy. Beyond surrogate models, we finally incorporate the machine learning algorithms to the state-of-the-art Shapley sensitivity analysis and present a ranking of the entry parameters based on their contributions to ground damages.

SMD↗

Selecting durable building envelope systems with machine learning assisted hygrothermal simulations database

Hygrothermal simulations provide insight into the energy performance and moisture durability of building envelope components under dynamic conditions. The inputs required for hygrothermal simulations are extensive, and carrying out simulations and analyses requires expert knowledge. An expert system, the Building Science Advisor (BSA), has been developed to predict the performance and select the energy-efficient and durable building envelope systems for different climates. The BSA consists of decision rules based on expert opinions and thousands of parametric simulation results for selected wall systems. The number of potential wall systems results in millions, too many to simulate all of them. We present how machine learning can help predict durability data, such as mold growth, while minimizing the number of simulations needed to run. The simulation results are used for training and validation of machine learning tools for predicting wall durability. We tested Artificial Neural Network (ANN) and Gradient Boosted Decision Trees (GBDT) for their applicability and model accuracy. Models developed with both methods showed adequate prediction performance (root mean square error of 0.195 and 0.209, respectively). Finally, we introduce how the information supports guidance for envelope design via an easy-to-use web-based tool that does not require the end-user to run hygrothermal simulations.

Salonvaara, Mikael↗

A novel machine learning based identification of potential adopter of rooftop solar photovoltaics

With the proliferation of rooftop solar photovoltaic installations, there is a need to proactively predict consumer potential for solar photovoltaic adoption, for improved electric utility planning and operation. Traditional analytical modeling approaches are limited to a few survey features and a larger part of the survey would remain untouched by the decision model. This article presents a novel, data-driven modeling approach that strategically prunes a large set of consumer profile features using a machine learning framework to train a model for predicting potential solar adoption. The approach utilizes the Gradient Boosting Decision Tree model through a Light Gradient Boosting framework that improves significantly over the poor prediction accuracy of the existing approaches. Model training using focal-loss based supervision is used to overcome the difficulty in identifying the potential adopters that is inherent in conventional data-driven models. In addition, to overcome possible data sparsity in a limited survey sample, a Generative Adversarial Network is presented to create synthetic user samples and its effectiveness on model performance is assessed. A Bayesian optimization approach is used to systematically arrive at the hyperparameters of the proposed model. Validation of the presented approach on a survey data collected by the National Rural Electric Cooperative Association in Virginia in 2018 demonstrates the excellent predictive capability of the machine learning based approach to modeling solar adoption reliably.

14 SOLAR ENERGY↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Predictive understanding of the surface tension and velocity of sound in ionic liquids using machine learning

Knowledge of the physical properties of ionic liquids (ILs), such as the surface tension and speed of sound, is important for both industrial and research applications. Unfortunately, technical challenges and costs limit exhaustive experimental screening efforts of ILs for these critical properties. Previous work has demonstrated that the use of quantum-mechanics-based thermochemical property prediction tools, such as the conductor-like screening model for real solvents, when combined with machine learning (ML) approaches, may provide an alternative pathway to guide the rapid screening and design of ILs for desired physiochemical properties. However, the question of which machine-learning approaches are most appropriate remains. In the present study, we examine how different ML architectures, ranging from tree-based approaches to feed-forward artificial neural networks, perform in generating nonlinear multivariate quantitative structure–property relationship models for the prediction of the temperature- and pressure-dependent surface tension of and speed of sound in ILs over a wide range of surface tensions (16.9–76.2 mN/m) and speeds of sound (1009.7–1992 m/s). The ML models are further interrogated using the powerful interpretation method, shapley additive explanations. We find that several different ML models provide high accuracy, according to traditional statistical metrics. The decision tree-based approaches appear to be the most accurate and precise, with extreme gradient-boosting trees and gradient-boosting trees being the best performers. However, our results also indicate that the promise of using machine-learning to gain deep insights into the underlying physics driving structure–property relationships in ILs may still be somewhat premature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An interpretable machine learning framework to understand bikeshare demand before and during the COVID-19 pandemic in New York City

In recent years, bikesharing systems have become increasingly popular as affordable and sustainable micromobility solutions. Advanced mathematical models such as machine learning are required to generate good forecasts for bikeshare demand. Here, this study proposes a machine learning modeling framework to estimate hourly demand in a large-scale bikesharing system. Two Extreme Gradient Boosting models were developed: one using data from before the COVID-19 pandemic (March 2019 to February 2020) and the other using data from during the pandemic (March 2020 to February 2021). Furthermore, a model interpretation framework based on SHapley Additive exPlanations was implemented. Based on the relative importance of the explanatory variables considered in this study, share of female users and hour of day were the two most important explanatory variables in both models. However, the month variable had higher importance in the pandemic model than in the pre-pandemic model.

99 GENERAL AND MISCELLANEOUS↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning Analysis of Hydrologic Exchange Flows and Transit Time Distributions in a Large Regulated River

Hydrologic exchange between river channels and adjacent subsurface environments is a key process that influences water quality and ecosystem function in river corridors. High-resolution numerical models were often used to resolve the spatial and temporal variations of exchange flows, which are computationally expensive. In this study, we adopt Random Forest (RF) and Extreme Gradient Boosting (XGB) approaches for deriving reduced order models of hydrologic exchange flows and associated transit time distributions, with integrated field observations (e.g., bathymetry) and hydrodynamic simulation data (e.g., river velocity, depth). The setup allows an improved understanding of the influences of various physical, spatial, and temporal factors on the hydrologic exchange flows and transit times. The predictors also contain those derived using hybrid clustering, leveraging our previous work on river corridor system hydromorphic classification. The machine learning-based predictive models are developed and validated along the Columbia River Corridor, and the results show that the top parameters are the thickness of the top geological formation layer, the flow regime, river velocity, and river depth; the RF and XGB models can achieve 70% to 80% accuracy and therefore are effective alternatives to the computational demanding numerical models of exchange flows and transit time distributions. Each machine learning model with its favorable configuration and setup have been evaluated. The transferability of the models to other river reaches and larger scales, which mostly depends on data availability, is also discussed.

97 MATHEMATICS AND COMPUTING↗

Automatic DDoS Attack Detection on SDNs: Preprint

Denial of Service (DoS) and Distributed Denial of Service (DDoS) attacks pose a serious threat to computing networks - especially to critical systems within the U.S. electrical grid. As attack mechanisms have increased in complexity and variety, more sophisticated detection mechanisms have become necessary to ensure network security. This paper explores the use of artificial intelligence to automate the process of detection and mitigation of DoS and DDoS attacks within the framework of Software-Defined Networking (SDN), to a high degree. Machine learning algorithms are trained to recognize DoS and DDoS attacks and are deployed in real-time to mitigate malicious network traffic. The results show a well-tuned gradient-boosted decision tree detecting DoS and DDoS attacks, as well as initial successful mitigation of attacks within an SDN framework.

cyber detection↗

Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM 2.5 ) Using Satellite Data Over Large Regions

Reconstructing the distribution of fine particulate matter (PM 2.5 ) in space and time, even far from ground monitoring sites, is an important exposure science contribution to epidemiologic analyses of PM 2.5 health impacts. Flexible statistical methods for prediction have demonstrated the integration of satellite observations with other predictors, yet these algorithms are susceptible to overfitting the spatiotemporal structure of the training datasets. We present a new approach for predicting PM 2.5 using machine-learning methods and evaluating prediction models for the goal of making predictions where they were not previously available. We apply extreme gradient boosting (XGBoost) modeling to predict daily PM 2.5 on a 1 x 1 km 2 resolution for a 13 state region in the Northeastern USA for the years 2000–2015 using satellite-derived aerosol optical depth and implement a recursive feature selection to develop a parsimonious model. We demonstrate excellent predictions of withheld observations but also contrast an RMSE of 3.11 μg/m 3 in our spatial cross-validation withholding nearby sites versus an overfit RMSE of 2.10 μg/m 3 using a more conventional random ten-fold splitting of the dataset. As the field of exposure science moves forward with the use of advanced machine-learning approaches for spatiotemporal modeling of air pollutants, our results show the importance of addressing data leakage in training, overfitting to spatiotemporal structure, and the impact of the predominance of ground monitoring sites in dense urban sub-networks on model evaluation. The strengths of our resultant modeling approach for exposure in epidemiologic studies of PM 2.5 include improved efficiency, parsimony, and interpretability with robust validation while still accommodating complex spatiotemporal relationships.

air pollution↗