Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boosting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Machine Learning Analysis of Impact of Western US Fires on Central US Hailstorms

Fires, including wildfires, harm air quality and essential public services like transportation, communication, and utilities. These fires can also influence atmospheric conditions, including temperature and aerosols, potentially affecting severe convective storms. Here, we investigate the remote impacts of fires in the western United States (WUS) on the occurrence of large hail (size: $\geqslant$ 2.54 cm) in the central US (CUS) over the 20-year period of 2001–20 using the machine learning (ML), Random Forest (RF), and Extreme Gradient Boosting (XGB) methods. The developed RF and XGB models demonstrate high accuracy (> 90%) and F1 scores of up to 0.78 in predicting large hail occurrences when WUS fires and CUS hailstorms coincide, particularly in four states (Wyoming, South Dakota, Nebraska, and Kansas). The key contributing variables identified from both ML models include the meteorological variables in the fire region (temperature and moisture), the westerly wind over the plume transport path, and the fire features (i.e., the maximum fire power and burned area). Importantly, the results confirm a linkage between WUS fires and severe weather in the CUS, corroborating the findings of our previous modeling study conducted on case simulations with a detailed physics model.

54 ENVIRONMENTAL SCIENCES↗

Harnessing Machine Learning to Predict MoS 2 Solid Lubricant Performance

Physical vapor deposited (PVD) molybdenum disulfide (MoS 2 ) solid lubricant coatings are an exemplar material system for machine learning methods due to small changes in process variables often causing large variations in microstructure and mechanical/tribological properties. Here, in this work, a gradient boosted regression tree machine learning method is applied to an existing experimental data set containing process, microstructure, and property information to create deeper insights into the process-structure–property relationships for molybdenum disulfide (MoS 2 ) solid lubricant coatings. The optimized and cross-validated models show good predictive capabilities for density, reduced modulus, hardness, wear rate, and initial coefficients of friction. The contribution of individual deposition variables (i.e., argon pressure, deposition power, target conditioning) on coating properties is highlighted through feature importance. The process-property relationships established herein show linear and non-linear relationships and highlight the influence of uncontrolled deposition variables (i.e., target conditioning) on the tribological performance.

MoS2↗

A Market Feedback Framework for Improved Estimates of the Arbitrage Value of Energy Storage Using Price-Taker Models

Price-taker (PT) models are often used to assess the potential value or revenue of energy arbitrage opportunities for energy storage in wholesale markets. But as greater amounts of energy storage are deployed on the grid, current PT models fail to predict the effects that energy storage itself can have on market prices. This can lead to an overestimation of the economic value of storage and an inability to capture price suppression. In this paper, we propose the use of a modified PT model to simulate the impact of increased storage deployment on energy prices and the resulting impact on revenue. Our method uses a gradient-boosting regressor to estimate the impact on prices, and we apply our method on historical price data from the PJM and California Independent System Operator wholesale markets. We use this approach to explore possible causes of electricity price suppression that occur from storage capacity additions, which is generally not possible with PT models.

arbitrage↗

Hybrid data-driven cement-stabilized soil design: An integration of machine learning, multi-objective optimization, and life cycle assessment

Soil stabilization is crucial in geotechnical engineering, yet conventional methods are often time-consuming, resource-intensive, and environmentally unsustainable. Despite growing interest in Machine Learning (ML) and optimization tools for mix design, few studies integrate these methods with decision-making techniques and environmental assessment to support practical implementation. This study proposes a hybrid data-driven framework for predicting strength, optimizing mix compositions, and evaluating environmental impacts via life cycle assessment of cement-stabilized soft soils. Six ML models were evaluated, and the top-performing eXtreme Gradient Boosting (XGB) model was further improved using the Grey Wolf Optimizer (GWO). The optimized XGB-GWO model, integrated with a polynomial cost function, served as the objective function in a multi-objective optimization problem solved via the Non-Dominated Sorting Genetic Algorithm II (NSGA-II), with final mix selection guided by the entropy-weighted TOPSIS method. Validation through a case study produced mix designs offering superior strength-cost trade-offs, with the optimal mix achieving 2243.2 kPa unconfined compressive strength and a 16.07 % reduction in carbon emissions compared to the highest-cost design. In conclusion, this study offers a sustainable, scalable approach to soil stabilization and supports informed decision-making in construction.

Life cycle assessment↗

Growing grasses in unprofitable areas of US Midwest croplands could increase species richness

The US has large potential to grow perennial energy crops, but because these crops are rarely grown in current agricultural landscapes, it is unclear how biodiversity may be affected. Over time, as agriculture has increased, many grassland species have declined. In addition, not all agricultural land is profitable for growing annual crops. Unprofitable areas were responsible for a loss of approximately $110 million USD per year from 2013 to 2016. Based on this, we want to know how converting less-profitable portions of agricultural fields to switchgrass, a native prairie grass, would influence species occurrence. To address this question, we developed an alternative landscape in which clustered corn/soy acres with a low return on investment (ROI) were replaced with grassland. We also developed and validated species distribution models to predict changes in species occurrence for 28 avian species in Iowa in response to landscape management. Furthermore, we compared results for three different models: Random forest (RF), Stochastic gradient boosting (GBM), and Neural network (Nnet) and found that all models performed well and predicted similar species distribution. Predicted species richness increased by 3.66% (RF), 2.79% (GBM), and 7.51% (Nnet) when we simulated a change in management for ~3% of Iowa's low ROI corn/soybean areas to grassland. If harvested, these areas could generate approximately 7.6 million dry tons/year of switchgrass for bioenergy, thereby increasing farmers earnings. Unprofitable areas tended to occur along streams, which suggest that incorporating partially harvested riparian buffers can benefit avian biodiversity, while improving water quality and reducing unnecessary costs for farmers.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Linear model decision trees as surrogates in optimization of engineering applications

Machine learning models are promising as surrogates in optimization when replacing difficult to solve equations or black-box type models. This work demonstrates the viability of linear model decision trees as piecewise-linear surrogates in decision-making problems. Linear model decision trees can be represented exactly in mixed-integer linear programming (MILP) and mixed-integer quadratic constrained programming (MIQCP) formulations. Furthermore, they can represent discontinuous functions, bringing advantages over neural networks in some cases. We present several formulations using transformations from Generalized Disjunctive Programming (GDP) formulations and modifications of MILP formulations for gradient boosted decision trees (GBDT). We then compare the computational performance of these different MILP and MIQCP representations in an optimization problem and illustrate their use on engineering applications. Importantly, we observe faster solution times for optimization problems with linear model decision tree surrogates when compared with GBDT surrogates using the Optimization and Machine Learning Toolkit (OMLT).

42 ENGINEERING↗

Charged particle reconstruction in CLAS12 using Machine Learning

In this work, we present studies of track parameter reconstruction from raw information in CLAS12 detector's Drift Chambers, using Machine Learning (ML). We study the resolution of tracks reconstructed with different types of ML models/algorithms, including Multi-Layer Perceptron (MLP), Extremely Randomized Trees (ERT) and Gradient Boosting Trees (GBT) using simulated data. We find that the resulting ML model is capable of reconstructing track parameters (particle momentum, and polar and azimuthal angles) with accuracy similar to Hit Based (HB) tracking code, but $150$ times faster. Moreover, physics reactions can be identified using the particles reconstructed by the neural network in real-time (with a rate of about $34~kHz$) during experimental data collection. The developed model can be used in numerous applications, such as triggering specific physics reactions in real-time, detector performance monitoring, and real-time detector calibration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning surrogate of physics-based building-stock simulator for end-use load forecasting

Building energy models are used to simulate heat and mass transfer and estimate end-use load in buildings. With the proliferation of solar photovoltaics on residential and commercial buildings, increasingly, buildings are expected to provide grid services, for which accurate and computationally efficient building energy simulations and end-use load prediction are imperative. Existing building energy simulation tools, however, have significant computational overhead that make them less practical in real-time deployment for optimization, design, uncertainty quantification and control in building energy management systems. Here this article presents a data-driven machine learning model based on light gradient boosting method (LightGBM) as a surrogate for a physics-based simulator for residential buildings to predict end-use load. The machine learning based surrogate model accounts for time-series related variables, seasonality and trend component of end-use load, and history of end-use load. The accuracy of the surrogate model is assessed on the prediction of the load profiles of 100 different houses in Cook County, Illinois, USA. The LightGBM surrogate model is shown to reduce the root-mean-squared error by 53% relative to a reference decision tree (DT) based model reported previously in the literature. Moreover, the model predicts the load spikes and high-ramp rate events throughout the year which are often the Achilles heel of other models in the literature. The machine learning based surrogate model is demonstrated to be computationally efficient, with a ten-fold reduction in the computational time compared to a physics-based building energy simulation, and suitable for uncertainty analysis and real-time control of building characteristics in response to uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Metal hydride composition-derived parameters as machine learning features for material design and H 2 storage

Though hydrogen is a promising energy carrier for a green future, many challenges persist. One is the difficulty in engineering storage solutions, with metal hydrides being a leading contender among solid-state strategies. To facilitate efficient searching of candidate materials, ridge regression, simple decision trees, random forest ensembles, and gradient boosting ensembles were employed to predict the energy of formation, with the random forest ensemble resulting in the lowest test set error. First, two public databases, Materials Project and HydPark, were searched for metal hydrides. Feature engineering was performed before the models were developed, resulting in electronegativity, density, atomic density, d-character, f-character, band gap, hydrogen weight fraction, magnetization, temperature, and pressure being retained. The models were then benchmarked by the lowest test error before a random forest ensemble was used to populate entries missing energy of formation. Furthermore, all were then scored by hydrogen storage capacity and energy of formation suitability. Readily available features including several derived from only the chemical formula which were found to be highly predictive. and so are promising for high-throughput screening of arbitrary novel hydride formulations and blends for thermodynamic feasibility.

25 ENERGY STORAGE↗

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗

"Hidden" hydrothermal technical potential & technoeconomics: Revealing permeability & fluids with more data

Historical hydrothermal estimates have largely relied on temperature or heat flow estimates ignoring the need for natural flowing fluids. More accurate hydrothermal estimates require some indication of permeability and fluids that naturally exist in the subsurface. This paper describes a novel approach that includes proxies of permeability and fluids in hydrothermal estimates by leveraging the relatively data-rich Great Basin. Specifically, nameplate capacities (megawatts) of operating geothermal plants, negative (0 megawatt) locations and 48 geophysical and geologic features are used to used in eXtreme Gradient Boosting (XGBoost) regression to make hydrothermal capacity predictions. Additionally, this work inputs the XGBoost-based hydrothermal predictions into the Renewable Energy Potential (reV) model to quantify technical capacity, its uncertainty and techno-economics. Compared to historical hydrothermal estimates, these predictions adhere to the 37 operating geothermal plants and negative locations. We present a method for subsampling the negative sites to bring the labels into balance that uses the geologic domain knowledge to proportionally represent negatives. Overall, the distributions of the hydrothermal technical capacity and the site levelized cost of energy are respectively much tighter, lower and more accurate than the previous estimates for the Great Basin, as they include geological and geophysical surrogates for permeability and fluids. Percentile (50th and 90th, median and high estimate, respectively) models provide bookends for these metrics.

13 HYDRO ENERGY↗

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage↗

Machine-learning and first-principles investigation of lightweight medium-entropy alloys for hydrogen-storage applications

The transition to a low-carbon economy demands efficient and sustainable energy-storage solutions, with hydrogen emerging as a promising clean-energy carrier and with metal hydrides recognized for their hydrogen-storage capacity. Here, we leverage machine learning (ML) to predict hydrogen-to-metal (H/M) ratios and solution energy by incorporating thermodynamic parameters and local lattice distortion (LLD) as key features. Our best-performing ML model provides improvements to H/M ratios and solution energies over a broad class of medium-entripy alloys (easily extendable to multi-principal-element alloys), such as Ti–Nb-X (X = Mo, Cr, Hf, Ta, V, Zr) and Co–Ni-X (X = Al, Mg, V). Ti–Nb–Mo alloys reveal compositional effects in H-storage behavior, in particular Ti, Nb, and V enhance H-storage capacity, while Mo reduces H/M and hydrogen weight percent by 40–50 %. We attributed results in molybdenum-rich alloys to slow hydrogen kinetics, as validated by our pressure-composition-temperature (PCT) isotherm experiments on pure Ti and Ti 5 Mo 95 alloys. Density functional theory (DFT) and molecular dynamics (MD) simulations also confirm that Ti and Nb promote H diffusion, whereas Mo hinders it, highlighting the interplay between electronic structure, lattice distortions, and hydrogen uptake. Notably, our Gradient Boosting Regression model identifies LLD as a critical factor in H/M predictions. Here, to aid material selection, we present two periodic tables illustrating elemental effects on (a) H 2 wt% and (b) solution energy, derived from ML, and provide a reference for identifying alloying elements that enhance hydrogen solubility and storage.

08 HYDROGEN↗

A large-scale comparison of Artificial Intelligence and Data Mining (AI&DM) techniques in simulating reservoir releases over the Upper Colorado Region

In recent years, the Artificial Intelligence and Data Mining (AI&DM) models have become popular tools in assisting various aspects of reservoir operation. However, the practical uses are still rarely reported. Comparison experiment of many AI&DM models over a large number of reservoir cases is particularly valuable to help reservoir operators first examine the usefulness and transferability of different AI&DM models, and then identify the most stable and reliable AI&DM model in assist of various decision-making processes. In this study, a total of 12 AI&DM models with different parameterizations and simulation scenarios are comprehensively tested out and compared in simulating the controlled reservoir outflows of 33 reservoir cases over the Upper Colorado Region, United States. Results show that the Random Forecast and the Long-Short-Term-Memory model could consistently derive the best statistical performance than other models under the baseline simulation scenario. The employed AI&DM models could obtain satisfactory statistical interquartile ranges (25–75%) between [0.6–0.9], [0.3–0.8], and [0.2–0.8], for CORR, NSE, and KGE measurements, respectively, and [1.5–6.5], [–15 to 20], and [0.5–8.5] for the normalized RMSE, PBIAS and RSR measurements, respectively. Results also show Multi-Layer Perceptron model and Extreme Gradient Boosting Tree Algorithm produced more stable and superior performance than other models under more complex input scenarios. We also found that the performance of different AI&DM models are closely relevant to the reservoir elevations, sizes, and functionalities. Discussions were made about the sensitivity of AI&DM models’ parameterizations and the key advantages of AI&DM models over the rule-based reservoir models. We further identify that the main advantage of AI&DM models is the flexibility in designing input structures, whereas the rule-based simulation model is rather limited. Future studies were suggested regarding the best way reservoir operators and researchers could use, select, and apply different AI&DM models in simulating reservoir releases under different natural and modeling environments. Finally, this comparison study also serves as a reference and a piece of groundwork for further promoting the practical uses of AI&DM models in assisting reservoir operation.

54 ENVIRONMENTAL SCIENCES↗

The Global LAnd Surface Satellite (GLASS) evapotranspiration product Version 5.0: Algorithm development and preliminary validation

An accurate estimation of spatially and temporally continuous global terrestrial evapotranspiration (ET) is essential in the assessment of surface energy, water and carbon cycles. The Global LAnd Surface Satellite (GLASS) ET product Version 4.0 (v4.0) based on the Bayesian model averaging (BMA) method was generated to estimate global terrestrial ET. However, certain uncertainty for the GLASS ET product v4.0 limits its application. In this study, we introduced the deep neural networks (DNN) merging framework to improve terrestrial ET estimation for GLASS ET product Version 5.0 (v5.0) generation by integrating five satellite-derived ET products [Moderate Resolution Imaging Spectroradiometer (MODIS) ET product (MOD16), Shuttleworth–Wallace dual-source ET product (SW), Priestley–Taylor-based ET product (PT-JPL), modified satellite-based Priestley–Taylor ET product (MS-PT) and simple hybrid ET product (SIM)]. We compared the performance of DNN method against other merging methods, including GLASS ET algorithm v4.0 (BMA), the gradient boosting regression tree (GBRT) method and the random forest (RF) method, based on 195 global eddy covariance (EC) flux towers covering observations from 2000 through 2015. Validations indicated that the DNN had the highest accuracy among four merging methods across different land cover types, yielding the highest average determination coefficients (R 2 , 0.62), root-mean-squared-error (RMSE, 24.1 W/m 2 ) and Kling–Gupta efficiency (KGE, 0.77) with a of 99% confidence interval. Compared with GLASS ET algorithm v4.0, the DNN improved on the R 2 by approximately 7% (p < 0.01) and the KGE by 10%. Based on the DNN, we then generated 8-day GLASS ET product v5.0 globally with a 1 km spatial resolution from 2001 to 2015 driven by GLASS vegetation and surface net radiation (R n ) datasets and Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA2) datasets. Finally, this global terrestrial ET product provides a valuable dataset for monitoring regional and global water resources and environmental changes.

54 ENVIRONMENTAL SCIENCES↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗