Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boost machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗

Mapping tree height in complex terrain of northern China using ultra-high-resolution images

Tree height is a key parameter for estimating forest biomass and carbon sequestration. In recent years, notable progress has been made in mapping tree height using satellite imagery. However, existing tree height products show low accuracy in mountainous and complex terrains, and few studies typically addressed tree height estimations in mountain areas. This study examines the Mentougou district of Beijing, China, characterized by complex terrain and mountainous landscapes. We analyzed two methods for estimating tree height: one using only spectral features and another combining spectral features with topographic factors (elevation, slope, aspect). We used 3-m resolution PlanetScope 8-band multispectral imagery, with 710 field-measured individual tree heights averaged to obtain 471 pixel-level tree height values as ground-truth, to develop tree height prediction models using eXtreme Gradient Boosting (XGBoost), Random Forest (RF), and Gradient Boosting Machine (GBM) models. The results show that the XGBoost model consistently presented the highest accuracy for both methods evaluated. Specifically, the XGBoost model that combined spectral data with elevation and slope variables with an R² of 0.75 and an RMSE of 2.69 m. Using the XGBoost model, we generated the tree height map for the Mentougou area at 3 m resolution, showing tree heights ranging from 0.5 to 30.4 m, and the model’s prediction error standard deviations ranged from 2.50 to 4.71 m, indicating reliable performance across varied terrain. Additionally, we compared and evaluated the global tree height products, identifying limitations in the accuracy within complex terrains. This study demonstrates the potential for accurately predicting tree heights by combining high-resolution multispectral satellites with a terrain factor modeling approach.

Complex terrain↗

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage↗

A Near-Real-Time Model for Predicting Electricity Disruptions in Texas During Winter Storms

There has been an increase in extreme weather events, posing a threat to power grid systems, potentially influenced by factors such as population growth, changes in ecosystems, land cover, and land use in the service area, as well as the growth of certain vegetation types. This research seeks to develop a predictive model to mitigate potential damages caused by future winter storms. This research utilizes the Light Gradient Boosting Machine (LightGBM), incorporating the number of power outages experienced at the county level, geographic details, weather information, and lagged outage and lagged weather data. The developed models were broadly divided into two groups, with six models in each group - one group without optimization and another with optimization, totaling 12 trained models. For model optimization, Bayesian optimization was employed using Root Mean Squared Error (RMSE) as the objective function. In results, when comparing Group 2 (the optimized group) with Group 1 (the non-optimized group), it was found that optimization did not always lead to a reduction in RMSE and Mean Absolute Error (MAE). However, in terms of Mean Directional Accuracy (MDA), while all results in Group 1 were below the baseline accuracy of 0.33, all results in Group 2 exceeded 0.33, with some cases showing an increase of more than three times the baseline. The results indicated that, in the optimized model group, Population and Pressure were the most influential factors when using current weather data and geographical information. When using lagged data, lagged recorded outages and lagged Pressure emerged as the most significant factors. Among the 12 developed models, the L-1-2-O model showed the lowest RMSE and MAE, as well as the highest accuracy, with values of 390.62 households and 168.13 households, respectively. To normalize the RMSE and MAE values, each metric was divided by the average number of households among the counties in Texas. For the L-1-2-O model, the scaled RMSE was 0.88% and the scaled MAE was 0.38%. In terms of MDA, which indicates the accuracy of the prediction direction, the L-1-O model achieved the highest score of 0.41. Although this study focused on Texas, which suffered the greatest impact from the winter storms in 2021, with additional validation, the methodology used in this research could be applied to other regions.

Lee, Jangjae [Texas A & M Univ., College Station, ↗

Estimation of the Surface Fluxes for Heat and Momentum in Unstable Conditions with Machine Learning and Similarity Approaches for the LAFE Data Set

Abstract Measurements of three flux towers operated during the land atmosphere feedback experiment (LAFE) are used to investigate relationships between surface fluxes and variables of the land–atmosphere system. We study these relations by means of two machine learning (ML) techniques: multilayer perceptrons (MLP) and extreme gradient boosting (XGB). We compare their flux derivation performance with Monin–Obukhov similarity theory (MOST) and a similarity relationship using the bulk Richardson number (BRN). The ML approaches outperform MOST and BRN. Best agreement with the observations is achieved for the friction velocity. For the sensible heat flux and even more so for the latent heat flux, MOST and BRN deviate from the observations while MLP and XGB yield more accurate predictions. Using MOST and BRN for latent heat flux, the root mean square errors (RMSE) are 107 Wm $$^{-2}$$ - 2 and 121 Wm $$^{-2}$$ - 2 , respectively, as well as the intercepts of the regression lines are $$\approx 110$$ ≈ 110 Wm $$^{-2}$$ - 2 . For the ML methods, the RMSEs reduce to 31 Wm $$^{-2}$$ - 2 for MLP and 33 Wm $$^{-2}$$ - 2 for XGB as well as the intercepts to just 4 Wm $$^{-2}$$ - 2 for MLP and $$-1$$ - 1 Wm $$^{-2}$$ - 2 for XGB with slopes of the regression lines close to 1, respectively. These results indicate significant deficiencies of MOST and BRN, particularly for the derivation of the latent heat flux. In fact, in contrast to the established theories, feature importance weighting demonstrates that the ML methods base their improved derivations on net radiation, the incoming and outgoing shortwave radiations, the air temperature gradient, and the available water contents, but not on the water vapor gradient. The results imply that further studies of surface fluxes and other turbulent variables with ML techniques provide great promise for deriving advanced flux parameterizations and their implementation in land–atmosphere system models.

54 ENVIRONMENTAL SCIENCES↗

Machine learning for fundamental spectroscopic and thermodynamic data of actinides and lanthanides

Accurately modeling optical spectra with absolute radiometric intensities is vital for nuclear forensics applications that depend on characterizing optical emissions from energetic nuclear phenomena. This requires precise knowledge of the individual atomic transition probabilities, known as Einstein A-coefficients, for each emission line. Obtaining these values theoretically or experimentally is often impractical due to the complex electronic structures and the number of transitions involved in atoms relevant to nuclear applications. In this study, we explore the use of machine learning to predict the Einstein A coefficients for atomic transitions. Seven models were evaluated that ranged from deep learning to decision tree algorithms, and found that gradient boosting performed best, specifically the Extreme Gradient Boosting (XGB) architecture, achieving a precision of 86% across transitions of 36 elements. Furthermore, the model was cross-validated using published transition probabilities reported in the literature and applied to estimate Pu plasma temperatures from a previous experiment conducted at Savannah River National Laboratory.

Atomic spectroscopy↗

Medium Energy Electron Flux in Earth's Outer Radiation Belt (MERLIN): A Machine Learning Model

The radiation belts of the Earth, filled with energetic electrons, comprise complex and dynamic systems that pose a significant threat to satellite operation. While various models of electron flux both for low and relativistic energies have been developed, the behavior of medium energy (120–600 keV) electrons, especially in the MEO region, remains poorly quantified. At these energies, electrons are driven by both convective and diffusive transport, and their prediction usually requires sophisticated 4D modeling codes. In this paper, we present an alternative approach using the Light Gradient Boosting (LightGBM) machine learning algorithm. The Medium Energy electRon fLux In Earth's outer radiatioN belt (MERLIN) model takes as input the satellite position, a combination of geomagnetic indices and solar wind parameters including the time history of velocity, and does not use persistence. MERLIN is trained on >15 years of the GPS electron flux data and tested on more than 1.5 years of measurements. Tenfold cross validation yields that the model predicts the MEO radiation environment well, both in terms of dynamics and amplitudes o f flux. Evaluation on the test set shows high correlation between the predicted and observed electron flux (0.8) and low values of absolute error. The MERLIN model can have wide space weather applications, providing information for the scientific community in the form of radiation belts reconstructions, as well as industry for satellite mission design, nowcast of the MEO environment, and surface charging analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimized Machine Learning Model for Predicting Groundwater Contamination

The use of physical models to predict groundwater contaminant movement remains technically challenging due to the complexity of the phenomena, the heterogeneity of key parameters in nature, and the presence of poorly defined interactive and feedback processes. New approaches to address these challenges are needed. In this study, we evaluate various Artificial Intelligence (AI)-based approaches to understand a hexavalent chromium (Cr(VI)) plumes located on the U.S. Department of Energy’s (DOE) Hanford Site in Richland, WA. The groundwater monitoring dataset used in this study included data from the 100 Area along the Columbia River and included data collected between 2010 to 2019. This study investigates the most prominent contaminant, Cr(VI), with the Extreme Gradient Boosting (XGBoost) machine learning model. The XGBoost models were compared with optimized versions using an Empirical Bayes Search Cross-Validation technique for better prediction. The optimized XGBoost model yielded an R^2 value of 0.99 on the training set and 0.85 on the testing set, whereas XGBoost without optimization yielded a value of 0.83 on the training set and 0.85 on the testing set. This paper provides an overview of a computational method for groundwater contamination modeling that shows promise for improving current remediation efforts.

Mazumdar, Hirak↗

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation↗

Building a landslide hazard indicator with machine learning and land surface models

The U.S. Pacific Northwest has a history of frequent and occasionally deadly landslides caused by various factors. Using a multivariate, machine-learning approach, we combined a Pacific Northwest Landslide Inventory with a 36-year gridded hydrologic dataset from the National Climate Assessment – Land Data Assimilation System to produce a landslide hazard indicator (LHI) on a daily 0.125-degree grid. The LHI identified where and when landslides were most probable over the years 1979–2016, addressing issues of bias and completeness that muddy the analysis of multi-decadal landslide inventories. The seasonal cycle was strong along the west coast, with a peak in the winter, but weaker east of the Cascade Range. This lagging indicator can fill gaps in the observational record to identify the seasonality of landslides over a large spatiotemporal domain and show how landslide hazard has responded to a changing climate.

XGBoost↗

fSCAml

A repository for predicting fractional snow covered area using a gradient boosted trees machine learning.

Crumley, Ryan L [Self-Employed]↗

Rapid Spaceborne Mapping of Wildfire Retardant Drops for Active Wildfire Management

Aerial application of fire retardant is a critical tool for managing wildland fire spread. Retardant applications are carefully planned to maximize fire line effectiveness, improve firefighter safety, protect high-value resources and assets, and limit environmental impact. However, topography, wind, visibility, and aircraft orientation can lead to differences between planned drop locations and the actual placement of the retardant. Information on the precise placement and areal extent of the dropped retardant can provide wildland fire managers with key information to (1) adaptively manage event resources, (2) assess the effectiveness of retardant slowing or stopping fire spread, (3) document location in relation to ecologically sensitive areas; and perform or validate cost-accounting for drop services. This study uses Sentinel-2 satellite data and commonly used machine learning classifiers to test an automated approach for detecting and mapping retardant application. We show that a multiclass model (retardant, burned, unburned, and cloud artifact classes) outperforms a single-class retardant model and that image differencing (post-application minus pre-application) outperforms single-image models. Compared to the random forest and support vector machine, the gradient boosting model performed the best with an overall accuracy of 0.88 and an F1 Score of 0.76 for fire retardant, though results were comparable for all three models. Our approach maps the full areal extent of the dropped retardant within minutes of image availability, rather than linear representations currently mapped by aerial GPS surveys. The development of this capability allows for the rapid assessment of retardant effectiveness and documentation of placement in relation to sensitive environments.

54 ENVIRONMENTAL SCIENCES↗

Median bed-material sediment particle size across rivers in the contiguous US

Abstract. Bed-material sediment particle size data, particularly the median sediment particle size (D50), are critical for understanding and modeling riverine sediment transport. However, sediment particle size observations are primarily available at individual sites. Large-scale modeling and assessment of riverine sediment transport are limited by the lack of continuous regional maps of bed-material sediment particle size. We hence present a map of D50 over the contiguous US in a vector format that corresponds to approximately 2.7 million river segments (i.e., flowlines) in the National Hydrography Dataset Plus (NHDPlus) dataset. We develop the map in four steps: (1) collect and process the observed D50 data from 2577 U.S. Geological Survey stations or U.S. Army Corps of Engineers sampling locations; (2) collocate these data with the NHDPlus flowlines based on their geographic locations, resulting in 1691 flowlines with collocated D50 values; (3) develop a predictive model using the eXtreme Gradient Boosting (XGBoost) machine learning method based on the observed D50 data and the corresponding climate, hydrology, geology, and other attributes retrieved from the NHDPlus dataset; and (4) estimate the D50 values for flowlines without observations using the XGBoost predictive model. We expect this map to be useful for various purposes, such as research in large-scale river sediment transport using model- and data-driven approaches, teaching environmental and earth system sciences, planning and managing floodplain zones, etc. The map is available at https://doi.org/10.5281/zenodo.4921987 (Li et al., 2021a).

54 ENVIRONMENTAL SCIENCES↗

Micro-structural features and material properties impact on adhesive metal joints via computational modeling and machine learning

The quality of structural bonding in practical applications depends on various factors arising from materials, pre-processing conditions, and manufacturing. Understanding how these factors influence bonding performance and determining their relative importance are of significant interest. Thus, this study evaluates the effects of microstructural features and material properties on the structural strength of adhesively-bonded metal joints at the submillimeter scale, utilizing a combination of Finite Element Modeling (FEM) and Machine Learning (ML) with Gradient Boosting Regression (GBR). The microstructural features include adhesive thickness, internal voids within the adhesive, adherend-adhesive interfacial voids, void size and volume fraction, and surface roughness. The material properties include the constitutive behavior of the adhesive, as well as the adherend-adhesive interfacial strength and fracture energy. The changes in structural strength and morphologies of the bonded metal structures with respect to different microstructural features and material properties were clarified by FEM. By further leveraging ML-GBR, the sequence of importance of these factors affecting bonding performance across various scenarios was summarized. This work provides valuable insights into the development of improved structural bonding for adhesive joints in industries such as automotive , aerospace, and beyond.

36 MATERIALS SCIENCE↗

Fair Bagging Boosting Models [SWR-24-38]

Fair Bagging Boosting Models is a software implementation of a framework for building, measuring bias and correcting bias in 3 popular forest machine learning models: gradient boosted trees (GBT), random forest (RF), and XGBoost models, using the XGBoost library. The framework takes advantage of the flexibility in XGBoost library to represent gradient boosted tree and random forest models, as well as the ability to use custom loss function.

Ugirumurera, Juliette↗

Harnessing Machine Learning to Predict MoS 2 Solid Lubricant Performance

Physical vapor deposited (PVD) molybdenum disulfide (MoS 2 ) solid lubricant coatings are an exemplar material system for machine learning methods due to small changes in process variables often causing large variations in microstructure and mechanical/tribological properties. Here, in this work, a gradient boosted regression tree machine learning method is applied to an existing experimental data set containing process, microstructure, and property information to create deeper insights into the process-structure–property relationships for molybdenum disulfide (MoS 2 ) solid lubricant coatings. The optimized and cross-validated models show good predictive capabilities for density, reduced modulus, hardness, wear rate, and initial coefficients of friction. The contribution of individual deposition variables (i.e., argon pressure, deposition power, target conditioning) on coating properties is highlighted through feature importance. The process-property relationships established herein show linear and non-linear relationships and highlight the influence of uncontrolled deposition variables (i.e., target conditioning) on the tribological performance.

MoS2↗

A bacterial sensor taxonomy across earth ecosystems for machine learning applications

Microbial communities have evolved to colonize all ecosystems of the planet, from the deep sea to the human gut. Microbes survive by sensing, responding, and adapting to immediate environmental cues. This process is driven by signal transduction proteins such as histidine kinases, which use their sensing domains to bind or otherwise detect environmental cues and “transduce” signals to adjust internal processes. We hypothesized that an ecosystem’s unique stimuli leave a sensor “fingerprint,” able to identify and shed insight on ecosystem conditions. To test this, we collected 20,712 publicly available metagenomes from Host-associated, Environmental, and Engineered ecosystems across the globe. We extracted and clustered the collection’s nearly 18M unique sensory domains into 113,712 similar groupings with MMseqs2. We built gradient-boosted decision tree machine learning models and found we could classify the ecosystem type (accuracy: 87%) and predict the levels of different physical parameters (R2 score: 83%) using the sensor cluster abundance as features. Feature importance enables identification of the most predictive sensors to differentiate between ecosystems which can lead to mechanistic interpretations if the sensor domains are well annotated. To demonstrate this, a machine learning model was trained to predict patient’s disease state and used to identify domains related to oxygen sensing present in a healthy gut but missing in patients with abnormal conditions. Moreover, since 98.7% of identified sensor domains are uncharacterized, importance ranking can be used to prioritize sensors to determine what ecosystem function they may be sensing. Furthermore, these new predictive sensors can function as targets for novel sensor engineering with applications in biotechnology, ecosystem maintenance, and medicine.

97 MATHEMATICS AND COMPUTING↗