Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boost machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Recent increases in annual, seasonal, and extreme methane fluxes driven by changes in climate and vegetation in boreal and temperate wetland ecosystems

Climate warming is expected to increase global methane (CH 4 ) emissions from wetland ecosystems. Although in situ eddy covariance (EC) measurements at ecosystem scales can potentially detect CH 4 flux changes, most EC systems have only a few years of data collected, so temporal trends in CH 4 remain uncertain. Here, we use established drivers to hindcast changes in CH 4 fluxes (FCH 4 ) since the early 1980s. We trained a machine learning (ML) model on CH 4 flux measurements from 22 [methane-producing sites] in wetland, upland, and lake sites of the FLUXNET-CH 4 database with at least two full years of measurements across temperate and boreal biomes. The gradient boosting decision tree ML model then hindcasted daily FCH 4 over 1981-2018 using meteorological reanalysis data. We found that, mainly driven by rising temperature, half of the sites (n = 11) showed significant increases in annual, seasonal, and extreme FCH 4 , with increases in FCH 4 of ca. 10% or higher found in the fall from 1981–1989 to 2010–2018. The annual trends were driven by increases during summer and fall, particularly at high-CH 4 -emitting fen sites dominated by aerenchymatous plants. We also found that the distribution of days of extremely high FCH 4 (defined according to the 95th percentile of the daily FCH 4 values over a reference period) have become more frequent during the last four decades and currently account for 10–40% of the total seasonal fluxes. The share of extreme FCH 4 days in the total seasonal fluxes was greatest in winter for boreal/taiga sites and in spring for temperate sites, which highlights the increasing importance of the non-growing seasons in annual budgets. Our results shed light on the effects of climate warming on wetlands, which appears to be extending the CH 4 emission seasons and boosting extreme emissions.

54 ENVIRONMENTAL SCIENCES↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Fast Evaluation of Aircraft Icing Severity Using Machine Learning Based on XGBoost

Aircraft icing represents a serious hazard in aviation which has caused a number of fatal accidents over the years. In addition, it can lead to substantial increase in drag and weight, thus reducing the aerodynamics performance of the airplane. The process of ice accretion on a solid surface is a complex interaction of aerodynamic and environmental variables. The complex relationship makes machine learning-based methods an attractive alternative to traditional numerical simulation-based approaches. In this study, we introduce a purely data-driven approach to find the complex pattern between different flight conditions and aircraft icing severity prediction. The supervised learning algorithm Extreme Gradient Boosting (XGBoost) is applied to establish the prediction framework which makes prediction based on any set of observations. The input flight conditions for the proposed prediction framework are liquid water content, droplet diameter and exposure time. The proposed approach is demonstrated in three cases: maximum ice thickness prediction, icing area prediction and icing severity level evaluation. Performance comparison studies and error analysis are also conducted to verify the effectiveness and performance of the proposed method. Results show that the proposed method has reasonable capability in evaluating aircraft icing severity.

Li, Sibo (ORCID:000000021705844X)↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

A Machine Learning-based Reliability Evaluation Model for Integrated Power-Gas Systems

This article proposes a hybrid machine learning method for the reliability evaluation of integrated power-gas systems (IPGS) under the uncertain component failure probability distributions. The Random Forest (RF) method is designed to select important features to solve the insufficient quantity of data and the curse of dimensionality problems. The Extreme Gradient Boosting (XGBoost) regression algorithm is developed to quantify the relationship between the uncertain parameters and reliability metrics. Moreover, a ten-fold cross-validation method is employed to further improve the accuracy of the regression model. Simulation results on three test systems show that the proposed method can achieve high accuracy for the reliability evaluation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

Characterize traction–separation relation and interfacial imperfections by data-driven machine learning models

Abstract Interfacial mechanical properties are important in composite materials and their applications, including vehicle structures, soft robotics, and aerospace. Determination of traction–separation (T–S) relations at interfaces in composites can lead to evaluations of structural reliability, mechanical robustness, and failures criteria. Accurate measurements on T–S relations remain challenging, since the interface interaction generally happens at microscale. With the emergence of machine learning (ML), data-driven model becomes an efficient method to predict the interfacial behaviors of composite materials and establish their mechanical models. Here, we combine ML, finite element analysis (FEA), and empirical experiments to develop data-driven models that characterize interfacial mechanical properties precisely. Specifically, eXtreme Gradient Boosting (XGBoost) multi-output regressions and classifier models are harnessed to investigate T–S relations and identify the imperfection locations at interface, respectively. The ML models are trained by macroscale force–displacement curves, which can be obtained from FEA and standard mechanical tests. The results show accurate predictions of T–S relations ( R 2 = 0.988) and identification of imperfection locations with 81% accuracy. Our models are experimentally validated by 3D printed double cantilever beam specimens from different materials. Furthermore, we provide a code package containing trained ML models, allowing other researchers to establish T–S relations for different material interfaces.

97 MATHEMATICS AND COMPUTING↗

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

Photon classification with Gradient Boosted Trees at CLAS12

Dihadron semi-inclusive deep inelastic scattering (SIDIS) of 10.6 GeV longitudinally polarized electrons off the proton has been measured using the CLAS12 detector at Jefferson Lab. Two separate channels, π + π 0 and π - π 0 , were analyzed, requiring the reconstruction of diphoton pairs. Here, in this analysis, we addressed the problem of false neutral particles being reconstructed by CLAS12's event builder, polluting the otherwise physical combinatorial background underneath the π 0 peak. A photon classifier using a Gradient Boosted Trees (GBTs) architecture was trained with Monte Carlo simulations to reduce the amount of background π 0 's. We show that the nearest-neighbor features learned by the model lead to a substantial increase in signal vs. background discrimination compared to previous CLAS12 π^0 analyses. The machine learning approach recovers several times more dihadron statistics for the dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

A Comparison of Infectious Disease Forecasting Methods across Locations, Diseases, and Time

Accurate infectious disease forecasting can inform efforts to prevent outbreaks and mitigate adverse impacts. This study compares the performance of statistical, machine learning (ML), and deep learning (DL) approaches in forecasting infectious disease incidences across different countries and time intervals. We forecasted three diverse diseases: campylobacteriosis, typhoid, and Q-fever, using a wide variety of features (n = 46) from public datasets, e.g., landscape, climate, and socioeconomic factors. We compared autoregressive statistical models to two tree-based ML models (extreme gradient boosted trees [XGB] and random forest [RF]) and two DL models (multi-layer perceptron and encoder–decoder model). The disease models were trained on data from seven different countries at the region-level between 2009–2017. Forecasting performance of all models was assessed using mean absolute error, root mean square error, and Poisson deviance across Australia, Israel, and the United States for the months of January through August of 2018. The overall model results were compared across diseases as well as various data splits, including country, regions with highest and lowest cases, and the forecasted months out (i.e., nowcasting, short-term, and long-term forecasting). Overall, the XGB models performed the best for all diseases and, in general, tree-based ML models performed the best when looking at data splits. There were a few instances where the statistical or DL models had minutely smaller error metrics for specific subsets of typhoid, which is a disease with very low case counts. Feature importance per disease was measured by using four tree-based ML models (i.e., XGB and RF with and without region name as a feature). The most important feature groups included previous case counts, region name, population counts and density, mortality causes of neonatal to under 5 years of age, sanitation factors, and elevation. This study demonstrates the power of ML approaches to incorporate a wide range of factors to forecast various diseases, regardless of location, more accurately than traditional statistical approaches.

59 BASIC BIOLOGICAL SCIENCES↗

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Hybrid data-driven cement-stabilized soil design: An integration of machine learning, multi-objective optimization, and life cycle assessment

Soil stabilization is crucial in geotechnical engineering, yet conventional methods are often time-consuming, resource-intensive, and environmentally unsustainable. Despite growing interest in Machine Learning (ML) and optimization tools for mix design, few studies integrate these methods with decision-making techniques and environmental assessment to support practical implementation. This study proposes a hybrid data-driven framework for predicting strength, optimizing mix compositions, and evaluating environmental impacts via life cycle assessment of cement-stabilized soft soils. Six ML models were evaluated, and the top-performing eXtreme Gradient Boosting (XGB) model was further improved using the Grey Wolf Optimizer (GWO). The optimized XGB-GWO model, integrated with a polynomial cost function, served as the objective function in a multi-objective optimization problem solved via the Non-Dominated Sorting Genetic Algorithm II (NSGA-II), with final mix selection guided by the entropy-weighted TOPSIS method. Validation through a case study produced mix designs offering superior strength-cost trade-offs, with the optimal mix achieving 2243.2 kPa unconfined compressive strength and a 16.07 % reduction in carbon emissions compared to the highest-cost design. In conclusion, this study offers a sustainable, scalable approach to soil stabilization and supports informed decision-making in construction.

Life cycle assessment↗

Predicting variable gene content in Escherichia coli using conserved genes

Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%–90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943–0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins” was accurately predicted (F1 = 0.902 [0.898–0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876–0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced descriptor identification and mechanism understanding for catalytic activity using a data-driven framework: revealing the importance of interactions between elementary steps

We report accurate identification of descriptors for catalytic activities has long been essential to the in-depth understanding of catalysis and recently to set the basis for catalyst screening. However, commonly used methods suffer from low accuracy in predictability. This study reports an enhanced approach to accurately identify the descriptors from a kinetic dataset using a machine learning (ML) surrogate model. CO hydrogenation to methanol over Cu-based catalysts was taken as a case study. Our model captures not only the contribution from individual elementary steps but also the interaction between relevant steps within a reaction network, which was found to be essential for high accuracy. As a result, six effective descriptors are identified, which are accurate enough to ensure the trained gradient boosted regression (GBR) model for good prediction of the methanol turnover frequency (TOF) over metal (M)-doped Cu(111) model surfaces (M = Au, Cu, Pd, Pt, Ni). More importantly, going beyond the purely mathematical ML model, the catalytic role of each identified descriptor can be revealed by using model-agnostic interpretation tools, which enhances the insight into the promoting effect of alloying. The trained GBR model outperforms the conventional derivative-based methods in terms of both the predictability and the mechanism understanding. It opens alternative possibilities toward accurate descriptor-based rational catalyst optimization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Plasma confinement state classification via FPP relevant microwave diagnostics

We present a parsimonious and robust machine learning approach for identifying plasma confinement states in fusion power plants (FPPs) where reliable identification of the low-confinement and high-confinement regimes is critical for safe and efficient operation. Unlike research-oriented devices, FPPs must operate with a severely constrained set of diagnostics. To address this challenge, we demonstrate that a minimalist model, using only electron cyclotron emission (ECE) signals, can achieve accurate and reliable state classification. ECE provides electron temperature profiles without the engineering or survivability issues of in-vessel probes, making it a primary candidate for FPP-relevant diagnostics. Our framework employs ECE as input, extracts features using radial basis functions, and applies a gradient boosting classifier, achieving a test accuracy of 96% (correct predictions). Robustness analysis and feature importance analyzes confirm the approach’s reliability. These results demonstrate that state-of-the-art performance is attainable from a restricted diagnostic set, paving the way for minimalist yet resilient plasma control architectures for FPPs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗