Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gradient boosting trees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

The Circular Velocity Curve of the Milky Way from 5–25 kpc Using Luminous Red Giant Branch Stars

We present a sample of 254,882 luminous red giant branch (LRGB) stars selected from the APOGEE and LAMOST surveys. By combining photometric and astrometric information from the Two Micron All Sky Survey and Gaia survey, the precise distances of the sample stars are determined by a supervised machine-learning algorithm: the gradient-boosted decision trees. To test the accuracy of the derived distances, member stars of globular clusters (GCs) and open clusters are used. The tests by cluster member stars show a precision of about 10% with negligible zero-point offsets, for the derived distances of our sample stars. The final sample covers a large volume of the Galactic disk(s) and halo of 0 < R < 30 kpc and |Z| ≤ 15 kpc. The rotation curve (RC) of the Milky Way across the radius of 5 ≲ R ≲ 25 kpc has been accurately measured with ~54,000 stars of the thin disk population selected from the LRGB sample. The derived RC shows a weak decline along R with a gradient of -1.83 ± 0.02 (stat.) ± 0.07 (sys.) km s -1 kpc -1 , in excellent agreement with the results measured by previous studies. The circular velocity at the solar position, yielded by our RC is 234.04 ± 0.08 (stat.) ± 1.36 (sys.) km s -1 , again in great consistency with other independent determinations. From the newly constructed RC, as well as constraints from other data, we have constructed a mass model for our Galaxy, yielding a mass of the dark matter halo of M 200 = (8.05 ± 1.15) × 10 11 M ⊙ with a corresponding radius of R 200 = 192.37 ± 9.24 kpc and a local dark matter density of 0.39 ± 0.03 GeV cm -3 .

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Machine Learning to Predict MoS 2 Solid Lubricant Performance

Physical vapor deposited (PVD) molybdenum disulfide (MoS 2 ) solid lubricant coatings are an exemplar material system for machine learning methods due to small changes in process variables often causing large variations in microstructure and mechanical/tribological properties. Here, in this work, a gradient boosted regression tree machine learning method is applied to an existing experimental data set containing process, microstructure, and property information to create deeper insights into the process-structure–property relationships for molybdenum disulfide (MoS 2 ) solid lubricant coatings. The optimized and cross-validated models show good predictive capabilities for density, reduced modulus, hardness, wear rate, and initial coefficients of friction. The contribution of individual deposition variables (i.e., argon pressure, deposition power, target conditioning) on coating properties is highlighted through feature importance. The process-property relationships established herein show linear and non-linear relationships and highlight the influence of uncontrolled deposition variables (i.e., target conditioning) on the tribological performance.

MoS2↗

The Global LAnd Surface Satellite (GLASS) evapotranspiration product Version 5.0: Algorithm development and preliminary validation

An accurate estimation of spatially and temporally continuous global terrestrial evapotranspiration (ET) is essential in the assessment of surface energy, water and carbon cycles. The Global LAnd Surface Satellite (GLASS) ET product Version 4.0 (v4.0) based on the Bayesian model averaging (BMA) method was generated to estimate global terrestrial ET. However, certain uncertainty for the GLASS ET product v4.0 limits its application. In this study, we introduced the deep neural networks (DNN) merging framework to improve terrestrial ET estimation for GLASS ET product Version 5.0 (v5.0) generation by integrating five satellite-derived ET products [Moderate Resolution Imaging Spectroradiometer (MODIS) ET product (MOD16), Shuttleworth–Wallace dual-source ET product (SW), Priestley–Taylor-based ET product (PT-JPL), modified satellite-based Priestley–Taylor ET product (MS-PT) and simple hybrid ET product (SIM)]. We compared the performance of DNN method against other merging methods, including GLASS ET algorithm v4.0 (BMA), the gradient boosting regression tree (GBRT) method and the random forest (RF) method, based on 195 global eddy covariance (EC) flux towers covering observations from 2000 through 2015. Validations indicated that the DNN had the highest accuracy among four merging methods across different land cover types, yielding the highest average determination coefficients (R 2 , 0.62), root-mean-squared-error (RMSE, 24.1 W/m 2 ) and Kling–Gupta efficiency (KGE, 0.77) with a of 99% confidence interval. Compared with GLASS ET algorithm v4.0, the DNN improved on the R 2 by approximately 7% (p < 0.01) and the KGE by 10%. Based on the DNN, we then generated 8-day GLASS ET product v5.0 globally with a 1 km spatial resolution from 2001 to 2015 driven by GLASS vegetation and surface net radiation (R n ) datasets and Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA2) datasets. Finally, this global terrestrial ET product provides a valuable dataset for monitoring regional and global water resources and environmental changes.

54 ENVIRONMENTAL SCIENCES↗

Quantum Chemistry-Driven Machine Learning Approach for the Prediction of the Surface Tension and Speed of Sound in Ionic Liquids

Ionic liquids (ILs) have unique solvent properties and have thus garnered significant interest. However, exhaustive experimental determination of the physicochemical properties of ILs is unrealistic due to the large structural diversity of anions and cations, their high cost, the requirements of elevated temperature and pressure, and the time required. To circumvent these experimental costs, computational approaches to accurately calculate these properties have emerged. Here in the present study, we present a demonstration of two machine learning (ML) models for the prediction of two critical IL physical properties, the surface tension and the speed of sound, across a wide range of temperatures and pressures. The models make use of molecular descriptors derived from the COSMO-RS, a quantum chemical-based model. The ML models show excellent agreement with experimental observations, with an R2 value of 0.96–0.99 and RMSE of 1.71 mN/m and 16.12 m/s for the surface tension and speed of sound, respectively. This work paves the way for the development of COSMO-RS-informed ML models for the prediction of IL properties which can help to further optimize and accelerate technology development for ILs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mapping Glacier Basal Sliding Applying Machine Learning

During the RESOLVE project (“High-resolution imaging in subsurface geophysics: development of a multi-instrument platform for interdisciplinary research”), continuous surface displacement and seismic array observations were obtained on Glacier d’Argentière in the French Alps for 35 days in May 2018. The data set is used to perform a detailed study of targeted processes within the highly dynamic cryospheric environment. In particular, the physical processes controlling glacial basal motion are poorly understood and remain challenging to observe directly. Especially in the Alpine region for temperate based glaciers where the ice rapidly responds to changing climatic conditions and thus, processes are strongly intermittent in time and heterogeneous in space. Spatially dense seismic and Global Positioning System (GPS) measurements are analyzed applying machine learning to gain insight into the processes controlling glacial motions of Glacier d’Argentière. Using multiple bandpass-filtered copies of the continuous seismic waveforms, we compute energy-based features, develop a matched field beamforming catalog and include meteorological observations. Features describing the data are analyzed with a gradient boosting decision tree model to directly estimate the GPS displacements from the seismic noise. We posit that features of the seismic noise provide direct access to the dominant parameters that drive displacement on the highly variable and unsteady surface of the glacier. The machine learning model infers daily fluctuations and longer term trends. The results show on-ice displacement rates are strongly modulated by activity at the base of the glacier. The techniques presented provide a new approach to study glacial basal sliding and discover its full complexity.

58 GEOSCIENCES↗

A machine learning approach for efficient multi-dimensional integration

Many physics problems involve integration in multi-dimensional space whose analytic solution is not available. The integrals can be evaluated using numerical integration methods, but it requires a large computational cost in some cases, so an efficient algorithm plays an important role in solving the physics problems. We propose a novel numerical multi-dimensional integration algorithm using machine learning (ML). After training a ML regression model to mimic a target integrand, the regression model is used to evaluate an approximation of the integral. Then, the difference between the approximation and the true answer is calculated to correct the bias in the approximation of the integral induced by ML prediction errors. Because of the bias correction, the final estimate of the integral is unbiased and has a statistically correct error estimation. Three ML models of multi-layer perceptron, gradient boosting decision tree, and Gaussian process regression algorithms are investigated. The performance of the proposed algorithm is demonstrated on six different families of integrands that typically appear in physics problems at various dimensions and integrand difficulties. The results show that, for the same total number of integrand evaluations, the new algorithm provides integral estimates with more than an order of magnitude smaller uncertainties than those of the VEGAS algorithm in most of the test cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Selecting durable building envelope systems with machine learning assisted hygrothermal simulations database

Hygrothermal simulations provide insight into the energy performance and moisture durability of building envelope components under dynamic conditions. The inputs required for hygrothermal simulations are extensive, and carrying out simulations and analyses requires expert knowledge. An expert system, the Building Science Advisor (BSA), has been developed to predict the performance and select the energy-efficient and durable building envelope systems for different climates. The BSA consists of decision rules based on expert opinions and thousands of parametric simulation results for selected wall systems. The number of potential wall systems results in millions, too many to simulate all of them. We present how machine learning can help predict durability data, such as mold growth, while minimizing the number of simulations needed to run. The simulation results are used for training and validation of machine learning tools for predicting wall durability. We tested Artificial Neural Network (ANN) and Gradient Boosted Decision Trees (GBDT) for their applicability and model accuracy. Models developed with both methods showed adequate prediction performance (root mean square error of 0.195 and 0.209, respectively). Finally, we introduce how the information supports guidance for envelope design via an easy-to-use web-based tool that does not require the end-user to run hygrothermal simulations.

Salonvaara, Mikael↗

Identification of low-momentum muons in the CMS detector using multivariate techniques in proton-proton collisions at $\sqrt{s}$ = 13.6 TeV

“Soft” muons with a transverse momentum below 10 GeV are featured in many processes studied by the CMS experiment, such as decays of heavy-flavor hadrons or rare tau lepton decays. Maximizing the selection efficiency for these muons, while simultaneously suppressing backgrounds from long-lived light-flavor hadron decays, is therefore important for the success of the CMS physics program. Multivariate techniques have been shown to deliver better muon identification performance than traditional selection techniques. To take full advantage of the large data set currently being collected during Run 3 of the CERN LHC, a new multivariate classifier based on a gradient-boosted decision tree has been developed. It offers a significantly improved separation of signal and background muons compared to a similar classifier used for the analysis of the Run 2 data. The performance of the new classifier is evaluated on a data set collected with the CMS detector in 2022 and 2023, corresponding to an integrated luminosity of 62 fb -1 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evaluating Recursive Blind Forecast Against API and Baseline: A Puerto Rican Case Study on Solar Irradiance for Normal and Extreme Weather

This paper leverages ongoing work in a community microgrid in Adjuntas, Puerto Rico to forecast global horizontal irradiance (GHI) and compare performance in normal and extreme weather. Given a positive correlation of 0.98 between GHI and PV power, forecasting GHI can be an effective, indirect forecast of photovoltaic (PV) power, especially in microgrids where the end-users, owners, operators, or other stakeholders are reluctant to share data for training or validation due to privacy and security concerns. A recursive one-shot (termed as "blind") forecast is, hence, formulated, wherein a gradient-boosted regression tree (GBR) is built to forecast GHI for a 7-day horizon in normal weather, and a 2-day horizon in extreme weather. To demonstrate its resilience, the architecture is trained on normal and hurricane weather GHI from 2002-2022. It is generalized on February 9-16, 2023, and on the landfall of Hurricane Nicole (Nov 4-5, 2022), respectively. Forecasts from GBR are compared against that from a satellite-based API resource and three baselines: persistence, averaging, and exponential smoothing. Results show GBR and persistence outperform sophisticated API in both types of weather for this case study.

Sundararajan, Aditya↗

Recent increases in annual, seasonal, and extreme methane fluxes driven by changes in climate and vegetation in boreal and temperate wetland ecosystems

Climate warming is expected to increase global methane (CH 4 ) emissions from wetland ecosystems. Although in situ eddy covariance (EC) measurements at ecosystem scales can potentially detect CH 4 flux changes, most EC systems have only a few years of data collected, so temporal trends in CH 4 remain uncertain. Here, we use established drivers to hindcast changes in CH 4 fluxes (FCH 4 ) since the early 1980s. We trained a machine learning (ML) model on CH 4 flux measurements from 22 [methane-producing sites] in wetland, upland, and lake sites of the FLUXNET-CH 4 database with at least two full years of measurements across temperate and boreal biomes. The gradient boosting decision tree ML model then hindcasted daily FCH 4 over 1981-2018 using meteorological reanalysis data. We found that, mainly driven by rising temperature, half of the sites (n = 11) showed significant increases in annual, seasonal, and extreme FCH 4 , with increases in FCH 4 of ca. 10% or higher found in the fall from 1981–1989 to 2010–2018. The annual trends were driven by increases during summer and fall, particularly at high-CH 4 -emitting fen sites dominated by aerenchymatous plants. We also found that the distribution of days of extremely high FCH 4 (defined according to the 95th percentile of the daily FCH 4 values over a reference period) have become more frequent during the last four decades and currently account for 10–40% of the total seasonal fluxes. The share of extreme FCH 4 days in the total seasonal fluxes was greatest in winter for boreal/taiga sites and in spring for temperate sites, which highlights the increasing importance of the non-growing seasons in annual budgets. Our results shed light on the effects of climate warming on wetlands, which appears to be extending the CH 4 emission seasons and boosting extreme emissions.

54 ENVIRONMENTAL SCIENCES↗

A bacterial sensor taxonomy across earth ecosystems for machine learning applications

Microbial communities have evolved to colonize all ecosystems of the planet, from the deep sea to the human gut. Microbes survive by sensing, responding, and adapting to immediate environmental cues. This process is driven by signal transduction proteins such as histidine kinases, which use their sensing domains to bind or otherwise detect environmental cues and “transduce” signals to adjust internal processes. We hypothesized that an ecosystem’s unique stimuli leave a sensor “fingerprint,” able to identify and shed insight on ecosystem conditions. To test this, we collected 20,712 publicly available metagenomes from Host-associated, Environmental, and Engineered ecosystems across the globe. We extracted and clustered the collection’s nearly 18M unique sensory domains into 113,712 similar groupings with MMseqs2. We built gradient-boosted decision tree machine learning models and found we could classify the ecosystem type (accuracy: 87%) and predict the levels of different physical parameters (R2 score: 83%) using the sensor cluster abundance as features. Feature importance enables identification of the most predictive sensors to differentiate between ecosystems which can lead to mechanistic interpretations if the sensor domains are well annotated. To demonstrate this, a machine learning model was trained to predict patient’s disease state and used to identify domains related to oxygen sensing present in a healthy gut but missing in patients with abnormal conditions. Moreover, since 98.7% of identified sensor domains are uncharacterized, importance ranking can be used to prioritize sensors to determine what ecosystem function they may be sensing. Furthermore, these new predictive sensors can function as targets for novel sensor engineering with applications in biotechnology, ecosystem maintenance, and medicine.

97 MATHEMATICS AND COMPUTING↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗