Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gradient boosting trees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

Insights into the origin of halo mass profiles from machine learning

ABSTRACT The mass distribution of dark matter haloes is the result of the hierarchical growth of initial density perturbations through mass accretion and mergers. We use an interpretable machine-learning framework to provide physical insights into the origin of the spherically-averaged mass profile of dark matter haloes. We train a gradient-boosted-trees algorithm to predict the final mass profiles of cluster-sized haloes, and measure the importance of the different inputs provided to the algorithm. We find two primary scales in the initial conditions (ICs) that impact the final mass profile: the density at approximately the scale of the haloes’ Lagrangian patch RL ($R\sim 0.7\, R_L$) and that in the large-scale environment (R ∼ 1.7 RL). The model also identifies three primary time-scales in the halo assembly history that affect the final profile: (i) the formation time of the virialized, collapsed material inside the halo, (ii) the dynamical time, which captures the dynamically unrelaxed, infalling component of the halo over its first orbit, (iii) a third, most recent time-scale, which captures the impact on the outer profile of recent massive merger events. While the inner profile retains memory of the ICs, this information alone is insufficient to yield accurate predictions for the outer profile. As we add information about the haloes’ mass accretion history, we find a significant improvement in the predicted profiles at all radii. Our machine-learning framework provides novel insights into the role of the ICs and the mass assembly history in determining the final mass profile of cluster-sized haloes.

79 ASTRONOMY AND ASTROPHYSICS↗

Composite Qdrift-product formulas for quantum and classical simulations in real and imaginary time

Recent study has shown that it can be advantageous to implement a composite channel that partitions the Hamiltonian H for a given simulation problem into subsets A and B such that H = A + B , where the terms in A are simulated with a Trotter-Suzuki channel and the B terms are randomly sampled via the Qdrift algorithm. Here we extend Qdrift and composite product formulas to imaginary time, formulating candidate classical algorithms for quantum Monte Carlo calculations. We upper bound the induced Schatten- 1 → 1 norm on both imaginary-time Qdrift and composite channels. Another recent result demonstrated that simulations of lattice Hamiltonians containing geometrically local interactions can be improved using a Lieb-Robinson argument to decompose H into subsets that contain only terms supported on that subset of the lattice. Here, we provide a quantum algorithm by unifying this result with the composite approach into “local composite channels” and we upper bound the diamond distance. We provide exact numerical simulations of algorithmic cost by counting the number of gates of the form e − i H j t and e − H j β to meet a certain error tolerance ε . In doing so, we optimize the partitioning into sets A and B using gradient boosted tree models from machine learning. These numerical studies are important given that product formulas have been historically known to outperform analytic upper bounds. We show constant factor advantages for a variety of interesting Hamiltonians, the maximum of which is a ≈ 20 -fold speedup that occurs in the simulation of Jellium. Published by the American Physical Society 2024

Pocrnic, Matthew (ORCID:0000000203089376)↗

Machine Learning Benchmarks for the Classification of Equivalent Circuit Models from Electrochemical Impedance Spectra

Analysis of Electrochemical Impedance Spectroscopy (EIS) data for electrochemical systems often consists of defining an Equivalent Circuit Model (ECM) using expert knowledge and then optimizing the model parameters to deconvolute various resistance, capacitive, inductive, or diffusion responses. For small data sets, this procedure can be conducted manually; however, it is not feasible to manually define a proper ECM for extensive data sets with a wide range of EIS responses. Automatic identification of an ECM would substantially accelerate the analysis of large sets of EIS data. We showcase machine learning methods to classify the ECMs of 9,300 impedance spectra provided by QuantumScape for the BatteryDEV hackathon. The best-performing approach is a gradient-boosted tree model utilizing a library to automatically generate features, followed by a random forest model using the raw spectral data. A convolutional neural network using boolean images of Nyquist representations is presented as an alternative, although it achieves a lower accuracy. We publish the data and open source the associated code. The approaches described in this article can serve as benchmarks for further studies. A key remaining challenge is the identifiability of the labels, underlined by the model performances and the comparison of misclassified spectra.

25 ENERGY STORAGE↗

fSCAml

A repository for predicting fractional snow covered area using a gradient boosted trees machine learning.

Crumley, Ryan L [Self-Employed]↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

CoSHA: Code for Stellar Properties Heuristic Assignment—for the MaStar Stellar Library

We introduce CoSHA: a Code for Stellar properties Heuristic Assignment. In order to estimate the stellar properties, CoSHA implements a Gradient Tree Boosting algorithm to label each star across the parameter space (T eff , $\mathrm{log}g$, [Fe/H], and [α/Fe]). We use CoSHA to estimate the stellar atmospheric parameters of 22,000 unique stars in the MaNGA Stellar Library (MaStar). To quantify the reliability of our approach, we run internal tests, using both the Göttingen Stellar Library (a theoretical library) and the first data release of MaStar, and external tests, by comparing the resulting distributions in the parameter space with the APOGEE estimates of the same properties. In summary, our parameter estimates span the ranges T eff = [2900, 12,000] K, $\mathrm{log}g=[-0.5,5.6]$, [Fe/H] = [-3.74, 0.81], and αM = [-0.22, 1.17]. We report internal (external) uncertainties of the properties of ${\sigma }_{{T}_{\mathrm{eff}}}\sim 43(240)$ K, ${\sigma }_{\mathrm{log}g}\sim 0.2(0.4)$, σ [Fe/H] ~ 0.16(0.24), and σ [α/Fe] ~ 0.09(0.08). These uncertainties are comparable to those of other methods with similar objectives. Despite the fact that CoSHA is not aware of the spatial distributions of these physical properties in the Milky Way, we are able to recover the main trends known in the literature. The catalog of physical properties for MaStar can be accessed online.

79 ASTRONOMY AND ASTROPHYSICS↗

Failure prediction and estimation of failure parameters

Machine-learning methods and apparatus are disclosed to determine frictional state or other parameters in an earthquake zone or other failing medium, using acoustic emission, seismic waves, or other detectable indicators of microscopic processes. Predictions of future failures are demonstrated in different regimes. A classifier is trained using time series of acoustic emission data along with historic data of frictional state or failure events. In disclosed examples, random forests and gradient boost trees are used, and grid-search or EGO procedures are used for hyperparameter tuning. Once trained, the classifier can be applied to testing or live data in order to assess a frictional state, assess seismic hazard, or make predictions regarding a future failure event. The technology has been developed in a double direct shear apparatus, but can be widely applied to seismic faults, other terrestrial failures, or failures in man-made structures. Variations are disclosed.

Johnson, Paul Allan↗

Measurements of Beam Spin Asymmetries in p+p0 and p´p0 Dihadron Production at CLAS12

Semi-Inclusive Deep Inelastic Scattering (SIDIS) is a powerful experimental tool for studying the internal structure and dynamics of the proton, revealing how quarks and gluons are distributed and interact within it. SIDIS describes a process where an elec tron scatters off one of the constituent quarks within the proton, causing it to undergo hadronization, creating multiple hadrons in the final state. Through factorization, the full process can be split into probabilistic components: one which describes the internal structure of the proton using Parton Distribution Functions (PDFs), and another which describes the hadronization process using Fragmentation Functions (FFs). These functions are non-perturbative quantities of Quantum Chromodynamics (QCD), meaning they cannot be calculated directly from first principles and must instead be extracted from experimental measurements. Acommon approach for accessing PDFs and FFs using SIDIS is to measure asymmetries. In this context, asymmetries correspond to subtle differences in the angular distribution of outgoing particles that arise when the spin orientation of the incoming beam or target is reversed. Because many of these effects only appear when spin is involved, they isolate specific, nuanced properties of the proton’s spin-structure that are otherwise hidden in spin averaged measurements. In practice, they show up as specific azimuthal modulations (e.g., sin ¿R, sin(¿h ´ ¿R)), whose amplitudes isolate convolutions of PDFs and FFs at leading and subleading twist. Non-zero asymmetries of these angular distributions can be traced back to unique combinations of PDFs and FFs, offering a way to probe them directly. In this work, we measure SIDIS by analyzing high energy electron-proton scattering events using the CLAS12 detector at Jefferson Lab. This study focuses on subset of SIDIS referred to as dihadron SIDIS, where pairs of hadrons — here p+p0 and p´p0 — are observed. We analyzed these dihadrons using detector data collected during Fall 2018 and Spring 2019, where longitudinally polarized electrons from the CEBAF accelerator were incident on a liquid hydrogen target. A photon classifier using a Gradient Boosted Trees (GBTs) architecture was trained using Monte Carlo simulations to reduce the amount of iv false combinatorial background p0’s. When deployed on experimental data, the model in creases our dihadron statistics by up to five-fold compared to previous CLAS12 p0 analyses. This work reports the first measurements of beam spin asymmetries for p+p0 and p´p0 dihadron production in SIDIS. The measured asymmetries offer new insights to the spin-dependent structure and dynamics within the proton, as well as the spin-dependent properties of quark fragmentation. Non-zero twist-3 sin¿R amplitudes are observed, pro viding sensitivity to the subleading twist PDF e(x). The PDF e(x) encodes quark-gluon correlations within the proton — a property that is otherwise inaccessible at leading twist. Additionally, this work measured significant twist-2 modulations carried by sin(¿h ´ ¿R) and sin(2¿h ´2¿R), providing experimental access to the helicity dihadron fragmentation function (DiFF) GK 1 . Because there is no equivalent quark helicity-dependent FF in single pion SIDIS, the DiFF GK 1 offers a unique lens into novel spin-dependent fragmentation. For instance, the twist-2 modulations observed in this study are enhanced by vector mesons created during fragmentation — a behavior predicted by phenomenological models. This study broadens our understanding of dihadron fragmentation, revealing new details about the flavor and charge dependence of hadronization.

Matousek, Gregory [Duke Univ., Durham, NC (United ↗

A novel machine learning based identification of potential adopter of rooftop solar photovoltaics

With the proliferation of rooftop solar photovoltaic installations, there is a need to proactively predict consumer potential for solar photovoltaic adoption, for improved electric utility planning and operation. Traditional analytical modeling approaches are limited to a few survey features and a larger part of the survey would remain untouched by the decision model. This article presents a novel, data-driven modeling approach that strategically prunes a large set of consumer profile features using a machine learning framework to train a model for predicting potential solar adoption. The approach utilizes the Gradient Boosting Decision Tree model through a Light Gradient Boosting framework that improves significantly over the poor prediction accuracy of the existing approaches. Model training using focal-loss based supervision is used to overcome the difficulty in identifying the potential adopters that is inherent in conventional data-driven models. In addition, to overcome possible data sparsity in a limited survey sample, a Generative Adversarial Network is presented to create synthetic user samples and its effectiveness on model performance is assessed. A Bayesian optimization approach is used to systematically arrive at the hyperparameters of the proposed model. Validation of the presented approach on a survey data collected by the National Rural Electric Cooperative Association in Virginia in 2018 demonstrates the excellent predictive capability of the machine learning based approach to modeling solar adoption reliably.

14 SOLAR ENERGY↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Tree-based algorithms for weakly supervised anomaly detection

Weakly supervised methods have emerged as a powerful tool for model-agnostic anomaly detection at the Large Hadron Collider (LHC). While these methods have shown remarkable performance on specific signatures such as dijet resonances, their application in a more model-agnostic manner requires dealing with a larger number of potentially noisy input features. In this paper, we show that using boosted decision trees as classifiers in weakly supervised anomaly detection gives superior performance compared to deep neural networks. Boosted decision trees are well known for their effectiveness in tabular data analysis. Our results show that they not only offer significantly faster training and evaluation times, but they are also robust to a large number of noisy input features. By using advanced gradient boosted decision trees in combination with ensembling techniques and an extended set of features, we significantly improve the performance of weakly supervised methods for anomaly detection at the LHC. This advance is a crucial step toward a more model-agnostic search for new physics. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Investigating boosted decision trees as a guide for inertial confinement fusion design

Inertial confined fusion experiments at the National Ignition Facility have recently entered a new regime approaching ignition. Improved modeling and exploration of the experimental parameter space were essential to deepening our understanding of the mechanisms that degrade and amplify the neutron yield. The growing prevalence of machine learning in fusion studies opens a new avenue for investigation. Here in this paper, we have applied the Gradient-Boosted Decision Tree machine-learning architecture to further explore the parameter space and find correlations with the neutron yield, a key performance indicator. We find reasonable agreement between the measured and predicted yield, with a mean absolute percentage error on a randomly assigned test set of 35.5%. This model finds the characteristics of the laser pulse to be the most influential in prediction, as well as the hohlraum laser entrance hole diameter and an enhanced capsule fabrication technique. We used the trained model to scan over the design space of experiments from three different campaigns to evaluate the potential of this technique to provide design changes that could improve the resulting neutron yield. While these data-driven model cannot predict ignition without examples of ignited shots in the training set, it can be used to indicate that an unseen shot design will at least be in the upper range of previously observed neutron yields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Interpretable boosted-decision-tree analysis for the Majorana Demonstrator

The Majorana Demonstrator is a leading experiment searching for neutrinoless double-beta decay with high purity germanium detectors (HPGe). Machine learning provides a new way to maximize the amount of information provided by these detectors, but the data-driven nature makes it less interpretable compared to traditional analysis. An interpretability study reveals the machine's decision-making logic, allowing us to learn from the machine to feedback to the traditional analysis. In this work, we have presented the first machine learning analysis of the data from the Majorana Demonstrator; this is also the first interpretable machine learning analysis of any germanium detector experiment. Two gradient boosted decision tree models are trained to learn from the data, and a game-theory-based model interpretability study is conducted to understand the origin of the classification power. By learning from data, this analysis recognizes the correlations among reconstruction parameters to further enhance the background rejection performance. By learning from the machine, this analysis reveals the importance of new background categories to reciprocally benefit the standard Majorana analysis. This model is highly compatible with next-generation germanium detector experiments like LEGEND since it can be simultaneously trained on a large number of detectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Linear model decision trees as surrogates in optimization of engineering applications

Machine learning models are promising as surrogates in optimization when replacing difficult to solve equations or black-box type models. This work demonstrates the viability of linear model decision trees as piecewise-linear surrogates in decision-making problems. Linear model decision trees can be represented exactly in mixed-integer linear programming (MILP) and mixed-integer quadratic constrained programming (MIQCP) formulations. Furthermore, they can represent discontinuous functions, bringing advantages over neural networks in some cases. We present several formulations using transformations from Generalized Disjunctive Programming (GDP) formulations and modifications of MILP formulations for gradient boosted decision trees (GBDT). We then compare the computational performance of these different MILP and MIQCP representations in an optimization problem and illustrate their use on engineering applications. Importantly, we observe faster solution times for optimization problems with linear model decision tree surrogates when compared with GBDT surrogates using the Optimization and Machine Learning Toolkit (OMLT).

42 ENGINEERING↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗