Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Decision tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

An automated approach to the design of decision tree classifiers

The classification of large dimensional data sets arising from the merging of remote sensing data with more traditional forms of ancillary data is considered. Decision tree classification, a popular approach to the problem, is characterized by the property that samples are subjected to a sequence of decision rules before they are assigned to a unique class. An automated technique for effective decision tree design which relies only on apriori statistics is presented. This procedure utilizes a set of two dimensional canonical transforms and Bayes table look-up decision rules. An optimal design at each node is derived based on the associated decision table. A procedure for computing the global probability of correct classfication is also provided. An example is given in which class statistics obtained from an actual LANDSAT scene are used as input to the program. The resulting decision tree design has an associated probability of correct classification of .76 compared to the theoretically optimum .79 probability of correct classification associated with a full dimensional Bayes classifier. Recommendations for future research are included.

Argentiero, P.↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

Disentangling error structures of precipitation datasets using decision trees

Characterizing error structures in precipitation products not only facilitates their proper applications for scientific and practical purposes but also helps improve their retrieval algorithms and processing methods. Despite the fact that multiple precipitation products have been assessed in the literature, factors that affect their error structures remain inadequately addressed. By interpreting 60 binary decision trees, this study disentangles the error characteristics of precipitation products in terms of their spatiotemporal patterns and geographical factors. Three independent precipitation products - two satellite-based and one reanalysis datasets: the Integrated Multi-satellitE Retrievals for GPM (Global Precipitation Measurement) late run (IMERG-L), Soil Moisture to Rain-Advanced SCATterometer (SM2RAIN-ASCAT), and the Modern-Era Retrospective analysis for Research and Applications, Version 2 uncorrected precipitation output (MERRA2-UC), are evaluated across the contiguous United States from 2010 to 2019. Here, the ground-based Stage IV precipitation dataset is used as the ground truth. Results indicate that the MERRA2-UC outperforms the IMERG-L and SM2RAIN-ASCAT with higher accuracy and more stable interannual patterns for the analysis period. Decision trees cross-assess three spatiotemporal factors and find that the underestimation of MERRA2-UC occurs in the east of the Rocky Mountains, and SM2RAIN-ASCAT underestimates precipitation over high latitudes, especially in winter. Additionally, the decision tree method ascribes system errors to nine different geographical characteristics, of which the distance to the coast, soil type, and DEM are the three dominant features. On the other hand, the land cover type, topography position index, and aspect are three relatively weak factors.

54 ENVIRONMENTAL SCIENCES↗

Biomass for Carbon Removal and Storage (BiCRS) Counterfactual Decision Tree

Counterfactual is the term used to describe a "business-as-usual" scenario which used as a baseline to compare against a new project, allowing the calculation of net impacts for a life cycle analysis (LCA). The choice of counterfactual is critical for determining the results from LCA and must be carefully justified to ensure a fair and accurate comparison. Using forest residues as an example, this decision tree illustrates decision points to be considered for sustainable biomass sourcing and provides a framework for estimating the carbon emissions or storage under the "business-as-usual” scenarios for biomass otherwise destined for use in Biomass for Carbon Removal and Storage (BiCRS) projects.

09 BIOMASS FUELS↗

MODIS Snow Cover Mapping Decision Tree Technique: Snow and Cloud Discrimination

Accurate mapping of snow cover continues to challenge cryospheric scientists and modelers. The Moderate-Resolution Imaging Spectroradiometer (MODIS) snow data products have been used since 2000 by many investigators to map and monitor snow cover extent for various applications. Users have reported on the utility of the products and also on problems encountered. Three problems or hindrances in the use of the MODIS snow data products that have been reported in the literature are: cloud obscuration, snow/cloud confusion, and snow omission errors in thin or sparse snow cover conditions. Implementation of the MODIS snow algorithm in a decision tree technique using surface reflectance input to mitigate those problems is being investigated. The objective of this work is to use a decision tree structure for the snow algorithm. This should alleviate snow/cloud confusion and omission errors and provide a snow map with classes that convey information on how snow was detected, e.g. snow under clear sky, snow tinder cloud, to enable users' flexibility in interpreting and deriving a snow map. Results of a snow cover decision tree algorithm are compared to the standard MODIS snow map and found to exhibit improved ability to alleviate snow/cloud confusion in some situations allowing up to about 5% increase in mapped snow cover extent, thus accuracy, in some scenes.

Riggs, George A.↗

Nanosecond machine learning regression with deep boosted decision trees in FPGA for high energy physics

We present a novel application of the machine learning / artificial intelligence method called boosted decision trees to estimate physical quantities on field programmable gate arrays (FPGA). The software package fwXmachina features a new architecture called parallel decision paths that allows for deep decision trees with arbitrary number of input variables. It also features a new optimization scheme to use different numbers of bits for each input variable, which produces optimal physics results and ultraefficient FPGA resource utilization. Problems in high energy physics of proton collisions at the Large Hadron Collider (LHC) are considered. Estimation of missing transverse momentum (E T miss ) at the first level trigger system at the High Luminosity LHC (HL-LHC) experiments, with a simplified detector modeled by Delphes, is used to benchmark and characterize the firmware performance. The firmware implementation with a maximum depth of up to 10 using eight input variables of 16-bit precision gives a latency value of $\mathcal{O}$(10) ns, independent of the clock speed, and $\mathcal{O}$(0.1)% of the available FPGA resources without using digital signal processors.

Instruments & Instrumentation↗

Using Boosted Decision Trees to Select High Quality Measurements in the Mu2e Experiment at Fermilab

This thesis presents the implementation and evaluation of a Boosted Decision Tree (BDT) model to improve the selection of high-quality track measurements in the Mu2e experiment at Fermilab. The Mu2e experiment is a high-energy physics experiments seeking to observe a rare theoretical physics process known as Charged Lepton Flavor Violation. A significant challenge faced by the Mu2e experiment are so-called background events, which are events whose data mimics that of the rare physics process the experiment seeks to observe. Without a mechanism to reduce background, it would be impossible to know whether Charged Lepton Flavor Violation occurred or not. To this end, high-quality track measurements must be distinguished from low-quality track measurements. A track can be conceived of as the reconstructed path of a particle that traveled through the Mu2e detector. In addition to other data, data about such tracks is stored using a C++-based framework, specific to the domain of high-energy physics, known as ROOT. A boosted decision tree model was trained using ROOT’s Toolkit For Multivariate Analysis by leveraging variables ancillary to track quality. In evaluation, the BDT achieves a ROC-AUC of 0.927 in discriminating good-quality tracks from poor-quality tracks. Such a score is indicative of both strong discrimination and strong generalization. Subsequently, it is shown that applying a BDT-based quality cut to the distribution of particle momenta significantly enhances the signal-to-background distinction for signal electrons, paving the way for improved sensitivity to Charged Lepton Flavor Violation.

Mullany, Brendan T. [Drew U.] (ORCID:0009000818888↗

The decision tree classifier - Design and potential

A new classifier has been developed for the computerized analysis of remote sensor data. The decision tree classifier is essentially a maximum likelihood classifier using multistage decision logic. It is characterized by the fact that an unknown sample can be classified into a class using one or several decision functions in a successive manner. The classifier is applied to the analysis of data sensed by Landsat-1 over Kenosha Pass, Colorado. The classifier is illustrated by a tree diagram which for processing purposes is encoded as a string of symbols such that there is a unique one-to-one relationship between string and decision tree.

Hauska, H.↗

Decision Tree for Variable Selection vs. Impact on Durability for Biomass and Biochar Burial Pathways [Slides]

Quantifying durability for lower-TRL BiCRS pathways has been challenging as limited data are available from real-world projects and long-term experiments, resulting in an overall lack of scientific consensus. We develop a decision tree that aims to summarize the current scientific understanding and state-of-the-art project experience. The decision tree can be used to (1) guide the selection of key variables and evaluate their relative impact on durability, (2) identify data and knowledge gaps for future research.

09 BIOMASS FUELS↗

Three-dimensional object recognition using similar triangles and decision trees

A system, TRIDEC, that is capable of distinguishing between a set of objects despite changes in the objects' positions in the input field, their size, or their rotational orientation in 3D space is described. TRIDEC combines very simple yet effective features with the classification capabilities of inductive decision tree methods. The feature vector is a list of all similar triangles defined by connecting all combinations of three pixels in a coarse coded 127 x 127 pixel input field. The classification is accomplished by building a decision tree using the information provided from a limited number of translated, scaled, and rotated samples. Simulation results are presented which show that TRIDEC achieves 94 percent recognition accuracy in the 2D invariant object recognition domain and 98 percent recognition accuracy in the 3D invariant object recognition domain after training on only a small sample of transformed views of the objects.

Spirkovska, Lilly↗

Using decision-tree classifier systems to extract knowledge from databases

One difficulty in applying artificial intelligence techniques to the solution of real world problems is that the development and maintenance of many AI systems, such as those used in diagnostics, require large amounts of human resources. At the same time, databases frequently exist which contain information about the process(es) of interest. Recently, efforts to reduce development and maintenance costs of AI systems have focused on using machine learning techniques to extract knowledge from existing databases. Research is described in the area of knowledge extraction using a class of machine learning techniques called decision-tree classifier systems. Results of this research suggest ways of performing knowledge extraction which may be applied in numerous situations. In addition, a measurement called the concept strength metric (CSM) is described which can be used to determine how well the resulting decision tree can differentiate between the concepts it has learned. The CSM can be used to determine whether or not additional knowledge needs to be extracted from the database. An experiment involving real world data is presented to illustrate the concepts described.

St.clair, D. C.↗

Learning from examples - Generation and evaluation of decision trees for software resource analysis

A general solution method for the automatic generation of decision (or classification) trees is investigated. The approach is to provide insights through in-depth empirical characterization and evaluation of decision trees for software resource data analysis. The trees identify classes of objects (software modules) that had high development effort. Sixteen software systems ranging from 3,000 to 112,000 source lines were selected for analysis from a NASA production environment. The collection and analysis of 74 attributes (or metrics), for over 4,700 objects, captured information about the development effort, faults, changes, design style, and implementation style. A total of 9,600 decision trees were automatically generated and evaluated. The trees correctly identified 79.3 percent of the software modules that had high development effort or faults, and the trees generated from the best parameter combinations correctly identified 88.4 percent of the modules on the average.

Selby, Richard W.↗

Development of heavy-duty vehicle representative driving cycles via decision tree regression

Previously, researchers who developed representative driving cycles mainly focused on light-duty vehicles and only considered vehicle speed and related derivations. In this paper, we propose a novel approach to develop representative cycles for heavy-duty vehicles. By implementing decision tree regression (DTR) to the Fleet DNA on-road vehicle data, a broader set of metrics, such as engine power and fuel consumption, can be used for more robust cycle development. Additionally, the influence of each metric on the regression target is also accounted for by a weighted number derived through the DTR to enhance the representativenss of the developed cycle. As case studies, we applied the proposed method to five heavy-duty vocations (drayage, long haul, regional haul, local delivery, and transit bus) and derived the most representative cycle, as well as four extreme cycles (maximal energy consumption, maximal power-weighted work, maximal fraction of high speed, and minimal fuel economy) to advance the related alternative powertrain design.

33 ADVANCED PROPULSION SYSTEMS↗

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees↗

Decision Tree Regression to Identify Representative Road Sections for Evaluating Performance of Connected and Automated Class 8 Tractors

Currently, connected and autonomous vehicle (CAV) technology is being developed for Class 8 tractor trucks aimed at improved safety and fuel economy and reduced CO2 emissions. Despite extensive efforts conducted across the world, the reported efficiency gains were varied from different research groups, raising concerns about the fidelity of models, the performance of control, and the effectiveness of the experimental validation. One root cause for this variation stems from the fact that the efficiency gain obtained from the CAV is sensitive to real-world conditions, including surrounding traffic and road grade. This study presents an approach aimed at identifying representative public road sections and facilitating CAV research from this perspective. By employing the decision tree regression (DTR) method to the Fleet DNA database, the most representative road sections can be identified. High-level metrics and detailed information of the derived road sections are also illustrated and discussed, which demonstrate their representativeness and the effectiveness of the approach. Meanwhile, the capability of this approach can be easily extended by integrating specific constraints into the DTR algorithm. As an example, a specific representative road section with an aggressive road grade profile was also provided via this approach.

27 ARPA - Advanced Research Projects Agency-Energy↗

Nanosecond anomaly detection with decision trees and real-time application to exotic Higgs decays

Abstract We present an interpretable implementation of the autoencoding algorithm, used as an anomaly detector, built with a forest of deep decision trees on FPGA, field programmable gate arrays. Scenarios at the Large Hadron Collider at CERN are considered, for which the autoencoder is trained using known physical processes of the Standard Model. The design is then deployed in real-time trigger systems for anomaly detection of unknown physical processes, such as the detection of rare exotic decays of the Higgs boson. The inference is made with a latency value of 30 ns at percent-level resource usage using the Xilinx Virtex UltraScale+ VU9P FPGA. Our method offers anomaly detection at low latency values for edge AI users with resource constraints.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗