Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Decision tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Predicting Ground Delay Program at an Airport Based on Meteorological Conditions

In this paper, we present two supervised-learning models, logistic regression and decision tree, to predict occurrence of ground delay program at an airport based on meteorological conditions and scheduled traffic demand. Such predictive capabilities can help the Federal Aviation Administration traffic managers and airline dispatchers to prepare mitigation strategies to reduce the impact of adverse weather. The models are applied to predict ground delay program occurrence at two major U.S. airports: Newark Liberty Intl. and San Francisco Intl. airports. The logistic regression model estimates the probability that a ground delay program will occur during a given hour. Decision tree, on the other hand, classifies an hour as a ground delay program or not based on the input variables. Results indicate that both models perform significantly better than a purely random prediction of ground delay program occurrence at the two airports. The logistic regression model performs better than the decision tree model. The degree to which various input variables impact the probability of ground delay program vary between the two airports. While the enroute convective weather is a dominant factor causing ground delay programs at New York airports, poor visibility and low cloud ceiling caused by marine stratus are major drivers of ground delay programs at San Francisco Intl. airport.

traffic flow management↗

Data fusion with artificial neural networks (ANN) for classification of earth surface from microwave satellite measurements

A data fusion system with artificial neural networks (ANN) is used for fast and accurate classification of five earth surface conditions and surface changes, based on seven SSMI multichannel microwave satellite measurements. The measurements include brightness temperatures at 19, 22, 37, and 85 GHz at both H and V polarizations (only V at 22 GHz). The seven channel measurements are processed through a convolution computation such that all measurements are located at same grid. Five surface classes including non-scattering surface, precipitation over land, over ocean, snow, and desert are identified from ground-truth observations. The system processes sensory data in three consecutive phases: (1) pre-processing to extract feature vectors and enhance separability among detected classes; (2) preliminary classification of Earth surface patterns using two separate and parallely acting classifiers: back-propagation neural network and binary decision tree classifiers; and (3) data fusion of results from preliminary classifiers to obtain the optimal performance in overall classification. Both the binary decision tree classifier and the fusion processing centers are implemented by neural network architectures. The fusion system configuration is a hierarchical neural network architecture, in which each functional neural net will handle different processing phases in a pipelined fashion. There is a total of around 13,500 samples for this analysis, of which 4 percent are used as the training set and 96 percent as the testing set. After training, this classification system is able to bring up the detection accuracy to 94 percent compared with 88 percent for back-propagation artificial neural networks and 80 percent for binary decision tree classifiers. The neural network data fusion classification is currently under progress to be integrated in an image processing system at NOAA and to be implemented in a prototype of a massively parallel and dynamically reconfigurable Modular Neural Ring (MNR).

Lure, Y. M. Fleming↗

A procedure for rule extraction from a Self-Organising plasma disruption predictor for JET

In a previous paper, a Self-Organizing Map had proven to be able to identify the regions of the plasma operative space characterizing the pre-disruptive phase at JET without relying on any a priori information. One of the strengths of this disruption predictor lies in its inherent self-organization capability. The Self-Organizing Map discovers non-trivial relationships and captures the complicated interplay of device diagnostics on the internal plasma states directly from the experimental data. Moreover, the provided model allows the visualization of high-dimensional plasma parameters and facilitates easy interrogation of the model to understand the reasons behind its correlations. In this paper, an additional step is taken towards the interpretability of models for predicting disruptions by training a Decision Tree to classify the plasma states according to the interpretation provided by the Self-Organizing Map (stable or at high risk of disruptions). The Decision tree provides a set of rules which describe the transition of the plasma towards the pre-disruptive phase as visualized in the Self-Organizing Map. The obtained rules for the database explored in the study identify four regions in the map, two of which are at risk of disruption. These regions correspond to partitions of a 3D space based on the peaking factors of the core and divertor radiation, as well as the Locked Mode. The agreement between the Self-Organizing Map answers and the rules supplied by the Decision Tree is confirmed by the comparison of the performance exhibited by the two models in the prediction of disruptions.

Setzu, Samuele [Univ. of Cagliari, Monserrato, Cag↗

Tree-based algorithms for weakly supervised anomaly detection

Weakly supervised methods have emerged as a powerful tool for model-agnostic anomaly detection at the Large Hadron Collider (LHC). While these methods have shown remarkable performance on specific signatures such as dijet resonances, their application in a more model-agnostic manner requires dealing with a larger number of potentially noisy input features. In this paper, we show that using boosted decision trees as classifiers in weakly supervised anomaly detection gives superior performance compared to deep neural networks. Boosted decision trees are well known for their effectiveness in tabular data analysis. Our results show that they not only offer significantly faster training and evaluation times, but they are also robust to a large number of noisy input features. By using advanced gradient boosted decision trees in combination with ensembling techniques and an extended set of features, we significantly improve the performance of weakly supervised methods for anomaly detection at the LHC. This advance is a crucial step toward a more model-agnostic search for new physics. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

The analysis of rapidly developing fog at the Kennedy Space Center

This report documents fog precursors and fog climatology at Kennedy Space Center (KSC) Florida from 1986 to 1990. The major emphasis of this report focuses on rapidly developing fog events that would affect the less than 7-statute mile visibility rule for End-Of-Mission (EOM) Shuttle landing at KSC (Rule 4-64(A)). The Applied Meteorology Unit's (AMU's) work is to: develop a data base for study of fog associated weather conditions relating to violations of this landing constraint; develop forecast techniques or rules-of-thumb to determine whether or not current conditions are likely to result in an acceptable condition at landing; validate the forecast techniques; and transition techniques to operational use. As part of the analysis the fog events were categorized as either advection, pre-frontal or radiation. As a result of these analyses, the AMU developed a fog climatological data base, identified fog precursors and developed forecaster tools and decision trees. The fog climatological analysis indicates that during the fog season (October to April) there is a higher risk for a visibility violation at KSC during the early morning hours (0700 to 1200 UTC), while 95 percent of all fog events have dissipated by 1600 UTC. A high number of fog events are characterized by a westerly component to the surface wind at KSC (92 percent) and 83 percent of the fog events had fog develop west of KSC first (up to 2 hours). The AMU developed fog decision trees and forecaster tools that would help the forecaster identify fog precursors up to 12 hours in advance. Using the decision trees as process tools ensures the important meteorological data are not overlooked in the forecast process. With these tools and a better understanding of fog formation in the local KSC area, the Shuttle weather support forecaster should be able to give the Launch and Flight Directors a better KSC fog forecast with more confidence.

Wheeler, Mark M.↗

A Model Tree Generator (MTG) Framework for Simulating Hydrologic Systems: Application to Reservoir Routing

Data-driven algorithms have been widely used as effective tools to mimic hydrologic systems. Unlike black-box models, decision tree algorithms offer transparent representations of systems and reveal useful information about the underlying process. A popular class of decision tree models is model tree (MT), which is designed for predicting continuous variables. Most MT algorithms employ an exhaustive search mechanism and a pre-defined splitting criterion to generate a piecewise linear model. However, this approach is computationally intensive, and the selection of the splitting criterion can significantly affect the performance of the generated model. These drawbacks can limit the application of MTs to large datasets. To overcome these shortcomings, a new flexible Model Tree Generator (MTG) framework is introduced here. MTG is equipped with several modules to provide a flexible, efficient, and effective tool for generating MTs. The application of the algorithm is demonstrated through simulation of controlled discharge from several reservoirs across the Contiguous United States (CONUS).

54 ENVIRONMENTAL SCIENCES↗

Avatar Tools

Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Siefert, Christopher↗

A comparison of machine learning methods to classify radioactive elements using prompt-gamma-ray neutron activation data

The detection of illicit radiological materials is critical to establishing a robust second line of defence in nuclear security. Neutron-capture prompt-gamma activation analysis (PGAA) can be used to detect multiple radioactive materials across the entire Periodic Table. However, long detection times and a high rate of false positives pose a significant hindrance in the deployment of PGAA-based systems to identify the presence of illicit substances in nuclear forensics. In the present work, six different machine-learning algorithms were developed to classify radioactive elements based on the PGAA energy spectra. The model performance was evaluated using standard classification metrics and trend curves with an emphasis on comparing the effectiveness of algorithms that are best suited for classifying imbalanced datasets. We analyse the classification performance based on Precision, Recall, F1-score, Specificity, Confusion matrix, ROC-AUC curves, and Geometric Mean Score (GMS) measures. The tree-based algorithms (Decision Trees, Random Forest and AdaBoost) have consistently outperformed Support Vector Machine and K-Nearest Neighbours. Based on the results presented, AdaBoost is the preferred classifier to analyse data containing PGAA spectral information due to the high recall and minimal false negatives reported in the minority class.

97 MATHEMATICS AND COMPUTING↗

Entropy removal of medical diagnostics

Shannon entropy is a core concept in machine learning and information theory, particularly in decision tree modeling. To date, no studies have extensively and quantitatively applied Shannon entropy in a systematic way to quantify the entropy of clinical situations using diagnostic variables (true and false positives and negatives, respectively). Decision tree representations of medical decision-making tools can be generated using diagnostic variables found in literature and entropy removal can be calculated for these tools. This concept of clinical entropy removal has significant potential for further use to bring forth healthcare innovation, such as quantifying the impact of clinical guidelines and value of care and applications to Emergency Medicine scenarios where diagnostic accuracy in a limited time window is paramount. This analysis was done for 623 diagnostic tools and provided unique insights into their utility. For studies that provided detailed data on medical decision-making algorithms, bootstrapped datasets were generated from source data to perform comprehensive machine learning analysis on these algorithms and their constituent steps, which revealed a novel and thorough evaluation of medical diagnostic algorithms.

97 MATHEMATICS AND COMPUTING↗

Algorithm For A Self-Growing Neural Network

CID3 algorithm simulates self-growing neural network. Constructs decision trees equivalent to hidden layers of neural network. Based on ID3 algorithm, which dynamically generates decision tree while minimizing entropy of information. CID3 algorithm generates feedforward neural network by use of either crisp or fuzzy measure of entropy.

Cios, Krzysztof J.↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

An improved classification tree analysis of high cost modules based upon an axiomatic definition of complexity

Identification of high cost modules has been viewed as one mechanism to improve overall system reliability, since such modules tend to produce more than their share of problems. A decision tree model was used to identify such modules. In this current paper, a previously developed axiomatic model of program complexity is merged with the previously developed decision tree process for an improvement in the ability to identify such modules. This improvement was tested using data from the NASA Software Engineering Laboratory.

Tian, Jianhui↗

QuantifyML: How good is my machine learning model?

This paper presents an approach, QuantifyML, which employs model counting to assess the learnability and robustness of machine learning models. Typically the efficacy of machine learning models is determined by computing their accuracy statistically on test data sets. However, this may be misleading, if the test data is not representative of the problem that is being studied. Further, two different models may have the same accuracy on a given data set, measured statistically, but may be very different in their behavior on unseen data. Also, models with high accuracy could have poor adversarial robustness. In QuantifyML, our goal is to precisely quantify the extent to which machine learning models have learned and generalized from the given data. In QuantifyML, a trained model is translated into a C program, which is fed to the CBMC model checking tool to produce a formula in Conjunctive Normal Form (CNF), which in turn is analyzed with state-of-the-art model counters to efficiently obtain precise counts w.r.t different outputs. QuantifyML enables i) evaluating the learnability of models by comparing the counts for the outputs to ground truth, expressed as logical predicates (if available), ii) comparing the performance of different models that may be built with different machine learning algorithms (e.g., decision-trees vs. neural networks), and iii) quantifying the robustness of trained models around given inputs. Our evaluation demonstrates these applications of QuantifyML on decision trees and neural networks trained to learn relational properties of graphs, for which we know the ground truth, and to perform image classification, for which we do not have the ground truth, but we can quantify local robustness.

Deep Neural Networks↗

Prognostic analysis of high-flow nasal cannula therapy and non-invasive ventilation in mild to moderate hypoxemia patients and construction of a machine learning model for 48-h intubation prediction—a retrospective analysis of the MIMIC database

Background This study aims to investigate the clinical outcome between high-flow nasal cannula (HFNC) and non-invasive ventilation (NIV) therapy in mild to moderate hypoxemic patients on the first ICU day and to develop a predictive model of 48-h intubation. Methods The study included adult patients from the MIMIC III and IV databases who first initiated HFNC or NIV therapy due to mild to moderate hypoxemia (100 < PaO2/FiO2 ≤ 300). The 48-h and 30-day intubation rates were compared using cross-sectional and survival analysis. Nine machine learning and six ensemble algorithms were deployed to construct the 48-h intubation predictive models, of which the optimal model was determined by its prediction accuracy. The top 10 risk and protective factors were identified using the Shapley interpretation algorithm. Result A total of 123,042 patients were screened, of which, 673 were from the MIMIC IV database for ventilation therapy comparison (HFNC n = 363, NIV n = 310) and 48-h intubation predictive model construction (training dataset n = 471, internal validation set n = 202) and 408 were from the MIMIC III database for external validation. The NIV group had a lower intubation rate (23.1% vs. 16.1%, p = 0.001), ICU 28-day mortality (18.5% vs. 11.6%, p = 0.014), and in-hospital mortality (19.6% vs. 11.9%, p = 0.007) compared to the HFNC group. Survival analysis showed that the total and 48-h intubation rates were not significantly different. The ensemble AdaBoost decision tree model (internal and external validation set AUROC 0.878, 0.726) had the best predictive accuracy performance. The model Shapley algorithm showed Sequential Organ Failure Assessment (SOFA), acute physiology scores (APSIII), the minimum and maximum lactate value as risk factors for early failure and age, the maximum PaCO 2 and PH value, Glasgow Coma Scale (GCS), the minimum PaO 2 /FiO 2 ratio, and PaO 2 value as protective factors. Conclusion NIV was associated with lower intubation rate and ICU 28-day and in-hospital mortality. Further survival analysis reinforced that the effect of NIV on the intubation rate might partly be attributed to the other impact factors. The ensemble AdaBoost decision tree model may assist clinicians in making clinical decisions, and early organ function support to improve patients’ SOFA, APSIII, GCS, PaCO 2 , PaO 2 , PH, PaO 2 /FiO 2 ratio, and lactate values can reduce the early failure rate and improve patient prognosis.

Fu, Wei↗

Prediction of Weather Impacted Airport Capacity using Ensemble Learning

Ensemble learning with the Bagging Decision Tree (BDT) model was used to assess the impact of weather on airport capacities at selected high-demand airports in the United States. The ensemble bagging decision tree models were developed and validated using the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) data and weather forecast at these airports. The study examines the performance of BDT, along with traditional single Support Vector Machines (SVM), for airport runway configuration selection and airport arrival rates (AAR) prediction during weather impacts. Testing of these models was accomplished using observed weather, weather forecast, and airport operation information at the chosen airports. The experimental results show that ensemble methods are more accurate than a single SVM classifier. The airport capacity ensemble method presented here can be used as a decision support model that supports air traffic flow management to meet the weather impacted airport capacity in order to reduce costs and increase safety.

Weather impact↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

traffic flow management↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

weather↗