Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Decision tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Interpretable boosted-decision-tree analysis for the Majorana Demonstrator

The Majorana Demonstrator is a leading experiment searching for neutrinoless double-beta decay with high purity germanium detectors (HPGe). Machine learning provides a new way to maximize the amount of information provided by these detectors, but the data-driven nature makes it less interpretable compared to traditional analysis. An interpretability study reveals the machine's decision-making logic, allowing us to learn from the machine to feedback to the traditional analysis. In this work, we have presented the first machine learning analysis of the data from the Majorana Demonstrator; this is also the first interpretable machine learning analysis of any germanium detector experiment. Two gradient boosted decision tree models are trained to learn from the data, and a game-theory-based model interpretability study is conducted to understand the origin of the classification power. By learning from data, this analysis recognizes the correlations among reconstruction parameters to further enhance the background rejection performance. By learning from the machine, this analysis reveals the importance of new background categories to reciprocally benefit the standard Majorana analysis. This model is highly compatible with next-generation germanium detector experiments like LEGEND since it can be simultaneously trained on a large number of detectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating boosted decision trees as a guide for inertial confinement fusion design

Inertial confined fusion experiments at the National Ignition Facility have recently entered a new regime approaching ignition. Improved modeling and exploration of the experimental parameter space were essential to deepening our understanding of the mechanisms that degrade and amplify the neutron yield. The growing prevalence of machine learning in fusion studies opens a new avenue for investigation. Here in this paper, we have applied the Gradient-Boosted Decision Tree machine-learning architecture to further explore the parameter space and find correlations with the neutron yield, a key performance indicator. We find reasonable agreement between the measured and predicted yield, with a mean absolute percentage error on a randomly assigned test set of 35.5%. This model finds the characteristics of the laser pulse to be the most influential in prediction, as well as the hohlraum laser entrance hole diameter and an enhanced capsule fabrication technique. We used the trained model to scan over the design space of experiments from three different campaigns to evaluate the potential of this technique to provide design changes that could improve the resulting neutron yield. While these data-driven model cannot predict ignition without examples of ignited shots in the training set, it can be used to indicate that an unseen shot design will at least be in the upper range of previously observed neutron yields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Offshore application of landslide susceptibility mapping using gradient-boosted decision trees: a Gulf of Mexico case study

Abstract Among natural hazards occurring offshore, submarine landslides pose a significant risk to offshore infrastructure installations attached to the seafloor. With the offshore being important for current and future energy production, there is a need to anticipate where future landslide events are likely to occur to support planning and development projects. Using the northern Gulf of Mexico (GoM) as a case study, this paper performs Landslide Susceptibility Mapping (LSM) using a gradient-boosted decision tree (GBDT) model to characterize the spatial patterns of submarine landslide probability over the United States Exclusive Economic Zone (EEZ) where water depths are greater than 120 m. With known spatial extents of historic submarine landslides and a Geographic Information System (GIS) database of known topographical, geomorphological, geological, and geochemical factors, the resulting model was capable of accurately forecasting potential locations of sediment instability. Results of a permutation modelling approach indicated that LSM accuracy is sensitive to the number of unique training locations with model accuracy becoming more stable as the number of training regions was increased. The influence that each input feature had on predicting landslide susceptibility was evaluated using the SHapely Additive exPlanations (SHAP) feature attribution method. Areas of high and very high susceptibility were associated with steep terrain including salt basins and escarpments. This case study serves as an initial assessment of the machine learning (ML) capabilities for producing accurate submarine landslide susceptibility maps given the current state of available natural hazard-related datasets and conveys both successes and limitations.

Dyer, Alec S. (ORCID:0000000219813904)↗

On the Investigation of Phase Fault Classification in Power Grid Signals: A Case Study for Support Vector Machines, Decision Tree and Random Forest

In monitoring the power grid, an ability to differentiate between fault types is essential to ensuring electrical safety. Accordingly, this study introduces a fault detection and classification method by considering different machine learning (ML) and feature extraction (FE) methods combinations. Specifically, the proposed method is established in two classification layers; the first layer determines the fault, and the second layer distinguishes the type of fault. Based on the proposed system model, this study seeks to determine the influential data attributes in a power grid signal using FE methods, including fast Fourier transform, power spectral density (PSD), auto-correlation, and wavelet transform (WT). A cross-comparison of the effectiveness of the Support Vector Machine (SVM), Decision Tree (DT), and Random Forest (RF) is also performed to accomplish the classification layers of the proposed method. The designed algorithm is analyzed under the various combinations of FE and ML methods, and outcomes are presented by considering the trade-off between computational complexity and prediction accuracy. The results reveal that the RF-based ML algorithm shows the most accurate classification performance with PSD, and the most time-saving of the models is the DT WT. Also, SVM emerges superior on a subsequent test of the simulated models on real-world signals.

Galbraith, Kelli↗

Utilizing waste heat in wastewater treatment plants for water desalination: Modeling and Multi-Objective optimization of a Multi-Effect desalination system using Decision Tree Regression and Pelican optimization algorithm

This paper examines the feasibility of using waste heat from wastewater treatment plants (WWTPs) for water desalination. A model was developed to utilize waste heat from the gensets at As Samra WWTP in Jordan, using real data and TRNSYS® software to calculate available waste heat. The desalination process was then modeled with ASPEN PLUS® software, focusing on multi-effect desalination (MED). Both series and parallel configurations for the MED system were compared. The study investigated the effects of system feeding flow rate, feeding pressure, and heat input on productivity, performance ratio, and recovery ratio. The study also introduces a novel optimization technique combining machine learning and modern optimization algorithms to maximize system productivity and performance. Initially, a decision tree regression (DTR) model is developed to establish relationships between key independent variables (flow rate, feed pressure, and heat input) and dependent variables (productivity, performance ratio, and recovery ratio). The Pelican Optimization Algorithm (POA) is then used to identify the optimal values of the independent variables for maximum productivity and performance. The results show that using a series configuration yields a system productivity of 3984.2 kg/hr, a performance ratio of 3.78, and a recovery ratio of 0.991 at a feed flow rate of 4000 kg/hr, feed pressure of 3 bars, and heat input of 719 kW. Optimal productivity (4421 kg/hr), performance ratio (3.81), and recovery ratio (0.851) are achieved at a feed flow rate of 5166 kg/hr, feed pressure of 3.2 bars, and heat input of 794 kW. In conclusion, the techno-economic assessment indicates a levelized cost of water of 1.63 USD/m 3 for parallel configurations and 1.65 USD/m 3 for series configurations, with a payback period of less than two years.

42 ENGINEERING↗

Integrating environmental understanding into freshwater floatovoltaic deployment using an effects hierarchy and decision trees

Abstract In an era of looming land scarcity and environmental degradation, the development of low carbon energy systems without adverse impacts on land and land-based resources is a global challenge. ‘Floatovoltaic’ energy systems—comprising floating photovoltaic (PV) panels over water—are an appealing source of low carbon energy as they spare land for other uses and attain greater electricity outputs compared to land-based systems. However, to date little is understood of the impacts of floatovoltaics on the hosting water body. Anticipating changes to water body processes, properties and services owing to floatovoltaic deployment represents a critical knowledge gap that may result in poor societal choices and water body governance. Here, we developed a theoretically-derived hierarchical effects framework for the assessment of floatovoltaic impacts on freshwater water bodies, emphasising ecological interactions. We describe how the presence of floatovoltaic systems may dramatically alter the air-water interface, with subsequent implications for surface meteorology, air-water fluxes and physical, chemical and biological properties of the recipient water body. We apply knowledge from this framework to delineate three response typologies—‘ magnitude’ , those for which the direction and magnitude of effect can be predicted; ‘ direction’ , those for which only the direction of effect can be predicted; and ‘ uncertain’ , those for which the response cannot be predicted—characterised by the relative importance of levels in the effects hierarchy. Illustrative decision trees are developed for an example water body response within each typology, specifically, evaporative water loss, cyanobacterial biomass, and phosphorus release from bed sediments, and implications for ecosystem services, including climate regulation, are discussed. Finally, the potential to use the new understanding of likely ecosystem perturbations to direct floatovoltaic design innovations and identify future research priorities is outlined, showcasing how inter-sectoral collaboration and environmental science can inform and optimise this low carbon, land-sparing renewable energy for ecosystem gains.

Environmental Sciences & Ecology↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗