Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “decision tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Operational Forecasting of Induced Seismicity (CRADA Final Report)

This was a collaborative effort between Lawrence Livermore National Security, LLC ("LLNS"), as manager and operator of Lawrence Livermore National Laboratory ("LLNL"), The Regents of the University of California, as manager and operator of Lawrence Berkeley National Laboratory (Collectively, Contractors) and Nanometrics, Inc. ("Participant"), to develop a toolkit called "Operational Forecasting of Induced Seismicity (ORION)" that includes a decision tree method for operational forecasting of induced seismicity rates related to fluid disposal operations.

58 GEOSCIENCES↗

Connectivity Troubleshooting Guide for Advanced Electricity Meters

This guide is designed to walk users through a troubleshooting process for disconnected advanced electricity meters. The guide centers on two connectivity troubleshooting decision trees related to power issues and network connectivity issues, respectively. Additionally, it contains additional background information such as definitions, diagrams, and checklists that can help a user prepare to troubleshoot meter connectivity issues.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Baseline Hypothetical Facility for the Production of 131 I and 99 Mo using Activation Targets

This report describes a hypothetical facility for production of medical radioisotopes via activation under the Proliferation Resistance and Optimization (PRO-X) program. The facility uses neutron activation of non-special nuclear material (SNM) to produce the medical isotopes 131 I and 99 Mo at a throughput of 60 Ci/week of 131 I and 5 Ci/week of 99 Mo. The hypothetical design was carried out using a 10 MWt research reactor. The precursors used for the activation process were TeO2 for 131 I and MoO 3 for 99 Mo. The processes are performed in 3 hot cells used for target receipt, extraction, purification low specific activity (LSA) generator introduction, and packaging. A fourth hotcell is used for waste processing. The hot cell processing area takes up a footprint of 15.4 m 2 with the total footprint of the facility, including space for administrative offices, non-rad labs, quality assurance, and radiation buffer areas set at 763 m 2 . Waste is produced at a weekly rate of 257.8 g low activity solid waste and 8032.7 mL of low activity liquid waste, 8032 mL of which is water. This baseline hypothetical facility for production of medical isotopes via activation was then compared and contrasted to the hypothetical facility for production of medical isotopes via fission products to show the differences in approach for the two production modes. The two production modes had several highlighted differences including the overall facility and hot cell layout, the type and amount of waste produced by the respective facilities, and economic factors impacting production mode. Finally, a decision tree for which production mode might be more beneficial for an entrant into medical isotope production was developed based on the differences examined and the desired output of medical isotopes desired by the entrant.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Machine Learning–Guided Boolean Matrix Inference for Real-Time O-RAN Conflict Detection

Open Radio Access Networks (O-RAN) are emerging, software-driven cellular architectures that promote flexibility by enabling components from different vendors to interoperate. Multiple control applications called xApps can independently adjust network parameters in near real time, often without awareness of each other's actions. This creates a system highly prone to unintended conflicts and performance degradation due to the inherent complexity of such openness. To model such systems and ultimately prevent or mitigate xApp conflicts, it is essential to understand the dynamic relationships between xApps (A), the control parameters they adjust (P), and the resulting KPI responses (K). While the mappings from A to P and from K to A can often be derived from xApp specifications, the relationship from P to K is typically hidden within the system’s dynamics and must be inferred from observed data. We propose a novel data-driven Boolean inference framework that uncovers the hidden P?K dependencies using machine learning and interpretable rule induction. Continuous parameters and KPIs are first binarized using decision tree classifiers, and a binary influence matrix L is then inferred by solving Boolean matrix equations over time. This compact representation improves interpretability and enables real-time tracking of dynamically evolving parameter-KPI dependencies. We demonstrate the effectiveness of our method in a realistic mobile handover scenario, where it accurately recovers the underlying logic and enables proactive conflict detection.

42 - ENGINEERING↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Prediction of Dielectric Constant in Series of Polymers by Quantitative Structure-Property Relationship (QSPR)

This work is devoted to the investigation of dielectric permittivity which is influenced by electronic, ionic, and dipolar polarization mechanisms, contributing to the material’s capacity to store electrical energy. In this study, an extended dataset of 86 polymers was analyzed, and two quantitative structure–property relationship (QSPR) models were developed to predict dielectric permittivity. From an initial set of 1273 descriptors, the most relevant ones were selected using a genetic algorithm, and machine learning models were built using the Gradient Boosting Regressor (GBR). In contrast to Multiple Linear Regression (MLR)- and Partial Least Squares (PLS)-based models, the gradient boosting models excel in handling nonlinear relationships and multicollinearity, iteratively optimizing decision trees to improve accuracy without overfitting. The developed GBR models showed high R2 coefficients of 0.938 and 0.822, for the training and test sets, respectively. An Accumulated Local Effect (ALE) technique was applied to assess the relationship between the selected descriptors—eight for the GB_A model and six for the GB_B model, and their impact on target property. ALE analysis revealed that descriptors such as TDB09m had a strong positive effect on permittivity, while MLOGP2 showed a negative effect. These results highlight the effectiveness of the GBR approach in predicting the dielectric properties of polymers, offering improved accuracy and interpretability.

Ascencio-Medina, Estefania↗

Distinguishing Orbiting and Infalling Dark Matter Particles with Machine Learning

Dark matter halos are typically defined as spheres that enclose some overdensity, but these sharp, somewhat arbitrary boundaries introduce nonphysical artifacts such as backsplash halos, pseudo-volution, and an incomplete accounting of halo mass. A more physically motivated alternative is to define halos as the collection of particles that are physically orbiting within their potential well. However, existing methods to classify particles as orbiting or infalling suffer from trade-offs between accuracy, computational cost, and generalizability across cosmologies. We present an efficient, yet accurate, supervised machine learning approach using decision trees. The classification is based on only the particle radii and velocities at two epochs. Compared to detailed analysis of particle trajectories, we find that our model matches the classification of 97% of particles. Consequently, we are able to quickly and accurately reproduce the density profiles of the orbiting and infalling components out to many virial radii. We demonstrate that our model generalizes to a significantly different cosmology that lies outside the training data set. We make publicly available both our final model and the code to train similar models.

79 ASTRONOMY AND ASTROPHYSICS↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

MicroBooNE investigations on the photon interpretation of the MiniBooNE low energy excess

The MicroBooNE experiment is a liquid argon time projection chamber with 85-ton active volume at Fermilab, operated from 2015 to 2020 to collect neutrino data from Fermilab's Booster Neutrino Beam. One of MicroBooNE's physics goals is to investigate possible explanations of the low-energy excess observed by the MiniBooNE experiment in $\nu_{\mu}\rightarrow \nu_{e}$ neutrino oscillation measurements. MicroBooNE has performed searches to test hypothetical interpretations of the MiniBooNE low-energy excess, including the underestimation of the photon background or instrinic $\nu_{e}$ background. This thesis presents MicroBooNE's searches for two neutral current (NC) single-photon production processes that contribute to the photon background of the MiniBooNE measurement: NC $\Delta$ resonance production followed by $\Delta$ radiative decay: $\Delta \rightarrow N\gamma$, and NC coherent single-photon production. Both searches take advantage of boosted decision trees to yield efficient background rejection, and a high-statistic NC $\pi^0$ measurement to constrain dominant background, and make use of MicroBooNE's first three years of data. The NC $\Delta \rightarrow N\gamma$ measurement yielded a bound on the $\Delta$ radiative decay process at 2.3 times the predicted nominal rate at 90\% confidence level(C.L.), disfavoring a candidate photon interpretation of the MiniBooNE low-energy excess as a factor of 3.18 times the nominal NC Δ radiative decay rate at the 94.8\% C.L. The NC coherent single photon measurement leads to the world's first experimental limit on the cross-section of this process below 1 GeV, of $1.49 \times 10^{-41} \text{cm}^2$ at 90\% C.L., corresponding to 24.0 times the nominal prediction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-channel, multi-template event reconstruction for SuperCDMS data using machine learning

SuperCDMS SNOLAB uses kilogram-scale germanium and silicon detectors to search for dark matter. Each detector has Transition Edge Sensors (TESs) patterned on the top and bottom faces of a large crystal substrate, with the TESs electrically grouped into six phonon readout channels per face. Noise correlations are expected among a detector's readout channels, in part because the channels and their readout electronics are located in close proximity to one another. Moreover, owing to the large size of the detectors, energy deposits can produce vastly different phonon propagation patterns depending on their location in the substrate, resulting in a strong position dependence in the readout-channel pulse shapes. Both of these effects can degrade the energy resolution and consequently diminish the dark matter search sensitivity of the experiment if not accounted for properly. We present a new algorithm for pulse reconstruction, mathematically formulated to take into account correlated noise and pulse shape variations. This new algorithm fits N readout channels with a superposition of M pulse templates simultaneously - hence termed the N$\times$M filter. We describe a method to derive the pulse templates using principal component analysis (PCA) and to extract energy and position information using a gradient boosted decision tree (GBDT). We show that these new N$\times$M and GBDT analysis tools can reduce the impact from correlated noise sources while improving the reconstructed energy resolution for simulated mono-energetic events by more than a factor of three and for the 71Ge K-shell electron-capture peak recoils measured in a previous version of SuperCDMS called CDMSlite to $<$ 50 eV from the previously published value of $\sim$100 eV. These results lay the groundwork for position reconstruction in SuperCDMS with the N$\times$M outputs.

Albakry, M. F. [British Columbia U.; TRIUMF]↗

Searching for Neutrino Tridents in the NOvA Near Detector

This dissertation presents a search for neutrino trident production in the NOvA near detector through the coherent ``dimuon" channel: $\nu_\mu +\hspace{1pt}\text{X} \rightarrow \nu_\mu + \mu^- + \mu^+ +\hspace{1pt}\text{X}$. Trident production is a rare, purely electroweak process with sensitivity to physics beyond the Standard Model. The theoretical background, motivation for studying the process, and previous experimental measurements are reviewed. The analysis uses data collected by the NOvA near detector (ND) from Fermilab's Neutrinos at the Main Injector (NuMI) beam between November 2014 and February 2024, corresponding to an exposure of $25.5\times 10^{20}$ protons on target. The ND is a segmented tracking calorimeter located 800~m from the beam target, receiving neutrinos with a mean energy of 2~GeV. A multi-pass background reduction strategy is implemented, including the development of a novel dimuon-specific tracking technique. Trident candidates are identified using a boost ed decision tree classifier trained on simulated signal and background events. Limited background Monte Carlo statistics necessitate the use of functional fits to sideband data, which are extrapolated to estimate backgrounds in the signal region. The unblinded data contain 9 trident-like events, with an estimated background of 5.66 $\pm$ 5.15 events. This yields a best fit estimate of 3.34 tridents compared to the Standard Model prediction of 4.66. A profiled Feldman-Cousins method is used to determine a 90\% confidence interval of [0,9.1] on the number of signal events, corresponding to an upper limit of 1.95$\times$ the Standard Model prediction. This result represents the lowest energy search for trident events to date, and the first experimental contribution to the process in 27 years.

Bowles, Reed Scott [Indiana U.]↗

Baryon Number Violation Search

Understanding the fundamental forces and symmetries of nature has long been a central goal of particle physics. While the Standard Model (SM) provides a successful framework, it does not guarantee the conservation of baryon number B or lepton number L, thus motivating searches for their violation. Proton decay, a fundamental process violating B, has been at the forefront of experimental searches for decades.The discovery of the weak neutral current in 1973 unified the electromagnetic and weak forces and inspired the creation of Grand Unified Theories (GUTs) that also unify the strong force. In 1974, the first-ever GUT, proposed by Georgi and Glashow, naturally predicted proton decay. This prediction led to an experimental push to validate these theories, and a large underground detector boom was born. Initially designed for proton decay searches, these detectors later proved invaluable to neutrino physics.Although no evidence for proton decay has yet been observed, next-generation large detectors, such as the Deep Underground Neutrino Experiment (DUNE), offer the opportunity to improve on current experimental limits. Utilizing its Liquid Argon Time Projection Chamber (LArTPC) technology, DUNE is positioned to probe rare processes such as proton decay with increased sensitivity.This thesis presents a sensitivity study for the dominant proton decay mode predicted by Supersymmetric GUTs, p → K+ν, utilizing machine learning approaches. Two methods are explored in this thesis: a Boosted Decision Tree (BDT) analysis and a Graphical Neural Network (GNN) analysis with NuGraph. A lifetime limit of 5.36 ± 0.69 × 1033 years for 400 kt-yrs is found using the BDT, while the GNN achieves a lifetime limit of 6.19±1.26×1033 years for 400-kt-yrs. The NuGraph result offers better sensitivity compared to the current limit set by Super-Kamiokande of 5.90 × 1033 while the BDT result offers a slightly lower sensitivity.Additionally, this thesis discusses cross-section work, a first-ever foray into proton decay and atmospheric neutrinos in a vertical drift (VD) DUNE detector, and extensive hardware contributions to the DUNE Far Detector (FD) 1 Module-0, ProtoDUNE-2, which serves as a testbed for the final detector design and installation.

Stokes, Tyler D. [Louisiana State U.] (ORCID:00000↗

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.

Jha, Piyush [Georgia Tech., Atlanta; Georgia Tech]↗

Weakly supervised anomaly detection with event-level variables

We introduce a new topology for weakly supervised anomaly detection searches, diobject plus X. In this topology, one looks for a resonance decaying to two standard model particles produced in association with other anomalous event activity (X). This additional activity is used for classification. We demonstrate how anomaly detection techniques which have been developed for dijet searches focusing on jet substructure anomalies can be applied to event-level anomaly detection in this topology. To robustly capture event-level features of multiparticle kinematics, we employ new physically motivated variables derived from the geometric structure of a collision’s phase space manifold. As a proof of concept, we explore the application of this approach to several benchmark signals in the di-𝜏 and di-𝜇 plus X final states. We demonstrate that our anomaly detection approach can reach discovery-level significances for signals that would be missed in a conventional bump-hunt approach.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Feature Engineering and Ensemble Methods for Imbalanced ICS Intrusion Detection: Pipeline Audit and Constrained Evaluation

Industries are becoming increasingly connected and are more vulnerable to cyberattacks due to the widened attack surface. Industrial Control Systems (ICS) are among the most critical sectors that malicious actors can target, as such attacks can cause significant operational disruption and physical damage. It is imperative to detect such attacks as early as possible. This paper evaluates constraint-conditioned optimistic performance estimates for traditional ML models in ICS intrusion detection (i.e., estimates obtained under contiguous, non-shuffled temporal evaluation without test-set alteration, but with pre-split feature engineering that may introduce temporal leakage, due to dataset constraints). Our findings are threefold. First, we quantify how iterative feature engineering affects tree-based ensemble performance and examine how pipeline decisions (split strategy, sampling scope, and cleaning policy) can inflate or reduce reported IDS results under constraint-bound evaluation. Second, we compare intrinsic class-imbalance handling across ensemble models. Third, under our current pipeline constraints (including pre-split feature engineering), CatBoost achieves the best performance on Water Storage Tank (accuracy: 0.9831, class-1 F1: 0.9682), while Light- GBM achieves the best performance on Gas Pipeline (accuracy: 0.9618, class-1 F1: 0.9086).

97 MATHEMATICS AND COMPUTING↗

A Behavior Tree Approach for Battery-Aware Inspection of Large Structures Using Drones

Electric multi-rotor drones have been used to inspect several structures, including large buildings and dams. In these inspections, energy consumption is a concern. To prevent the drone from running out of battery, commercial drones usually come back to their home position when the battery level reaches a minimum threshold. The pilots then need to replace the battery and use their own experience to restart the inspection mission approximately from where it ended before the drone returned home. Instead of relying on the human operator, in this paper, we automate this process using behavior trees, which is an effective way to perform autonomous mission control and supervision. By integrating battery management strategies into a behavior tree framework, this paper demonstrates the drone’s adaptive and resilient decision-making when confronted with limited power constraints. We implemented our methodology using a commercial drone and tested the proposed ideas in a photogrammetry-based inspection task.

42 ENGINEERING↗