Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boosting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Rapid Spaceborne Mapping of Wildfire Retardant Drops for Active Wildfire Management

Aerial application of fire retardant is a critical tool for managing wildland fire spread. Retardant applications are carefully planned to maximize fire line effectiveness, improve firefighter safety, protect high-value resources and assets, and limit environmental impact. However, topography, wind, visibility, and aircraft orientation can lead to differences between planned drop locations and the actual placement of the retardant. Information on the precise placement and areal extent of the dropped retardant can provide wildland fire managers with key information to (1) adaptively manage event resources, (2) assess the effectiveness of retardant slowing or stopping fire spread, (3) document location in relation to ecologically sensitive areas; and perform or validate cost-accounting for drop services. This study uses Sentinel-2 satellite data and commonly used machine learning classifiers to test an automated approach for detecting and mapping retardant application. We show that a multiclass model (retardant, burned, unburned, and cloud artifact classes) outperforms a single-class retardant model and that image differencing (post-application minus pre-application) outperforms single-image models. Compared to the random forest and support vector machine, the gradient boosting model performed the best with an overall accuracy of 0.88 and an F1 Score of 0.76 for fire retardant, though results were comparable for all three models. Our approach maps the full areal extent of the dropped retardant within minutes of image availability, rather than linear representations currently mapped by aerial GPS surveys. The development of this capability allows for the rapid assessment of retardant effectiveness and documentation of placement in relation to sensitive environments.

54 ENVIRONMENTAL SCIENCES↗

Can Simple Machine Learning Tools Extend and Improve Temperature-Based Methods to Infer Streambed Flux?

Temperature-based methods have been developed to infer 1D vertical exchange flux between a stream and the subsurface. Current analyses rely on fitting physically based analytical and numerical models to temperature time series measured at multiple depths to infer daily average flux. These methods have seen wide use in hydrologic science despite strong simplifying assumptions including a lack of consideration of model structural error or the impacts of multidimensional flow or the impacts of transient streambed hydraulic properties. We performed a “perfect-model experiment” investigation to examine whether regression trees, with and without gradient boosting, can extract sufficient information from model-generated subsurface temperature time series, with and without added measurement error, to infer the corresponding exchange flux time series at the streambed surface. Using model-generated, synthetic data allowed us to assess the basic limitations to the use of machine learning; further examination of real data is only warranted if the method can be shown to perform well under these ideal conditions. We also examined whether the inherent feature importance analyses of tree-based machine learning methods can be used to optimize monitoring networks for exchange flux inference.

54 ENVIRONMENTAL SCIENCES↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cloud drop number concentrations over the western North Atlantic Ocean: seasonal cycle, aerosol interrelationships, and other influential factors

Cloud drop number concentrations (N d ) over the western North Atlantic Ocean (WNAO) are generally highest during the winter (DJF) and lowest in summer (JJA), in contrast to aerosol proxy variables (aerosoloptical depth, aerosol index, surface aerosol mass concentrations, surface cloud condensation nuclei (CCN) concentrations) that generally peak inspring (MAM) and JJA with minima in DJF. Using aircraft, satellite remote sensing, ground-based in situ measurement data, and reanalysis data, we characterize factors explaining the divergent seasonal cycles and furthermore probe into factors influencing N d on seasonal timescales. The results can be summarized well by features most pronounced in DJF, including features associated with cold-air outbreak (CAO) conditions such as enhanced values of CAO index, planetary boundary layer height (PBLH),low-level liquid cloud fraction, and cloud-top height, in addition to winds aligned with continental outflow. Data sorted into high- and low-N d days in each season, especially in DJF, revealed that all of these conditions were enhanced on the high-N d days, including reduced sea level pressure and stronger wind speeds. Although aerosols may be more abundant in MAM and JJA, the conditions needed to activate those particles into cloud droplets are weaker than in colder months, which is demonstrated by calculations of the strongest (weakest) aerosol indirect effects in DJF (JJA) based on comparing N d to perturbations in four different aerosol proxy variables (total and sulfate aerosol optical depth, aerosol index, surface mass concentration of sulfate). We used three machine learning models and up to 14 input variables to infer about most influential factors related to N d for DJF and JJA, with the best performance obtained with gradient-boosted regression tree (GBRT) analysis. The model results indicated that cloud fraction was the most important input variable, followed by some combination (depending on season) of CAO index and surface mass concentrations of sulfate and organic carbon. Future work is recommended to further understand aspects uncovered here such as impacts of free tropospheric aerosol entrainment on clouds, degree of boundary layer coupling, wet scavenging, and giant CCN effects on aerosol–N d relationships, updraft velocity, and vertical structure of cloud properties such as adiabaticity that impact the satellite estimation of N d .

54 ENVIRONMENTAL SCIENCES↗

Best estimate of the planetary boundary layer height from multiple remote sensing measurements

Remote sensing measurements have been widely used to estimate the planetary boundary layer height (PBLHT). Each remote sensing approach offers unique strengths and faces different limitations. In this study, we use machine learning (ML) methods to produce a best-estimate PBLHT (PBLHT-BE-ML) by integrating four PBLHT estimates derived from remote sensing measurements at the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) observatory. Three ML models – random forest (RF) classifier, RF regressor, and light gradient-boosting machine (LightGBM) – were trained on a dataset from 2017 to 2023 that included radiosonde, various remote sensing PBLHT estimates, and atmospheric meteorological conditions. Evaluations indicated that PBLHT-BE-ML from all three models improved alignment with the PBLHT derived from radiosonde data (PBLHT-SONDE), with LightGBM demonstrating the highest accuracy under both stable and unstable boundary layer conditions. Feature analysis revealed that the most influential input features at the SGP site were the PBLHT estimates derived from (a) potential temperature profiles retrieved using Raman lidar (RL) and atmospheric emitted radiance interferometer (AERI) measurements (PBLHT-THERMO), (b) vertical velocity variance profiles from Doppler lidar (PBLHT-DL), and (c) aerosol backscatter profiles from micropulse lidar (PBLHT-MPL). The trained models were then used to predict PBLHT-BE-ML at a temporal resolution of 10 min, effectively capturing the diurnal evolution of PBLHT and its significant seasonal variations, with the largest diurnal variation observed over summer at the SGP site. We applied these trained models to data from the ARM Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) field campaign (EPC), where the PBLHT-BE-ML, particularly with the LightGBM model, demonstrated improved accuracy against PBLHT-SONDE. Analyses of model performance at both the SGP and EPC sites suggest that expanding the training dataset to include various surface types, such as ocean and ice-covered areas, could further enhance ML model performance for PBLHT estimation across varied geographic regions.

Zhang, Damao [Pacific Northwest National Laborator↗

Median bed-material sediment particle size across rivers in the contiguous US

Abstract. Bed-material sediment particle size data, particularly the median sediment particle size (D50), are critical for understanding and modeling riverine sediment transport. However, sediment particle size observations are primarily available at individual sites. Large-scale modeling and assessment of riverine sediment transport are limited by the lack of continuous regional maps of bed-material sediment particle size. We hence present a map of D50 over the contiguous US in a vector format that corresponds to approximately 2.7 million river segments (i.e., flowlines) in the National Hydrography Dataset Plus (NHDPlus) dataset. We develop the map in four steps: (1) collect and process the observed D50 data from 2577 U.S. Geological Survey stations or U.S. Army Corps of Engineers sampling locations; (2) collocate these data with the NHDPlus flowlines based on their geographic locations, resulting in 1691 flowlines with collocated D50 values; (3) develop a predictive model using the eXtreme Gradient Boosting (XGBoost) machine learning method based on the observed D50 data and the corresponding climate, hydrology, geology, and other attributes retrieved from the NHDPlus dataset; and (4) estimate the D50 values for flowlines without observations using the XGBoost predictive model. We expect this map to be useful for various purposes, such as research in large-scale river sediment transport using model- and data-driven approaches, teaching environmental and earth system sciences, planning and managing floodplain zones, etc. The map is available at https://doi.org/10.5281/zenodo.4921987 (Li et al., 2021a).

54 ENVIRONMENTAL SCIENCES↗

ELM2.1-XGBfire1.0: improving wildfire prediction by integrating a machine learning fire model in a land surface model

Wildfires have shown increasing trends in both frequency and severity across the contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth system models (ESMs). Alternatively, fire models based on machine learning (ML), which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ELM2.1-XGBFire1.0) that integrates an eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran–C–Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001–2019, the ELM2.1-XGBFire1.0 outperforms process-based fire models in terms of spatial distribution and seasonal variations. The ELM2.1-XGBFire1.0 has proven to be a new tool for studying vegetation–fire interactions and, more importantly, enables seamless exploration of climate–fire feedback, working as an active component of E3SM.

54 ENVIRONMENTAL SCIENCES↗

Failure prediction and estimation of failure parameters

Machine-learning methods and apparatus are disclosed to determine frictional state or other parameters in an earthquake zone or other failing medium, using acoustic emission, seismic waves, or other detectable indicators of microscopic processes. Predictions of future failures are demonstrated in different regimes. A classifier is trained using time series of acoustic emission data along with historic data of frictional state or failure events. In disclosed examples, random forests and gradient boost trees are used, and grid-search or EGO procedures are used for hyperparameter tuning. Once trained, the classifier can be applied to testing or live data in order to assess a frictional state, assess seismic hazard, or make predictions regarding a future failure event. The technology has been developed in a double direct shear apparatus, but can be widely applied to seismic faults, other terrestrial failures, or failures in man-made structures. Variations are disclosed.

Johnson, Paul Allan↗

Automatic DDoS Attack Detection on SDNs: Preprint

Denial of Service (DoS) and Distributed Denial of Service (DDoS) attacks pose a serious threat to computing networks - especially to critical systems within the U.S. electrical grid. As attack mechanisms have increased in complexity and variety, more sophisticated detection mechanisms have become necessary to ensure network security. This paper explores the use of artificial intelligence to automate the process of detection and mitigation of DoS and DDoS attacks within the framework of Software-Defined Networking (SDN), to a high degree. Machine learning algorithms are trained to recognize DoS and DDoS attacks and are deployed in real-time to mitigate malicious network traffic. The results show a well-tuned gradient-boosted decision tree detecting DoS and DDoS attacks, as well as initial successful mitigation of attacks within an SDN framework.

cyber detection↗

A Methodology for Simulating Supercritical CO2 Heat Transfer Experiments Using Machine Learning Models

To support the growth of supercritical carbon dioxide (sCO2) power cycles in the energy industry, this study seeks to train a machine learning model to mirror experimental data to predict new heat transfer data. To do this experimental data was amassed, one preliminary set comprised of 16 test results, and an expanded version comprised of 38 test results. With the goal of predicting experimental apparatus temperatures and pressures, several iterations of models were tested investigating the impact of model hyper-parameters, data inclusion, and data pre-processing on model performance. A total of 15 variations cumulatively of Gaussian Process Regressors, Gradient Boosting Regressors, and Multi-Layer Perceptrons were trained and validated on the preliminary set, and the best algorithm of each class was re-trained on the expanded set. These were compared based on test/train R^2 , test/train mean absolute error (MAE), and validation MAE, to identify the successfulness of these models. It was shown temperatures could be predicted within just a few degrees, showing the potential of this approach. Future research has been identified with approaches to improve pressure and temperature predictions going forward.

Grabowski, Owen↗

Event-Based Energy Impact Tracking and Forecasting with Limited Measurements for Rooftop Units

Packaged air conditioning units and heat pumps, also known as rooftop units (RTUs), are responsible for almost 133 billion kWh of electricity usage annually on site for space cooling U.S. commercial buildings. In addition, the use of heat pumps is a trend we expect to accelerate as buildings transition from fossil fuel-based heating to electricity as a key step for decarbonizing the U.S. commercial buildings sector. However, the operation conditions and energy use of RTUs and heat pumps are usually not well monitored as they are not commonly integrated with building automation systems and lack exposed sensing and control points. To fill this gap, this paper proposes a framework for tracking and forecasting energy impacts resulting from degradation of performance and improved performance for unit servicing using limited data. The proposed framework makes use of a constrained dataset, specifically measurements of the outdoor air temperature and the power demand of individual RTUs, to track and forecast changes in energy use associated with changes in performance over various temporal horizons ranging from days to weeks. Following the detection of an RTU fault, performance degradation, or performance improvement, the framework employs a prediction model to assess the cumulative energy impact. We demonstrate the effectiveness of the method with field-collected data for servicing and degradation examples and compare the predicting accuracy of Gradient Boosting Decision Tree (GBDT) Regression models to Support Vector Regression and Linear Regression models. The results show that GBDT achieved the best accuracy for time-series validation datasets for the servicing and degradation cases, and the prediction model was able to track the cumulative energy impacts of events. The proposed framework can inform building owners of the cumulative change in energy usage of RTUs associated with performance degradation, performance improvement, or a fault.

packaged air conditioners, packaged heat pumps, ro↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif↗

Multi-channel, multi-template event reconstruction for SuperCDMS data using machine learning

SuperCDMS SNOLAB uses kilogram-scale germanium and silicon detectors to search for dark matter. Each detector has Transition Edge Sensors (TESs) patterned on the top and bottom faces of a large crystal substrate, with the TESs electrically grouped into six phonon readout channels per face. Noise correlations are expected among a detector's readout channels, in part because the channels and their readout electronics are located in close proximity to one another. Moreover, owing to the large size of the detectors, energy deposits can produce vastly different phonon propagation patterns depending on their location in the substrate, resulting in a strong position dependence in the readout-channel pulse shapes. Both of these effects can degrade the energy resolution and consequently diminish the dark matter search sensitivity of the experiment if not accounted for properly. We present a new algorithm for pulse reconstruction, mathematically formulated to take into account correlated noise and pulse shape variations. This new algorithm fits N readout channels with a superposition of M pulse templates simultaneously - hence termed the N$\times$M filter. We describe a method to derive the pulse templates using principal component analysis (PCA) and to extract energy and position information using a gradient boosted decision tree (GBDT). We show that these new N$\times$M and GBDT analysis tools can reduce the impact from correlated noise sources while improving the reconstructed energy resolution for simulated mono-energetic events by more than a factor of three and for the 71Ge K-shell electron-capture peak recoils measured in a previous version of SuperCDMS called CDMSlite to $<$ 50 eV from the previously published value of $\sim$100 eV. These results lay the groundwork for position reconstruction in SuperCDMS with the N$\times$M outputs.

Albakry, M. F. [British Columbia U.; TRIUMF]↗

Measurements of Beam Spin Asymmetries in p+p0 and p´p0 Dihadron Production at CLAS12

Semi-Inclusive Deep Inelastic Scattering (SIDIS) is a powerful experimental tool for studying the internal structure and dynamics of the proton, revealing how quarks and gluons are distributed and interact within it. SIDIS describes a process where an elec tron scatters off one of the constituent quarks within the proton, causing it to undergo hadronization, creating multiple hadrons in the final state. Through factorization, the full process can be split into probabilistic components: one which describes the internal structure of the proton using Parton Distribution Functions (PDFs), and another which describes the hadronization process using Fragmentation Functions (FFs). These functions are non-perturbative quantities of Quantum Chromodynamics (QCD), meaning they cannot be calculated directly from first principles and must instead be extracted from experimental measurements. Acommon approach for accessing PDFs and FFs using SIDIS is to measure asymmetries. In this context, asymmetries correspond to subtle differences in the angular distribution of outgoing particles that arise when the spin orientation of the incoming beam or target is reversed. Because many of these effects only appear when spin is involved, they isolate specific, nuanced properties of the proton’s spin-structure that are otherwise hidden in spin averaged measurements. In practice, they show up as specific azimuthal modulations (e.g., sin ¿R, sin(¿h ´ ¿R)), whose amplitudes isolate convolutions of PDFs and FFs at leading and subleading twist. Non-zero asymmetries of these angular distributions can be traced back to unique combinations of PDFs and FFs, offering a way to probe them directly. In this work, we measure SIDIS by analyzing high energy electron-proton scattering events using the CLAS12 detector at Jefferson Lab. This study focuses on subset of SIDIS referred to as dihadron SIDIS, where pairs of hadrons — here p+p0 and p´p0 — are observed. We analyzed these dihadrons using detector data collected during Fall 2018 and Spring 2019, where longitudinally polarized electrons from the CEBAF accelerator were incident on a liquid hydrogen target. A photon classifier using a Gradient Boosted Trees (GBTs) architecture was trained using Monte Carlo simulations to reduce the amount of iv false combinatorial background p0’s. When deployed on experimental data, the model in creases our dihadron statistics by up to five-fold compared to previous CLAS12 p0 analyses. This work reports the first measurements of beam spin asymmetries for p+p0 and p´p0 dihadron production in SIDIS. The measured asymmetries offer new insights to the spin-dependent structure and dynamics within the proton, as well as the spin-dependent properties of quark fragmentation. Non-zero twist-3 sin¿R amplitudes are observed, pro viding sensitivity to the subleading twist PDF e(x). The PDF e(x) encodes quark-gluon correlations within the proton — a property that is otherwise inaccessible at leading twist. Additionally, this work measured significant twist-2 modulations carried by sin(¿h ´ ¿R) and sin(2¿h ´2¿R), providing experimental access to the helicity dihadron fragmentation function (DiFF) GK 1 . Because there is no equivalent quark helicity-dependent FF in single pion SIDIS, the DiFF GK 1 offers a unique lens into novel spin-dependent fragmentation. For instance, the twist-2 modulations observed in this study are enhanced by vector mesons created during fragmentation — a behavior predicted by phenomenological models. This study broadens our understanding of dihadron fragmentation, revealing new details about the flavor and charge dependence of hadronization.

Matousek, Gregory [Duke Univ., Durham, NC (United ↗

Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification

Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.

Sun, Jingbo↗

Learning-based demand-supply-coupled charging station location problem for electric vehicle demand management

We present a learning-based, demand-supply-coupled optimization model for the charging station location problem (CSLP), aiming to integrate the concept of electric vehicle (EV) charging demand management into the planning of charging infrastructures. In stage one, a gradient boosting-based learning model is developed to predict the charging demand of a charging station based on 15 defined features. Next, in stage two, a demand–supply-coupled CSLP model is developed to optimize the total charging usage rates of both existing and newly selected charging stations. We design a gradient-based stochastic spatial search algorithm to solve the proposed model. A case study with 6-year charging event data from Kansas City Missouri is performed. Results show that the proposed method can generate satisfactory charging demand predictions, and can increase charging usage rates by 14%, outperforming two benchmark approaches. Furthermore, the results of this research are poised to guide agencies in identifying optimal locations for new charging stations.

33 ADVANCED PROPULSION SYSTEMS↗