Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boosting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Fault location in High Voltage Multi-terminal dc Networks Using Ensemble Learning

Precise location of faults for large distance power transmission networks is essential for faster repair and restoration process. High Voltage direct current (HVdc) networks using modular multi-level converter (MMC) technology has found its prominence for interconnected multi-terminal networks. This allows for large distance bulk power transmission at lower costs. However, they cope with the challenge of dc faults. Fast and efficient methods to isolate the network under dc faults have been widely studied and investigated. After successful isolation, it is essential to precisely locate the fault. The post-fault voltage and current signatures are a function of multiple factors and thus accurately locating faults on a multi-terminal network is challenging. In this paper, we discuss a novel data-driven ensemble learning based approach for accurate fault location. Here we utilize the eXtreme Gradient Boosting (XGB) method for accurate fault location. The sensitivity of the proposed algorithm to measurement noise, fault location, resistance and current limiting inductance are performed on a radial three-terminal MTdc network designed in Power System Computer Aided Design (PSCAD)/Electromagnetic Transients including dc (EMTdc).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Estimation of the Surface Fluxes for Heat and Momentum in Unstable Conditions with Machine Learning and Similarity Approaches for the LAFE Data Set

Abstract Measurements of three flux towers operated during the land atmosphere feedback experiment (LAFE) are used to investigate relationships between surface fluxes and variables of the land–atmosphere system. We study these relations by means of two machine learning (ML) techniques: multilayer perceptrons (MLP) and extreme gradient boosting (XGB). We compare their flux derivation performance with Monin–Obukhov similarity theory (MOST) and a similarity relationship using the bulk Richardson number (BRN). The ML approaches outperform MOST and BRN. Best agreement with the observations is achieved for the friction velocity. For the sensible heat flux and even more so for the latent heat flux, MOST and BRN deviate from the observations while MLP and XGB yield more accurate predictions. Using MOST and BRN for latent heat flux, the root mean square errors (RMSE) are 107 Wm $$^{-2}$$ - 2 and 121 Wm $$^{-2}$$ - 2 , respectively, as well as the intercepts of the regression lines are $$\approx 110$$ ≈ 110 Wm $$^{-2}$$ - 2 . For the ML methods, the RMSEs reduce to 31 Wm $$^{-2}$$ - 2 for MLP and 33 Wm $$^{-2}$$ - 2 for XGB as well as the intercepts to just 4 Wm $$^{-2}$$ - 2 for MLP and $$-1$$ - 1 Wm $$^{-2}$$ - 2 for XGB with slopes of the regression lines close to 1, respectively. These results indicate significant deficiencies of MOST and BRN, particularly for the derivation of the latent heat flux. In fact, in contrast to the established theories, feature importance weighting demonstrates that the ML methods base their improved derivations on net radiation, the incoming and outgoing shortwave radiations, the air temperature gradient, and the available water contents, but not on the water vapor gradient. The results imply that further studies of surface fluxes and other turbulent variables with ML techniques provide great promise for deriving advanced flux parameterizations and their implementation in land–atmosphere system models.

54 ENVIRONMENTAL SCIENCES↗

Micro-structural features and material properties impact on adhesive metal joints via computational modeling and machine learning

The quality of structural bonding in practical applications depends on various factors arising from materials, pre-processing conditions, and manufacturing. Understanding how these factors influence bonding performance and determining their relative importance are of significant interest. Thus, this study evaluates the effects of microstructural features and material properties on the structural strength of adhesively-bonded metal joints at the submillimeter scale, utilizing a combination of Finite Element Modeling (FEM) and Machine Learning (ML) with Gradient Boosting Regression (GBR). The microstructural features include adhesive thickness, internal voids within the adhesive, adherend-adhesive interfacial voids, void size and volume fraction, and surface roughness. The material properties include the constitutive behavior of the adhesive, as well as the adherend-adhesive interfacial strength and fracture energy. The changes in structural strength and morphologies of the bonded metal structures with respect to different microstructural features and material properties were clarified by FEM. By further leveraging ML-GBR, the sequence of importance of these factors affecting bonding performance across various scenarios was summarized. This work provides valuable insights into the development of improved structural bonding for adhesive joints in industries such as automotive , aerospace, and beyond.

36 MATERIALS SCIENCE↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning Analysis of Hydrologic Exchange Flows and Transit Time Distributions in a Large Regulated River

Hydrologic exchange between river channels and adjacent subsurface environments is a key process that influences water quality and ecosystem function in river corridors. High-resolution numerical models were often used to resolve the spatial and temporal variations of exchange flows, which are computationally expensive. In this study, we adopt Random Forest (RF) and Extreme Gradient Boosting (XGB) approaches for deriving reduced order models of hydrologic exchange flows and associated transit time distributions, with integrated field observations (e.g., bathymetry) and hydrodynamic simulation data (e.g., river velocity, depth). The setup allows an improved understanding of the influences of various physical, spatial, and temporal factors on the hydrologic exchange flows and transit times. The predictors also contain those derived using hybrid clustering, leveraging our previous work on river corridor system hydromorphic classification. The machine learning-based predictive models are developed and validated along the Columbia River Corridor, and the results show that the top parameters are the thickness of the top geological formation layer, the flow regime, river velocity, and river depth; the RF and XGB models can achieve 70% to 80% accuracy and therefore are effective alternatives to the computational demanding numerical models of exchange flows and transit time distributions. Each machine learning model with its favorable configuration and setup have been evaluated. The transferability of the models to other river reaches and larger scales, which mostly depends on data availability, is also discussed.

97 MATHEMATICS AND COMPUTING↗

Modeling household online shopping demand in the U.S.: a machine learning approach and comparative investigation between 2009 and 2017

Despite the rapid growth of online shopping and research interest in the relationship between online and in-store shopping, national-level modeling and investigation of the demand for online shopping with a prediction focus remain limited in the literature. Here, this paper differs from prior work and leverages two recent releases of the U.S. National Household Travel Survey (NHTS) data for 2009 and 2017 to develop machine learning (ML) models, specifically gradient boosting machine (GBM), for predicting household-level online shopping purchases. The NHTS data allow for not only conducting nationwide investigation but also at the level of households, which is more appropriate than at the individual level given the connected consumption and shopping needs of members in a household. We follow a systematic procedure for model development including employing Recursive Feature Elimination algorithm to select input variables (features) in order to reduce the risk of model overfitting and increase model explainability. Among several ML models, GBM is found to yield the best prediction accuracy. Extensive post-modeling investigation is conducted in a comparative manner between 2009 and 2017, including quantifying the importance of each input variable in predicting online shopping demand, and characterizing value-dependent relationships between demand and the input variables. In doing so, two latest advances in machine learning techniques, namely Shapley value-based feature importance and Accumulated Local Effects plots, are adopted to overcome inherent drawbacks of the popular techniques in current ML modeling. The modeling and investigation are performed at the national level, with a number of findings obtained. The models developed and insights gained can be used for online shopping-related freight demand generation and may also be considered for evaluating the potential impact of relevant policies on online shopping demand.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Utilization of Synthetic Near-Infrared Spectra via Generative Adversarial Network to Improve Wood Stiffness Prediction

Near-infrared (NIR) spectroscopy is widely used as a nondestructive evaluation (NDE) tool for predicting wood properties. When deploying NIR models, one faces challenges in ensuring representative training data, which large datasets can mitigate but often at a significant cost. Machine learning and deep learning NIR models are at an even greater disadvantage because they typically require higher sample sizes for training. In this study, NIR spectra were collected to predict the modulus of elasticity (MOE) of southern pine lumber (training set = 573 samples, testing set = 145 samples). To account for the limited size of the training data, this study employed a generative adversarial network (GAN) to generate synthetic NIR spectra. The training dataset was fed into a GAN to generate 313, 573, and 1000 synthetic spectra. The original and enhanced datasets were used to train artificial neural networks (ANNs), convolutional neural networks (CNNs), and light gradient boosting machines (LGBMs) for MOE prediction. Overall, results showed that data augmentation using GAN improved the coefficient of determination (R 2 ) by up to 7.02% and reduced the error of predictions by up to 4.29%. ANNs and CNNs benefited more from synthetic spectra than LGBMs, which only yielded slight improvement. All models showed optimal performance when 313 synthetic spectra were added to the original training data; further additions did not improve model performance because the quality of the datapoints generated by GAN beyond a certain threshold is poor, and one of the main reasons for this can be the size of the initial training data fed into the GAN. LGBMs showed superior performances than ANNs and CNNs on both the original and enhanced training datasets, which highlights the significance of selecting an appropriate machine learning or deep learning model for NIR spectral-data analysis. The results highlighted the positive impact of GAN on the predictive performance of models utilizing NIR spectroscopy as an NDE technique and monitoring tool for wood mechanical-property evaluation. Further studies should investigate the impact of the initial size of training data, the optimal number of generated synthetic spectra, and machine learning or deep learning models that could benefit more from data augmentation using GANs.

59 BASIC BIOLOGICAL SCIENCES↗

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Assessment of Machine Learning Wall Modeling Approaches for Large Eddy Simulation of Gas Turbine Film Cooling Flows: An a Priori Study

Here, in this work, a priori analysis of machine learning (ML) strategies is carried out with the goal of data-driven wall modeling for large eddy simulation (LES) of gas turbine film cooling flows. High-fidelity flow datasets are extracted from wall-resolved LES (WRLES) of flow over a flat plate interacting with the coolant flow supplied by a single row of 7-7-7 shaped cooling holes inclined at 30 degrees with the flat plate at different blowing ratios (BR). The WRLES are performed using the high-order Nek5000 spectral element computational fluid dynamics (CFD) solver. Light gradient boosting machine (LightGBM) is employed as the ML algorithm for the data-driven wall model. Parametric tests are conducted to systematically assess the influence of a wide range of input flow features (velocity components, velocity gradients, pressure gradients, and fluid properties) on the accuracy of ML wall model with respect to prediction of wall shear stress. In addition, the use of spatial stencil and time delay is also explored within the ML wall modeling framework. It is shown that features associated with gradients of the streamwise and spanwise velocity components have a major impact on the prediction fidelity of wall model, while the effect of gradients of wall-normal velocity component is found to be negligible. Moreover, adding flow feature information from an x-y-z spatial stencil significantly improves the ML model accuracy and generalizability compared to just using local flow features from the matching location. Overall, highest prediction accuracy is achieved when both spatial stencil and time delay features are incorporated within the data-driven wall modeling paradigm.

33 ADVANCED PROPULSION SYSTEMS↗

Tree-based algorithms for weakly supervised anomaly detection

Weakly supervised methods have emerged as a powerful tool for model-agnostic anomaly detection at the Large Hadron Collider (LHC). While these methods have shown remarkable performance on specific signatures such as dijet resonances, their application in a more model-agnostic manner requires dealing with a larger number of potentially noisy input features. In this paper, we show that using boosted decision trees as classifiers in weakly supervised anomaly detection gives superior performance compared to deep neural networks. Boosted decision trees are well known for their effectiveness in tabular data analysis. Our results show that they not only offer significantly faster training and evaluation times, but they are also robust to a large number of noisy input features. By using advanced gradient boosted decision trees in combination with ensembling techniques and an extended set of features, we significantly improve the performance of weakly supervised methods for anomaly detection at the LHC. This advance is a crucial step toward a more model-agnostic search for new physics. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Investigating boosted decision trees as a guide for inertial confinement fusion design

Inertial confined fusion experiments at the National Ignition Facility have recently entered a new regime approaching ignition. Improved modeling and exploration of the experimental parameter space were essential to deepening our understanding of the mechanisms that degrade and amplify the neutron yield. The growing prevalence of machine learning in fusion studies opens a new avenue for investigation. Here in this paper, we have applied the Gradient-Boosted Decision Tree machine-learning architecture to further explore the parameter space and find correlations with the neutron yield, a key performance indicator. We find reasonable agreement between the measured and predicted yield, with a mean absolute percentage error on a randomly assigned test set of 35.5%. This model finds the characteristics of the laser pulse to be the most influential in prediction, as well as the hohlraum laser entrance hole diameter and an enhanced capsule fabrication technique. We used the trained model to scan over the design space of experiments from three different campaigns to evaluate the potential of this technique to provide design changes that could improve the resulting neutron yield. While these data-driven model cannot predict ignition without examples of ignited shots in the training set, it can be used to indicate that an unseen shot design will at least be in the upper range of previously observed neutron yields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Interpretable boosted-decision-tree analysis for the Majorana Demonstrator

The Majorana Demonstrator is a leading experiment searching for neutrinoless double-beta decay with high purity germanium detectors (HPGe). Machine learning provides a new way to maximize the amount of information provided by these detectors, but the data-driven nature makes it less interpretable compared to traditional analysis. An interpretability study reveals the machine's decision-making logic, allowing us to learn from the machine to feedback to the traditional analysis. In this work, we have presented the first machine learning analysis of the data from the Majorana Demonstrator; this is also the first interpretable machine learning analysis of any germanium detector experiment. Two gradient boosted decision tree models are trained to learn from the data, and a game-theory-based model interpretability study is conducted to understand the origin of the classification power. By learning from data, this analysis recognizes the correlations among reconstruction parameters to further enhance the background rejection performance. By learning from the machine, this analysis reveals the importance of new background categories to reciprocally benefit the standard Majorana analysis. This model is highly compatible with next-generation germanium detector experiments like LEGEND since it can be simultaneously trained on a large number of detectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Circular Velocity Curve of the Milky Way from 5–25 kpc Using Luminous Red Giant Branch Stars

We present a sample of 254,882 luminous red giant branch (LRGB) stars selected from the APOGEE and LAMOST surveys. By combining photometric and astrometric information from the Two Micron All Sky Survey and Gaia survey, the precise distances of the sample stars are determined by a supervised machine-learning algorithm: the gradient-boosted decision trees. To test the accuracy of the derived distances, member stars of globular clusters (GCs) and open clusters are used. The tests by cluster member stars show a precision of about 10% with negligible zero-point offsets, for the derived distances of our sample stars. The final sample covers a large volume of the Galactic disk(s) and halo of 0 < R < 30 kpc and |Z| ≤ 15 kpc. The rotation curve (RC) of the Milky Way across the radius of 5 ≲ R ≲ 25 kpc has been accurately measured with ~54,000 stars of the thin disk population selected from the LRGB sample. The derived RC shows a weak decline along R with a gradient of -1.83 ± 0.02 (stat.) ± 0.07 (sys.) km s -1 kpc -1 , in excellent agreement with the results measured by previous studies. The circular velocity at the solar position, yielded by our RC is 234.04 ± 0.08 (stat.) ± 1.36 (sys.) km s -1 , again in great consistency with other independent determinations. From the newly constructed RC, as well as constraints from other data, we have constructed a mass model for our Galaxy, yielding a mass of the dark matter halo of M 200 = (8.05 ± 1.15) × 10 11 M ⊙ with a corresponding radius of R 200 = 192.37 ± 9.24 kpc and a local dark matter density of 0.39 ± 0.03 GeV cm -3 .

79 ASTRONOMY AND ASTROPHYSICS↗

Unraveling the Correlation between Raman and Photoluminescence in Monolayer MoS 2 through Machine‐Learning Models

Abstract 2D transition metal dichalcogenides (TMDCs) with intense and tunable photoluminescence (PL) have opened up new opportunities for optoelectronic and photonic applications such as light‐emitting diodes, photodetectors, and single‐photon emitters. Among the standard characterization tools for 2D materials, Raman spectroscopy stands out as a fast and non‐destructive technique capable of probing material's crystallinity and perturbations such as doping and strain. However, a comprehensive understanding of the correlation between photoluminescence and Raman spectra in monolayer MoS 2 remains elusive due to its highly nonlinear nature. Here, the connections between PL signatures and Raman modes are systematically explored, providing comprehensive insights into the physical mechanisms correlating PL and Raman features. This study's analysis further disentangles the strain and doping contributions from the Raman spectra through machine‐learning models. First, a dense convolutional network (DenseNet) to predict PL maps by spatial Raman maps is deployed. Moreover, a gradient boosted trees model (XGBoost) with Shapley additive explanation (SHAP) to bridge the impact of individual Raman features in PL features is applied. Last, a support vector machine (SVM) to project PL features on Raman frequencies is adopted. This work may serve as a methodology for applying machine learning to characterizations of 2D materials.

Lu, Ang‐Yu↗