Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “local clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Attention Enabled Multi-Agent DRL for Decentralized Volt-VAR Control of Active Distribution System Using PV Inverters and SVCs

This paper proposes attention enabled multi-agent deep reinforcement learning (MADRL) framework for active distribution network decentralized Volt-VAR control. Using the unsupervised clustering, the whole distribution system can be decomposed into several sub-networks according to the voltage and reactive power sensitivity relationships. Then, the distributed control problem of each sub-network is modeled as Markov games and solved by the improved MADRL algorithm, where each sub-network is modeled as an adaptive agent. An attention mechanism is developed to help each agent focus on specific information that is mostly related to the reward. All agents are centrally trained offline to learn the optimal coordinated Volt-VAR control strategy and executed in a decentralized manner to make online decisions with only local information. Compared with other distributed control approaches, the proposed method can effectively deal with uncertainties, achieve fast decision makings, and significantly reduce the communication requirements. Comparison results with model-based and other data-driven methods on IEEE 33-bus and 123-bus systems demonstrate the benefits of the proposed approach.

distribution network↗

Morphological descriptors of nanoparticles: The link between atomistic structures and x-ray absorption spectra

Understanding and quantifying the morphology of nanoparticles are essential for linking their atomic structure to diverse applications and verifying theoretical models. While experimental information on the structure of nanoparticles in the size range below ∼5 nm can be extracted from x-ray absorption spectroscopy using a small number of descriptors—most commonly coordination numbers—developing an understanding of morphology descriptors from experimental data remains a challenge. Here, in this study, we introduce NanoGene, a genetic algorithm-based method for generating structurally diverse nanoparticle models guided by user-defined descriptors. We establish correlations among structural, size-related, and morphological descriptors and demonstrate how experimentally accessible parameters, such as coordination numbers, can be leveraged to infer otherwise inaccessible ones, such as the generalized coordination number or particle oblateness. Principal component and clustering analyses reveal the relative importance of descriptors, with the number of atoms emerging as a key discriminant of the nanoparticle structure. By providing both the methodology and an extensive dataset of nanoparticle geometries, this work offers a practical foundation for descriptor-based analysis and interpretation of experimental observations, bridging the gap between local atomic coordinates and global morphological characterization.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Reaction–drift–diffusion models from master equations: application to material defects

We present a general method to produce well-conditioned continuum reaction–drift–diffusion equations directly from master equations on a discrete, periodic state space. We assume the underlying data to be kinetic Monte Carlo models (i.e. continuous-time Markov chains) produced from atomic sampling of point defects in locally periodic environments, such as perfect lattices, ordered surface structures or dislocation cores, possibly under the influence of a slowly varying external field. Our approach also applies to any discrete, periodic Markov chain. Here, the analysis identifies a previously omitted non-equilibrium drift term, present even in the absence of external forces, which can compete in magnitude with the reaction rates, thus being essential to correctly capture the kinetics. To remove fast modes which hinder time integration, we use a generalized Bloch relation to efficiently calculate the eigenspectrum of the master equation. A well conditioned continuum equation then emerges by searching for spectral gaps in the long wavelength limit, using an established kinetic clustering algorithm to define a proper reduced, Markovian state space.

36 MATERIALS SCIENCE↗

Short- and medium-range orders in Al90Tb10 glass and their relation to the structures of competing crystalline phases

Molecular dynamics simulations using an interatomic potential developed by artificial neural network deep machine learning are performed to study the local structural order in Al90Tb10 metallic glass. We show that more than 80% of the Tb-centered clusters in Al90Tb10 glass have short-range order (SRO) with their 17 first coordination shell atoms stacked in a ‘3661’ or ‘15551’ sequence. Medium-range order (MRO) in Bergman-type packing extended out to the second and third coordination shells is also clearly observed. Analysis of the network formed by the ‘3661’ and ‘15551’ clusters show that ~82% of such SRO units share their faces or vertexes, while only ~6% of neighboring SRO pairs are interpenetrating. Such a network topology is consistent with the Bergman-type MRO around the Tb-centers. Moreover, crystal structure searches using genetic algorithm and the neural network interatomic potential reveal several low-energy metastable crystalline structures in the composition range close to Al 90 Tb 10 . Some of these crystalline structures have the ‘3661’ SRO while others have the ‘15551’ SRO. While the crystalline structures with the ‘3661’ SRO also exhibit the MRO very similar to that observed in the glass, the ones with the ‘15551’ SRO have very different atomic packing in the second and third shells around the Tb centers from that of the Bergman-type MRO observed in the glassy phase.

36 MATERIALS SCIENCE↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Artificial intelligence-empowered cellular morphometric risk score improves prognostic stratification of cutaneous squamous cell carcinoma

Abstract Background Risk stratification of cutaneous squamous cell carcinoma (cSCC) is essential for managing patients. Objectives To determine if artificial intelligence and machine learning might help to stratify patients with cSCC by risk using more than solely clinical and histopathological factors. Methods We retrieved a retrospective cohort of 104 patients whose cSCCs had been excised with clear margins. Clinical and histopathological risk factors were evaluated. Haematoxylin and eosin-stained slides were scanned and analysed by an algorithm based on the stacked predictive sparse decomposition technique. Cellular morphometric biomarkers (CMBs) were identified via machine learning and used to derive a cellular morphometric risk score (CMRS) that classified cSCCs into clusters of differential prognoses. Concordance analysis, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and accuracy were calculated and compared with results obtained with the Brigham and Women’s Hospital (BWH) staging system. The performance of the combination of the BWH staging system and the CMBs was also analysed. Results There were no differences among the CMRS groups in terms of clinical and histopathological risk factors and T-stage assignment, but there were significant differences in prognosis. Combining the CMRS with BWH staging systems increased distinctiveness and improved prognostic performance. C-indices were 0.91 local recurrence and 0.91 for nodal metastasis when combining the two approaches. The NPV was 94.41% and 96.00%, the PPV was 36.36% and 41.67%, and accuracy reached 86.75% and 89.16%, respectively, with the combined approach. Conclusions CMRS is helpful for cSCC risk stratification beyond classic clinical and histopathological risk features. Combining the information from the CMRS and the BWH staging system offers outstanding prognostic performance for patients with high-risk cSCC.

Pérez-Baena, Manuel J.↗

The impact of urban configuration types on urban heat islands, air pollution, CO 2 emissions, and mortality in Europe: a data science approach

The world is becoming increasingly urbanized. As cities around the world continue to grow, it is important for urban planners and policymakers to understand how different urban configuration patterns affect the environment and human health. We aimed at identifying European urban configuration types, based on the Local Climate Zones categories and street design variables from Open Street Map, and evaluating their association with motorized traffic flows, Surface Urban Heat Island (SUHI) intensities, tropospheric nitrogen dioxide (NO 2 ), CO 2 per capita emissions and age-standardized mortality. We considered 946 European cities from 31 countries for the analysis defined in the 2018 Urban Audit database, of which 919 European cities were analysed. Data were collected at a 250 m × 250 m grid cell resolution. We divided all cities into five concentric rings based on the Burgess concentric urban planning model and calculated the mean values of all variables for each ring. First, to identify distinct urban configuration types, we applied the Uniform Manifold Approximation and Projection for Dimension Reduction method, followed by the k-means clustering algorithm. Next, statistical differences in exposures (including SUHI) and mortality between the resulting urban configuration types were evaluated using a Kruskal–Wallis test followed by a post-hoc Dunn's test. We identified four distinct urban configuration types characterising European cities: compact high density (n=246), open low-rise medium density (n=245), open low-rise low density (n=261), and green low density (n=167). Compact high density cities were a small size, had high population densities, and a low availability of natural areas. In contrast, green low-density cities were a large size, had low population densities, and a high availability of natural areas and cycleways. The open low-rise medium and low-density cities were a small to medium size with medium to low population densities and low to moderate availability of green areas. Motorised traffic flows and NO 2 exposure were significantly higher in compact high density and open low rise medium density cities when compared with green low density and open low-rise low density cities. Additionally, green low-density cities had a significantly lower SUHI effect compared with all other urban configuration types. Per person CO 2 emissions were significantly lower in compact high density cities compared with green low density cities. Lastly, green low density cities had significantly lower mortality rates when compared with all other urban configuration types. Our findings indicate that, although the compact city model is more sustainable, European compact cities still face challenges related to poor environmental quality and health. Our results have notable implications for urban and transport planning policies in Europe and contribute to the ongoing discussion on which city models can bring the greatest benefits for the environment, climate, and health.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Noise-induced barren plateaus in variational quantum algorithms

Abstract Variational Quantum Algorithms (VQAs) may be a path to quantum advantage on Noisy Intermediate-Scale Quantum (NISQ) computers. A natural question is whether noise on NISQ devices places fundamental limitations on VQA performance. We rigorously prove a serious limitation for noisy VQAs, in that the noise causes the training landscape to have a barren plateau (i.e., vanishing gradient). Specifically, for the local Pauli noise considered, we prove that the gradient vanishes exponentially in the number of qubits n if the depth of the ansatz grows linearly with n . These noise-induced barren plateaus (NIBPs) are conceptually different from noise-free barren plateaus, which are linked to random parameter initialization. Our result is formulated for a generic ansatz that includes as special cases the Quantum Alternating Operator Ansatz and the Unitary Coupled Cluster Ansatz, among others. For the former, our numerical heuristics demonstrate the NIBP phenomenon for a realistic hardware noise model.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward scalable quantum computations of atomic nuclei

We solve the nuclear two-body and three-body bound states via quantum simulations of pionless effective field theory on a lattice in position space. While the employed lattice remains small, the usage of local Hamiltonians including two- and three-body forces ensures that the number of Pauli terms scales linearly with increasing numbers of lattice sites. We use an adaptive ansatz grown from unitary coupled cluster theory to parametrize the ground states of the deuteron and 3 He, compute their corresponding energies, and analyze the scaling of the required computational resources. Our quantum simulations reproduce exact benchmarks for 2 H and 3 He within 100 keV, requiring at most 30 layers in the ansatz and thus resulting in modest circuit depths. Additionally, we find the number of shots required to reach a given precision scales linearly in the lattice size and more mildly in the system size. Furthermore, based on the agreement with exact benchmarks and mild scaling, we conclude that this can be an efficient, scalable approach for quantum computations of nuclear ground states, particularly to prepare initial states for quantum phase estimation or other filtering algorithms.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab↗

LoVoCCS. I. Survey Introduction, Data Processing Pipeline, and Early Science Results

We present the Local Volume Complete Cluster Survey (LoVoCCS; we pronounce it as "low-vox" or "law-vox," with stress on the second syllable), an NSF's National Optical-Infrared Astronomy Research Laboratory survey program that uses the Dark Energy Camera to map the dark matter distribution and galaxy population in 107 nearby (0.03 < z < 0.12) X-ray luminous ([0.1–2.4 keV] L X500 > 10 44 erg s –1 ) galaxy clusters that are not obscured by the Milky Way. The survey will reach Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) Year 1–2 depth (for galaxies r = 24.5, i = 24.0, signal-to-noise ratio (S/N) > 20; u = 24.7, g = 25.3, z = 23.8, S/N > 10) and conclude in ~2023 (coincident with the beginning of LSST science operations), and will serve as a zeroth-year template for LSST transient studies. We process the data using the LSST Science Pipelines that include state-of-the-art algorithms and analyze the results using our own pipelines, and therefore the catalogs and analysis tools will be compatible with the LSST. We demonstrate the use and performance of our pipeline using three X-ray luminous and observation-time complete LoVoCCS clusters: A3911, A3921, and A85. A3911 and A3921 have not been well studied previously by weak lensing, and we obtain similar lensing analysis results for A85 to previous studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Unraveling the structural stability and the electronic structure of ThO 2 clusters

Unraveling the correlations between the geometry, the relative energy and the electronic structure of metal oxide nanostructures is crucial for a better control of their size, shape and properties. Here, we investigated these correlations for stoichiometric thorium dioxide clusters ranging from ThO 2 to Th 8 O 16 using a chemically-driven geometry search algorithm in combination with state-of-the-art first principles calculations. This strategy allows us to homogeneously screen the potential energy surface of actinide oxide clusters for the first time. It is found that the presence of peroxo and superoxo groups tends to increase the total energy of the system by at least 3.5 eV and 7 eV, respectively. For the larger clusters, the presence of terminal oxygen atoms increases the energy by about 0.5 eV. Regarding the electronic structure, it is found that the HOMO–LUMO gap is larger in systems containing only bridging oxygen atoms (~2–3.5 eV) than for systems containing oxo groups (~1–3 eV), peroxo groups (~0–2 eV), and superoxo groups (~0–1 eV). Furthermore, while the LUMO is always dominated by thorium orbitals, the composition of the HOMO changes in the presence or the absence of oxo, peroxo and/or superoxo groups: in the presence of peroxo groups, it is dominated by thorium orbitals, in all other cases, it is dominated by oxygen orbitals, and is rather localized in the presence of terminal oxo or superoxo groups. These correlations are of great interest for synthesizing clusters with tailored properties, especially for applications in the field of nuclear energy and heterogeneous catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DELVE Milky Way Satellite Galaxy Census. I. Satellite Population and Survey Selection Function in DES, DELVE, and Pan-STARRS

The properties of Milky Way satellite galaxies have important implications for galaxy formation, reionization, and the fundamental physics of dark matter. However, the population of Milky Way satellites includes the faintest known galaxies, and current observations are incomplete. To understand the impact of observational selection effects on the known satellite population, we perform rigorous, quantitative estimates of the Milky Way satellite galaxy detection efficiency in three wide-field survey datasets: the Dark Energy Survey Year 6, the DECam Local Volume Exploration Data Release 3, and the Pan-STARRS1 Data Release 1. Together, these surveys cover ∼13,600 deg 2 to g ∼ 24.0 and ∼27,700 deg 2 to g ∼ 22.5, spanning ∼91% of the high-Galactic-latitude sky (∣b∣ ≥ 15°). We apply multiple detection algorithms over the combined footprint and recover 49 known satellites above a strict census detection threshold. To characterize the sensitivity of our census, we run our detection algorithms on a large set of simulated galaxies injected into the survey data, which allows us to develop models that predict the detectability of satellites as a function of their properties. We then fit an empirical model to our data and infer the luminosity function, radial distribution, and size–luminosity relation of Milky Way satellite galaxies. Our empirical model predicts a total of $265^{+79}_{-47}$ satellite galaxies with −20 ≤ M V ≤ 0, half-light radii of 15 ≤ r 1/2 , (pc) ≤ 3000, and galactocentric distances of 10 ≤ D GC (kpc) ≤ 300. We also identify a mild anisotropy in the angular distribution of the observed galaxies, at a significance of ∼2σ, which can be attributed to the clustering of satellites associated with the LMC.

Tan, Chin Yi [Univ. of Chicago, IL (United States)↗

Automated identification of dominant physical processes

The identification of processes that locally and approximately dominate dynamical system behavior has enabled significant advances in understanding and modeling nonlinear differential dynamical systems. Conventional methods of dominant process identification involve piecemeal and ad hoc (non-rigorous, informal) scaling analyses to identify dominant balances of governing equation terms and to delineate the spatiotemporal boundaries (boundaries in space and/or time) of each dominant balance. For the first time, we present an objective global measure of the fit of dominant balances to observations, which is desirable for automation, and was previously undefined. Furthermore, we propose a formal definition of the dominant balance identification problem in the form of an optimization problem. Here, we show that the optimization can be performed by various machine learning algorithms, enabling the automatic identification of dominant balances. Our method is algorithm agnostic and it eliminates reliance upon expert knowledge to identify dominant balances which are not known beforehand.

42 ENGINEERING↗

Resilient Autonomous Wind Farms: Preprint

With the advent of an increasing number of control strategies that seek to optimize wind turbine performance on a farm-level, taking account of individual wind turbine information to achieve wind farm-level objectives has become an increasingly important goal. Methods for controlling wind turbines on an individual and farm level have seen significant development, and an abundance of new implementations for gathering and using data from turbines have created potential for novel control mechanisms which can further optimize the performance and delivery characteristics of a wind farm. A key element of making these wind farms more efficient is to develop reliable algorithms that use local sensor information that is already being collected, such as supervisory control and data acquisition (SCADA) data, local meteorological stations, and nearby radars/sodars/lidars. Making use of information from all wind turbines in a wind farm can enable such approaches as determining the atmospheric conditions across the farm, improving fault-finding, and enabling more efficient overall control of farm-wide optimizations through mechanisms such as wake-steering. However, these approaches typically involve a centralized communications and control center. In order to ensure the resilient operation of the farm, it is necessary to develop an approach which distributes the calculation and communication amongst multiple nodes throughout the farm. In this fashion, a redundant, robust, and secure network can be created, which can tolerate faults in calculation, communication, and even external attacks which seek to disrupt the operation of the wind farm. This paper introduces the use of the Raft Byzantine Fault Tolerance algorithm in the implementation of autonomous control of a wind farm. This implementation will allow for fault tolerance for malfunctioning nodes, sensors, transmitters, and connectors. This approach is equally extensible to account for malicious actors. It will be shown to achieve overall consensus, provided the number of faults/malicious nodes is less than 3$n$+1, where $n$ is the number of turbine cluster faults which may occur, and to be robust in the face of multiple arbitrary faults.

autonomous↗

Excited States of Crystalline Point Defects with Multireference Density Matrix Embedding Theory

Accurate and affordable methods to characterize the electronic structure of solids are important for targeted materials design. Embedding-based methods provide an appealing balance in the trade-off between cost and accuracy-particularly when studying localized phenomena. Here, we use the density matrix embedding theory (DMET) algorithm to study the electronic excitations in solid-state defects with a restricted open-shell Hartree-Fock (ROHF) bath and multireference impurity solvers, specifically, complete active space self-consistent field (CASSCF) and n-electron valence state second-order perturbation theory (NEVPT2). In this work, we apply the method to investigate the electronic excitations in an oxygen vacancy (OV) on a MgO(100) surface and find absolute deviations within 0.05 eV between DMET using the CASSCF/NEVPT2 solver, denoted as CAS-DMET/NEVPT2-DMET, and the nonembedded CASSCF/NEVPT2 approach. Next, we establish the practicality of DMET by extending it to larger supercells for the OV defect and a neutral silicon vacancy in diamond where the use of nonembedded CASSCF/NEVPT2 is extremely expensive.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine-learning based approach to examine ecological processes influencing the diversity of riverine dissolved organic matter composition

Dissolved organic matter (DOM) assemblages in freshwater rivers are formed from mixtures of simple to complex compounds that are highly variable across time and space. These mixtures largely form due to the environmental heterogeneity of river networks and the contribution of diverse allochthonous and autochthonous DOM sources. Most studies are, however, confined to local and regional scales, which precludes an understanding of how these mixtures arise at large, e.g., continental, spatial scales. The processes contributing to these mixtures are also difficult to study because of the complex interactions between various environmental factors and DOM. Here we propose the use of machine learning (ML) approaches to identify ecological processes contributing toward mixtures of DOM at a continental-scale. We related a dataset that characterized the molecular composition of DOM from river water and sediment with Fourier-transform ion cyclotron resonance mass spectrometry to explanatory physicochemical variables such as nutrient concentrations and stable water isotopes ( 2 H and 18 O). Using unsupervised ML, distinctive clusters for sediment and water samples were identified, with unique molecular compositions influenced by environmental factors like terrestrial input and microbial activity. Sediment clusters showed a higher proportion of protein-like and unclassified compounds than water clusters, while water clusters exhibited a more diversified chemical composition. We then applied a supervised ML approach, involving a two-stage use of SHapley Additive exPlanations (SHAP) values. In the first stage, SHAP values were obtained and used to identify key physicochemical variables. These parameters were employed to train models using both the default and subsequently tuned hyperparameters of the Histogram-based Gradient Boosting (HGB) algorithm. The supervised ML approach, using HGB and SHAP values, highlighted complex relationships between environmental factors and DOM diversity, in particular the existence of dams upstream, precipitation events, and other watershed characteristics were important in predicting higher chemical diversity in DOM. Our data-driven approach can now be used more generally to reveal the interplay between physical, chemical, and biological factors in determining the diversity of DOM in other ecosystems.

54 ENVIRONMENTAL SCIENCES↗