Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Different methods of estimating riverbed sediment grain size diverge at the basin scale

Introduction: The distribution of sediment grain size in streams and rivers is often quantified by the median grain size (D50), a key metric for understanding and predicting hydrologic and biogeochemical function of streams and rivers. Manual D50 measurements are time-consuming and ignore larger grains, while approaches to model D50 based on catchment characteristics may over-generalize and miss site-scale heterogeneity. Machine learning-enabled object detection methods like You Only Look Once (YOLO) provides an alternative that enables estimation of D50 that is faster than manual measurements and more site-specific than predictions based on catchment characteristics. Methods: To understand the potential role of object detection methods for improving understanding of D50, we compared D50 estimates made manually, predicted from catchment characteristics, and using a YOLO-enabled approach across the Yakima River Basin. Results: We found distinct differences between methods for D50 averages and variability, and relationships between D50 estimates and basin characteristics. Discussion: We discuss the advantages and limitations of object detection methods versus current methods, and explore potential future directions to combine D50 methods to better estimate spatiotemporal variation of D50, and improve incorporation into basin-scale models.

grain size distribution↗

Machine learning coupled multi-scale modeling for redox flow batteries

The reaction distribution in macro or device-scale has been studied for redox flow batteries. The reaction distribution on electrode pore-scale structure however is not well understood, lacking especially on how the reaction distribution on the pore-scale may impact the overall performance of a flow battery. This study introduces for the first time a framework of a multi-scale model that provides understanding of the relationship between the pore-scale electrode structure reaction and the device-scale electrochemical reaction uniformity within the flow battery. A reduced order model is constructed based on 128 pore-scale simulations, which provide a quantitative relationship between the battery operation conditions (inlet velocity, current density, inlet concentration) and the surface reaction uniformity for the pore-scale sample. The multi-scale framework upscales this pore-scale surface reaction uniformity to device-scale combined uniformity. Based on the multi-scale model, a time-varying optimization of the inlet velocity is established, leading to significant reduction on pump power consumption with targeted surface reaction uniformity. The multi-scale model establishes the critical link between the micro-structure of a flow battery component and its performance at the macro-scale, therefore providing rationale for further operational or material optimization.

flow batteries, machine learning, multi-scale mode↗

Machine learning assisted bayesian inference of mix and hot-spot conditions in NIF implosions

Experiments on the National Ignition Facility (NIF) have provided clear evidence of ablator material mixing into the Hot-Spot, leading to degraded performance. However, inferring the amount of mix and Hot-Spot conditions from typical experimental observations (e.g. x-ray spectra and images) is highly challenging. Here, we have developed an analysis method that utilizes machine learning assisted Bayesian inference to find the probability distributions of the Hot-Spot and mix conditions. This approach uses a neural network, trained on an idealized 2-dimensional representation of the Hot-Spot and mix distribution, and Bayesian inference to find the statistical distributions of Hot-Spot conditions that provide a match with observations. We have tested this method with synthetic data from simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The impact of the $\text{WHIM}$ on the $\text{IGM}$ thermal state determined from the low- z Lyman $\alpha$ forest

At z ≲ 1, shock heating caused by large-scale velocity flows and possibly violent feedback from galaxy formation, converts a significant fraction of the cool gas (T ~ 10 4 K) in the intergalactic medium (IGM) into warm–hot phase (WHIM) with T > 10 5 K, resulting in a significant deviation from the previously tight power-law IGM temperature–density relationship, T = T 0 (P/$\bar{p}$) γ-1 ⁠. This study explores the impact of the WHIM on measurements of the low-z IGM thermal state, [T 0 , γ], based on the b–NH 1 distribution of the Ly α forest. Exploiting a machine learning-enabled simulation-based inference method trained on Nyx hydrodynamical simulations, we demonstrate that [T 0 , γ] can still be reliably measured from the b–NH 1 distribution at z = 0.1, notwithstanding the substantial WHIM in the IGM. To investigate the effects of different feedback, we apply this inference methodology to mock spectra derived from the IllustrisTNG and Illustris simulations at z = 0.1. The results suggest that the underlying [T 0 , γ] of both simulations can be recovered with biases as low as |Δlog(T 0 /K)| ≲ 0.05 dex, |Δγ| ≲ 0.1, smaller than the precision of a typical measurement. Given the large differences in the volume-weighted WHIM fractions between the three simulations (Illustris 38 percent, IllustrisTNG 10 percent, and Nyx 4 per cent), we conclude that the b–N H1 distribution is not sensitive to the WHIM under realistic conditions. Finally, we investigate the physical properties of the detectable Ly α absorbers, and discover that although their T and Δ distributions remain mostly unaffected by feedback, they are correlated with the photoionization rate used in the simulation.

79 ASTRONOMY AND ASTROPHYSICS↗

An Efficient Bayesian Approach to Learning Droplet Collision Kernels: Proof of Concept Using “Cloudy,” a New n -Moment Bulk Microphysics Scheme

The small-scale microphysical processes governing the formation of precipitation particles cannot be resolved explicitly by cloud resolving and climate models. Instead, they are represented by microphysics schemes that are based on a combination of theoretical knowledge, statistical assumptions, and fitting to data (“tuning”). Historically, tuning was done in an ad hoc fashion, leading to parameter choices that are not explainable or repeatable. Recent work has treated it as an inverse problem that can be solved by Bayesian inference. The posterior distribution of the parameters given the data—the solution of Bayesian inference—is found through computationally expensive sampling methods, which require over $\mathcal{O}$(10 5 ) evaluations of the forward model; this is prohibitive for many models. We present a proof of concept of Bayesian learning applied to a new bulk microphysics scheme named “Cloudy,” using the recently developed Calibrate-Emulate-Sample (CES) algorithm. Cloudy models collision-coalescence and collisional breakup of cloud droplets with an adjustable number of prognostic moments and with easily modifiable assumptions for the cloud droplet mass distribution and the collision kernel. The CES algorithm uses machine learning tools to accelerate Bayesian inference by reducing the number of forward evaluations needed to $\mathcal{O}$(10 2 ). It also exhibits a smoothing effect when forward evaluations are polluted by noise. In a suite of perfect-model experiments, we show that CES enables computationally efficient Bayesian inference of parameters in Cloudy from noisy observations of moments of the droplet mass distribution. In an additional imperfect-model experiment, a collision kernel parameter is successfully learned from output generated by a Lagrangian particle-based microphysics model.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning in Power System Operations: Training Data

Reliability and stability of the electric grid today has depended upon operations of the grid which include the protective relay. Today, the electricity sector faces new challenges with the shift of generation resource characteristics away from the traditional “big iron” generation to inverter-based resources (IBR) which shift the physics and assumption used in grid operation and protection. These changing conditions represent new challenges for protective relays (identification of faults) and increased challenges for protection engineers (correct settings and configuration, reduction of mis-operations), both issues recognized in research and industry. Finding new approaches to reduce mis-operations in relaying and new approaches to fault identification is critical to grid operations. Using today’s modern technology of embedded systems, edge computing, machine learning (ML), and communications we can help address challenges and augment and improve on existing power system operations methodologies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

GridDS: Data Science Toolkit for Energy Grid Data

According to the U.S. Energy Information Administration (EIA), the demand for energy is expected to increase 50% by the year 20501. While energy standards, such as the Institute of Electrical and Electronics Engineers (IEEE) Standard 1547, (Basso 2015) and monitoring with wide area management systems (WAMS) (Liu 2017, Zhou 2016) have enabled large scale data collection and storage, the application of this data in mitigating costs associated with increased consumer demand is an ongoing focus for energy research. This ubiquitous data collection presents a promising opportunity for machine learning and data science to improve efficiency of distributed energy resources (DERs). The GridDS software toolkit is designed to leverage advanced metering infrastructure (AMI), outage management systems data (OMS), Supervisory control Data Acquisition (SCADA), and geographic information systems (GIS) to forecast future energy demands and detect incipient grid failures. GridDS is a python software library designed to be modular and generalizable to data recorded by DERs. In adapting to disparate datasets recorded by various WAMS, GridDS provides a range of unique functionality not presently implemented in current WAMS which have highly specific software infrastructure by design. GridDS functionality ranges from data specification and preparation, to training and validation for state of the art machine learning, to interactive data visualization. For data intake, GridDS combines: Pandera: a library for creating data specifications. TimeScaleDB: a postgresSQL database infrastructure for efficient storage of timeseries data. Dataset class: A custom dataset class / interface that ensures modularity between a range of synthetic and live recorded datasets. Is

Ladd, Alexander↗

Normalizing Flows for Microscopic Many-Body Calculations: An Application to the Nuclear Equation of State

We report that normalizing flows are a class of machine learning models used to construct a complex distribution through a bijective mapping of a simple base distribution. We demonstrate that normalizing flows are particularly well suited as a Monte Carlo integration framework for quantum many-body calculations that require the repeated evaluation of high-dimensional integrals across smoothly varying integrands and integration regions. As an example, we consider the finite-temperature nuclear equation of state. An important advantage of normalizing flows is the ability to build highly expressive models of the target integrand, which we demonstrate enables precise evaluations of the nuclear free energy and its derivatives. Furthermore, we show that a normalizing flow model trained on one target integrand can be used to efficiently calculate related integrals when the temperature, density, or nuclear force is varied. This work will support future efforts to build microscopic equations of state for numerical simulations of supernovae and neutron star mergers that employ state-of-the-art nuclear forces and many-body methods.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Preliminary Results on Bayesian Inverse UQ for OECD/NEA WPNCS Subgroup 14 Benchmark Exercise for Error Recovery and Experimental Coverage

The Organization for Economic Cooperation and Development (OECD) Working Party on Nucelar Criticality Safety (WPNCS) has proposed a benchmark exercise representative of neutronic behavior in criticality experiments. Here, the goal is to develop confidence in data assimilation techniques used to adjust nuclear data. Participants are given synthetic experimental models with associated measured data and asked to estimate the model parameters given the model and measurements as well as provide predictions for separate application models. In this work, we performed data assimilation using Bayesian inverse Uncertainty Quantification (UQ) with machine learning surrogate models to produce posterior parameter distributions for the requested parameters and posterior predictive distributions for the requested responses. Several experimental models are shown to insufficiently inform the posterior parameter distributions for the applications involved. However, given sufficient experimental data, posterior parameter estimates yielded reduced uncertainty in the response predictions of interest while covering the experimental data.

Bayesian Inference↗

Facility Cybersecurity Framework Best Practices

Federal facilities are increasingly adopting automation and connecting to the Internet creating an energy-internet-of-things environment that converges operational technology (OT) and information technology (IT). Today's buildings increasingly weave together networked sensors and cyber and physical systems that enable data to be collected, aggregated, exchanged, stored and monetized in new ways. Building technological advances have created new energy technology, services, markets and value creation opportunities (e.g. transactive energy, two-way grid communications, machine learning, and increased use of renewable and distributed energy resources). But as larger data sets are being exchanged at faster speeds between an increasing number of OT systems, it becomes more difficult to protect the security of the data lifecycle and the physical equipment it interacts with. These challenges are especially difficult to overcome because the economic and environmental gain (interoperability, big data, social networks and ubiquitous information sharing) are driving these prominent trends in the digital age. Often cybersecurity is an afterthought. The U.S. Department of Energy’s (DOE) Federal Energy Management Program (FEMP) funded the Pacific Northwest National Laboratory (PNNL) to develop various cybersecurity tools, trainings, and reports to aid federal facility managers – and other building owners and operators – in better applying frameworks and lessons learned from the National Institute of Standards and Technology (NIST) Cybersecurity Framework (CSF), risk management framework (RMF), DOE’s cybersecurity capability maturity model (C2M2), and a wide variety of industry best practices and guidance documents (i.e., NIST 800 series, Department of Defense United Facilities Criteria). This set of tools, collectively known as the FEMP Facility-Related Control System Cyber Toolkit (FRCS Cyber Toolkit)2, is focused on cybersecurity concerns from facility-related control systems and other operational technology (OT), such as industrial control systems (ICS). The FRCS Cyber Toolkit can be applied across six of the sixteen critical infrastructure sectors designated by the Department of Homeland Security, including government facilities, healthcare and public health, commercial facilities (e.g., public assembly, offices, lodging), financial services (e.g., banking and insurance), emergency services (e.g., fire and police stations), and information technology. With increasingly converged IT and OT systems, it is crucial to address OT cybersecurity considerations and assess how the seam of these two systems could impact the overall cybersecurity posture of a facility. The objective of this report is to provide an overview of the best possible method to use FRCS Cyber Toolkit (section 2.0) and distilled cybersecurity best practices for the federal facilities to address growing non-linear cyber threats (section 3.0). Recommendations in this document are aggregated from several NIST and other documents (see Appendix A for additional details).

97 MATHEMATICS AND COMPUTING↗

Facility Cybersecurity Framework Best Practices Version 2.0

Federal facilities are increasingly adopting automation and connecting to the Internet creating an energy-internet-of-things environment that converges operational technology (OT) and information technology (IT). Today's buildings increasingly weave together networked sensors and cyber and physical systems that enable data to be collected, aggregated, exchanged, stored and monetized in new ways. Building technological advances have created new energy technology, services, markets and value creation opportunities (e.g. transactive energy, two-way grid communications, machine learning, and increased use of renewable and distributed energy resources). But as larger data sets are being exchanged at faster speeds between an increasing number of OT systems, it becomes more difficult to protect the security of the data lifecycle and the physical equipment it interacts with. These challenges are especially difficult to overcome because the economic and environmental gain (interoperability, big data, social networks and ubiquitous information sharing) are driving these prominent trends in the digital age. Often cybersecurity is an afterthought.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Facility Cybersecurity Framework Best Practices Version 2.0

Federal facilities are increasingly adopting automation and connecting to the Internet creating an energy-internet-of-things environment that converges operational technology (OT) and information technology (IT). Today's buildings increasingly weave together networked sensors and cyber and physical systems that enable data to be collected, aggregated, exchanged, stored and monetized in new ways. Building technological advances have created new energy technology, services, markets and value creation opportunities (e.g. transactive energy, two-way grid communications, machine learning, and increased use of renewable and distributed energy resources). But as larger data sets are being exchanged at faster speeds between an increasing number of OT systems, it becomes more difficult to protect the security of the data lifecycle and the physical equipment it interacts with. These challenges are especially difficult to overcome because the economic and environmental gain (interoperability, big data, social networks and ubiquitous information sharing) are driving these prominent trends in the digital age. Often cybersecurity is an afterthought.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Machine learning analysis reveals relationship between pomacentrid calls and environmental cues

Sound production rates of fishes can be used as an indicator for coral reef health, providing an opportunity to utilize long-term acoustic recordings to assess environmental change. As acoustic datasets become more common, computational techniques need to be developed to facilitate analysis of the massive data files produced by long-term monitoring. Machine learning techniques demonstrate an advantage in the identification of fish sounds over manual sampling approaches. Here we evaluated the ability of convolutional neural networks to identify and monitor call patterns for pomacentrids (damselfishes) in a tropical reef region of the western Pacific. A stationary hydrophone was deployed for 39 mo (2014-2018) in the National Park of American Samoa to continuously record the local marine acoustic environment. A neural network was trained—achieving 94% identification accuracy of pomacentrids—to demonstrate the applicability of machine learning in fish acoustics and ecology. The distribution of sound production was found to vary on diel and interannual timescales. Additionally, the distribution of sound production was correlated with wind speed, water temperature, tidal amplitude, and sound pressure level. This research has broad implications for state-of-the-art acoustic analysis and promises to be an efficient, scalable asset for ecological research, environmental monitoring, and conservation planning.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

Combining synchrotron X-ray diffraction, mechanistic modeling and machine learning for in situ subsurface temperature quantification during laser melting

Laser melting, such as that encountered during additive manufacturing, produces extreme gradients of temperature in both space and time, which in turn influence microstructural development in the material. Qualification and model validation of the process itself and the resulting material necessitate the ability to characterize these temperature fields. However, well established means to directly probe the material temperature below the surface of an alloy while it is being processed are limited. To address this gap in characterization capabilities, a novel means is presented to extract subsurface temperature-distribution metrics, with uncertainty, from in situ synchrotron X-ray diffraction measurements to provide quantitative temperature evolution data during laser melting. Temperature-distribution metrics are determined using Gaussian process regression supervised machine-learning surrogate models trained with a combination of mechanistic modeling (heat transfer and fluid flow) and X-ray diffraction simulation. The trained surrogate model uncertainties are found to range from 5 to 15% depending on the metric and current temperature. The surrogate models are then applied to experimental data to extract temperature metrics from an Inconel 625 nickel superalloy wall specimen during laser melting. The maximum temperatures of the solid phase in the diffraction volume through melting and cooling are found to reach the solidus temperature as expected, with the mean and minimum temperatures found to be several hundred degrees less. The extracted temperature metrics near melting are determined to be more accurate because of the lower relative levels of mechanical elastic strains. However, uncertainties for temperature metrics during cooling are increased due to the effects of thermomechanical stress.

36 MATERIALS SCIENCE↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

A versatile machine learning workflow for high-throughput analysis of supported metal catalyst particles

Accurate and efficient characterization of nanoparticles (NPs), particularly regarding particle size distribution, is essential for advancing our understanding of their structure-property relationship and facilitating their design for various applications. In this study, we introduce a novel two-stage artificial intelligence (AI)-driven workflow for NP analysis that leverages prompt engineering techniques from state-of-the-art single-stage object detection and large-scale vision transformer (ViT) architectures. This methodology is applied to transmission electron microscopy (TEM) and scanning TEM (STEM) images of heterogeneous catalysts, enabling high-resolution, high-throughput analysis of particle size distributions for supported metal catalyst NPs. The model's performance in detecting and segmenting NPs is validated across diverse heterogeneous catalyst systems, including various metals (Ru, Cu, PtCo, and Pt), supports (silica (SiO 2 ), γ-alumina (γ-Al 2 O 3 ), and carbon black), and particle diameter size distributions with mean and standard deviations ranging from 1.6 ± 0.2 nm to 9.7 ± 4.6 nm. The proposed machine learning (ML) methodology achieved an average F1 overlap score of 0.91 ± 0.01 and demonstrated the ability to disentangle overlapping NPs anchored on catalytic support materials. The segmentation accuracy is further validated using the Hausdorff distance and robust Hausdorff distance metrics, with the 90th percent of the robust Hausdorff distance showing errors within 0.4 ± 0.1 nm to 1.4 ± 0.6 nm. In conclusion, our AI-assisted NP analysis workflow demonstrates robust generalization across diverse datasets and can be readily applied to similar NP segmentation tasks without requiring costly model retraining.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗