Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hydrogen Evolution on Electrode‐Supported Pt n Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

Abstract We report the size‐dependent activity and stability of supported Pt 1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with size‐dependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (≈2 H/Pt). An ensemble‐based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. The stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential‐ and adsorbate‐coverage‐dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

Zhang, Zisheng↗

Hydrogen Evolution on Electrode–Supported Ptn Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

We report the size-dependent activity and stability of supported Pt1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with sizedependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (~2 H/Pt). An ensemble-based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. Furthermore, the stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential- and adsorbate-coverage-dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cryptate binding energies towards high throughput chelator design: metadynamics ensembles with cluster–continuum solvation

A tiered forcefield/semiempirical/meta-GGA pipeline together with a thermodynamic scheme designed with error cancellation in mind was developed to calculate binding energies of [2.2.2] cryptate complexes of mono- and divalent cations. Stable complexes of Na, K, Rb, Ca, Zn and Pb were generated, revealing consistent cation–N lengths but highly variable cation–O lengths and an amine stacking mechanism potentially augmenting the cation size selectivity. Metadynamics, used for searching the high-dimensional potential energy surface, together with a cluster–continuum model for affordable – yet accurate – solvation modeling, enabled the discovery of more stable geometries than those previously reported. Similar solvation energy curve shapes for lone vs. coordinated ions enabled rapid solvation convergence via the cancellation of errors stemming from finite cluster sizes. In conclusion, an R 2 of 0.850 vs. experimental aqueous binding energies was obtained, validating this scheme as the backbone of a high-throughput workflow for chelator design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unravelling the orbits of cluster galaxy populations according to their dominant gas ionization source

ABSTRACT We investigate the kinematical and dynamical properties of cluster galaxy populations classified according to their dominant source of gas ionization, namely: star-forming (SF) galaxies, optical active galactic nuclei (AGNs), mixed SF plus AGN ionization (transition objects, T), and quiescent (Q) galaxies. We stack 8892 member galaxies from 336 relaxed galaxy clusters to build an ensemble cluster and estimate the observed projected profiles of numerical density and velocity dispersion, $\sigma _P(R)$, of each galaxy population. The MAMPOSSt code and the Jeans equations inversion technique are used to constrain the velocity anisotropy profiles of the galaxy populations in both parametric and non-parametric ways. We find that Q (SF) galaxies display the lowest (highest) typical cluster-centric distances and velocity dispersion values. Transition galaxies are more concentrated and tend to exhibit lower velocity dispersion values than SF galaxies. Galaxies that host an optical AGN are as concentrated as Q galaxies but display velocity dispersion values similar to those of the SF population. MAMPOSSt is able to find equilibrium solutions that successfully recover the observed $\sigma _P(R)$ profile only for the Q, T, and AGN populations. We find that the orbits of all populations are consistent with isotropy in the inner regions, becoming increasingly radial with the distance from the cluster centre. These results suggest that Q galaxies are in equilibrium within their clusters, while SF galaxies have more recently arrived in the cluster environment. Finally, the T and AGN populations appear to be in an intermediate dynamical state between those of the SF and Q populations.

Valk, Greique A. (ORCID:0009000827731299)↗

A Novel Data Segmentation Method for Data-driven Phase Identification

This paper presents a smart meter phase identification algorithm for two cases: meter-phase-label-known and meter-phase-label-unknown. To improve the identification accuracy, a data segmentation method is proposed to exclude data segments that are collected when the voltage correlation between smart meters on the same phase is weakened. Then, using the selected data segments, a hierarchical clustering method is used to calculate the correlation distances and cluster the smart meters. If the phase labels are unknown, a Connected-Triple-based Similarity (CTS) method is adapted to further improve the phase identification accuracy of the ensemble clustering method. The methods are developed and tested on both synthetic and real feeder data sets. Here, simulation results show that the proposed phase identification algorithm outperforms the state-of-the-art methods in both accuracy and robustness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Grafting nanometer metal/oxide interface towards enhanced low-temperature acetylene semi-hydrogenation

Metal/oxide interface is of fundamental significance to heterogeneous catalysis because the seemingly “inert” oxide support can modulate the morphology, atomic and electronic structures of the metal catalyst through the interface. The interfacial effects are well studied over a bulk oxide support but remain elusive for nanometer-sized systems like clusters, arising from the challenges associated with chemical synthesis and structural elucidation of such hybrid clusters. We hereby demonstrate the essential catalytic roles of a nanometer metal/oxide interface constructed by a hybrid Pd/Bi 2 O 3 cluster ensemble, which is fabricated by a facile stepwise photochemical method. The Pd/Bi 2 O 3 cluster, of which the hybrid structure is elucidated by combined electron microscopy and microanalysis, features a small Pd-Pd coordination number and more importantly a Pd-Bi spatial correlation ascribed to the heterografting between Pd and Bi terminated Bi 2 O 3 clusters. The intra-cluster electron transfer towards Pd across the as-formed nanometer metal/oxide interface significantly weakens the ethylene adsorption without compromising the hydrogen activation. As a result, a 91% selectivity of ethylene and 90% conversion of acetylene can be achieved in a front-end hydrogenation process with a temperature as low as 44 °C.

36 MATERIALS SCIENCE↗

Unleashing the Power of Industrial Big Data through Scalable Manual Labeling

Big Data plays a central role in the remarkable results achieved by Machine Learning (ML) and especially Deep Learning (DL) in the recent years. However, the difficulty in obtaining a reasonable amount of labeled samples limits ML/DL application in various domains, including industrial equipment and system monitoring. In this paper the need for methods that turn manual labeling into a scalable process is highlighted. A real world problem is analyzed for which weak supervision methods, successfully employed in other domains, did not produce acceptable results. An alternative approach based on clustering ensembles is described and tested, achieving good performance.

Paes Leao, Bruno↗

Distilling Knowledge from Ensembles of Cluster-Constrained-Attention Multiple-Instance Learners for Whole Slide Image Classification

The peculiar nature of whole slide imaging (WSI), digitizing conventional glass slides to obtain multiple high resolution images which capture microscopic details of a patient’s histopathological features, has garnered increased interest from the computer vision research community over the last two decades. Given the unique computational space and time complexity inherent to gigapixel-size whole slide image data, researchers have proposed novel machine learning algorithms to aid in the performance of diagnostic tasks in clinical pathology. One effective algorithm represents a Whole slide image as a bag of smaller image patches, which can be represented as low-dimension image patch embeddings. Weakly supervised deep-learning methods, such as cluster-constrained-attention multiple instance learning (CLAM), have shown promising results when combined with image patch embeddings. While traditional ensemble classifiers yield improved task performance, such methods come with a steep cost in model complexity. Through knowledge distillation, it is possible to retain some performance improvements from an ensemble, while minimizing costs to model complexity. In this work, we implement a weakly supervised ensemble using clustering-constrained-attention multiple-instance learners (CLAM), which uses attention and instance-level clustering to identify task salient regions and feature extraction in whole slides. By applying logit-based and attention-based knowledge distillation, we show it is possible to retain some performance improvements resulting from the ensemble at zero cost to model complexity.

Alamudun, Folami↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Analytical and EZmock covariance validation for the DESI 2024 results

The estimation of uncertainties in cosmological parameters is an important challenge in Large-Scale-Structure (LSS) analyses. For standard analyses such as Baryon Acoustic Oscillations (BAO) and Full-Shape two approaches are usually considered. First: analytical estimates of the covariance matrix use Gaussian approximations and (nonlinear) clustering measurements to estimate the matrix, which allows a relatively fast and computationally cheap way to generate matrices that adapt to an arbitrary clustering measurement. On the other hand, sample covariances are an empirical estimate of the matrix based on an ensemble of clustering measurements from fast and approximate simulations. While more computationally expensive due to the large amount of simulations and volume required, these allow us to take into account systematics that are impossible to model analytically. In this work we compare these two approaches in order to enable DESI's key analyses. We find that the configuration space analytical estimate performs satisfactorily in BAO analyses and its flexibility in terms of input clustering makes it the fiducial choice for DESI's 2024 BAO analysis. On the contrary, the analytical computation of the covariance matrix in Fourier space does not reproduce the expected measurements in terms of Full-Shape analyses, which motivates the use of a corrected mock covariance for DESI's 2024 Full Shape analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Ensemble Effects on Hydroxide Bond Dissociation Free Energies in Polyoxovanadate Clusters

Understanding structure-property relationships is foundational to numerous modern chemistries, such as proton-coupled electron transfer (PCET). However, an experimentally measured property is the result of the behavior from an ensemble of molecules. Neglecting ensemble effects, especially under complex chemical environments, may obfuscate these relationships and lead to discrepancies between theory and experiment. In this work, we demonstrate the impact of configurational entropy and local chemical environments on hydroxide bond dissociation free energies [BDFE- (O−H)] for a set of polyoxovanadate nanoclusters, at ambient conditions. The O−H bond strengths are investigated via density functional theory (DFT) coupled with statistical thermodynamic analysis and bilinear modeling, and compared with previous experimental results on the same systems, namely electrochemical solutions of: [V 6 O 13−x (OH) x (TRIOL R ) 2 ] −2 (x = 2, 4, 6; R = NO 2 , Me) and [V 6 O 11−x (OMe) 2 (OH) x (TRIOL NO 2 ) 2 ] −2 (x = 2, 4). Interestingly, we find that ensemble effects, even at room temperature, can account for a significant portion of the BDFE(O−H) trend with the degree of reduction via H atom binding, which cannot be fully captured by single-structure, static DFT calculations. Moreover, we find that the ensemble effects may be replicated statistically, requiring only enumeration of energetically accessible H-binding sites. With the ensemble effects resolved, we present a simple bilinear model to reconcile remaining biases between experiment and ensemble-informed theory, which corelate with clusterspecific electronic environment differences. The bilinear model achieves outstanding accuracy vs experiments with a root-mean squared error of 0.4 kcal/mol. Finally, based on the physicochemical characteristics of hydrogen interaction with polyoxometalates, we present a simple methodology that captures the BDFE(O−H) trend while dramatically reducing required DFT calculations by 98% and achieving accuracy within 1 kcal/mol. Overall, this work elucidates the roles and structural origins of configurational entropy and chemical effects on polyoxometalate hydroxide bond energies, with potential applicability to various atomically precise metal oxide systems. Importantly, it introduces models for rapid and highly accurate property calculations in connection with experiments.

Cluster chemistry↗

Catalytic Activity of an Ensemble of Sites for CO 2 Hydrogenation to Methanol on a ZrO 2 -on-Cu Inverse Catalyst

The significant increase in CO 2 emissions from heavy fossil fuel utilization has raised serious concerns, highlighting the need for effective methods to convert CO 2 into value-added chemicals. Here, in this work, we report a computational investigation on the catalytic activity of ZrO 2 -on-Cu inverse catalysts for CO 2 hydrogenation to methanol, considering highly dispersed ZrO 2 trimers on Cu (111). Such clusters present a large ensemble of formate-containing configurations, Zr 3 O n (OH) m (OCHO) l , making the evaluation of the catalytic activity very challenging. We found that the sites on the various catalyst configurations exhibit markedly different activities for formate hydrogenation, despite their similar free energy and composition. To understand these differences in reactivity, we examined the structural and electronic nature of the low free-energy catalyst configurations and identified that the energy of the lowest unoccupied orbital of the reacting formate, modified by its binding with the catalytic site, is a descriptor for the reaction energy of the formate hydrogenation step. From there, we screened an ensemble of catalyst structures using this descriptor to predict highly active metastable catalyst configurations and computed the reaction pathways and transition states for formate hydrogenation. From this investigation, we distinguished reactive from nonreactive sites and formate species on the ZrO 2 /Cu inverse catalyst based on structural and electronic features. We showed that rare metastable configurations control the activity. Additionally, an efficient method for examining the reactivity of a large number of coexisting catalyst structures was developed.

catalysts↗