Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Quantum annealing for jet clustering with thrust

Quantum computing holds the promise of substantially speeding up computationally expensive tasks, such as solving optimization problems over a large number of elements. In high-energy collider physics, quantum-assisted algorithms might accelerate the clustering of particles into jets. In this study, we benchmark quantum annealing strategies for jet clustering based on optimizing a quantity called “thrust” in electron-positron collision events. Here, we find that quantum annealing yields similar performance to exact classical approaches and classical heuristics, after tuning the annealing parameters. Without tuning, comparable performance can be obtained through a hybrid quantum/classical approach.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Resource-Efficient Chemistry on Quantum Computers with the Variational Quantum Eigensolver and The Double Unitary Coupled-Cluster approach

Applications of quantum simulation algorithms to obtain electronic energies of molecules on noisy intermediate-scale quantum (NISQ) devices require careful consideration of resources describing the complex electron correlation effects. In modeling second-quantized problems, the biggest challenge confronted is that the number of qubits scales linearly with the size of molecular basis. This poses a significant limitation on the size of the basis sets and the number of correlated electrons included in quantum simulations of chemical processes. To address this issue and to enable more realistic simulations on NISQ computers, we employ the double unitary coupled-cluster (DUCC) method to effectively downfold correlation effects into the reduced-size orbital space, commonly referred to as the active space. Using downfolding techniques, we demonstrate that properly constructed effective Hamiltonians can capture the effect of the whole orbital space in small-size active spaces. Combining the downfolding pre-processing technique with the Variational Quantum Eigensolver, we solve for the ground-state energy of H2 and Li2 in the cc-pVTZ basis using the DUCC-reduced active spaces. We compare these results to full configuration-interaction and high-level coupled-cluster reference calculations.

quantum computing, variational quantum solver, cou↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A possibilistic approach to clustering

Fuzzy clustering has been shown to be advantageous over crisp (or traditional) clustering methods in that total commitment of a vector to a given class is not required at each image pattern recognition iteration. Recently fuzzy clustering methods have shown spectacular ability to detect not only hypervolume clusters, but also clusters which are actually 'thin shells', i.e., curves and surfaces. Most analytic fuzzy clustering approaches are derived from the 'Fuzzy C-Means' (FCM) algorithm. The FCM uses the probabilistic constraint that the memberships of a data point across classes sum to one. This constraint was used to generate the membership update equations for an iterative algorithm. Recently, we cast the clustering problem into the framework of possibility theory using an approach in which the resulting partition of the data can be interpreted as a possibilistic partition, and the membership values may be interpreted as degrees of possibility of the points belonging to the classes. We show the ability of this approach to detect linear and quartic curves in the presence of considerable noise.

Krishnapuram, Raghu↗

Application of High-Dimensional Fuzzy K-Means Cluster Analysis to CALIOP/CALIPSO Version 4.1 Cloud-Aerosol Discrimination

This study applies fuzzy k-means (FKM) cluster analyses to a subset of the parameters reported in the CALIPSO lidar level 2 data products in order to classify the layers detected as either clouds or aerosols. The results obtained are used to assess the reliability of the cloud–aerosol discrimination (CAD) scores reported in the version 4.1 release of the CALIPSO data products. FKM is an unsupervised learning algorithm, whereas the CALIPSO operational CAD algorithm (COCA) takes a highly supervised approach. Despite these substantial computational and architectural differences, our statistical analyses show that the FKM classifications agree with the COCA classifications for more than 94 % of the cases in the troposphere. This high degree of similarity is achieved because the lidar-measured signatures of the majority of the clouds and the aerosols are naturally distinct, and hence objective methods can independently and effectively separate the two classes in most cases. Classification differences most often occur in complex scenes (e.g., evaporating water cloud filaments embedded in dense aerosol) or when observing diffuse features that occur only intermittently (e.g., volcanic ash in the tropical tropopause layer). The two methods examined in this study establish overall classification correctness boundaries due to their differing algorithm uncertainties. In addition to comparing the outputs from the two algorithms, analysis of sampling, data training, performance measurements, fuzzy linear discriminants, defuzzification, error propagation, and key parameters in feature type discrimination with the FKM method are further discussed in order to better understand the utility and limits of the application of clustering algorithms to space lidar measurements. In general, we find that both FKM and COCA classification uncertainties are only minimally affected by noise in the CALIPSO measurements, though both algorithms can be challenged by especially complex scenes containing mixtures of discrete layer types. Our analysis results show that attenuated backscatter and color ratio are the driving factors that separate water clouds from aerosols; backscatter intensity, depolarization, and mid-layer altitude are most useful in discriminating between aerosols and ice clouds; and the joint distribution of backscatter intensity and depolarization ratio is critically important for distinguishing ice clouds from water clouds.

Zeng, Shan↗

Qubit coupled cluster singles and doubles variational quantum eigensolver ansatz for electronic structure calculations

Variational quantum eigensolver (VQE) for electronic structure calculations is believed to be one major potential application of near term quantum computing. Among all proposed VQE algorithms, the unitary coupled cluster singles and doubles excitations (UCCSD) VQE ansatz has achieved high accuracy and received a lot of research interest. However, the UCCSD VQE based on fermionic excitations needs extra terms for the parity when using Jordan–Wigner transformation. Here we introduce a new VQE ansatz based on the particle preserving exchange gate to achieve qubit excitations. The proposed VQE ansatz has gate complexity up-bounded to O ( n 4 ) for all-to-all connectivity where n is the number of qubits of the Hamiltonian. Numerical results of simple molecular systems such as BeH 2 , H 2 O, N 2 , H 4 and H 6 using the proposed VQE ansatz gives very accurate results within errors about 10 -3 Hartree.

Physics↗

The Localized Active Space Method with Unitary Selective Coupled Cluster

Here, we introduce a hybrid quantum-classical algorithm, the localized active space unitary selective coupled cluster singles and doubles (LAS-USCCSD) method. Derived from the localized active space unitary coupled cluster (LAS-UCCSD) method, LAS-USCCSD first performs a classical LASSCF calculation, then selectively identifies the most important parameters (cluster amplitudes used to build the multireference UCC ansatz) for restoring interfragment interaction energy using this reduced set of parameters with the variational quantum eigensolver method. We benchmark LAS-USCCSD against LAS-UCCSD by calculating the total energies of (H 2 ) 2 , (H 2 ) 4 , and trans-butadiene, and the magnetic coupling constant for a bimetallic compound [Cr 2 (OH) 3 (NH 3 ) 6 ] 3+ . For these systems, we find that LAS-USCCSD reduces the number of required parameters and thus the circuit depth by at least 1 order of magnitude, an aspect which is important for the practical implementation of multireference hybrid quantum-classical algorithms like LAS-UCCSD on near-term quantum computers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Knowledge Network-Based Approach to Facilitate Annotation of Clinical Pathway Component Clusters

Mining electronic health records (EHRs) to identify contextually related clinical concept clusters that tend to co-occur temporarily and consistently could improve data-driven clinical pathway (CP) construction. However, the automatic extraction of contextually related clinical concept clusters contains a vast amount of irrelevant information. Hence, this paper proposes a knowledge network-enabled literature-based discovery (LBD) approach to remove noise from clusters. The authors used published literature to filter spurious concepts from the clusters and used data from the US Department of Veterans Affairs’s major depressive disorder (MDD) cohort of Operation Enduring Freedom/Operation Iraqi Freedom (OEF/OIF) for their experimentation. The approach was applied to 2,967 clusters extracted from the MDD OEF/OIF database. The experimental results demonstrate that the proposed approach can filter 94% of the irrelevant information. Moreover, the authors applied various network mining algorithms to analyze the clusters and demonstrated that LBD, along with network mining techniques, is a useful method for finding accurate contextually related clinical concept clusters. This could help domain researchers perform advanced analytics in CPs.

Hasan, S M Shamimul↗

Hiperclust

This software leverages transfer learning to analyze atom probe tomography (APT) data. It is trained on synthetic data and then applies this knowledge to predict the optimal number of clusters for a given APT dataset. Initially, the software used preliminary clustering to estimate the general structure of the data. Based on this, it provides suggestions for key parameters like minimum cluster size and minimum number of points. These parameters are critical for algorithms like HDBSCAN, ensuring accurate cluster formation without the need for trial-and-error testing. The software runs on High-Performance computing (HPC) systems, enabling fast, scalable analysis of large APT datasets, ultimately saving time and improving the reliability of clustering outcomes.

Tang, Yalei [Idaho National Laboratory (INL), Idah↗

On Efficient Multigrid Methods for Materials Processing Flows with Small Particles

Multiscale modeling of materials requires simulations of multiple levels of structural hierarchy. The computational efficiency of numerical methods becomes a critical factor for simulating large physical systems with highly desperate length scales. Multigrid methods are known for their superior efficiency in representing/resolving different levels of physical details. The efficiency is achieved by employing interactively different discretizations on different scales (grids). To assist optimization of manufacturing conditions for materials processing with numerous particles (e.g., dispersion of particles, controlling flow viscosity and clusters), a new multigrid algorithm has been developed for a case of multiscale modeling of flows with small particles that have various length scales. The optimal efficiency of the algorithm is crucial for accurate predictions of the effect of processing conditions (e.g., pressure and velocity gradients) on the local flow fields that control the formation of various microstructures or clusters.

Thomas, James↗

Transitory sensitivity in automatic chemical kinetic mechanism analysis

Abstract Detailed chemical kinetic mechanisms are necessary for resolving many important chemical processes. As the chemistry of smaller molecules has become better grounded and quantum chemistry calculations have become cheaper, kineticists have become interested in constructing progressively larger kinetic mechanisms to model increasingly complex chemical processes. These large kinetic mechanisms prove incredibly difficult to refine and time‐consuming to interpret. Traditional sensitivity analysis on a large mechanism can range from inconvenient to practically impossible without special techniques to reduce the computational cost. We first present a new time‐local sensitivity analysis we term transitory sensitivity analysis. Transitory sensitivity analysis is demonstrated in an example to accurately identify traditionally sensitive reactions at an 18,000x speed up over traditional sensitivities. By fusing transitory sensitivity analysis with more traditional time‐local branching, pathway, and cluster analyses, we develop an algorithm for efficient automatic mechanism analysis. This automatic mechanism analysis at a time point is able to identify the reactions a target is most sensitive to using transitory sensitivity analysis and then propose hypotheses why the reaction might be sensitive using branching, pathway, and cluster analyses. We implement these algorithms within the reaction mechanism simulator (RMS) package, which enables us to report the automatic mechanism analysis results in highly readable text formats and in molecular flux diagrams.

Johnson, Matthew S.↗

Copacabana: a probabilistic membership assignment method for galaxy clusters

Cosmological analyses using galaxy clusters in optical/near-infrared photometric surveys require robust characterization of their galaxy content. Precisely determining which galaxies belong to a cluster is crucial. In this paper, we present the COlor Probabilistic Assignment of Clusters And BAyesiaN Analysis (Copacabana) algorithm. Copacabana computes membership probabilities for all galaxies within an aperture centred on the cluster using photometric redshifts, colours, and projected radial probability density functions. We use simulations to validate Copacabana and we show that it achieves up to 89 per cent membership accuracy with a mild dependence on photometric redshift uncertainties and choice of aperture size. We find that the precision of the photometric redshifts has the largest impact on the determination of the membership probabilities followed by the choice of the cluster aperture size. We also quantify how much these uncertainties in the membership probabilities affect the stellar mass–cluster mass scaling relation, a relation that directly impacts cosmology. Using the sum of the stellar masses weighted by membership probabilities (⁠μ * ⁠) as the observable, we find that Copacabana can reach an accuracy of 0.06 dex in the measurement of the scaling relation at low redshift for a Legacy Survey of Space and Time type survey. These results indicate the potential of Copacabana and μ * to be used in cosmological analyses of optically selected clusters in the future.

79 ASTRONOMY AND ASTROPHYSICS↗

Hierarchical median narrow band for level set segmentation of cervical cell nuclei

This paper presents a novel hierarchical nuclei segmentation algorithm for isolated and overlapping cervical cells based on a narrow band level set implementation. Our method applies a new multiscale analysis algorithm to estimate the number of clusters in each image region containing cells, which turns into the input to a narrow band level set algorithm. We assess the nuclei segmentation results on three public cervical cell image databases. Overall, our segmentation method outperformed six state-of-the-art methods concerning the number of correctly segmented nuclei and the Dice coefficient reached values equal to or higher than 0.90. We also carried out classification experiments using features extracted from our segmentation results and the proposed pipeline achieved the highest average accuracy values equal to 0.89 and 0.77 for two-class and three-class problems, respectively. Furthermore, these results demonstrated the suitability of the proposed segmentation algorithm to integrate decision support systems for cervical cell screening.

47 OTHER INSTRUMENTATION↗

Skew-Symmetric adjacency matrices for clustering directed graphs

Graph clustering methods often critically rely on the symmetry of graph matrices. Developing analogous methods for digraphs often proves more challenging, because digraph matrices are typically asymmetric and not orthogonally diagonalizable. However, researchers have recently proposed several complex-valued Hermitian digraph matrices. In particular, one such representation has been utilized as an input to an algorithm for finding imbalanced cuts. In this work, we establish an algebraic relationship between this matrix and an associated real-valued matrix. We show using this real-valued matrix for imbalanced cut-finding algorithms is not only sufficient but advantageous. Our algorithm uses less memory and asymptotically less computation while provably preserving solution quality. We also show our method can be easily implemented using standing computational building blocks, possesses better numerical properties, and loans itself to a natural interpretation via an objective function relaxation argument. We empirically demonstrate these advantages on real world data sets and show how our algorithm can uncover meaningful cluster structure.

Hayashi, Koby↗

Optical selection bias and projection effects in stacked galaxy cluster weak lensing

ABSTRACT Cosmological constraints from current and upcoming galaxy cluster surveys are limited by the accuracy of cluster mass calibration. In particular, optically identified galaxy clusters are prone to selection effects that can bias the weak lensing mass calibration. We investigate the selection bias of the stacked cluster lensing signal associated with optically selected clusters, using clusters identified by the redMaPPer algorithm in the Buzzard simulations as a case study. We find that at a given cluster halo mass, the residuals of redMaPPer richness and weak lensing signal are positively correlated. As a result, for a given richness selection, the stacked lensing signal is biased high compared with what we would expect from the underlying halo mass probability distribution. The cluster lensing selection bias can thus lead to overestimated mean cluster mass and biased cosmology results. We show that the lensing selection bias exhibits a strong scale dependence and is approximately 20–60 per cent for ΔΣ at large scales. This selection bias largely originates from spurious member galaxies within ±20–60 $h^{-1}\, \rm Mpc$ along the line of sight, highlighting the importance of quantifying projection effects associated with the broad redshift distribution of member galaxies in photometric cluster surveys. While our results qualitatively agree with those in the literature, accurate quantitative modelling of the selection bias is needed to achieve the goals of cluster lensing cosmology and will require synthetic catalogues covering a wide range of galaxy–halo connection models.

79 ASTRONOMY AND ASTROPHYSICS↗

Free spectator nucleons in ultracentral relativistic heavy-ion collisions as a probe of neutron skin

Besides the yield ratio of free spectator neutrons produced in ultracentral 96 Zr + 96 Zr to 96 Ru + 96 Ru collisions, we propose that the yield ratio N n /N p of free spectator neutrons to protons in a single collision system at the BNL Relativistic Heavy Ion Collider or the CERN Large Hadron Collider can be a more sensitive probe of the neutron-skin thickness Δrnp and the slope parameter L of the symmetry energy. Here, the idea is demonstrated based on the proton and neutron density distributions of colliding nuclei obtained from Skyrme-Hartree-Fock-Bogolyubov calculations, and on a Glauber model that provides information of spectator matter. The final spectator particles are produced from direct emission, clusterization by a minimum spanning tree algorithm or a Wigner function approach, and deexcitation of heavy clusters by gemini. A larger Δr np associated with a larger L value increases the isospin asymmetry of spectator matter and thus leads to a larger N n /N p , especially in ultracentral collisions where the multiplicity of free nucleons are free from the uncertainties of cluster formation and deexcitation. We have further shown that the double ratio of N n /N p in isobaric collision systems or in collisions by isotopes helps to cancel the detecting efficiency for protons. Effects from nuclear deformation and electromagnetic excitation are studied, and they are found to be subdominant compared to the expected sensitivity to Δr np .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Cobweb/3: A portable implementation

An algorithm is examined for data clustering and incremental concept formation. An overview is given of the Cobweb/3 system and the algorithm on which it is based, as well as the practical details of obtaining and running the system code. The implementation features a flexible user interface which includes a graphical display of the concept hierarchies that the system constructs.

Mckusick, Kathleen↗