Learning Associations between Features and Clusters: An Interpretable Deep Clustering Method
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Understanding flow traffic patterns in networks, such as the Internet or service provider networks, is crucial to improving their design and building them robustly. However, as networks grow and become more complex, it is increasingly cumbersome and challenging to study how the many flow patterns, sizes and the continually changing source-destination pairs in the network evolve with time. Here, we present Netostat, a visualization-based network analysis tool that uses visual representation and a mathematics framework to study and capture flow patterns, using graph theoretical methods such as clustering, similarity and difference measures. Netostat generates an interactive graph of all traffic patterns in the network, to isolate key elements that can provide insights for traffic engineering. We present results for U.S. and European research networks, ESnet and GEANT, demonstrating network state changes, to identify major flow trends, potential points of failure, and bottlenecks.
Feature extraction and data compression of LANDSAT data is accomplished by BCCA program which reduces costs associated with transmitting, storing, distributing, and interpreting multispectral image data. Algorithm uses spatially local clustering to extract features from image data to describe spectral characteristics of data set. Approach requires only simple repetitive computations, and parallel processing can be used for very high data rates. Program is written in FORTRAN IV for batch execution and has been implemented on SEL 32/55.
Here, we introduce a hybrid quantum-classical algorithm, the localized active space unitary selective coupled cluster singles and doubles (LAS-USCCSD) method. Derived from the localized active space unitary coupled cluster (LAS-UCCSD) method, LAS-USCCSD first performs a classical LASSCF calculation, then selectively identifies the most important parameters (cluster amplitudes used to build the multireference UCC ansatz) for restoring interfragment interaction energy using this reduced set of parameters with the variational quantum eigensolver method. We benchmark LAS-USCCSD against LAS-UCCSD by calculating the total energies of (H 2 ) 2 , (H 2 ) 4 , and trans-butadiene, and the magnetic coupling constant for a bimetallic compound [Cr 2 (OH) 3 (NH 3 ) 6 ] 3+ . For these systems, we find that LAS-USCCSD reduces the number of required parameters and thus the circuit depth by at least 1 order of magnitude, an aspect which is important for the practical implementation of multireference hybrid quantum-classical algorithms like LAS-UCCSD on near-term quantum computers.
We present a polynomial-scaling algorithm for the localized active space unitary selective coupled cluster singles and doubles (LAS-USCCSD) method. In this approach, cluster excitations are selected based on a threshold ϵ determined by the absolute gradients of the LAS-UCCSD energy with respect to cluster amplitudes. Using the generalized Wick’s theorem for multireference wave functions, we derive the gradient expression as a polynomial function of one-, two-, and three-body reduced density matrices and 1- and 2-electron integrals, valid for any multireference wave function. The resulting gradient implementation exhibits a memory scaling of 𝒪(N 6 ), with N spin orbitals in the combined active space of all fragments. The variational quantum eigensolver is used to optimize the selected cluster excitations on a quantum simulator. Furthermore, by plotting the energy error, defined as the difference between the LAS-USCCSD and corresponding CASCI energies, against the inverse cluster amplitude selection threshold (ϵ –1 ) for polyene chains containing 2 to 5 π-bond units, we establish a relationship between the energy error and the threshold. To further validate the accuracy of LAS-USCCSD, we computed the cis–trans isomerization energy of stilbene (a 20-qubit system) and the magnetic coupling constant of the tris-hydroxo-bridged chromium dimer [Cr 2 (OH) 3 (NH 3 ) 6 ] 3+ (evaluated as both 12- and 20-qubit systems) using the Qiskit-Qulacs simulator. Assessing such examples is important to determine the practical feasibility of quantum simulations for chemically realistic systems. Toward this goal, with the LAS-USCCSD algorithm we estimated the quantum resources required for simulating an active space of (30e,22o) in [Cr 2 (OH) 3 (NH 3 ) 6 ] 3+ , a size that remains beyond the reach of current quantum simulators for accurate treatment.
State preparation for quantum algorithms is crucial for achieving high accuracy in quantum chemistry and competing with classical algorithms. The localized active space–unitary coupled cluster (LAS–UCC) algorithm iteratively loads a fragment-based multireference wave function onto a quantum computer. Here, in this study, we compare two state preparation methods, quantum phase estimation (QPE) and direct initialization (DI), for each fragment. We test the two state preparation methods on three systems, ranging from a model system, a set of interacting hydrogen molecules, to more realistic chemical problems, like the C–C double bond breaking in transbutadiene and the spin ladder in a bimetallic system. We analyze the impact of QPE parameters, such as the number of ancilla qubits and Trotter steps, on the prepared state. We find a trade-off between the methods, where DI requires fewer resources for smaller fragments, while QPE is more efficient for larger fragments. Our resource estimates highlight the benefits of system fragmentation in state preparation for subsequent quantum chemical calculations. These findings have broad applications for preparing multireference quantum chemical wave functions on quantum circuits that can be used for realistic chemical applications.
Abstract Droplet-level interactions in clouds are often parameterized by a modified gamma fitted to a “global” droplet size distribution. Do “local” droplet size distributions of relevance to microphysical processes look like these average distributions? This paper describes an algorithm to search and classify characteristic size distributions within a cloud. The approach combines hypothesis testing, specifically, the Kolmogorov–Smirnov (KS) test, and a widely used class of machine learning algorithms for identifying clusters of samples with similar properties: density-based spatial clustering of applications with noise (DBSCAN) is used as the specific example for illustration. The two-sample KS test does not presume any specific distribution, is parameter free, and avoids biases from binning. Importantly, the number of clusters is not an input parameter of the DBSCAN-type algorithms but is independently determined in an unsupervised fashion. As implemented, it works on an abstract space from the KS test results, and hence spatial correlation is not required for a cluster. The method is explored using data obtained from the Holographic Detector for Clouds (HOLODEC) deployed during the Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) field campaign. The algorithm identifies evidence of the existence of clusters of nearly identical local size distributions. It is found that cloud segments have as few as one and as many as seven characteristic size distributions. To validate the algorithm’s robustness, it is tested on a synthetic dataset and successfully identifies the predefined distributions at plausible noise levels. The algorithm is general and is expected to be useful in other applications, such as remote sensing of cloud and rain properties. Significance Statement A typical cloud can have billions of drops spread over tens or hundreds of kilometers in space. Keeping track of the sizes, positions, and interactions of all of these droplets is impractical, and, as such, information about the relative abundance of large and small drops is typically quantified with a “size distribution.” Droplets in a cloud interact locally, however, so this work is motivated by the question of whether the cloud droplet size distribution is different in different parts of a cloud. A new method, based on hypothesis testing and machine learning, determines how many different size distributions are contained in a given cloud. This is important because the size distribution describes processes such as cloud droplet growth and light transmission through clouds.
By combining a local sky-fitting algorithm with a Fourier point-spread-function matching technique, nova outbursts have been searched for inside 54 of the globular clusters contained on the Ciardullo et al. (1987 and 1990) H-alpha survey frames of M 31. Over a mean effective survey time of about 2.0 years, no cluster exhibited a magnitude increase indicative of a nova explosion. If the cataclysmic variables (CVs) contained within globular clusters are similar to those found in the field, then these data imply that the overdensity of CVs within globulars is at least several times less than that of the high-luminosity X-ray sources. If tidal capture is responsible for the high density of hard binaries within globulars, then the probability of capturing condensed objects inside globular clusters may depend strongly on the mass of the remnant.
A major limitation of additive manufacturing (AM) processes is that local conditions of material deposition frequently lead to unintentional heterogeneities in microstructure and properties within a single component, despite nominally uniform process conditions. Up to now, there has been no way to a priori determine the distribution of these heterogeneities, requiring expensive trial-and-error approaches to fabrication, testing, and characterization. Here, a physics-based framework for creating a digital representation of the laser powder bed fusion (PBF) process is proposed to predict the variation in solidification behavior that leads to heterogeneous microstructures in an as-built part. By leveraging in situ process data stored in the part’s digital thread, the scan path and process parameters were input into a heat transfer model which predicted solidification data at the melt pool scale. A two-step unsupervised clustering algorithm was used to first cluster the local solidification conditions (12.5µm 3 voxels) and then to cluster the regional behavior on the scale of multiple scan passes and print layers (250µm 3 super-voxels). This process was used to identify regions with similar solidification characteristics for multiple locations in a Stainless Steel 316-L component. The corresponding as-built part was sectioned and characterized using electron backscatter diffraction (EBSD). Quantitative analysis of the pole figures confirmed that the predicted regions of heterogeneity in the solidification conditions corresponded with differences in the observed microstructure. In conclusion, this work shows a viable path for estimating the microstructural heterogeneity for additively manufactured parts to either limit microstructural variation throughout a part or to enable functionality-based variation of the microstructure.
This research is focused on the identification of cracking mechanisms for cement paste using acoustic emission data, recorded from compression and notched four-point bending tests. A procedure is developed for analyzing the data by employing an agglomerative hierarchical clustering method, an artificial neural network, and a ray-tracing source location algorithm. An agglomerative hierarchical clustering method is utilized to cluster the AE data from a compression test using frequency-dependent features. A neural network is trained using the compression test data and applied to the AE data emitted during the four-point bending test. The clustered data from the four-point bending test is localized using a ray-tracing algorithm. Based on the occurrence and locations of the clustered events and signal feature analyses, potential cracking mechanisms are identified and assigned.
Abstract Detailed chemical kinetic mechanisms are necessary for resolving many important chemical processes. As the chemistry of smaller molecules has become better grounded and quantum chemistry calculations have become cheaper, kineticists have become interested in constructing progressively larger kinetic mechanisms to model increasingly complex chemical processes. These large kinetic mechanisms prove incredibly difficult to refine and time‐consuming to interpret. Traditional sensitivity analysis on a large mechanism can range from inconvenient to practically impossible without special techniques to reduce the computational cost. We first present a new time‐local sensitivity analysis we term transitory sensitivity analysis. Transitory sensitivity analysis is demonstrated in an example to accurately identify traditionally sensitive reactions at an 18,000x speed up over traditional sensitivities. By fusing transitory sensitivity analysis with more traditional time‐local branching, pathway, and cluster analyses, we develop an algorithm for efficient automatic mechanism analysis. This automatic mechanism analysis at a time point is able to identify the reactions a target is most sensitive to using transitory sensitivity analysis and then propose hypotheses why the reaction might be sensitive using branching, pathway, and cluster analyses. We implement these algorithms within the reaction mechanism simulator (RMS) package, which enables us to report the automatic mechanism analysis results in highly readable text formats and in molecular flux diagrams.
The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.
Protection against dc faults is one of the main technical hurdles faced when operating converter-based HVdc systems. Protection becomes even more challenging for multi-terminal dc (MTdc) systems with more than two terminals/converter stations. In this paper, a hybrid primary fault detection algorithm for MTdc systems is proposed to detect a broad range of failures. Sensor measurements, i.e., line currents and dc reactor voltages measured at local terminals, are first processed by a top-level context clustering algorithm. For each cluster, the best fault detector is selected among a detector pool according to a rule resulting from a learning algorithm. The detector pool consists of several existing detection algorithms, each performing differently across fault scenarios. The proposed hybrid primary detection algorithm: i) offers superior performance compared to an individual detector through a data-driven approach; ii) detects all major fault types including pole-to-pole (P2P), pole-to-ground (P2G), and external dc faults; iii) identifies faults with various fault locations and impedances; iv) is more robust to noisy sensor measurements compared to existing methods; v) does not require exhaustive simulation and sampling for training the model. Performance and effectiveness of the proposed algorithm are evaluated and verified based on time-domain simulations in the PSCAD/EMTDC software environment. The results confirm satisfactory operation, accuracy, and detection speed of the proposed algorithm under various fault scenarios.
We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.
This work focuses on incorporating pairwise constraints into a spectral clustering algorithm. A new constrained spectral clustering method is proposed, as well as an active constraint acquisition technique and a heuristic for parameter selection. We demonstrate that our constrained spectral clustering method, CSC, works well when the data exhibits what we term local proximity structure.
The building stock in the United States (U.S.) varies significantly as a function of several macro variables such as: climate, building type, vintage, and density. These variables change across the U.S. and can also significantly impact energy usage of the individual buildings and overall stock. For example, the square foot density and building type varies by several orders of magnitude from Manhattan to the eastern plains of Colorado. The diversity in energy use of the building stock of different areas of the U.S. is significant, and as a result, analyses that require localized results need to consider the relevant geography and the current makeup of the building stock. This document discusses the development and implementation of a stock clustering algorithm that produces a technically rigorous, consistent, and repeatable collection of geographies which are used as the basis for localized analysis. This framework considers the impact of built environment density, diversity, and climate in creating groupings of counties that create a far more nuanced analysis framework than national averages. Clustering of counties together represents a similarity of building characteristics and climate zone.
Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.
In this article, a parallel algorithm which applies Givens rotations to selectively annihilate k(k + 1)/2 nonzero elements from two k x n(k not more than n) upper trapezoidal submatrices is described. The new algorithm is suitable for implementation on either a pair of directly connected local-memory processors or two clusters of multiple tightly-coupled processors. Analyses show that in both cases the proposed algorithms achieve optimal speed-up by balancing the work load distribution and masking interprocessor or intercluster communication by computation if k is much small than n. In the context of solving large scale least squares problems, this submatrix merging step is repetitively needed during the entire computation and, furthermore, there are usually many pairs of such submatrices to be merged with each submatrix stored in the memory of a processor or a cluster of processors. The proposed algorithm can be applied to each pair of submatrices concurrently, and thus parallelizes an important step in solving the least squares problems.