Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dynamic clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automatic Traffic Queue-End Identification using Location-Based Waze User Reports

Traffic queues, especially queues caused by non-recurrent events such as incidents, are unexpected to high-speed drivers approaching the end of queue (EOQ) and become safety concerns. Though the topic has been extensively studied, the identification of EOQ has been limited by the spatial-temporal resolution of traditional data sources. This study explores the potential of location-based crowdsourced data, specifically Waze user reports. It presents a dynamic clustering algorithm that can group the location-based reports in real time and identify the spatial-temporal extent of congestion as well as the EOQ. The algorithm is a spatial-temporal extension of the density-based spatial clustering of applications with noise (DBSCAN) algorithm for real-time streaming data with an adaptive threshold selection procedure. Here, the proposed method was tested with 34 traffic congestion cases in the Knoxville, Tennessee area of the United States. It is demonstrated that the algorithm can effectively detect spatial-temporal extent of congestion based on Waze report clusters and identify EOQ in real-time. The Waze report-based detection are compared to the detection based on roadside sensor data. The results are promising: The EOQ identification time of Waze is similar to the EOQ detection time of traffic sensor data, with only 1.1 min difference on average. In addition, Waze generates 1.9 EOQ detection points every mile, compared to 1.8 detection points generated by traffic sensor data, suggesting the two data sources are comparable in respect of reporting frequency. The results indicate that Waze is a valuable complementary source for EOQ detection where no traffic sensors are installed.

99 GENERAL AND MISCELLANEOUS↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Investigation of plasmon relaxation mechanisms using nonadiabatic molecular dynamics

Hot carriers generated from the decay of plasmon excitation can be harvested to drive a wide range of physical or chemical processes. However, their generation efficiency is limited by the concomitant phonon-induced relaxation processes by which the energy in excited carriers is transformed into heat. However, simulations of dynamics of nanoscale clusters are challenging due to the computational complexity involved. Here, in this paper, we adopt our newly developed Trajectory Surface Hopping (TSH) nonadiabatic molecular dynamics algorithm to simulate plasmon relaxation in Au 20 clusters, taking the atomistic details into account. The electronic properties are treated within the Linear Response Time-Dependent Tight-binding Density Functional Theory (LR-TDDFTB) framework. The relaxation of plasmon due to coupling to phonon modes in Au 20 beyond the Born–Oppenheimer approximation is described by the TSH algorithm. The numerically efficient LR-TDDFTB method allows us to address a dense manifold of excited states to ensure the inclusion of plasmon excitation. Starting from the photoexcited plasmon states in Au 20 cluster, we find that the time constant for relaxation from plasmon excited states to the lowest excited states is about 2.7 ps, mainly resulting from a stepwise decay process caused by low-frequency phonons of the Au 20 cluster. Furthermore, our simulations show that the lifetime of the phonon-induced plasmon dephasing process is ~10.4 fs and that such a swift process can be attributed to the strong nonadiabatic effect in small clusters. Our simulations demonstrate a detailed description of the dynamic processes in nanoclusters, including plasmon excitation, hot carrier generation from plasmon excitation dephasing, and the subsequent phonon-induced relaxation process.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Predicting images for the dynamics of stellar clusters ( π-DOC ): a deep learning framework to predict mass, distance, and age of globular clusters

ABSTRACT Dynamical mass estimates of simple systems such as globular clusters (GCs) still suffer from up to a factor of 2 uncertainty. This is primarily due to the oversimplifications of standard dynamical models that often neglect the effects of the long-term evolution of GCs. Here, we introduce a new approach to measure the dynamical properties of GCs, based on the combination of a deep-learning framework and the large amount of data from direct N-body simulations. Our algorithm, π-DOC (Predicting Images for the Dynamics Of stellar Clusters) is composed of two convolutional networks, trained to learn the non-trivial transformation between an observed GC luminosity map and its associated mass distribution, age, and distance. The training set is made of V-band luminosity and mass maps constructed as mock observations from N-body simulations. The tests on π-DOC demonstrate that we can predict the mass distribution with a mean error per pixel of 27 per cent, and the age and distance with an accuracy of 1.5 Gyr and 6 kpc, respectively. In turn, we recover the shape of the mass-to-light profile and its global value with a mean error of 12 per cent, which implies that we efficiently trace mass segregation. A preliminary comparison with observations indicates that our algorithm is able to predict the dynamical properties of GCs within the limits of the training set. These encouraging results demonstrate that our deep-learning framework and its forward modelling approach can offer a rapid and adaptable tool competitive with standard dynamical models.

Chardin, Jonathan↗

Distribution Cutoff for Clusters near the Gel Point

The mechanical and dynamic properties of developing networks near the gel point are susceptible to the distribution of clusters coexisting with percolating networks. The distribution of cluster numbers follows a broad power law, wrapped by a cutoff function that rapidly decays at a characteristic size. The form of the cutoff function has been speculated based on known results from lattice percolation and, in certain cases, solved. We obtained this cutoff function from simulated dynamic clusters of polymeric precursor chains using a hybrid Monte Carlo algorithm. The results obtained from three different precursor chain lengths are consistent with each other and are consistent with the expectation from lattice percolation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electromagnetic Transient (EMT) Simulation Algorithm for Evaluation of Photovoltaic (PV) Generation Systems

Use of inverter-based resources facilitating renewable energy resources such as photovoltaic (PV) generation is increasing rapidly with decreasing costs and reduced emissions associated. To accommodate such rapid growth of inverter-based resources like PV systems, electromagnetic transient (EMT) simulation models of both PV systems and grids are required to analyze the interaction of PVs in the grid (like the post-event analysis). In addition, the EMT simulation would help with the planning of future power grid with a large number of PVs as well as other inverter-based distributed generation systems. In this paper, the EMT simulation models of PV systems and grids are developed based on the differential algebraic equations (DAEs) representing their EMT dynamics. Furthermore, advanced simulation algorithms including numerical stiffness-based hybrid discretization, DAE clustering and aggregation, multi-order integration, and matrix splitting approaches are applied to accelerate the EMT simulation. The proposed algorithm was applied to 125 PV inverters within 52-bus medium-voltage (MV) distribution grid.

Choi, Jongchan↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

Possibilities and Limitations of Kinematically Identifying Stars from Accreted Ultra-faint Dwarf Galaxies

Abstract The Milky Way has accreted many ultra-faint dwarf galaxies (UFDs), and stars from these galaxies can be found throughout our Galaxy today. Studying these stars provides insight into galaxy formation and early chemical enrichment, but identifying them is difficult. Clustering stellar dynamics in 4D phase space ( E , L z , J r , J z ) is one method of identifying accreted structure that is currently being utilized in the search for accreted UFDs. We produce 32 simulated stellar halos using particle tagging with the Caterpillar simulation suite and thoroughly test the abilities of different clustering algorithms to recover tidally disrupted UFD remnants. We perform over 10,000 clustering runs, testing seven clustering algorithms, roughly twenty hyperparameter choices per algorithm, and six different types of data sets each with up to 32 simulated samples. Of the seven algorithms, HDBSCAN most consistently balances UFD recovery rates and cluster realness rates. We find that, even in highly idealized cases, the vast majority of clusters found by clustering algorithms do not correspond to real accreted UFD remnants and we can generally only recover 6% of UFDs remnants at best. These results focus exclusively on groups of stars from UFDs, which have weak dynamic signatures compared to the background of other stars. The recoverable UFD remnants are those that accreted recently, z accretion ≲ 0.5. Based on these results, we make recommendations to help guide the search for dynamically linked clusters of UFD stars in observational data. We find that real clusters generally have higher median energy and J r , providing a way to help identify real versus fake clusters. We also recommend incorporating chemical tagging as a way to improve clustering results.

79 ASTRONOMY AND ASTROPHYSICS↗

Ab Initio Study of the Beryllium Isotopes 7 Be to 12 Be

We present a systematic ab initio study of the low-lying states in beryllium isotopes from 7 Be to 12 Be using nuclear lattice effective field theory with the N 3 ⁢LO interaction. Our calculations achieve good agreement with experimental data for energies, radii, and electromagnetic properties. We introduce a novel, model-independent method to quantify nuclear shapes, uncovering a distinct pattern in the interplay between positive and negative parity states across the isotopic chain. By combining Monte Carlo sampling of the many-body density operator with a novel nucleon-grouping algorithm, the prominent two-center cluster structures, the emergence of one-neutron halo, complex nuclear molecular dynamics such as 𝜋 orbital and 𝜎 orbital, emerge naturally.

binding energy & masses↗

Substructure in the stellar halo near the Sun: I. Data-driven clustering in integrals-of-motion space

Context. Merger debris is expected to populate the stellar haloes of galaxies. In the case of the Milky Way, this debris should be apparent as clumps in a space defined by the orbital integrals of motion of the stars. Aims. Our aim is to develop a data-driven and statistics-based method for finding these clumps in integrals-of-motion space for nearby halo stars and to evaluate their significance robustly. Methods. We used data from Gaia EDR3, extended with radial velocities from ground-based spectroscopic surveys, to construct a sample of halo stars within 2.5 kpc from the Sun. We applied a hierarchical clustering method that makes exhaustive use of the single linkage algorithm in three-dimensional space defined by the commonly used integrals of motion energy E, together with two components of the angular momentum, L z and L ⊥ . To evaluate the statistical significance of the clusters, we compared the density within an ellipsoidal region centred on the cluster to that of random sets with similar global dynamical properties. By selecting the signal at the location of their maximum statistical significance in the hierarchical tree, we extracted a set of significant unique clusters. By describing these clusters with ellipsoids, we estimated the proximity of a star to the cluster centre using the Mahalanobis distance. Additionally, we applied the HDBSCAN clustering algorithm in velocity space to each cluster to extract subgroups representing debris with different orbital phases. Results. Our procedure identifies 67 highly significant clusters (> 3σ), containing 12% of the sources in our halo set, and 232 subgroups or individual streams in velocity space. In total, 13.8% of the stars in our data set can be confidently associated with a significant cluster based on their Mahalanobis distance. Inspection of the hierarchical tree describing our data set reveals a complex web of relations between the significant clusters, suggesting that they can be tentatively grouped into at least six main large structures, many of which can be associated with previously identified halo substructures, and a number of independent substructures. This preliminary conclusion is further explored in a companion paper, in which we also characterise the substructures in terms of their stellar populations. Conclusions. Our method allows us to systematically detect kinematic substructures in the Galactic stellar halo with a data-driven and interpretable algorithm. The list of the clusters and the associated star catalogue are provided in two tables available at the CDS.

79 ASTRONOMY AND ASTROPHYSICS↗

A k-mer based approach for classifying viruses without taxonomy identifies viral associations in human autism and plant microbiomes

Viruses are an underrepresented taxa in the study and identification of microbiome constituents; however, they play an essential role in health, microbiome regulation, and transfer of genetic material. Only a few thousand viruses have been isolated, sequenced, and assigned a taxonomy, which limits the ability to identify and quantify viruses in the microbiome. Additionally, the vast diversity of viruses represents a challenge for classification, not only in constructing a viral taxonomy, but also in identifying similarities between a virus’ genotype and its phenotype. However, the diversity of viral sequences can be leveraged to classify their sequences in metagenomic and metatranscriptomic samples, even if they do not have a taxonomy. To identify and quantify viruses in transcriptomic and genomic samples, we developed a dynamic programming algorithm for creating a classification tree out of 715,672 metagenome viruses. To create the classification tree, we clustered proportional similarity scores generated from the k-mer profiles of each of the metagenome viruses to create a database of metagenomic viruses. The resulting Kraken2 database of the metagenomic viruses can be found here: https://www.osti.gov/biblio/1615774 and is compatible with Kraken2. We then integrated the viral classification database with databases created with genomes from NCBI for use with ParaKraken (a parallelized version of Kraken provided in Supplemental Zip 1), a metagenomic/transcriptomic classifier. To illustrate the breadth of our utility for classifying metagenome viruses, we analyzed data from a plant metagenome study identifying genotypic and compartment specific differences between two Populus genotypes in three different compartments. We also identified a significant increase in abundance of eight viral sequences in post mortem brains in a human metatranscriptome study comparing Autism Spectrum Disorder patients and controls. We also show the potential accuracy for classifying viruses by utilizing both the JGI and NCBI viral databases to identify the uniqueness of viral sequences. Finally, we validate the accuracy of viral classification with NCBI databases containing viruses with taxonomy to identify pathogenic viruses in known COVID-19 and cassava brown streak virus infection samples. Our method represents the compulsory first step in better understanding the role of viruses in the microbiome by allowing for a more complete identification of sequences without taxonomy. Better classification of viruses will improve identifying associations between viruses and their hosts as well as viruses and other microbiome members. Despite the lack of taxonomy, this database of metagenomic viruses can be used with any tool that utilizes a taxonomy, such as Kraken, for accurate classification of viruses.

59 BASIC BIOLOGICAL SCIENCES↗

A scalable variational method for estimating the latent infection-rate field of an outbreak

In this paper, we explore whether the infection-rate of a disease can serve as a robust monitoring variable in epidemiological surveillance algorithms. The infection-rate is dependent on population mixing patterns that do not vary erratically day-to-day; in contrast, daily case-counts used in contemporary surveillance algorithms are corrupted by reporting errors. The technical challenge lies in estimating the latent infection-rate from case-counts. Here we devise a Bayesian method to estimate the infection-rate across multiple adjoining areal units, and then use it, via an anomaly detector, to discern a change in epidemiological dynamics. We extend an existing model for estimating the infection-rate in an areal unit by incorporating a Markov random field model, so that we may estimate infection-rates across multiple areal units, while preserving spatial correlations observed in the epidemiological dynamics. To carry out the high-dimensional Bayesian inverse problem, we develop an implementation of mean-field variational inference specific to the infection model and integrate it with the random field model to incorporate correlations across counties. The method is tested on estimating the COVID-19 infection-rates across all 33 counties in New Mexico using data from the summer of 2020, and then employing them to detect the arrival of the Fall 2020 COVID-19 wave. We perform the detection using a temporal algorithm that is applied county-by-county. We also show how the infection-rate field can be used to cluster counties with similar epidemiological dynamics.

60 APPLIED LIFE SCIENCES↗

Combined Machine Learning and Molecular Dynamics Reveal Two States of Hydration of a Single Functional Group of Cationic Polymeric Brushes

The state of hydration of a macromolecular system regulates a plethora of different properties of such a system. In this article, we develop a novel machine learning (ML) approach, based on the unsupervised clustering algorithm, for probing the hydration behavior of the {N(CH 3 ) 3 } + functional group of the PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] polyelectrolyte (PE) brush system. The PE brushes and the brush-supported water molecules and counterions (chloride ions) are first described using all-atom molecular dynamics (MD) simulations. The simulation data is subsequently used in our ML framework to identify that (1) the {N(CH 3 ) 3 } + functional groups of the PMETAC brushes have two distinct hydration states with one state (state 1) being characterized by less structured water molecules and the other state (state 2) being characterized by more structured water molecules and (2) an enhancement in the brush grafting density leads to the progressive dissapparenace of state 2. An increase in the grafting density increases the number of chloride counterions in a given volume around the {N(CH 3 ) 3 } + functional group and increases the number of shared water molecules between the {N(CH 3 ) 3 } + and Cl - . The chloride counterions are associated with a hydration layer with much less structured water molecules. Therefore, with an increase in the grafting density, an increase in the percentage of shared water molecules leads to the prevalence of the hydration state [of the {N(CH 3 ) 3 } + moiety] with less structured water molecules. Finally, we explain how the present findings are commensurate with two key previous related results, namely a significantly large chloride ion mobility inside the PMETAC brush layer and the {N(CH 3 ) 3 } + -Cl - average distance remaining independent of the PMETAC brush grafting density. Furthermore, we anticipate that the combined ML-MD-simulation approach proposed in this study can be adapted to probe other soft matter systems to reveal new insights of the underlying mechanisms of emergent phenomenon.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗