NSPACE - An unsupervised clustering algorithm based on discretized marginal distributions
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We introduce a general type of optimization algorithm which infers data models relating two different but intertwined types of information about each of a set of objects.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.
Cluster Compression Algorithm (CCA) preprocesses Landsat image data immediately following satellite data sensor (receiver). Data are reduced by extracting pertinent image features and compressing this result into concise format for transmission to ground station. This results in narrower transmission bandwidth, increased data-communication efficiency, and reduced computer time in reconstructing and analyzing image. Similar technique could be applied to other types of recorded data to cut costs of transmitting, storing, distributing, and interpreting complex information.
ABSTRACT Galaxy clusters enable unique opportunities to study cosmology, dark matter, galaxy evolution, and strongly lensed transients. We here present a new cluster-finding algorithm, CluMPR (Clusters from Masses and Photometric Redshifts), that exploits photometric redshifts (photo-z’s) as well as photometric stellar mass measurements. CluMPR uses a 2D binary search tree to search for overdensities of massive galaxies with similar redshifts on the sky and then probabilistically assigns cluster membership by accounting for photo-z uncertainties. We leverage the deep DESI Legacy Survey grzW1W2 imaging over one-third of the sky to create a catalogue of $\sim 300\, 000$ galaxy cluster candidates out to z = 1, including tabulations of member galaxies and estimates of each cluster’s total stellar mass. Compared to other methods, CluMPR is particularly effective at identifying clusters at the high end of the redshift range considered (z = 0.75–1), with minimal contamination from low-mass groups. These characteristics make it ideal for identifying strongly lensed high-redshift supernovae and quasars that are powerful probes of cosmology, dark matter, and stellar astrophysics. As an example application of this cluster catalogue, we present a catalogue of candidate wide-angle strongly lensed quasars in Appendix C. The nine best candidates identified from this sample include two known lensed quasar systems and a possible changing-look lensed QSO with SDSS spectroscopy. All code and catalogues produced in this work are publicly available (see Data Availability).
This research focuses on a new neural network scene classification technique. The task is to identify scene elements in Advanced Very High Resolution Radiometry (AVHRR) data from three scene types: polar, desert and smoke from biomass burning in South America (smoke). The ultimate goal of this research is to design and implement a computer system which will identify the clouds present on a whole-Earth satellite view as a means of tracking global climate changes. Previous research has reported results for rule-based systems (Tovinkere et at 1992, 1993) for standard back propagation (Watters et at. 1993) and for a hierarchical approach (Corwin et al 1994) for polar data. This research uses a hierarchical neural network with don't care conditions and applies this technique to complex scenes. A hierarchical neural network consists of a switching network and a collection of leaf networks. The idea of the hierarchical neural network is that it is a simpler task to classify a certain pattern from a subset of patterns than it is to classify a pattern from the entire set. Therefore, the first task is to cluster the classes into groups. The switching, or decision network, performs an initial classification by selecting a leaf network. The leaf networks contain a reduced set of similar classes, and it is in the various leaf networks that the actual classification takes place. The grouping of classes in the various leaf networks is determined by applying an iterative clustering algorithm. Several clustering algorithms were investigated, but due to the size of the data sets, the exhaustive search algorithms were eliminated. A heuristic approach using a confusion matrix from a lightly trained neural network provided the basis for the clustering algorithm. Once the clusters have been identified, the hierarchical network can be trained. The approach of using don't care nodes results from the difficulty in generating extremely complex surfaces in order to separate one class from all of the others. This approach finds pairwise separating surfaces and forms the more complex separating surface from combinations of simpler surfaces. This technique both reduces training time and improves accuracy over the previously reported results. Accuracies of 97.47%, 95.70%, and 99.05% were achieved for the polar, desert and smoke data sets.
Explore the source record for details and available documents.
The Cluster Compression Algorithm (CCA), which was developed to reduce costs associated with transmitting, storing, distributing, and interpreting LANDSAT multispectral image data is described. The CCA is a preprocessing algorithm that uses feature extraction and data compression to more efficiently represent the information in the image data. The format of the preprocessed data enables simply a look-up table decoding and direct use of the extracted features to reduce user computation for either image reconstruction, or computer interpretation of the image data. Basically, the CCA uses spatially local clustering to extract features from the image data to describe spectral characteristics of the data set. In addition, the features may be used to form a sequence of scalar numbers that define each picture element in terms of the cluster features. This sequence, called the feature map, is then efficiently represented by using source encoding concepts. Various forms of the CCA are defined and experimental results are presented to show trade-offs and characteristics of the various implementations. Examples are provided that demonstrate the application of the cluster compression concept to multi-spectral images from LANDSAT and other sources.
A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.
The state of hydration of a macromolecular system regulates a plethora of different properties of such a system. In this article, we develop a novel machine learning (ML) approach, based on the unsupervised clustering algorithm, for probing the hydration behavior of the {N(CH 3 ) 3 } + functional group of the PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] polyelectrolyte (PE) brush system. The PE brushes and the brush-supported water molecules and counterions (chloride ions) are first described using all-atom molecular dynamics (MD) simulations. The simulation data is subsequently used in our ML framework to identify that (1) the {N(CH 3 ) 3 } + functional groups of the PMETAC brushes have two distinct hydration states with one state (state 1) being characterized by less structured water molecules and the other state (state 2) being characterized by more structured water molecules and (2) an enhancement in the brush grafting density leads to the progressive dissapparenace of state 2. An increase in the grafting density increases the number of chloride counterions in a given volume around the {N(CH 3 ) 3 } + functional group and increases the number of shared water molecules between the {N(CH 3 ) 3 } + and Cl - . The chloride counterions are associated with a hydration layer with much less structured water molecules. Therefore, with an increase in the grafting density, an increase in the percentage of shared water molecules leads to the prevalence of the hydration state [of the {N(CH 3 ) 3 } + moiety] with less structured water molecules. Finally, we explain how the present findings are commensurate with two key previous related results, namely a significantly large chloride ion mobility inside the PMETAC brush layer and the {N(CH 3 ) 3 } + -Cl - average distance remaining independent of the PMETAC brush grafting density. Furthermore, we anticipate that the combined ML-MD-simulation approach proposed in this study can be adapted to probe other soft matter systems to reveal new insights of the underlying mechanisms of emergent phenomenon.
Feature extraction and data compression of LANDSAT data is accomplished by BCCA program which reduces costs associated with transmitting, storing, distributing, and interpreting multispectral image data. Algorithm uses spatially local clustering to extract features from image data to describe spectral characteristics of data set. Approach requires only simple repetitive computations, and parallel processing can be used for very high data rates. Program is written in FORTRAN IV for batch execution and has been implemented on SEL 32/55.
When simulating a lattice system near its critical temperature, local algorithms for modeling the system’s evolution can introduce very large autocorrelation times into sampled data. Here, this critical slowing down places restrictions on the analysis that can be completed in a timely manner of the behavior of systems around the critical point. Because it is often desirable to study such systems around this point, a new algorithm must be introduced. Therefore, we turn to cluster algorithms, such as the Swendsen–Wang algorithm and the Wolff clustering algorithm. They incorporate global updates which generate new lattice configurations with little correlation to previous states, even near the critical point. We look to accelerate the rate at which these algorithm are capable of running by implementing and benchmarking a parallel implementation of each algorithm designed to run on GPUs under NVIDIA’s CUDA framework. A 17 and 90 fold increase in the computational rate was, respectively, experienced when measured against the equivalent algorithm implemented in serial code.
An insight into the characteristics which determine the performance of a clustering algorithm is presented. In order for the techniques which are examined to accurately cluster data, two conditions must be simultaneously satisfied. First the data must have a particular structure, and second the parameters chosen for the clustering algorithm must be correct. By examining the structure of the data from the Cl flight line, it is clear that no single set of parameters can be used to accurately cluster all the different crops. The effectiveness of either a noniterative or iterative clustering algorithm to accurately cluster data representative of the Cl flight line is questionable. Thus extensive a prior knowledge is required in order to use cluster analysis in its present form for applications like assisting in the definition of field boundaries and evaluating the homogeneity of a field. New or modified techniques are necessary for clustering to be a reliable tool.
Abstract The Milky Way has accreted many ultra-faint dwarf galaxies (UFDs), and stars from these galaxies can be found throughout our Galaxy today. Studying these stars provides insight into galaxy formation and early chemical enrichment, but identifying them is difficult. Clustering stellar dynamics in 4D phase space ( E , L z , J r , J z ) is one method of identifying accreted structure that is currently being utilized in the search for accreted UFDs. We produce 32 simulated stellar halos using particle tagging with the Caterpillar simulation suite and thoroughly test the abilities of different clustering algorithms to recover tidally disrupted UFD remnants. We perform over 10,000 clustering runs, testing seven clustering algorithms, roughly twenty hyperparameter choices per algorithm, and six different types of data sets each with up to 32 simulated samples. Of the seven algorithms, HDBSCAN most consistently balances UFD recovery rates and cluster realness rates. We find that, even in highly idealized cases, the vast majority of clusters found by clustering algorithms do not correspond to real accreted UFD remnants and we can generally only recover 6% of UFDs remnants at best. These results focus exclusively on groups of stars from UFDs, which have weak dynamic signatures compared to the background of other stars. The recoverable UFD remnants are those that accreted recently, z accretion ≲ 0.5. Based on these results, we make recommendations to help guide the search for dynamically linked clusters of UFD stars in observational data. We find that real clusters generally have higher median energy and J r , providing a way to help identify real versus fake clusters. We also recommend incorporating chemical tagging as a way to improve clustering results.
The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.