Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Understanding nanoscale structural distortions in Pb(Zr 0.2 Ti 0.8 )O 3 by utilizing X-ray nanodiffraction and clustering algorithm analysis

Hard X-ray nanodiffraction provides a unique nondestructive technique to quantify local strain and structural inhomogeneities at nanometer length scales. However, sample mosaicity and phase separation can result in a complex diffraction pattern that can make it challenging to quantify nanoscale structural distortions. In this work, a k-means clustering algorithm was utilized to identify local maxima of intensity by partitioning diffraction data in a three-dimensional feature space of detector coordinates and intensity. This technique has been applied to X-ray nanodiffraction measurements of a patterned ferroelectric PbZr 0.2 Ti 0.8 O 3 sample. The analysis reveals the presence of two phases in the sample with different lattice parameters. A highly heterogeneous distribution of lattice parameters with a variation of 0.02 Å was also observed within one ferroelectric domain. This approach provides a nanoscale survey of subtle structural distortions as well as phase separation in ferroelectric domains in a patterned sample.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An evaluation of ISOCLS and CLASSY clustering algorithms for forest classification in northern Idaho

Both the iterative self-organizing clustering system (ISOCLS) and the CLASSY algorithms were applied to forest and nonforest classes for one 1:24,000 quadrangle map of northern Idaho and the classification and mapping accuracies were evaluated with 1:30,000 color infrared aerial photography. Confusion matrices for the two clustering algorithms were generated and studied to determine which is most applicable to forest and rangeland inventories in future projects. In an unsupervised mode, ISOCLS requires many trial-and-error runs to find the proper parameters to separate desired information classes. CLASSY tells more in a single run concerning the classes that can be separated, shows more promise for forest stratification than ISOCLS, and shows more promise for consistency. One major drawback to CLASSY is that important forest and range classes that are smaller than a minimum cluster size will be combined with other classes. The algorithm requires so much computer storage that only data sets as small as a quadrangle can be used at one time.

Werth, L. F.↗

Classification of posture maintenance data with fuzzy clustering algorithms

Sensory inputs from the visual, vestibular, and proprioreceptive systems are integrated by the central nervous system to maintain postural equilibrium. Sustained exposure to microgravity causes neurosensory adaptation during spaceflight, which results in decreased postural stability until readaptation occurs upon return to the terrestrial environment. Data which simulate sensory inputs under various sensory organization test (SOT) conditions were collected in conjunction with Johnson Space Center postural control studies using a tilt-translation device (TTD). The University of West Florida applied the fuzzy c-meams (FCM) clustering algorithms to this data with a view towards identifying various states and stages of subjects experiencing such changes. Feature analysis, time step analysis, pooling data, response of the subjects, and the algorithms used are discussed.

Bezdek, James C.↗

The CLASSY clustering algorithm: Description, evaluation, and comparison with the iterative self-organizing clustering system (ISOCLS)

A clustering method, CLASSY, was developed, which alternates maximum likelihood iteration with a procedure for splitting, combining, and eliminating the resulting statistics. The method maximizes the fit of a mixture of normal distributions to the observed first through fourth central moments of the data and produces an estimate of the proportions, means, and covariances in this mixture. The mathematical model which is the basic for CLASSY and the actual operation of the algorithm is described. Data comparing the performances of CLASSY and ISOCLS on simulated and actual LACIE data are presented.

Lennington, R. K.↗

Minijet clustering algorithm using transverse-momentum seeds in high-energy nuclear collisions

We propose an algorithm to detect mini-jet clusters in high-energy nuclear collisions, by selecting a high-transverse-momentum (pT) particle as a seed and assigning a clustering radius (R) in the pseudorapidity and azimuthal-angle space. Our PYTHIA simulations for p+p collisions show that a scheme with a seeding p T of around 0.5 GeV/c and R of approximately 0.6 satisfactorily identifies mini-jet clusters. The correlation between clusters obtained in PYTHIA calculations using the algorithm exhibits the proper behavior of hard-scattering-like processes, suggesting its usefulness in isolating mini-jet-like clusters from non-hard-scattering soft processes when applied to actual nuclear-collision data, thereby allowing a closer examination of both the mini-jet and the soft mechanisms.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Classification of posture maintenance data with fuzzy clustering algorithms

Sensory inputs from the visual, vestibular, and proprioreceptive systems are integrated by the central nervous system to maintain postural equilibrium. Sustained exposure to microgravity causes neurosensory adaptation during spaceflight, which results in decreased postural stability until readaptation occurs upon return to the terrestrial environment. Data which simulate sensory inputs under various conditions were collected in conjunction with JSC postural control studies using a Tilt-Translation Device (TTD). The University of West Florida proposed applying the Fuzzy C-Means Clustering (FCM) Algorithms to this data with a view towards identifying various states and stages. Data supplied by NASA/JSC were submitted to the FCM algorithms in an attempt to identify and characterize cluster substructure in a mixed ensemble of pre- and post-adaptational TTD data. Following several unsuccessful trials with FCM using a full 11 dimensional data set, a set of two channels (features) were found to enable FCM to separate pre- from post-adaptational TTD data. The main conclusions are that: (1) FCM seems able to separate pre- from post-TTD subject no. 2 on the one trial that was used, but only in certain subintervals of time; and (2) Channels 2 (right rear transducer force) and 8 (hip sway bar) contain better discrimination information than other supersets and combinations of the data that were tried so far.

Bezdek, James C.↗

Displacement Analysis of Geothermal Field Based on PSInSAR And SOM Clustering Algorithms A Case Study of Brady Field, Nevada—USA

The availability of free and high temporal resolution satellite data and advanced SAR techniques allows us to analyze ground displacement cost-effectively. Our aim was to properly define subsidence and uplift areas to delineate a geothermal field and perform time-series analysis to identify temporal trends. A Persistent Scatterer Interferometry (PSI) algorithm was used to estimate vertical displacement in the Brady geothermal field located in Nevada by analyzing 70 Sentinel-1A Synthetic-Aperture Radar (SAR) images, between January 2017 and December 2019. To classify zones affected by displacement, an unsupervised Self-Organizing Map (SOM) algorithm was applied to classify points based on their behavior in time, and those clusters were used to determine subsidence, uplift, and stable regions automatically. Finally, time-series analysis was applied to the clustered data to understand the inflection dates. The maximum subsidence is –19 mm/yr with an average value of –6 mm/yr within the geothermal field. The maximum uplift is 14 mm/yr with an average value of 4 mm/yr within the geothermal field. The uplift occurred on the NE of the field, where the injection wells are located. On the other hand, subsidence is concentrated on the SW of the field where the production wells are located. The coupling of the PSInSAR and the SOM algorithms was shown to be effective in analyzing the direction and pattern of the displacements observed in the field.

54 ENVIRONMENTAL SCIENCES↗

Optimal Electrification Using Renewable Energies: Microgrid Installation Model with Combined Mixture k-Means Clustering Algorithm, Mixed Integer Linear Programming, and Onsset Method

Optimal planning and design of microgrids are priorities in the electrification of off-grid areas. Indeed, in one of the Sustainable Development Goals (SDG 7), the UN recommends universal access to electricity for all at the lowest cost. Several optimization methods with different strategies have been proposed in the literature as ways to achieve this goal. This paper proposes a microgrid installation and planning model based on a combination of several techniques. The programming language Python 3.10 was used in conjunction with machine learning techniques such as unsupervised learning based on K-means clustering and deterministic optimization methods based on mixed linear programming. These methods were complemented by the open-source spatial method for optimal electrification planning: onsset. Four levels of study were carried out. The first level consisted of simulating the model obtained with a cluster, which is considered based on the elbow and k-means clustering method as a case study. The second level involved sizing the microgrid with a capacity of 40 kW and optimizing all the resources available on site. The example of the different resources in the Togo case was considered. At the third level, the work consisted of proposing an optimal connection model for the microgrid based on voltage stability constraints and considering, above all, the capacity limit of the source substation. Finally, the fourth level involved a planning study of electrification strategies based mainly on microgrids according to the study scenario. The results of the first level of study enabled us to obtain an optimal location for the centroid of the cluster under consideration, according to the different load positions of this cluster. Then, the results of the second level of study were used to highlight the optimal resources obtained and proposed by the optimization model formulated based on the various technology costs, such as investment, maintenance, and operating costs, which were based on the technical limits of the various technologies. In these results, solar systems account for 80% of the maximum load considered, compared to 7.5% for wind systems and 12.5% for battery systems. Next, an optimal microgrid connection model was proposed based on the constraints of a voltage stability limit estimated to be 10% of the maximum voltage drop. The results obtained for the third level of study enabled us to present selective results for load nodes in relation to the source station node. Finally, the last results made it possible to plan electrification using different network technologies and systems in the short and long term. The case study of Togo was taken into account. The various results obtained from the different techniques provide the necessary leads for a feasibility study for optimal electrification of off-grid areas using microgrid systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Detection of Anomalies in Gamma Background Radiation Data with K-Means and Self-Organizing Map Clustering Algorithms (Consortium on Nuclear Security Technologies (CONNECT) Q1 Report)

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. The challenge is that spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

61 RADIATION PROTECTION AND DOSIMETRY↗

A Network-Based Algorithm for Clustering Multivariate Repeated Measures Data

The National Aeronautics and Space Administration (NASA) Astronaut Corps is a unique occupational cohort for which vast amounts of measures data have been collected repeatedly in research or operational studies pre-, in-, and post-flight, as well as during multiple clinical care visits. In exploratory analyses aimed at generating hypotheses regarding physiological changes associated with spaceflight exposure, such as impaired vision, it is of interest to identify anomalies and trends across these expansive datasets. Multivariate clustering algorithms for repeated measures data may help parse the data to identify homogeneous groups of astronauts that have higher risks for a particular physiological change. However, available clustering methods may not be able to accommodate the complex data structures found in NASA data, since the methods often rely on strict model assumptions, require equally-spaced and balanced assessment times, cannot accommodate missing data or differing time scales across variables, and cannot process continuous and discrete data simultaneously. To fill this gap, we propose a network-based, multivariate clustering algorithm for repeated measures data that can be tailored to fit various research settings. Using simulated data, we demonstrate how our method can be used to identify patterns in complex data structures found in practice.

Koslovsky, Matthew↗

Quantum cluster algorithm for data classification

Abstract We present a quantum algorithm for data classification based on the nearest-neighbor learning algorithm. The classification algorithm is divided into two steps: Firstly, data in the same class is divided into smaller groups with sublabels assisting building boundaries between data with different labels. Secondly we construct a quantum circuit for classification that contains multi control gates. The algorithm is easy to implement and efficient in predicting the labels of test data. To illustrate the power and efficiency of this approach, we construct the phase transition diagram for the metal-insulator transition of VO 2 , using limited trained experimental data, where VO 2 is a typical strongly correlated electron materials, and the metallic-insulating phase transition has drawn much attention in condensed matter physics. Moreover, we demonstrate our algorithm on the classification of randomly generated data and the classification of entanglement for various Werner states, where the training sets can not be divided by a single curve, instead, more than one curves are required to separate them apart perfectly. Our preliminary result shows considerable potential for various classification problems, particularly for constructing different phases in materials.

97 MATHEMATICS AND COMPUTING↗