Feature Selection Clustering and Prototype Placement for Turbulence Data Sets.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
The 45th Weather Squadron (45 WS) asked the Applied Meteorology Unit (AMU) to review the 30-year-old, lightning stand-off distances of 5 nautical miles (nmi) for applicability to today's operations. This was based on the realization that previous lightning strike distance studies did not match how 45 WS issues lightning warnings (Roeder, 2008). The previous lightning distance studies were from the point of origin of the lightning or the average starting location that would tend to be in the core of the thunderstorm. However, the 45 WS issues lightning warnings based on the edge of a preexisting lightning area. Before beginning the AMU project, it took several years to develop a method to calculate a distance distribution beyond a preexisting area (Roeder, 2015). The AMU pulled Lightning Detection and Ranging (LDAR) sensor data from 1/1/2013 to 12/31/2013. This dataset consisted of 37 million individual source data points from the LDAR sensors. Only sources within 50 km north, south, east or west of the LDAR grid center were included in the dataset. This limited the use of LDAR data to that with the greatest accuracy of source detection and increased data processing. Points were grouped into flashes based on spatial and temporal criteria. Based on the sensitivity analysis the AMU performed on the flash clustering algorithm, a time value of 0.3 seconds was found to model flashes adequately. Distance parameters were tested from 1,500 to 7,500 meters (m) in 500 m increments. Distance parameters of both 3,000 m and 4,000 m produced results in the plotting tool that were most representative of the physical behavior of lightning. Thus statistics were gathered for the most representative of these spatial and temporal criteria on the flash size and the polygon expansion distance in order to find the correct data distributions. The best fit curves for the LDAR polygon expansion frequency vs. distances for both the 3 kilometer (km) and 4 km distance threshold values were exponential decay functions and had R2 values of > 0.998, indicating good model fits. The equations of the best fit curves were then used to calculate a desired safety radius of 4 nmi for either 3 km or 4 km distance threshold criteria. The AMU analysis concludes the safe reduction of the 5 nmi lightning warning circles to 4 nmi should improve the operational impact by 36% if based on distance from the center of the property area being protected. If based on the edge of the property being protected, then the reduction is 4.5 nmi to 4 nmi and the operational impact is decreased by 21%. For the 6 nmi lightning warning circles, the recommended 4 nmi stand-off distance will result in a safe reduction of operational impact of 31% if based on the center of the area being protected, or 16% if based on the edge of the property.
Protection against dc faults is one of the main technical hurdles faced when operating converter-based HVdc systems. Protection becomes even more challenging for multi-terminal dc (MTdc) systems with more than two terminals/converter stations. In this paper, a hybrid primary fault detection algorithm for MTdc systems is proposed to detect a broad range of failures. Sensor measurements, i.e., line currents and dc reactor voltages measured at local terminals, are first processed by a top-level context clustering algorithm. For each cluster, the best fault detector is selected among a detector pool according to a rule resulting from a learning algorithm. The detector pool consists of several existing detection algorithms, each performing differently across fault scenarios. The proposed hybrid primary detection algorithm: i) offers superior performance compared to an individual detector through a data-driven approach; ii) detects all major fault types including pole-to-pole (P2P), pole-to-ground (P2G), and external dc faults; iii) identifies faults with various fault locations and impedances; iv) is more robust to noisy sensor measurements compared to existing methods; v) does not require exhaustive simulation and sampling for training the model. Performance and effectiveness of the proposed algorithm are evaluated and verified based on time-domain simulations in the PSCAD/EMTDC software environment. The results confirm satisfactory operation, accuracy, and detection speed of the proposed algorithm under various fault scenarios.
We propose a flexible framework for clustering hypergraph-structured data based on recently proposed random walks utilizing edge-dependent vertex weights. When incorporating edge-dependent vertex weights (EDVW), a weight is associated with each vertex-hyperedge pair, yielding a weighted incidence matrix of the hypergraph. Such weightings have been utilized in term-document representations of text data sets. We explain how random walks with EDVW serve to construct different hypergraph Laplacian matrices, and then develop a suite of clustering methods that use these incidence matrices and Laplacians for hypergraph clustering. Using 20Newsgroup, U.S. patent, Reuters' Corpus Volume 1, and genetics data sets, we compare the performance of these clustering algorithms experimentally against a variety of existing hypergraph clustering methods. We show that the proposed methods produce higher-quality clusters.
SAND2024-01234O TICC is a clustering algorithm that labels a sequence of data points according to numerical properties. This library is a Python implementation of the algorithm described in "Toeplitz Inverse Covariance-Based Clustering of Multivariate Time Series Data" (Hallac et al. 2017). It includes documentation, performance improvements, examples, and test coverage. This library allows users to automatically segment a series of multivariate data points according to their covariance—that is, the way the values at each data point are changing in relation to one another. This is useful for identifying periods in which a system is behaving. For example, if a sensor is measuring a car's velocity, steering wheel angle, braking and acceleration, TICC can determine when the car was stopped, beginning/exiting a turn, slowing or accelerating at an intersection, or driving on straight or curved roads. TICC can be applied to measure multiple quantities at known times. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525
Secondary crashes are crashes that occur as a result of the nonrecurrent congestion originating from primary crashes, and always have a greater impact on safety and traffic than a single crash. A better understanding of secondary crashes would benefit traffic incident management, and this requires accurate identification of secondary crashes. This study explores using crowdsourced Waze user reports to identify secondary crashes. Here, a network-based clustering algorithm is proposed to extract the primary crash cluster, including all user reports originating from the primary crash, and any crash that occurred within the cluster would be a secondary crash. This method works as a filter to select accurate primary–secondary relationships, thus precisely identifying secondary crashes. A case study is performed with crashes occurring from June to December 2019 on a 30-mi stretch of I-40 in Knoxville, TN. A static threshold method (crash duration and 10 mi) was used to preselect the potential primary–secondary crash pairs, and 75 out of 708 crashes were identified as potential secondary crashes. Based on the preselected primary–secondary crash pairs, 17 secondary crashes were obtained with the proposed method and the results were compared with one of the commonly used methods, the speed contour plot method. Though the proposed method captured fewer secondary crashes, it did identify several secondary crashes that could not be observed with the speed contour plot method. The results showed the applicability of the method and the potential of crowdsourced Waze user reports in secondary crash identification.
A new clustering algorithm is presented that is based on dimensional information. The algorithm includes an inherent feature selection criterion, which is discussed. Further, a heuristic method for choosing the proper number of intervals for a frequency distribution histogram, a feature necessary for the algorithm, is presented. The algorithm, although usable as a stand-alone clustering technique, is then utilized as a global approximator. Local clustering techniques and configuration of a global-local scheme are discussed, and finally the complete global-local and feature selector configuration is shown in application to a real-time adaptive classification scheme for the analysis of remote sensed multispectral scanner data.
Two unsupervised classification procedures were applied to ratioed and unratioed LANDSAT multispectral scanner data of an area of spatially complex vegetation and terrain. An objective accuracy assessment was undertaken on each classification and comparison was made of the classification accuracies. The two unsupervised procedures use the same clustering algorithm. By on procedure the entire area is clustered and by the other a representative sample of the area is clustered and the resulting statistics are extrapolated to the remaining area using a maximum likelihood classifier. Explanation is given of the major steps in the classification procedures including image preprocessing; classification; interpretation of cluster classes; and accuracy assessment. Of the four classifications undertaken, the monocluster block approach on the unratioed data gave the highest accuracy of 80% for five coarse cover classes. This accuracy was increased to 84% by applying a 3 x 3 contextual filter to the classified image. A detailed description and partial explanation is provided for the major misclassification. The classification of the unratioed data produced higher percentage accuracies than for the ratioed data and the monocluster block approach gave higher accuracies than clustering the entire area. The moncluster block approach was additionally the most economical in terms of computing time.
Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.
In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.
Frequency stability assessment is one critical aspect of power system security assessment. Traditional N-1 screening method is based on the simulations of a few typical daily and seasonal operation scenarios. However, the increasing integration of inverter-based renewables and the retirement of conventional synchronous generators result in decreasing system inertia and growing complexity of system operating conditions. Selecting a few typical operation scenarios cannot cover all operating conditions, and the time-domain simulation of all operation conditions requires tremendous time. This paper proposes a more efficient frequency stability assessment method based on deep learning. The affinity propagation clustering algorithm is used to divide the dataset into different clusters, so the selected dataset for training can cover the diversified operating conditions as much as possible. Also, feature normalization is applied to both the training dataset and testing dataset in order to remove any unnecessary bias. Especially, trained model based on full dataset normalization has bounded error in the prediction. The case study on the reduced 240-bus WECC system demonstrates that the proposed method can predict accurate frequency nadir with limited training dataset. The deep learning model using the revised feature normalization can predict more accurate frequency nadir than that using the traditional feature normalization and has very small maximum prediction error.
ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, it uses linear BVH for its low construction cost and sufficient quality. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase. ArborX also implements several advanced algorithms using geometric search, such as density-based clustering algorithm DBSCAN.
Clustering algorithms can identify groups in large data sets, such as star catalogs and hyperspectral images. In general, clustering methods cannot analyze items that have missing data values. Common solutions either fill in the missing values (imputation) or ignore the missing data (marginalization). Imputed values are treated as just as reliable as the truly observed data, but they are only as good as the assumptions used to create them. In contrast, we present a method for encoding partially observed features as a set of supplemental soft constraints and introduce the KSC algorithm, which incorporates constraints into the clustering process. In experiments on artificial data and data from the Sloan Digital Sky Survey, we show that soft constraints are an effective way to enable clustering with missing values.
We present a novel flexible bi-level spatiotemporal clustering algorithm to extract events based on their intensity and spatiotemporal structures. Our algorithm consists of using (i) a novel space-time k-means clustering to obtain spatiotemporally coherent intensity clusters, and (ii) a density-based spatial clustering of applications with noise (DBSCAN) to spatiotemporally section the intensity clusters into individual events. We discuss the development of the algorithm, the selection, tuning and meaning of the parameters within each step, as well as its validation. Finally, we apply the algorithm to a spatiotemporal drought index, standardized vapor pressure deficit drought index (SVDI), over the continental United States (US) from 1980–2021 and show that it captures historical drought events over the continental United States and their spatiotemporal extents.
Hyperspectral imagers produce very large images, with each pixel recorded at hundreds or thousands of different wavelengths. The ability to automatically generate summaries of these data sets enables several important applications, such as quickly browsing through a large image repository or determining the best use of a limited bandwidth link (e.g., determining which images are most critical for full transmission). Clustering algorithms can be used to generate these summaries, but traditional clustering methods make decisions based only on the information contained in the data set. In contrast, we present a new method that additionally leverages existing spectral libraries to identify materials that are likely to be present in the image target area. We find that this approach simultaneously reduces runtime and produces summaries that are more relevant to science goals.
The conditions under which a hybrid of clustering and canonical analysis for image classification produce optimum results were analyzed. The approach involves generation of classes by clustering for input to canonical analysis. The importance of the number of clusters input and the effect of other parameters of the clustering algorithm (ISOCLS) were examined. The approach derives its final result by clustering the canonically transformed data. Therefore the importance of number of clusters requested in this final stage was also examined. The effect of these variables were studied in terms of the average separability (as measured by transformed divergence) of the final clusters, the transformation matrices resulting from different numbers of input classes, and the accuracy of the final classifications. The research was performed with LANDSAT MSS data over the Hazleton/Berwick Pennsylvania area. Final classifications were compared pixel by pixel with an existing geographic information system to provide an indication of their accuracy.
We extend and enhance the concept of the Cloud Regimes (CRs) developed from two-dimensional joint histograms of cloud optical thickness and cloud top pressure from the Moderate Resolution Imaging Spectroradiometer (MODIS), by adding precipitation information in order to better understand cloud-precipitation relationships. Taking advantage of the high-resolution Integrated Multi-satellitE Retrievals for GPM (IMERG) precipitation dataset, cloud-precipitation “hybrid” regimes are derived by implementing the k-means clustering algorithm with advanced initialization and objective measures to determine the most optimal clusters. By expressing precipitation rates within 1-degree grid cell as histograms and making choices on the relative weight of cloud and precipitation, we could obtain several editions of hybrid cloud-precipitation regimes (CPRs), and examine their characteristics. In the deep tropics, when precipitation is weighted weakly, the cloud part of the hybrid entroids resembles the centroid of cloud-only regimes, but still tightens the cloud-precipitation relationship by decreasing the precipitation variability of each regime. As precipitation weight progressively increases, the shape of the cloudy part of the hybrid centroids becomes blunter, while the precipitation part of the centroids sharpens. In the case where cloud and precipitation are weighted equally, the CPRs representing high clouds with intermediate to heavy precipitation exhibit distinct features in the precipitation parts of the centroids, which allows us to project them onto the 30-minly IMERG domain. Such a projection can be used to overcome the temporal sparseness of MODIS cloud observations, which leads to great application potential for various convection-focused studies, including diurnal cycle analysis.
The clustering program ISOCLS developed at the Johnson Space Center, Houston, Texas, has been extensively used in the pattern analysis and classification of remote sensor data collected by aircraft and by the Earth Resources Technology Satellite ERTS-1. This paper discusses the theory behind this clustering algorithm. Several new ideas that have been incorporated in ISOCLS are discussed. Among these are the novel philosophy of operation behind the procedure, which assumes that a population (i.e., a class or a cluster) can be treated as the union of an appropriate number of subpopulations, and the termination of the clustering program by a 'chaining algorithm.' Finally, this paper reports the results of the application of ISOCLS to an investigation on rangeland vegetation mapping using ERTS-1 data.