Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

An application of cluster analysis for determining homogeneous subregions: The agroclimatological point of view

A stratification oriented to crop area and yield estimation problems was performed using an algorithm of clustering. The variables used were a set of agroclimatological characteristics measured in each one of the 232 municipalities of the State of Rio Grande do Sul, Brazil. A nonhierarchical cluster analysis was used and the pseudo F-statistics criterion was implemented for determining the "cut point" in the number of strata.

Parada, N. D. J.↗

What Determines the Number and the Timing of Pulses in Afternoon Precipitation in the Green Ocean Amazon (GoAmazon) Observations?

Using GoAmazon observations over the central Amazon, we investigate double-pulse events of afternoon precipitation for the first time. Relative humidity is the dominant factor in determining the onset timing and the number of pulses of precipitation. A moister free troposphere on single-pulse days features heavy precipitation in the early afternoon, which stabilizes the atmosphere and prohibits further convection. On early–onset double-pulse days, weaker precipitation from shallower convection in the first pulse leads to less consumption of atmospheric instability, making the environment favorable for a second pulse of stronger precipitation. A cluster tracking algorithm using the scanning radar reflectivity is developed to examine the relationship between precipitation pulses. Through explicitly tracking convective clusters that exist from the first to the second pulse, convective aggregation plays a more dominant role to the survival of clusters with larger final size, and on average around 20% of cluster survivals are due to natural growth.

54 ENVIRONMENTAL SCIENCES↗

VST ATLAS galaxy cluster catalogue I: cluster detection and mass calibration

Taking advantage of ∼4700 deg 2 optical coverage of the Southern sky offered by the VST ATLAS survey, we construct a new catalogue of photometrically selected galaxy groups and clusters using the orca cluster detection algorithm. The catalogue contains ∼22 000 detections with N 200 > 10 and ∼9000 with N 200 > 20. We estimate the photometric redshifts of the clusters using machine learning and find the redshift distribution of the sample to extend to z ∼ 0.7, peaking at z ∼ 0.25. We calibrate the ATLAS cluster mass-richness scaling relation using masses from the MCXC, Planck, ACT DR5, and SDSS redMaPPer cluster samples. We estimate the ATLAS sample to be > 95 per cent complete and > 85 per cent pure at z < 0.35 and in the M 200m >1 x 10 14 h -1 M ⊙ mass range. At z < 0.35, we also find the ATLAS sample to be more complete than redMaPPer, recovering a ~ 40 per cent higher fraction of Abell clusters. This higher sample completeness places the amplitude of the z < 0.35 ATLAS cluster mass function closer to the predictions of a ΛCDM model with parameters based on the Planck CMB analyses, compared to the mass functions of the other cluster samples. However, strong tensions between the observed ATLAS mass functions and models remain. We shall present a detailed cosmological analysis of the ATLAS cluster mass functions in paper II. In the future, optical counterparts to X-ray-detected eROSITA clusters can be identified using the ATLAS sample. The catalogue is also well suited for auxiliary spectroscopic target selection in 4MOST. The ATLAS cluster catalogue is publicly available at http://astro.dur.ac.uk/cosmology/vstatlas/cluster_catalogue/.

79 ASTRONOMY AND ASTROPHYSICS↗

Joint pattern recognition/data compression concept for ERTS multispectral imaging

This paper describes a new technique which jointly applies clustering and source encoding concepts to obtain data compression. The cluster compression technique basically uses clustering to extract features from the measurement data set which are used to describe characteristics of the entire data set. In addition, the features may be used to approximate each individual measurement vector by forming a sequence of scalar numbers which define each measurement vector in terms of the cluster features. This sequence, called the feature map, is then efficiently represented by using source encoding concepts. A description of a practical cluster compression algorithm is given and experimental results are presented to show trade-offs and characteristics of various implementations. Examples are provided which demonstrate the application of cluster compression to multispectral image data of the Earth Resources Technology Satellite.

Hilbert, E. E.↗

Spectroscopic Characterization of redMaPPer Galaxy Clusters with DESI

Optical galaxy cluster identification algorithms such as redMaPPer promise to enable an array of astrophysical and cosmological studies, but suffer from biases whereby galaxies in front of and behind a galaxy cluster are mistakenly associated with the primary cluster halo. These projection effects caused by irreducible photometric redshift uncertainty must be quantified to facilitate the use of optical cluster catalogues. We present measurements of galaxy cluster projection effects and velocity dispersion using spectroscopy from the Dark Energy Spectroscopic Instrument. Our findings are as follows: we confirm that the fraction of redMaPPer putative member galaxies mistakenly associated with cluster haloes is richness dependent, being more than twice as large at low richness than high richness; we present the first spectroscopic evidence of an increase in projection effects with increasing redshift, by as much as 25 per cent from $z\sim 0.1$ to $z\sim 0.2$; moreover, we find qualitative evidence for luminosity dependence in projection effects, with fainter galaxies being more commonly far behind clusters than their bright counterparts; finally, we fit the scaling relation between measured mean spectroscopic richness and velocity dispersion, finding an implied linear scaling between spectroscopic richness and halo mass. We discuss further directions for the application of spectroscopic data sets to improve use of optically selected clusters to test cosmological models.

clusters↗

Automated cloud screening of AVHRR imagery using split-and-merge clustering

Previous methods to segment clouds from ocean in AVHRR imagery have shown varying degrees of success, with nighttime approaches being the most limited. An improved method of automatic image segmentation, the principal component transformation split-and-merge clustering (PCTSMC) algorithm, is presented and applied to cloud screening of both nighttime and daytime AVHRR data. The method combines spectral differencing, the principal component transformation, and split-and-merge clustering to sample objectively the natural classes in the data. This segmentation method is then augmented by supervised classification techniques to screen clouds from the imagery. Comparisons with other nighttime methods demonstrate its improved capability in this application. The sensitivity of the method to clustering parameters is presented; the results show that the method is insensitive to the split-and-merge thresholds.

Gallaudet, Timothy C.↗

The PSZ-MCMF catalogue of Planck clusters over the DES region

ABSTRACT We present the first systematic follow-up of Planck Sunyaev–Zeldovich effect (SZE) selected candidates down to signal-to-noise (S/N) of 3 over the 5000 deg2 covered by the Dark Energy Survey. Using the MCMF cluster confirmation algorithm, we identify optical counterparts, determine photometric redshifts, and richnesses and assign a parameter, fcont, that reflects the probability that each SZE-optical pairing represents a random superposition of physically unassociated systems rather than a real cluster. The new PSZ-MCMF cluster catalogue consists of 853 MCMF confirmed clusters and has a purity of 90 per cent. We present the properties of subsamples of the PSZ-MCMF catalogue that have purities ranging from 90 per cent to 97.5 per cent, depending on the adopted fcont threshold. Halo mass estimates M500, redshifts, richnesses, and optical centres are presented for all PSZ-MCMF clusters. The PSZ-MCMF catalogue adds 589 previously unknown Planck identified clusters over the DES footprint and provides redshifts for an additional 50 previously published Planck-selected clusters with S/N>4.5. Using the subsample with spectroscopic redshifts, we demonstrate excellent cluster photo-z performance with an RMS scatter in Δz/(1 + z) of 0.47 per cent. Our MCMF based analysis allows us to infer the contamination fraction of the initial S/N>3 Planck-selected candidate list, which is ∼50 per cent. We present a method of estimating the completeness of the PSZ-MCMF cluster sample. In comparison to the previously published Planck cluster catalogues, this new S/N>3 MCMF confirmed cluster catalogue populates the lower mass regime at all redshifts and includes clusters up to z∼1.3.

79 ASTRONOMY AND ASTROPHYSICS↗

Cluster-cluster clustering

The cluster correlation function xi sub c(r) is compared with the particle correlation function, xi(r) in cosmological N-body simulations with a wide range of initial conditions. The experiments include scale-free initial conditions, pancake models with a coherence length in the initial density field, and hybrid models. Three N-body techniques and two cluster-finding algorithms are used. In scale-free models with white noise initial conditions, xi sub c and xi are essentially identical. In scale-free models with more power on large scales, it is found that the amplitude of xi sub c increases with cluster richness; in this case the clusters give a biased estimate of the particle correlations. In the pancake and hybrid models (with n = 0 or 1), xi sub c is steeper than xi, but the cluster correlation length exceeds that of the points by less than a factor of 2, independent of cluster richness. Thus the high amplitude of xi sub c found in studies of rich clusters of galaxies is inconsistent with white noise and pancake models and may indicate a primordial fluctuation spectrum with substantial power on large scales.

Barnes, J.↗

Low-level processing for real-time image analysis

A system that detects object outlines in television images in real time is described. A high-speed pipeline processor transforms the raw image into an edge map and a microprocessor, which is integrated into the system, clusters the edges, and represents them as chain codes. Image statistics, useful for higher level tasks such as pattern recognition, are computed by the microprocessor. Peak intensity and peak gradient values are extracted within a programmable window and are used for iris and focus control. The algorithms implemented in hardware and the pipeline processor architecture are described. The strategy for partitioning functions in the pipeline was chosen to make the implementation modular. The microprocessor interface allows flexible and adaptive control of the feature extraction process. The software algorithms for clustering edge segments, creating chain codes, and computing image statistics are also discussed. A strategy for real time image analysis that uses this system is given.

Eskenazi, R.↗

Smart Pixel Sensors for the HL-LHC

Large-scale particle physics experiments produce tens of terabytes of data every second. Innovative methods to manage the data rate at the HL-LHC, which expects to operate at 10x the luminosity of what the LHC was initially designed for, are needed. AI-on the chip provides a way to intelligently filter out low momentum clusters in the pixel detector. This will open up an opportunity to use the pixel detector for the first time in the CMS Level-1 trigger, and lead to increased sensitivity to new physics measurements and searches. We have taped out our first chip, which incorporates a $p_T$ filtering algorithm on an ASIC chip. Our initial $p_T$ filtering algorithm considers clusters that are tracked by CMS. We will report on ongoing studies seeking to enhance the performance of our filter by utilizing unsupervised learning on untracked clusters, thus increasing background rejection.

43 PARTICLE ACCELERATORS↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Automated Storm Tracking and the Lightning Jump Algorithm Using GOES-R Geostationary Lightning Mapper (GLM) Proxy Data

This study develops a fully automated lightning jump system encompassing objective storm tracking, Geostationary Lightning Mapper proxy data, and the lightning jump algorithm (LJA), which are important elements in the transition of the LJA concept from a research to an operational based algorithm. Storm cluster tracking is based on a product created from the combination of a radar parameter (vertically integrated liquid, VIL), and lightning information (flash rate density). Evaluations showed that the spatial scale of tracked features or storm clusters had a large impact on the lightning jump system performance, where increasing spatial scale size resulted in decreased dynamic range of the system's performance. This framework will also serve as a means to refine the LJA itself to enhance its operational applicability. Parameters within the system are isolated and the system's performance is evaluated with adjustments to parameter sensitivity. The system's performance is evaluated using the probability of detection (POD) and false alarm ratio (FAR) statistics. Of the algorithm parameters tested, sigma-level (metric of lightning jump strength) and flash rate threshold influenced the system's performance the most. Finally, verification methodologies are investigated. It is discovered that minor changes in verification methodology can dramatically impact the evaluation of the lightning jump system.

lightning jump↗

Generating Ground Reference Data for a Global Impervious Surface Survey

We are engaged in a project to produce a 30m impervious cover data set of the entire Earth for the years 2000 and 2010 based on the Landsat Global Land Survey (GLS) data set. The GLS data from Landsat provide an unprecedented opportunity to map global urbanization at this resolution for the first time, with unprecedented detail and accuracy. Moreover, the spatial resolution of Landsat is absolutely essential to accurately resolve urban targets such as buildings, roads and parking lots. Finally, with GLS data available for the 1975, 1990, 2000, and 2005 time periods, and soon for the 2010 period, the land cover/use changes due to urbanization can now be quantified at this spatial scale as well. Our approach works across spatial scales using very high spatial resolution commercial satellite data to both produce and evaluate continental scale products at the 30m spatial resolution of Landsat data. We are developing continental scale training data at 1m or so resolution and aggregating these to 30m for training a regression tree algorithm. Because the quality of the input training data are critical, we have developed an interactive software tool, called HSegLearn, to facilitate the photo-interpretation of high resolution imagery data, such as Quickbird or Ikonos data, into an impervious versus non-impervious map. Previous work has shown that photo-interpretation of high resolution data at 1 meter resolution will generate an accurate 30m resolution ground reference when coarsened to that resolution. Since this process can be very time consuming when using standard clustering classification algorithms, we are looking at image segmentation as a potential avenue to not only improve the training process but also provide a semi-automated approach for generating the ground reference data. HSegLearn takes as its input a hierarchical set of image segmentations produced by the HSeg image segmentation program [1, 2]. HSegLearn lets an analyst specify pixel locations as being either positive or negative examples, and displays a classification of the study area based on these examples. For our study, the positive examples are examples of impervious surfaces and negative examples are examples of non-impervious surfaces. HSegLearn searches the hierarchical segmentation from HSeg for the coarsest level of segmentation at which selected positive example locations do not conflict with negative example locations and labels the image accordingly. The negative example regions are always defined at the finest level of segmentation detail. The resulting classification map can be then further edited at a region object level using the previously developed HSegViewer tool [3]. After providing an overview of the HSeg image segmentation program, we provide a detailed description of the HSegLearn software tool. We then give examples of using HSegLearn to generate ground reference data and conclude with comments on the effectiveness of the HSegLearn tool.

Tilton, James C.↗

Uniformly Ordered Binary Decision Algorithm for Benchmark Experiment Correlations in Whisper Validation

When performing a validation exercise for determining the upper subcritical limit of a nuclear criticality safety application, an analyst should select and perform a statistical analysis on a population of benchmark experiments that are neutronically similar to the application. The size of this population should be sufficiently large such that the statistical analysis has a high degree of confidence that the bias plus bias uncertainty (calculational margin) has been accurately quantified. A complication arises because many benchmark experiments share common components, leading to correlations in their measured effective multiplication factors. Correlations between benchmark experiments within the population reduces its predictive power. This motivates the need for methods that consider benchmark experiment correlations and ensure adequate statistical significance of results. The Whisper code is a statistical analysis pack- age that incorporates nuclear data sensitivity coefficients from MCNP to assess benchmark experiment similarity and then performs an extreme-value analysis to estimate the bias plus bias uncertainty. The original methodology in Whisper does not consider the effect of benchmark experiment correlations when making this estimation, and this summary proposes the uniformly ordered binary decision algorithm to address this shortcoming. The original methodology in Whisper computes similarity coefficients ck for an application compared to all benchmark experiments in its library and develops weighting factors for a selected population proportional to the ck values. The methodology can be interpreted as statistically emulating a validation exercise for a particular application where the weighting factors may be viewed as the likelihood that an analyst would include a particular benchmark experiment within the population. The effective sample size of the population is the expected or mean number of benchmark experiments in the population. The uniformly ordered binary decision algorithm identifies clusters of correlated benchmark experiments within the population and then computes adjusted weighting factors based on the magnitude of the correlation coefficients within the cluster to compute a reduced effective sample size accounting for the lower information content because of correlations. Benchmark experiments within the cluster are ordered randomly with equal probability and probabilistic decisions are made as to whether a benchmark. experiment within the cluster should treated as redundant with a previous one; if two redundant benchmark experiments are included, then the conservative worst case bias plus bias uncertainty is used and the pair is counted as a single benchmark experiment in the population. Results are provided for HEU solutions in a research version of the Whisper software using benchmark experiment correlations provided by DICE, the Database for the International Criticality Safety Benchmark Evaluation Project (ICSBEP). These show that there can be a significant increase in the bias plus bias uncertainty because the effective sample size is reduced, and therefore the algorithm, needing to meet sample size requirements, expands the benchmark experiment population by accepting less similar benchmark experiments that would have otherwise not been included.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The fuzzy C spherical shells algorithm - A new approach

The fuzzy c spherical shells (FCSS) algorithm is specially designed to search for clusters that can be described by circular arcs or, more generally, by shells of hyperspheres. In this paper, a new approach to the FCSS algorithm is presented. This algorithm is computationally and implementationally simpler than other clustering algorithms that have been suggested for this purpose. An unsupervised algorithm which automatically finds the optimum number of clusters is also proposed. This algorithm can be used when the number of clusters is not known. It uses a cluster validity measure to identify good clusters, merges all compatible clusters, and eliminates spurious clusters to achieve the final result. Experimental results on several data sets are presented.

Krishnapuram, Raghu↗

Making the most of missing values : object clustering with partial data in astronomy

We demonstrate a clustering analysis algorithm, KSC, that a) uses all observed values and b) does not discard the partially observed objects. KSC uses soft constraints defined by the fully observed objects to assist in the grouping of objects with missing values. We present an analysis of objects taken from the Sloan Digital Sky Survey to demonstrate how imputing the values can be misleading and why the KSC approach can produce more appropriate results.

clustering↗

Geometric Interpretation of the Cluster Location Problem Part I: Theory

We present a new framing of the seismic location problem using principles drawn from differential geometry. Our interpretation relies upon the common assumption that travel times observed across a network are continuous, differentiable functions of source location. In consequence, travel‐time functions constitute a differentiable map between the source region and a Riemannian manifold. The manifold is said to be the image of the source region embedded in a generally high‐dimension travel‐time vector space. A cluster of events in the source region has an image of discrete points on the manifold, that, except in the simplest cases, cannot be viewed directly. However, it is possible to project the image of a cluster into a tangent space of the manifold for direct visualization. The projection operator can be computed directly from the data without a velocity model, but produces a distorted rendering of the cluster geometry. With a model we can predict the distortions and correct them to estimate cluster geometry. We develop these points with the simplest possible example, one for which direct visualization of the manifold is possible, using the example as an introduction to the relevant concepts from differential geometry in a familiar setting. The tangent space, a local linearization of the manifold, plays a key role. We develop a metric to estimate the limits of linearization, that is, to determine when the curvature of the manifold invalidates the linear assumption. We also examine the interplay of model error, inadequate network geometry, and pick error. We then generalize our results from the simple case to the general case of 3D source regions observed by general networks. Although we do suggest a new “project and correct” method for location, we do not develop it into a practical algorithm. In conclusion, our intention rather is to highlight new analytical methods grounded in differential geometry.

East Pacific Ocean Islands↗

Using action space clustering to constrain the recent accretion history of Milky Way-like galaxies

ABSTRACT In the currently favoured cosmological paradigm galaxies form hierarchically through the accretion of satellites. Since a satellite is less massive than the host, its stars occupy a smaller volume in action space. Actions are conserved when the potential of the host halo changes adiabatically, so stars from an accreted satellite would remain clustered in action space as the host evolves. In this paper, we identify recently disrupted accreted satellites in three Milky Way-like disc galaxies from the cosmological baryonic FIRE-2 simulations by tracking satellites through simulation snapshots. We try to recover these satellites by applying the cluster analysis algorithm Enlink to the orbital actions of accreted star particles in the z = 0 snapshot. Even with completely error-free mock data we find that only 35 per cent (14/39) satellites are well recovered while the rest (25/39) are poorly recovered (i.e. either contaminated or split up). Most (10/14 ∼70 per cent) of the well-recovered satellites have infall times <7.1 Gyr ago and total mass >4 × 108M⊙ (stellar mass more than 1.2 × 106 M⊙, although our upper mass limit is likely to be resolution dependent). Since cosmological simulations predict that stellar haloes include a population of in situ stars, we test our ability to recover satellites when the data include 10–50 per cent in situ contamination. We find that most previously well-recovered satellites stay well recovered even with 50 per cent contamination. With the wealth of 6D phase space data becoming available we expect that cluster analysis in action space will be useful in identifying the majority of recently accreted and moderately massive satellites in the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗