NASA NTRS · 20070019774
Clustering with Missing Values: No Imputation Required
Abstract
Clustering algorithms can identify groups in large data sets, such as star catalogs and hyperspectral images. In general, clustering methods cannot analyze items that have missing data values. Common solutions either fill in the missing values (imputation) or ignore the missing data (marginalization). Imputed values are treated as just as reliable as the truly observed data, but they are only as good as the assumptions used to create them. In contrast, we present a method for encoding partially observed features as a set of supplemental soft constraints and introduce the KSC algorithm, which incorporates constraints into the clustering process. In experiments on artificial data and data from the Sloan Digital Sky Survey, we show that soft constraints are an effective way to enable clustering with missing values.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wagstaff, Kiri. 2004-07-15. Clustering with Missing Values: No Imputation Required. https://ntrs.nasa.gov/citations/20070019774
Cite the original work for its findings. Save a collection to share your selection of sources.