Partition-based overlapping clustering using cluster's parameters and relations
Tanawat Limungkura, Peerapon Vateekul · 2017
Conventional clustering algorithms based on the assumption that a data point can be assigned to only a single cluster. In spite of, there are several types of data that a data point belongs to multiple categories and causes ground-truth clusters overlap. To handle this situation, several algorithms are proposed and referred as “overlapping clustering”. One of state-of-the-art partition-based overlapping clustering technique is “Non-exhaustive, Overlapping K-Means” or “NEO-K-Means” in short, which is an extension of K-Means clustering algorithm. Although NEO-K-Means works effectively for most real-world multi-category data. However, the process of assigning clusters focuses only on the minimum distance from the data to the centroid and ignores other essential parameters, such as, distance between clusters and the radius of clusters. Moreover, the number of estimation of data points on overlapping area is also still not accurate enough. These are huge drawbacks of NEO-K-Means that makes clustering accuracy lower than it should be. In this paper, we aim to overcome this limitation in NEO-K-Means by using radius of cluster and distance between clusters for assisting in estimating data and assigning clusters. The experimental results show that our method significantly outperforms NEO-K-Means on nine real multi-category data sets in terms of F1.