Circluster: Storing cluster shapes for clustering
Sajad Shirali-Shahreza, Soheil Hassas Yeganeh, Hassan Abolhassani, Jafar Alim Habibi · 2008
One of the important problems in knowledge discovery from data is clustering. Clustering is the problem of partitioning a set of data using unsupervised techniques. An important characteristic of a clustering technique is the shape of the cluster it can find. Clustering methods which are capable to find simple cluster shapes are usually fast but inaccurate for complex data sets. Ones capable to find complex cluster shapes are usually not fast but accurate. In this paper, we propose a simple clustering technique named circlusters. Circlusters are circles partitioned into different radius sectors. Circlusters can be used to create hybrid approaches with density based or partitioning based methods. We also propose a naive clustering method that is capable to find complex clusters in O(n). This method operates in two phases. In the first phase, circlusters are created to approximate the shape of the data set. In the second phase, connected circlusters are found to form the final clusters.