Optimal K-Means Clustering Method Using Silhouette Coefficient

L. Nitya Sai, M. Sai Shreya, A. Anjan Subudhi, B. Jaya Lakshmi, K.B. Madhuri · International Journal of Applied Research on Information Technology and Computing · 2017

K-Means is one partitional-based clustering algorithm that accepts K, a user defined parameter as input. Choosing optimal K value is an open issue. It is very difficult to choose K value as it depends on distribution of data points in the feature space. Euclidean distance is one of the distance metrics used to calculate similarity between the data points. By means of Silhouette coefficient (SC), the quality of the cluster can be measured based on the concept of cohesion and separation between clusters. It ranges from [-1,1] where 1 indicates best quality of cluster while -1 indicates poor quality. In this paper, SC is used to estimate optimal K-value with which clusters are formed using K-means clustering. Depending on the optimal K value, better clusters can be obtained in the result for a given data set.

Read the paper · More papers on PaperTik