Correct Number of Clusters (CNC) Description Length in Arbitrary Shape Clustering
Mahdi Shamsi, Faizan Rahman, Soosan Beheshti · 2019
One of the main challenges in clustering unlabeled data sets is determining the unknown correct number of clusters (CNC). K-means is a well known and widely used clustering algorithm in this context which requires the correct number of the clusters for proper performance. To address this problem, various validity indices approaches aim to optimize a desired criteria based on measuring the compactness of cluster and the separation between cluster. K-MACE algorithm is a validity index clustering method based on estimating the average error between the correct cluster center and the estimated cluster center for each data point. We propose a modified version of K-MACE that is based on minimizing the CNC codelength. The proposed theory handles clusters that are arbitrary shaped and/or nonlinearly separable. Simulation results confirms superiority of the proposed method over well known validity index methods in the sense of accurate CNC estimation as well as optimizing Adjusted Random Index (ARI) and the Normalized Variation Information (NVI) measures.