A Cluster Validity Indexing Method Based on Entropy for Solving Cluster Overlapping Problem

Lin Phen-Lan, Huang Ping-Hsuan, Huang Po-Whei · Frontiers in artificial intelligence and applications · 2015

Data clustering technique can be used in many fields, such as data mining, statistical data analysis, image analysis, pattern recognition, etc. Good clustering can result in computational reduction in related application programs; however, it is hard to achieve without knowing how many clusters that a data set should be partitioned, which is common in many applications. The way to find the optimal number of clusters is called cluster validity. In this paper, we proposed a new cluster validity indexing method that aims to solve cluster overlapping problem. Our method adapts the concept of cluster validity index defined as the ratio of compactness and separation and enhances it by integrating an entropy-based weight to the definition of separation so that the new weighted-separation of two overlapped clusters will be larger than that of two non-overlapped clusters, where the distance between the two cluster-centroids are the same. Experiments on six synthetic datasets comprising 3 to 10 clusters with some clusters overlapped each other demonstrate that our proposed method achieves 100% accuracy of validity index for all these datasets and is superior to all other compared methods.

Read the paper · More papers on PaperTik