17. Evaluation of Clustering Algorithms
Society for Industrial and Applied Mathematics eBooks · 2007
In the literature of data clustering, a lot of algorithms have been proposed for different applications and different sizes of data. But clustering a data set is an unsupervised process; there are no predefined classes and no examples that can show that the clusters found by the clustering algorithms are valid (Halkidi et al., 2002a). To compare the clustering results of different clustering algorithms, it is necessary to develop some validity criteria. Also, if the number of clusters is not given in the clustering algorithms, it is a highly nontrivial task to find the optimal number of clusters in the data set. To do this, we need some cluster validity methods. In this chapter, we will present various kinds of cluster validity methods appearing in the literature. 17.1 Introduction In general, there are three fundamental criteria to investigate the cluster validity: external criteria, internal criteria, and relative criteria (Jain and Dubes, 1988; Theodoridis and Koutroubas, 1999; Halkidi et al., 2002a). Some validity index criteria work well when the clusters are compact but do not work sufficiently (Halkidi et al., 2002a) if the clusters have arbitrary shape (applications in spatial data, biology data, etc.). Figure 17.1 summarizes some popular criteria. The first two approaches involve statistical testing, which is computationally expensive. The third, i.e., relative criteria, does not involve statistical testing.