Inference for similarity indices
Farag Shuweihdi, Charles C Taylor · 1999
Clustering methods are considered as unsupervised learning techniques as there are no predetermined subpopulations. Evaluating the results of clustering algorithms is the main topic of cluster validity. This paper tries to contribute to a better understanding and inference for the distribution, of such a case, of the Rand Index (Rand, 1971) under several conditions. Besides this, a bootstrapping test for testing the significance of observed value of the similarity indices is provided and compared with the p-values yielded from the density function of the Rand index. The comparison is made using simulated data. Let nij be the number of objects which belong to cluster i of partition P1 and cluster j of partition P2, i = 1, · · · , R, j = 1, · · · , C, with ni. = ∑C j nij , n.j = ∑R i nij , the size of the clusters in P1 , P2 respectively, and n = ∑R i ni. = ∑C j n.j, the total number of observations. After that the Rand index can be applied to this result; this index is given by