Verification of supervised clustering validity and its applications

Yan Zhan, Yuan Fang, Xizhao Wang · 2003

The clustering validity problem decides whether the clustering result of data set is reasonable. The validity functions frequently used for clustering are based on unsupervised clustering, whose central meaning is the confirmation of class number in unsupervised clustering. This paper puts forward a validity function for judging clustering; the numerical experiments prove its validity. For applications in k-nearest neighbor classification, we can view the class center as its representative if the data set satisfies the qualification of this function. In this way, we can reduce the iteration query in all the training data while we only need to compare the similarities between the testing data and each class center. When selecting k-neighbors, we can only choose the nearest neighbor (1-nearest neighbor) in order to achieve precise classification and avoid the trouble of looking for the k-value. This will reduce the query complexity greatly and improve the efficiency of the nearest neighbor algorithm.

Read the paper · More papers on PaperTik