Effects of Similarity Metrics on Document Clustering
Kazem Taghva, Rushikesh Veni · 2010
Document clustering or unsupervised document classification is an automated process of grouping documents with similar content. A typical technique uses a similarity function to compare documents. In the literature, many similarity functions such as dot product or cosine measures are proposed for the comparison operator. In these papers, we evaluate the effects of many similarity functions on k-mean clustering algorithm. Based on our analysis, we conclude that Chi-Square works best for the document collection with efficiency around 80% followed by Canberra and Euclidean distances with 70%. The results also indicate that the distance metrics like Bray-Curtis, Variational and Trigonometric function didn't produce good results.