Design and analysis of novel similarity measure for clustering and classification of high dimensional text documents

G. Suresh Reddy, T. V. Rajinikanth, A. Ananda Rao · 2014

The main idea of this research is to first design the similarity measure which can be used to of find the similarity between any two text documents and use the same to perform clustering. The similarity measure designed is analyzed to study the behavior in the best case, average case and worst case situations. The drawback of Euclidean, Cosine, Jaccard similarity measures are overcome using the proposed measure. The similarity measure is evaluated considering reuters-21578 dataset. The results show that the proposed measure overcomes other measures.

Read the paper · More papers on PaperTik