A New Normalized Similarity for Discriminating Similar Documents
Jeong-Hoon Ji, Chang-Koen Ryu, Gyun Woo, Hwan-Gue Cho · 2008
To find out similar document pairs from a set of documents, computing normalization similarities is inevitable because the sizes of documents are different from documents to documents. However, the normalized similarities proposed up to now are still unreliably sensitive to the size of programs compared. Due to this fact, most previously announced similarity detection tools have difficulties in determining the cutoff threshold to discriminate similar documents from a set of documents. In this paper, we propose a new normalized similarity based on Weibull distribution. To test the effectiveness of the new similarity measure, we applied it in detecting similar program pairs from a set of programs. According to the experiment, the new similarity measure showed very nice characteristics in discriminating the very similar program pairs from other pairs. Also, the proposed normalized similarity is effective in detecting similar documents written in natural languages.