Improved Implementation for Finding Text Similarities in Large Sets of Data - Notebook for PAN at CLEF 2011.

J. Grman, Rudolf Ravas · 2011

In this article we describe a new algorithm method for the detection of plagiarism. The method removes numerous limitations of our older method, which has been used as part of a complex information system for the detection of plagiarism. The method has been tested using multiple corpora mainly in Slovak language. With the PAN-09 and PAN-10 corpora it was of great advantage that we could compare our results with the results of other methods. The very good initial results gave us motivation to implement multiple algorithm and parameter improvements.

Read the paper · More papers on PaperTik