Quick Full-text Identification Algorithm for Document Set Plagiarism

Mingxiao Hu · Jisuanji gongcheng · 2010

In order to identify plagiarisms for local document set,this paper defines the document plagiarism distance as an approximate generalized edit distance based on returning number and skipping number,then uses this distance.After analyzing the sufficient conditions of satisfying triangle inequality or weak triangle inequality for the distance,it proposes an efficient full-text identification algorithm which can find out all ordered plagiarizing document pairs faithfully.Experimental results show that the algorithm improves the identifying efficiency by 3 times to 5 times meanwhile it does not lower the recall ratio.

Read the paper · More papers on PaperTik