Detecting Short Passages of Similar Text in Large Document Collections

Caroline Lyon, James A. Malcolm, Bob Dickerson · University of Hertfordshire Research Archive (University of Hertfordshire) · 2001

This paper presents a statistical method for fingerprinting text. In a large collection of independently written documents each text is associated with a fingerprint which should be different from all the others. If fingerprints are too close then it is suspected that passages of copied or similar text occur in two documents. Our method exploits the characteristic distribution of word trigrams, and measures to determine similarity are based on set theoretic principles. The system was developed using a corpus of broadcast news reports and has been successfully used to detect plagiarism in students' work. It can find small sections that are similar as well as those that are identical. The method is very simple and effective, but seems not to have been used before

Read the paper · More papers on PaperTik