Detecting Text Reuse with Modified and Weighted N-grams

Rao Muhammad Adeel Nawab, Mark Stevenson, Paul David Clough · Joint Conference on Lexical and Computational Semantics · 2012

Text reuse is common in many scenarios and documents are often based, at least in part, on existing documents. This paper reports an approach to detecting text reuse which identifies not only documents which have been reused verbatim but is also designed to identify cases of reuse when the original has been rewritten. The approach identifies reuse by comparing word n-grams in documents and modifies these (by substituting words with synonyms and deleting words) to identify when text has been altered. The approach is applied to a corpus of newspaper stories and found to outperform a previously reported method.

Read the paper · More papers on PaperTik