Semantic-based plagiarism detection

Radim Řehůřek · 2008

Summary: I aim to develop an efficient text plagiarism detection system. Currently used systems concentrate on copy detection, and as such are inca- pable of detecting finer cases of plagiarism, which include stealing of ideas rather than exact words. I want to employ Latent Semantic Indexing (LSI) to build semantic structure from a set of documents, and use this knowledge to guide plagiarism queries. LSI serves both as a tool for transforming text data into a smaller, conceptual space and consequently as a performance booster. However, the resulting document representation is very dense, in the sense that each concept is assigned a non-zero real valued number. This poses a problem to efficient querying, because the commonly used tech- nique of inverted index files is not applicable. Popular space-partitioning and data-partitioning indexing techniques also prove inadequate, due to their poor scalability with regard to the VS dimensionality. A choice of an improvement over linear scan called VA-File is considered. An improve- ment to internal system design clarity in the form of segmenting the doc- uments according to topics before further processing is introduced. This process is also hoped to improve retrieval performance. A novel combina- tion of these general methods into a system, their modifications and perfor- mance assessment is proposed to be the subject of my thesis.

Read the paper · More papers on PaperTik