Duplicate Detection in Documents and WebPages Using Improved Longest Common Subsequence and Documents Syntactical Structures
Mohamed Elhadi, Amjad Al-Tobi · 2009
This paper reports on experiments performed to investigate the use of a combined part of speech (POS) and an improved longest common subsequence (LCS) in the analysis and calculation of similarity between texts. The text's syntactical structures were used as a representation for documents. An improved LCS algorithm was applied to such a representation to compare and rank the documents according to the similarity of their representative string. The approach was applied in detecting duplicate documents within a corpus, and in the filtering of search engine results. Results obtained were encouraging.