Web Document Duplicate Removal Algorithm Based on Keyword Sequences

Wei Li, Jianyi Liu, Cong Wang · 2006

There are many identical documents across the Web. The effective duplicate removal has become one of the most important techniques to improve search engines. In this paper, we take both syntax information and semantic information into account, and put forward a Web document duplicate removal algorithm based on keyword sequences, which is called KSM (keyword sequences method). The main intuition behind KSM is as follows. The keyword sequences of the Web document can be used to depict its structure feature (syntax) and intension feature (semantics). By the comparison of keyword sequences between similar documents, we can judge whether there is information redundancy. The experimental results show that KSM can greatly reduce the probability of mistaking similar documents for identical documents while remarkably improving the resistance to document noises.

Read the paper · More papers on PaperTik