A new STC algorithm based on repeats

Hua Jiang · Microcomputer Information · 2009

Current de-duplication algorithms mainly focus on keywords de-duplication or semantic fingerprint de-duplication and may cause error when processing Web pages.This paper using the repeats as mapped sentences to make the suffix tree. Using the inverted index method to storage the data. Experiment results show that this method can find similar Web pages efficiently,this algorithm can reach a high precision in mono-language deletion of duplicated web pages, and this algorithm can also reach a maximum precision when it is applied to deletion of duplicated web pages.

Read the paper · More papers on PaperTik