Rearch on Large Scale Documents Deduplication Technique based on Simhash Algorithm

Yi Yu, Zijian Hu, Yuzhu Zhang · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2015

On the background of the deduplication needs of repeated documents in Internet, research the deduplication technique based on Simhash algorithm on large-scale documents.On the basis of taking the Simhash algorithm as core algorithm in duplicated documents detection, improve the procedure of achieving documents features of this algorithm.It takes the meaning and length of words as a consideration factor in measuring the weight of words.Aiming at the Simhash signature of a 64-bit, provide the document service of making a similarity comparison based on the full text and paragraphs.Through test data and analysis, this technique can guarantee the stable operation, 100 million documents can be memorized in each example.The average request response time is about 20 ms.The respon.setime will increase during the peak hour, but, in general, will not go over 100 ms.

Read the paper · More papers on PaperTik