Application of Bloom Filter for Duplicate URL Detection in a Web Crawler
Aveksha Kapoor, Vinay Arora · 2016
A web crawler is an important component for web based information retrieval system. In this paper, a detailed study of an open-source map-reduce based web crawler, Apache Nutch has been done. In addition, the bloom filter methodology for de-duplicated URL list generation and updating in each iteration of the crawl, has been implemented. The bloom based approach is compared with the existing URL fetch-time approach (inherent to Nutch) in terms of time. Further, two types of hashes-jenkins hash and murmur hash have been used in the bloom filter; their performance in terms of time, false positive has been compared with each other. It has also been shown that the usage of this additional data structure, bloom filter in the proposed approach is memory-efficient. The paper is an attempt to improve upon the time-efficiency of a map-reduce based crawler. Using bloom filter, based upon comparison against different parameters our results indicate the efficiency and effectiveness of the proposed approach.