Information Retrieval from the Web and Application of Migrating Crawler

Niraj Singhal, Rohit Agarwal, Ashutosh Dixit, Anukrati Sharma · 2011

Study reports that about 40% of current internet traffic and bandwidth consumption is due to the web crawlers that retrieve pages for indexing by the different search engines. As the size of the web continues to grow, searching it for useful information has become increasingly difficult. The centralized crawling techniques are unable to cope up with constantly growing web. In this paper it is presented that distributed crawling methods based on migrating crawlers are an essential tool for allowing such access that minimizes network utilization and also keeps up with document changes.

Read the paper · More papers on PaperTik