Bulk-Synchronous On-Line Crawling on Clusters of Computers

Mauricio Marı́n, Carolina Bonacic · 2008

This paper describes the design of a crawler devised to perform the periodic retrieval of Web documents for a search engine able to accept on-line updates in a concurrent manner. On-line updates comes in the form of insertions of new documents or update of existing ones, all of them mixed with the usual user queries. The search engine is bulk-synchronous which allows it to deal efficiently with the concurrency control problem. The crawler is also bulk- synchronous so that it can be integrated into the same P- processors cluster executing the search engine. This paper describes and evaluates the practical feasibility of such a crawler.

Read the paper · More papers on PaperTik