The anatomy of web crawlers
Shruti Sharma, Parul Gupta · 2015
World Wide Web (www) is the gigantic and richest source of information. To retrieve the information from this imperative resource, Search Engines are generally used. For this purpose these Search engines rely on massive collections of web pages that have been downloaded by web crawlers. A Web crawler is a program that traverses the web by following the ever changing, dense and distributed hyperlinked structure and thereafter storing downloaded pages in a large repository which is later indexed for efficient execution of user queries. Thus, web crawlers are becoming increasingly important. Various web crawling architectures have been proposed in recent years. In this paper a survey of different architectures of web crawlers along with their comparisons has been carried out that takes into account various important features like scalability, manageability, page refresh policy, politeness policy etc.