A Study on Design, Development and Deployment of Web Crawler Algorithms and Their Metrics

J. Arthy, K. Raja · 2024

Globally, the search engine is extremely important in reducing the difficulty of information exploration. An internet spider, bot, or program known as a web crawler is used by search engines to maintain an index and retrieve online sites in response to user queries. Web pages are really crawled in order to be explored, starting from the seed URL. In this study, the standard protocols of the literature survey were followed to describe the crawler mechanism also analyzed and compared several research publications to explain the crawler types, individual architectural types, policies, methodologies, metrics, techniques, and equipments for every category was covered. Additionally, the maximum crawler algorithms and the open-source crawler tools were listed and described. This paper gives insight into choosing the optimal crawler algorithm for search engine optimization and for various applications.

Read the paper · More papers on PaperTik