A Comprehensive Study on Performance of Serial and Parallel Web Crawlers

Elina Parajuli, K. Vimala Devi, Narayanan Prashanth, Ashutosh Bashyal · 2024

Web crawlers, also known as web spiders or web robots, are automated programs that systematically browse the World Wide Web, indexing web pages and gathering data for various purposes. This paper addresses the inefficiency of default Python crawlers, caused by the Global Interpreter Lock (GIL), in the context of web page collection for search engines. It aims to create a concurrent Python crawler and compare its performance with a single-threaded counterpart. The crawler will undergo three iterations: a sequential version, a concurrent version using thread pool executor, and an aiohttp version for concurrent requests. Performance metrics such as speed of retrieval, internal processing, and memory requirements will be analyzed. The goal is to develop an efficient Python crawler for continuous web crawling, providing insights for developers and users, and comparing Python's performance with another programming language.

Read the paper · More papers on PaperTik