Design of Four Web Crawlers based on Python

Pan Liu, Yihao Li, Xuankui Zheng, Shili Ai, Wenjie Zhang · 2022

The data on websites is an important source of data for both big data analysis and machine learning. Due to the limitation of data crawling on some websites, the general web crawler will be invalid. To facilitate the crawling of data in websites with different structures, this paper introduces four types of web crawlers. Then, based on some third party libraries developed for Python, the corresponding Python programs are designed respectively for these four web crawlers. This paper provides a technical guide for those researchers who want to construct web crawlers quickly.

Read the paper · More papers on PaperTik