A topic-specific Web crawler based on content and structure mining
Rong Qian, Kejun Zhang, Geng Zhao · 2013
This paper discusses a topic-specific intelligent Web crawler based on Web content and structure mining. The method takes advantage of the characteristics of the neural network and introduces the reinforcement learning to find the relativity between the crawled web pages and the topic. When calculating the correlation, we just select the important tags of HTML makeup of the Web page, to analyze the web page's content and structure. The experiments show that our method improves the efficiency and accuracy clearly.