Intelligent crawling based on rough set for web resource discovery
LingXia Hu · 2010
The rapid development of the Internet brings a new problem, which is how to rapidly and effectively retrieve needed web resource from vast number of web pages. The progress of machine learning techniques shows a new direction of solving this problem. In this paper, intelligent crawling algorithm based on rough set is proposed. The algorithm use the hypertext features behavior in order to perform topic specific resource discovery. Our experiment in this regard has provided better Harvest rate and better Target recall for focused crawling.