Intelligent crawling based on rough set for web resource discovery

LingXia Hu · 2010

The rapid development of the Internet brings a new problem, which is how to rapidly and effectively retrieve needed web resource from vast number of web pages. The progress of machine learning techniques shows a new direction of solving this problem. In this paper, intelligent crawling algorithm based on rough set is proposed. The algorithm use the hypertext features behavior in order to perform topic specific resource discovery. Our experiment in this regard has provided better Harvest rate and better Target recall for focused crawling.

Read the paper · More papers on PaperTik