Using Hyperlink Information to Improve Crawler's Searching Strategy

Wanli Zuo · Journal of Changchun Post and Telecommunication Institute · 2005

A crawler must face two problems when it searches pages in internet. One is that an internet search engine cannot contain entire Web pages due to huge volumes of data in internet. Because of the constraint of hardware resource, the other is that the Web pages stored in the internet search engine are limited. Crawling in Web space according to the strategy of the traditional breadth-first search, if a crawler respects the importance of every page equally, the quality of Web pages collected by the crawler is not high. The algorithm proposed in the present paper makes the best use of hyperlink information contained in the Web pages to the great extent, overcomes the limitations of the blind crawling strategy owned by the traditional breadth-first search. It is improved by using the hyperlink information contained in the Web pages on the base of the traditional breadth-first search. The experiment results show the Web pages crawled by using this algorithm those are relevant to a pre-defined set of topics are over 50%.

Read the paper · More papers on PaperTik