A popularity-based URL ordering algorithm for crawlers

Aravind Chandramouli, Susan Gauch, Joshua Eno · 2010

Uniform Resource Locator (URL) ordering algorithms are used by Web crawlers to determine the order in which to download pages from the Web. The current approaches for URL ordering based on link structure are expensive and/or miss many good pages, particularly in social network environments. In this paper, we present a novel URL ordering algorithm that exploits the access count information present in the Web logs on the individual Websites. In particular, we develop algorithms based on internal and external counts and by using this popularity information for URL ordering, we are able to retrieve high quality pages earlier in the crawl. We perform our experiments on two data sets using the Web logs from university and CiteSeer Websites and, on these data sets, we achieve a statistically significant improvement in the ordering of the high quality pages (as indicated by Google's PageRank) of 57.2% and 65.7% over that of a breadth-first search crawl.

Read the paper · More papers on PaperTik