Evaluation of a Graph-based Topical Crawler.
Aurel Cami, Narsingh Deo · 2006
Abstract – Topical (or, focused) crawlers have become important tools in dealing with the massiveness and dynamic nature of the World Wide Web. Guided by a data mining component that monitors and analyzes the boundary of the set of crawled pages, a focused crawler selectively seeks out pages on a pre-defined topic. Recent research indicates that both the textual content of web pages and the structural information enclosed in the Web graph need to be exploited in order to build high quality focused crawlers. While, a variety of text-based and graphbased measures of similarity that can direct a focused crawler toward relevant pages have been developed, much remains to be done toward formally evaluating and ranking the effectiveness of various focused crawling algorithms. Inspired by a recent and comprehensive evaluation framework for focused crawlers, we analyze the performance of a graph-based algorithm and compare it with two other algorithms: a breadth-first one and a textbased, best-first one. The results suggest that our graphbased algorithm is faster and only slightly less effective than the text-based, best-first algorithm, while significantly outperforming the breadth-first one.