Rank-Aware Crawling of Hidden Web sites

George Valkanas, Alexandros Ntoulas, Dimitrios Gunopulos · 2011

An ever-increasing amount of valuable information on the Web today is stored inside online databases and is accessible only after the users issue a query through a search interface. Such information is collectively called the“Hidden Web”and is mostly inaccessible by traditional search engine crawlers that scout the Web following links. Since the only way to access the Hidden Web pages is through the submission of queries to the Hidden Web sites, previous work [14, 18] has focused on how to automatically generate queries in order to incrementally retrieve and cover a Hidden Web site in depth, as much as possible. Forcertainapplicationshoweveritisnotnecessarytohave crawled a Hidden-Web site in-depth. For example, a metasearcher or a content aggregator will utilize only the top portion of the ranked result lists coming from the querying of a Hidden Web site instead of its full content. Hence, if we can crawl a Hidden Web site in breadth, i.e. download just the top results for all potential queries, we can enable such applications without the need for allocating resources for fully crawling a potentially huge Hidden Web site. In this paper we present algorithms for crawling a Hidden Web site by taking the ranking of the results into account. Since we do not know all potential queries that may be directed to the Web site in advance, we study how to approximate the site’s ranking function so that we can compute the top results based on the data collected so far. We provide a framework for performing ranking-aware Hidden Web crawling and we show experimental results on a real Web site demonstrating the performance of our methods.

Read the paper · More papers on PaperTik