A WEBIR Crawling Framework for Retrieving Highly Relevant Web Documents: Evaluation Based on Rank Aggregation and Result Merging Algorithms

Shashi Shekhar, K. V. Arya, Rohit Agarwal, Rakesh Kumar · 2011

Finding relevant information on the web is an ongoing problem. Commercial search engines like Google rely on sophisticated algorithms to index huge collection of web pages to make them accessible to user queries. Users, however, are still frequently overloaded with irrelevant results. The required information is available in replicated manner scattered in various disjoint databases. For effective web information retrieval, user need to consult several commercial search engines working on different architecture and principles. Rank aggregation and Result merging is the key component of a crawling mechanism used by the commercial search engines. Once the results from various search engines are collected, they need to be merged into a single unified ranked list. The effectiveness of any crawling mechanism is closely related to the rank aggregation and result merging algorithm it employs. In this paper, we investigate a variety of rank aggregation and result merging algorithms based on a wide range of available information. The effectiveness of these algorithms is then compared experimentally to our proposed crawling framework based on queries from the TREC Web track and 3 most popular general-purpose search engines. Our experiments yield two important results. First, simple result merging strategies can outperform Google, Yahoo and MSN Live. Second, Proposed Content Based Result Aggregation (CBRA) algorithm outperforms other existing content based merging algorithms based on full document content.

Read the paper · More papers on PaperTik