Result Merging in a Peer-to-Peer Web Search Engine
Sergey Chernov, Gerhard Weikum, Christian Zimmer · Max Planck Institute for Plasma Physics · 2005
A tremendous amount of information in the Internet requires powerful search engines. Currently, only the commercial centralized search engines like Google can process terabytes of Web documents. Such approaches fail in indexing the “Hidden Web” located in the intranets and local databases, and with an exponential growing of information volume the situation becomes even worse. Peer-to-Peer (P2P) systems can be pursued for extending the current search capabilities. The Minerva project is a Web search engine based on a P2P architecture. In this thesis, we investigate the effectiveness of the different result merging methods for the Minerva system. Each peer provides an efficient search engine for its own focused Web crawl. Each peer can pose a query against a number of selected peers; the selection is based on a database ranking algorithm. The best top-k results from several highly ranked peers are collected by the query initiator and merged into a single list. We address problem of the result merging. We select several merging methods, which are feasible for use in a heterogeneous, dynamic, distributed environment. The experimental framework for these methods was implemented and the effectiveness of the merging techniques was studied with the TREC Web data. The language modeling based ranking method produced the most robust and accurate results under the different conditions. We also proposed a new merging method, which incorporates the preference-based language model. The novelty of the method is that the preference-based language model is obtained from the pseudo-relevance feedback on the best peer in the database ranking. In every tested setup, the new method was at least as effective as the baseline or slightly better.