Using Past Queries for Resource Selection in Distributed Information Retrieval
Sulleyman Cetintas, Luo Si, Hao Yuan · Purdue e-Pubs (Purdue University System) · 2011
Federated text search provides a unified search interface for multiple search engines of distributed text information sources.Resource selection is an important component for federated text search, which selects a small number of information sources that contain the largest number of relevant documents for a user query.Most prior research of resource selection focused on selecting information sources by analyzing static information of available information sources that is sampled in the offline manner.On the other hand, most prior research ignored a large amount of valuable information like the results from past queries.This paper proposes a new resource selection technique (which is called qSim) that utilizes the search results of past queries for estimating the utilities of available information sources for a specific user query.The new algorithm calculates the query similarities between a specific query and all past queries, and then estimates the utilities of available information sources by the weighted combination of results of past queries with respect to the query similarities.The new resource selection algorithm is practical as it does not require relevance judgment of past queries and it only utilizes regression based results merging method to rank the results of past queries.Furthermore, a combined resource selection approach is proposed to integrate the two approaches of learning from past queries and using static sampled information.An extensive set of experiments demonstrate the effectiveness of the new resource selection algorithm of learning from past queries as well as the combined resource selection algorithm in several configurations.