A Framework for Efficient Document Ranking Using Order and Non Order Based Fitness Function

Hazra Imran, Aditi Sharan · 2010

Abstract—One central problem of information retrieval is to determine the relevance of documents with respect to the user information needs. The choice of similarity measure is crucial for improving search effectiveness of a retrieval system. Different similarity measures have been suggested to match the query and documents. This study investigates the use of Genetic Algorithm to increase the efficiency of information retrieval by defining a combined similarity measure. Genetic Algorithm has been used for learning weights of the components of the combined similarity measure. We have provided a weight-learning algorithm for the same. We have considered order based and non-order based fitness functions to evaluate the goodness of the solution. A non-order based fitness function is based on recall-precision values only. However, it has been observed that a better fitness function can be obtained if we also consider the order in which relevant documents are retrieved. This leads to an idea of order based fitness functions. We evaluated the efficacy of a genetic algorithm with various fitness functions. Further, we provide a framework for applying genetic algorithms to improve the retrieval efficiency by combing various similarity measures. The experiments have been carried out on TREC data collection. The results have been compared with various well-known similarity measures. Index Terms—Document retrieval, genetic algorithms, similarity measures, information retrieval, vector space Model. I.

Read the paper · More papers on PaperTik