An approach to information retrieval based on the language-modeling kernel
Vijay V. Raghavan, Ying Xie · 2004
In this dissertation, we propose a novel integrated information retrieval approach that provides a unified solution to challenging problems faced by existing, popular information retrieval models. The first problem relates to the vector space model. We found that different information needs in fact require different vector spaces to represent documents. However, the question of how to dynamically build optimal vector spaces that are tailored to users' different information needs remains unexplored. The second problem relates to the language modeling approach. It is conceptually difficult for the language modeling approach to utilize the advantages of machine learning techniques. In order to solve these problems, we designed a kernel function called the language-modeling kernel. This kernel function retains all the modeling benefits provided by the language modeling approach. Meanwhile, for each information need, it dynamically determines an optimal vector space, based on which machine learning algorithms, such as the support vector machine (SVM), can be applied to find an optimal decision boundary that separates the relevant documents from the non-relevant ones. Furthermore, an effective double learning strategy is proposed based on the language-modeling kernel. Large-scale experiments on standard test-beds show that our approach makes significant improvements over the state-of-art information retrieval methods.