Learning to Rank Using Semantic Features in Document Retrieval

Weixin Tian, Fuxi Zhu · 2009

This paper describes an approach to retrieve documents adopting machine learning method. The application of machine learning to document retrieval, which can be so called ldquolearning to rankrdquo, has been a hot research topic in the information retrieval and machine learning communities recently. One of the characters that discriminates this work from other studies is the use of semantic features while classifying the documents. Firstly, we extract the semantic structures from collection and index documents on these structures. Later, we construct a SVM classifier and generate it using pairwise training data. Lastly, we use this SVM to judge the relevance of documents according to given queries. In the extracting of semantic structures phrase, an analysis system that we developed previously is used. The analysis system is based on modifying relations (MR) and on the help of a knowledge base called MRKB. We give a general introduction about the system in this paper. The experiment is done on the benchmark dataset OHSUMED, and the experimental results shows that the proposed method outperforms other approaches based barely on string frequency.

Read the paper · More papers on PaperTik