Information Retrieval Using PU Learning Based Re-ranking

Chong Teng, Yanxiang He, Donghong Ji, Han Ren, Lingpeng Yang, Wei Xiong · 2015

In this paper, we describe our approach for information retrieval for question answering (IR4QA) on simple Chinese language of NTCIR-7 tasks. Firstly, we use both bi-grams and single Chinese characters as index units and use OKAPI BM25 as retrieval model. Secondly, we re-rank all documents ’ orders for the first retrieval documents. We focus mostly on the document re-ranking technique. We address probabilistically labeling relevant degree between the first retrieval documents and query topics. In other words, we want to know the probability of a document belongs to relevance/irrelevance class. We employ PU㧔positive and unlabeled 㧕 learning to solve this problem, and use Bayesian classifier and EM algorithm in process of computing the probability. Consequently, those relevant documents with high probability are updated rank. Lastly, we use re-ranked retrieved documents to do query expansion. Evaluation at NTCIR-7 shows that our group achieves 0.3862 and 0.3806 MAP based on pseudo-qrels and real qrels respectively.

Read the paper · More papers on PaperTik