On Penalising Late Arrival of Relevant Documents in Information Retrieval Evaluation with Graded Relevance

Tetsuya Sakai · 2007

Large-scale information retrieval evaluation efforts such as TREC and NTCIR have tended to adhere to binary-relevance evaluation metrics, even when graded relevance data were available. However, the NTCIR-6 Crosslingual Task has finally started adopting graded-relevance metrics, though only as additional metrics. This paper compares three existing graded-relevance metrics that were mentioned in the Call for Participation of the NTCIR-6 Crosslingual Task in terms of the ability to control how severely “late arrival ” of relevant documents should be penalised. We argue and demonstrate that Q-measure is more flexible than normalised Discounted Cumulative Gain and generalised Average Precision. We then suggest a brief guideline for conducting a reliable information retrieval evaluation with graded relevance. Keywords: Q-measure, nDCG, generalised average precision, rank correlation, bootstrap sensitivity.

Read the paper · More papers on PaperTik