Computing Maximized Effectiveness Distance for Recall-Based Metrics
Alistair Moffat · IEEE Transactions on Knowledge and Data Engineering · 2017
Given an effectiveness metric M(·), two ordered document rankings X1and X2generated by a score-based information retrieval activity, and relevance labels in regard to some subset (possibly empty) of the documents appearing in the two rankings, Tan and Clarke's Maximized Effectiveness Distance (MED) computes the greatest difference in metric score that can be achieved that is consistent with all provided information, crystallized via a set of relevance assignments to the unlabeled documents such that |M(X1) - M(X2)| is maximized. The closer the maximized effectiveness distance is to zero, the more similar X1and X2can be considered to be from the point of view of the metric M(·). Here, we consider issues that arise when Tan and Clarke's definitions are applied to recall-based metrics, notably normalized discounted cumulative gain (NDCG), and average precision (AP). In particular, we show that MED can be applied to NDCG without requiring an a priori assumption in regard to the total number of relevant documents; we also show that making such an assumption leads to different outcomes for both NDCG and average precision (AP) compared to when no such assumption is made.