Evaluating Relative Retrieval Effectiveness with Normalized Residual Gain

Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, Charles L. A. Clarke · 2024

Traditional search evaluation metrics, such as MRR and NDCG, focus on absolute measures of effectiveness. While they allow us to compare the absolute performance of one retrieval method to another, we do not know if systems with similar absolute performance achieve this performance by finding the same items, or by finding different items with similar relevance grades. To address this problem, several recent proposals have measured the relative performance of a retrieval method in the context of the results from one or more other methods. In this paper, we address theoretical limitations of these proposals and introduce a new metric called Normalized Residual Gain (NRG) that can be seen as an extension of the underlying absolute metric, rather than as an entirely new metric. Operating in the context of the results retrieved by one or more other methods, NRG adjusts gain values according to the browsing model of the absolute metric. Through testing over the MS MARCO dev small and TREC DL 2019 datasets, we find that higher absolute effectiveness does not necessarily correlate with a higher NRG score, which will vary depending on context. In particular, in the context of modern neural models, NRG suggests that a traditional BM25 ranker continues to find relevant items missed by even the best neural models.

Read the paper · More papers on PaperTik