Better Effectiveness Metrics for SERPs, Cards, and Rankings
Paul Thomas, Alistair Moffat, Peter Bailey, Falk Scholer, Nick Craswell · 2018
Offline metrics for IR evaluation are often derived from a user model that seeks to capture the interaction between the user and the ranking, conflating the interaction with a ranking of documents with the user's interaction with the search results page. A desirable property of any effectiveness metric is if the scores it generates over a set of rankings correlate well with the "satisfaction" or "goodness" scores attributed to those same rankings by a population of searchers.