Evaluating ADM on a Four-Level Relevance Scale Document Set from NTCIR

Vincenzo Della Mea, Luca Di Gaspero, Stefano Mizzaro · NTCIR · 2004

Most common effectiveness measures for Information Retrieval (IR) systems are based on the assumptions of binary relevance (either a document is relevant to a given query or it is not) and binary retrieval (either a document is retrieved or it is not). These assumptions are often questioned, since almost everybody agrees that relevance and retrieval are matter of degree (three or more categories, if not a continuum). However, the standard practice in IR systems evaluation remains based on the use of precision and recall (and related measures), thus hindering IR development and evaluation. We recently questioned these assumptions, and proposed a new measure named ADM (Average Distance Measure), in order to pass from binary to continuous relevance and retrieval [21, 6]. In this paper we describe the idea on which ADM is based, a conceptual analysis of this new measure, and the results of a first experimental validation on TREC data, which feature 2-levels human relevance judgments and IR systems that rank the retrieved documents. Furthermore, we present some experimental results on NTCIR-3 data, which feature 4-levels human relevance judgments and IR systems that assign a numeric score to the retrieved documents. Both conceptual and experimental results show that ADM might be potentially adequate, provided that IR system designers take some care in sys

Read the paper · More papers on PaperTik