Pairwise Crowd Judgments

Ziying Yang, Alistair Moffat, Andrew H. Turpin · 2018

Relevance judgments are conventionally formed by small numbers of experts using ordinal relevance scales defined by two or more relevance categories. Such judgments often contain many ties: documents in the same category that cannot be separated by relevance. Here we explore the use of crowd-sourcing and combined three-way relevance assessments using pairwise preference, absolute relevance, and relevance ratio, with forced choice testing and embedded quality control processes, seeking to reduce assessment ties, and to increase judgment consistency. In particular, the crowd-sourced judgments from these three approaches were normalized into numeric relevance scores, and compared against judgments arising via three previous techniques: NIST binary; Sormunen; and magnitude estimation. The relationship between generated judgment reliability and number of document pairs assessed was also explored, as was the effect that factors such as document length, topic difficulty, number of documents judged, and assessment time, have on assessment reliability. Lastly, we investigate the extent to which the methodology used to collect judgments affects the ability of an experiment to discriminate between IR systems.

Read the paper · More papers on PaperTik