Relevance assessment

Peter Bailey, Nick Craswell, Ian M. Soboroff, Paul Thomas, Arjen P. de Vries, Emine Yılmaz · 2008

We investigate to what extent people making relevance judgements for a reusable IR test collection are exchangeable. We consider three classes of judge: "gold standard" judges, who are topic originators and are experts in a particular information seeking task; "silver standard" judges, who are task experts but did not create topics; and "bronze standard" judges, who are those who did not define topics and are not experts in the task.

Read the paper · More papers on PaperTik