Exploring relevance assessment using crowdsourcing for faceted and ambiguous queries
Sri Devi Ravana, Parnia Samimi, Parisa Dabir Ashtyani · 2014
Relevance assessment usually generated by human experts that can be a time-consuming, difficult and potentially expensive process. Recently, crowdsourcing has been presented to be a fast and cheap method to make relevance assessments in a semi-automatic way. However, previous work on the limit of crowdsourcing in IR evaluation is still inadequate and need further investigation especially for varying nature of queries used during the Web search. In this study, we have observed the responses from the crowdsourced workers and experts in assessing the judgments, compared the agreement between the relevance judgments made by an expert assessor and crowdsourced worker, and finally explored the constancy in system ranking when these two sets of relevance judgments were used to score the systems. Two commercial search engines were compared using two different types of queries namely the faceted and ambiguous queries. In general, both set of judgments ranked the systems in the same order although the absolute systems scores vary slightly. However, the findings shows that the type of query used does influence the agreement between assessors and the system performance measure.