{Sequential crowdsourced labeling as an epsilon-greedy exploration in a Markov Decision Process}

Vikas Chandrakant Raykar, Priyanka Agrawal · 2014

Crowdsourcing marketplaces are widely used for curating large annotated datasets by col-lecting labels from multiple annotators. In such scenarios one has to balance the trade-off between the accuracy of the collected la-bels, the cost of acquiring these labels, and the time taken to finish the labeling task. With the goal of reducing the labeling cost, we introduce the notion of sequential crowd-sourced labeling, where instead of asking for all the labels in one shot we acquire labels from annotators sequentially one at a time. We model it as an epsilon-greedy exploration in a Markov Decision Process with a Bayesian decision theoretic utility function that incor-porates accuracy, cost and time. Experimen-tal results confirm that the proposed sequen-tial labeling procedure can achieve similar ac-curacy at roughly half the labeling cost and at any stage in the labeling process the algo-rithm achieves a higher accuracy compared to randomly asking for the next label. 1

Read the paper · More papers on PaperTik