Making the Most of Crowdsourced Document Annotations: Confused Supervised LDA

Paul Felt, Eric K. Ringger, Jordan Lee Boyd-Graber, Kevin D. Seppi · 2015

Corpus labeling projects frequently use low-cost workers from microtask marketplaces; however, these workers are often inexperienced or have misaligned incentives.Crowdsourcing models must be robust to the resulting systematic and nonsystematic inaccuracies.We introduce a novel crowdsourcing model that adapts the discrete supervised topic model sLDA to handle multiple corrupt, usually conflicting (hence "confused") supervision signals.Our model achieves significant gains over previous work in the accuracy of deduced ground truth.

Read the paper · More papers on PaperTik