Making the Most of Crowdsourced Document Annotations: Confused Supervised LDA
Paul Felt, Eric K. Ringger, Jordan Lee Boyd-Graber, Kevin D. Seppi · 2015
Corpus labeling projects frequently use low-cost workers from microtask marketplaces; however, these workers are often inexperienced or have misaligned incentives.Crowdsourcing models must be robust to the resulting systematic and nonsystematic inaccuracies.We introduce a novel crowdsourcing model that adapts the discrete supervised topic model sLDA to handle multiple corrupt, usually conflicting (hence "confused") supervision signals.Our model achieves significant gains over previous work in the accuracy of deduced ground truth.