Crowdsourcing for structured labeling with applications to protein folding
Jian Xun Peng, Qiang Liu, Alexander Ihler · 2013
Label aggregation is an important problem in crowdsourcing research. Besides simple classification and regression tasks, many crowdsourcing experiments require workers to return structured outputs as labels. These labels are often inherently constrained, making the aggregation task very difficult. We approach this problem by formulating it as an exemplar-based clustering problem and choosing the most representative cluster centers as the aggregated label. This approach naturally extends the majority voting method and incorporates workers’ abilities. We evaluate our method on the most recent protein structure prediction experiment CASP10, where the structured outputs are 3D coordinates of proteins, and find that our method outperforms both majority voting and the single best predictor. These results are very promising since in CASP10 all earlier aggregation methods performed worse than the single best predictor. 1.