Noise or additional information? Leveraging crowdsource annotation item agreement for natural language tasks.

Emily Jamison, Iryna Gurevych · 2015

In order to reduce noise in training data, most natural language crowdsourcing annotation tasks gather redundant labels and aggregate them into an integrated label, which is provided to the classifier.However, aggregation discards potentially useful information from linguistically ambiguous instances.For five natural language tasks, we pass item agreement on to the task classifier via soft labeling and low-agreement filtering of the training dataset.We find a statistically significant benefit from low item agreement training filtering in four of our five tasks, and no systematic benefit from soft labeling.

Read the paper · More papers on PaperTik