Detecting annotation noise in automatically labelled data

Ines Rehbein, Josef Ruppenhofer · 2017

We introduce a method for error detection in automatically annotated text, aimed at supporting the creation of high-quality language resources at affordable cost.Our method combines an unsupervised generative model with human supervision from active learning.We test our approach on in-domain and out-of-domain data in two languages, in AL simulations and in a real world setting.For all settings, the results show that our method is able to detect annotation errors with high precision and high recall.

Read the paper · More papers on PaperTik