Reducing Annotation Effort for Quality Estimation via Active Learning

Daniel Beck, Lucia Specia, Trevor Cohn · 2013

Quality estimation models provide feedback on the quality of machine translated texts. They are usually trained on humanannotated datasets, which are very costly due to its task-specific nature. We investigate active learning techniques to reduce the size of these datasets and thus annotation effort. Experiments on a number of datasets show that with as little as 25 % of the training instances it is possible to obtain similar or superior performance compared to that of the complete datasets. In other words, our active learning query strategies can not only reduce annotation effort but can also result in better quality predictors. 1

Read the paper · More papers on PaperTik