Agree to Disagree: Analysis of Inter-Annotator Disagreements in Human Evaluation of Machine Translation Output

Maja Popović · 2021

This work describes an analysis of interannotator disagreements in human evaluation of machine translation output.The errors in the analysed texts were marked by multiple annotators under guidance of different quality criteria: adequacy, comprehension, and an unspecified generic mixture of adequacy and fluency.Our results show that different criteria result in different disagreements, and indicate that a clear definition of quality criterion can improve the inter-annotator agreement.Furthermore, our results show that for certain linguistic phenomena which are not limited to one or two words (such as word ambiguity or gender) but span over several words or even entire phrases (such as negation or relative clause), disagreements do not necessarily represent "errors" or "noise" but are rather inherent to the evaluation process.On the other hand, for some other phenomena (such as omission or verb forms) agreement can be easily improved by providing more precise and detailed instructions to the evaluators.

Read the paper · More papers on PaperTik