Continuous Measurement Scales in Human Evaluation of Machine Translation

Yvette Graham, Timothy J. Baldwin, Alistair Moffat, Justin Zobel · Trinity's Access to Research Output (TARA) (Trinity College Dublin) · 2013

We explore the use of continuous rating scales for human evaluation in the context of machine translation evaluation, comparing two assessor-intrinsic quality-control techniques that do not rely on agreement with expert judgments. Experiments employing Amazon's Mechanical Turk service show that quality-control techniques made possible by the use of the continuous scale show dramatic improvements to intra-annotator agreement of up to +0.101 in the kappa coefficient, with inter-annotator agreement increasing by up to +0.144 when additional standardization of scores is applied.

Read the paper · More papers on PaperTik