Measurement of Progress in Machine Translation

Yvette Graham, Timothy J. Baldwin, Aaron Harwood, Alistair Moffat, Justin Zobel · 2012

Machine translation (MT) systems can only be improved if their performance can be reliably measured and compared. However, measurement of the quality of MT output is not straightforward, and, as we discuss in this paper, relies on correlation with inconsistent human judgments. Even when the question is captured via “is translation A better than translation B ” pairwise comparisons, empirical evidence shows that inter-annotator consistency in such experiments is not particularly high; for intra-judge consistency – computed by showing the same judge the same pair of candidate translations twice – only low levels of agreement are achieved. In this paper we review current and past methodologies for human evaluation of translation quality, and explore the ramifications of current practices for automatic MT evaluation. Our goal is to document how the methodologies used for collecting human judgments of machine translation quality have evolved; as a result, we raise key questions in connection with the low levels of judgment agreement and the lack of mechanisms for longitudinal evaluation. 1

Read the paper · More papers on PaperTik