Efficient Elicitation of Annotations for Human Evaluation of Machine Translation

Keisuke Sakaguchi, Matt Post, Benjamin Van Durme · 2014

A main output of the annual Workshop on Statistical Machine Translation (WMT) is a ranking of the systems that partici-pated in its shared translation tasks, pro-duced by aggregating pairwise sentence-level comparisons collected from human judges. Over the past few years, there have been a number of tweaks to the ag-gregation formula in attempts to address issues arising from the inherent ambigu-ity and subjectivity of the task, as well as weaknesses in the proposed models and the manner of model selection. We continue this line of work by adapt-ing the TrueSkillTM algorithm — an online approach for modeling the relative skills of players in ongoing competitions, such as Microsoft’s Xbox Live — to the hu-man evaluation of machine translation out-put. Our experimental results show that TrueSkill outperforms other recently pro-posed models on accuracy, and also can significantly reduce the number of pair-wise annotations that need to be collected by sampling non-uniformly from the space of system competitions. 1

Read the paper · More papers on PaperTik