Beyond Linguistic Equivalence. An Empirical Study of Translation Evaluation in a Translation Learner Corpus
Mihaela Vela, Anne-Kathrin Schumann, Andrea Wurm · 2014
The realisation that fully automatic translation in many settings is still far from producing output that is equal or superior to human translation has lead to an intense interest in translation evaluation in the MT community.However, research in this field, by now, has not only largely ignored the tremendous amount of relevant knowledge available in a closely related discipline, namely translation studies, but also failed to provide a deeper understanding of the nature of "translation errors" and "translation quality".This paper presents an empirical take on the latter concept, translation quality, by comparing human and automatic evaluations of learner translations in the KOPTE corpus.We will show that translation studies provide sophisticated concepts for translation quality estimation and error annotation.Moreover, by applying well-established MT evaluation scores, namely BLEU and Meteor, to KOPTE learner translations that were graded by a human expert, we hope to shed light on properties (and potential shortcomings) of these scores.