Enhanced Bilingual Evaluation Understudy
Krzysztof Wołk, Krzysztof P. Marasek · Lecture Notes on Information Theory · 2014
Our research extends the Bilingual Evaluation Understudy (BLEU) evaluation technique for statistical machine translation to make it more adjustable and robust.We intend to adapt it to resemble human evaluation more.We perform experiments to evaluate the performance of our technique against the primary existing evaluation methods.We describe and show the improvements it makes over existing methods as well as correlation to them.When human translators translate a text, they often use synonyms, different word orders or style, and other similar variations.We propose an SMT evaluation technique that enhances the BLEU metric to consider variations such as those. I. INTRODUCTIONTo make progress in Statistical Machine Translation (SMT), the quality of its results must be evaluated.It has been recognized for quite some time that using humans to evaluate SMT approaches is very expensive and timeconsuming.[1] As a result, human evaluation cannot keep up with the growing and continual need for SMT evaluation.This led to the recognition that the development of automated SMT evaluation techniques is critical.[1,2] Evaluation is particularly crucial for translation between diverse language pairs, such as Polish and English.Polish has complex declension, 7 cases, 15 gender forms, and complicated grammatical construction procedures.This leads to a very large Polish vocabulary and great complexity in data requirements for SMT.Meanwhile, the order of subjects, verbs, and objects is not important to determine the meaning of a Polish sentence.Instead, many variations of word order mean the same thing in this language.Unlike Polish, the English language does not have declensions.In addition, word order, esp. the Subject-Verb-Object (SVO) pattern, is absolutely crucial to determining the meaning of an English sentence.These differences in the Polish and English languages lead to great translation complexity.In addition, the lack of lexical data availability and phrase models only further complicates SMT between those languages.In [2] Reeder compiled an initial list of SMT evaluation metrics.Further research has led to the development of newer metrics.Prominent metrics include: Bilingual Evaluation Understudy (BLEU); the National Institute of Standards and Technology (NIST) metric; Translation Error Rate (TER), the Metric for Evaluation of Translation with Explicit Ordering (METEOR); Length Penalty, Precision, n-gram Position difference Penalty and Recall (LEPOR); and the Rankbased Intuitive Bilingual Evaluation Score (RIBES).This paper presents extensions to existing SMT evaluation metrics.Section 2 describes the existing evaluation techniques.Our enhanced method is discussed in Section 3. Section 4 describes the experiments we performed to compare our method with existing evaluation methods.Section 5 discusses the results and future potential research in this area.II.