Extending BLEU Evaluation Method with Linguistic Weight

Muyun Yang, Junguo Zhu, Jufeng Li, Lixin Wang, Haoliang Qi, Sheng Li, Liu Daxin · 2008

BLEU is one of the most popular metrics for automatic evaluation of machine translation quality. Focusing on its ignorance of different effects of various translation units upon translation quality, this paper extends proper weights to different words and n-grams in the framework of BLEU. The linear regression method is adopted to capture the human perception on translation quality via word types and n-gram length. Compared with other linguistic-rich metrics based on machine learning, the proposed approach is simple and largely preserves BLEUpsilas advantage of language independence. Experimental results indicate that this method brings a much better evaluation performance for both human translation and machine translation than original BLEU.

Read the paper · More papers on PaperTik