Tackling Sparse Data Issue in Machine Translation Evaluation

Ondřej Bojar, Kamil Kos, David Mareček · Meeting of the Association for Computational Linguistics · 2010

We illustrate and explain problems of n-grams-based machine translation (MT) metrics (e.g. BLEU) when applied to morphologically rich languages such as Czech. A novel metric SemPOS based on the deep-syntactic representation of the sentence tackles the issue and retains the performance for translation to English as well.

Read the paper · More papers on PaperTik