Reordering rules for phrase-based statistical machine translation.
Boxing Chen, Mauro Cettolo, Marcello Federico · 2006
This paper proposes the use of rules automatically ex-tracted from word aligned training data to model word reordering phenomena in phrase-based statistical machine translation. Scores computed from matching rules are used as additional feature functions in the rescoring stage of the automatic translation process from various languages to En-glish, in the ambit of a popular traveling domain task. Rules are defined either on Part-of-Speech or words. Part-of-Speech rules are extracted from and applied to Chinese, while lexicalized rules are extracted from and applied to Chi-nese, Japanese and Arabic. Both Part-of-Speech and lexicalized rules yield an ab-solute improvement of the BLEU score of 0.4-0.9 points without affecting the NIST score, on the Chinese-to-English translation task. On other language pairs which differ a lot in the word order, the use of lexicalized rules allows to observe significant improvements as well. 1.