Exploiting Document-Level Context for Data-Driven Machine Translation

Ralf D. Brown · 2008

This paper presents a method for exploiting document-level similarity between the docu-ments in the training corpus for a corpus-driven (statistical or example-based) machine translation system and the input documents it must translate. The method is simple to imple-ment, efficient (increases the translation time of an example-based system by only a few percent), and robust (still works even when the actual document boundaries in the input text are not known). Experiments on French-English and Arabic-English showed relative gains over the same system without using document-level similarity of up to 7.4 % and 5.4%, respectively, on the BLEU metric. 1

Read the paper · More papers on PaperTik