Inflating Training Data for Statistical Machine Translation using Unaligned Monolingual Data

Wei Yang, Zhongwen Zhao, Yves Lepage · 2015

In data-driven machine translation, parallel corpora are an extremely important resource. For language pairs that involve English, there exist many freely available bilingual or multilingual parallel corpora, especially for European languages. To improve the translation quality for less-resourced language pairs, such as Chinese–Japanese, larger and larger aligned training data are needed. The constitution of large bilingual corpora is not easy for less documented language pairs. In this paper, we show how to construct a Chinese–Japanese quasi-parallel corpus automatically by using analogical associations based on a small amount of parallel sentences and a reasonable amount of monolingual data. We perform SMT experiments in Chinese–Japanese and compare a baseline system and a system build by adding the quasi-parallel corpus. On the same test set, the translation quality significantly improved over the baseline system.

Read the paper · More papers on PaperTik