Using Comparable Corpora to Augment Statistical Machine Translation Models in Low Resource Settings

Ann Irvine · 2014

Previously, statistical machine translation (SMT) models have been estimated from parallel corpora, or pairs of translated sentences. In this thesis, we directly incorporate comparable corpora into the estimation of end-to-end SMT models. In contrast to parallel corpora, comparable corpora are pairs of monolingual corpora that have some cross-lingual similarities, for example topic or publication date, but that do not necessarily contain any direct translations. Comparable corpora are more readily available in large quantities than parallel corpora, which require significant human effort to compile. We use comparable corpora to estimate machine translation model parameters and show that doing so improves performance in settings where a limited amount of parallel data is available for training. The major contributions of this thesis are the following: ‚ We release ‘language packs ’ for 151 human languages, which include bilingual dictionaries, comparable corpora of Wikipedia document pairs, comparable cor-pora of time-stamped news text that we harvested from the web, and, for

Read the paper · More papers on PaperTik