Online Language Model adaptation via N-gram Mixtures for Statistical Machine Translation

Germán Sanchis-Trilles, Mauro Cettolo, Fondazione Bruno Kessler · 2010

The problem of language model adaptation in statistical machine translation is considered. A mixture of language models is employed, which is obtained by clustering the bilingual training data. Unsupervised clustering is guided by either the development or the test set. Different mixture weight estimation schemes are proposed and compared, at the level of either single or all source sentences. Experimental results show that, by training different specific language models weighted according to the actual input instead of using a single target language model, translation quality is improved, as measured by BLEU and TER. 1

Read the paper · More papers on PaperTik