Online Language Model Adaptation for Spoken Dialog Translation

Mauro Cettolo, Nicola Bertoldi, Marcello Federico · 2009

This paper focuses on the problem of language model adapta-tion in the context of Chinese-English cross-lingual dialogs, as set-up by the challenge task of the IWSLT 2009 Evalu-ation Campaign. Mixtures of n-gram language models are investigated, which are obtained by clustering bilingual train-ing data according to different available human annotations, respectively, at the dialog level, turn level, and dialog act level. For the latter case, clustering of IWSLT data was in fact induced through a comparable Italian-English parallel corpus provided with dialog act annotations. For the sake of adaptation, mixture weight estimation is performed either at the level of single source sentence or test set. Estimated weights are then transferred to the target language mixture model. Experimental results show that, by training different specific language models weighted according to the actual input instead of using a single target language model, signifi-cant gains in terms of perplexity and BLEU can be achieved. 1.

Read the paper · More papers on PaperTik