Linear Mixture Models for Robust Machine Translation
Marine Jacinthe Carpuat, Cyril Goutte, George Foster · 2014
As larger and more diverse parallel texts become available, how can we lever-age heterogeneous data to train robust machine translation systems that achieve good translation quality on various test domains? This challenge has been ad-dressed so far by repurposing techniques developed for domain adaptation, such as linear mixture models which combine estimates learned on homogeneous sub-domains. However, learning from large heterogeneous corpora is quite different from standard adaptation tasks with clear domain distinctions. In this paper, we show that linear mixture models can re-liably improve translation quality in very heterogeneous training conditions, even if the mixtures do not use any domain knowledge and attempt to learn generic models rather than adapt them to the tar-get domain. This surprising finding opens new perspectives for using mixture mod-els in machine translation beyond clear cut domain adaptation tasks. 1