Domain Adaptation in Statistical Machine Translation of User-Forum Data using Component Level Mixture Modelling.
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith · 2011
This paper reports experiments on adapting components of a Statistical Machine Trans-lation (SMT) system for the task of trans-lating online user-generated forum data from Symantec. Such data is monolingual, and differs from available bitext MT training re-sources in a number of important respects. For this reason, adaptation techniques are impor-tant to achieve optimal results. We investi-gate the use of mixture modelling to adapt our models for this specific task. Individual models, created from different in-domain and out-of-domain data sources, are combined us-ing linear and log-linear weighting methods for the different components of an SMT sys-tem. The results show a more profound effect of language model adaptation over translation model adaptation with respect to translation quality. Surprisingly, linear combination out-performs log-linear combination of the mod-els. The best adapted systems provide a sta-tistically significant improvement of 1.78 ab-solute BLEU points (6.85 % relative) and 2.73 absolute BLEU points (8.05 % relative) over the baseline system for English–German and English–French, respectively. 1