Parametric Weighting of Parallel Data for Statistical Machine Translation

Kashif Ur Rehman Shah, Loïc Barrault, Holger Schwenk · 2011

During the last years there is increas-ing interest in methods that perform some kind of weighting of heterogeneous paral-lel training data when building a statistical machine translation system. It was for in-stance observed that training data that is close to the period of the test data is more valuable than older data (Hardt and Elm-ing, 2010; Levenberg et al., 2010). In this paper we obtain such a weighting by resampling alignments using weights that decrease with the temporal distance of bi-texts to the test set. By these means, we can use all the available bitexts and still put an emphasis on the most recent one. The main idea of our approach is to use a parametric form or meta-weights for the weighting of the different parts of the bi-texts. This ensures that our approach has only few parameters to optimize. We re-port experimental results on the Europarl corpus, translating from French to En-glish and further verified it on the official WMT’11 task, translating from English to French. Our method achieves improve-ments of about 0.6 points BLEU on the test set with respect to a system trained on data without any weighting. 1

Read the paper · More papers on PaperTik