Statistical Machine Translation with Automatic Identification of Translationese
Naama Twitto, Noam Ordan, Shuly Wintner · 2015
Translated texts (in any language) are so markedly different from original ones that text classification techniques can be used to tease them apart.Previous work has shown that awareness to these differences can significantly improve statistical machine translation.These results, however, required meta-information on the ontological status of texts (original or translated) which is typically unavailable.In this work we show that the predictions of translationese classifiers are as good as meta-information.First, when a monolingual corpus in the target language is given, to be used for constructing a language model, predicting the translated portions of the corpus, and using only them for the language model, is as good as using the entire corpus.Second, identifying the portions of a parallel corpus that are translated in the direction of the translation task, and using only them for the translation model, is as good as using the entire corpus.We present results from several language pairs and various data sets, indicating that these results are robust and general.