German Compounds and Statistical Machine Translation. Can they get along?
Carla Parra Escartín, Stephan Peitz, Hermann Ney · 2014
This paper reports different experiments created to study the impact of using linguistics to preprocess German com-pounds prior to translation in Statistical Machine Translation (SMT). Compounds are a known challenge both in Machine Translation (MT) and Translation in gen-eral as well as in other Natural Language Processing (NLP) applications. In the case of SMT, German compounds are split into their constituents to decrease the number of unknown words and improve the re-sults of evaluation measures like the Bleu score. To assess to which extent it is neces-sary to deal with German compounds as a part of preprocessing in SMT systems, we have tested different compound splitters and strategies, such as adding lists of com-pounds and their translations to the train-ing set. This paper summarizes the re-sults of our experiments and attempts to yield better translations of German nom-inal compounds into Spanish and shows how our approach improves by up to 1.4 Bleu points with respect to the baseline. 1