Linguistic tuple segmentation in n-gram-based statistical machine translation
Adrià de Gispert, José Bernardo Mariño Acebal · 2006
Ngram-based Statistical Machine Translation relies on a standard Ngram language model of tuples to estimate the translation pro-cess. In training, this translation model requires a segmentation of each parallel sentence, which involves taking a hard decision on tuple segmentation when a word is not linked during word align-ment. This is especially critical when this word appears in the target language, as this hard decision is compulsory. In this paper we present a thorough study of this situation, comparing for the rst time each of the proposed techniques in two independent tasks, namely EnglishSpanish European Parlia-ment Proceedings large-vocabulary task and ArabicEnglish Basic Travel Expressions small-data task. In the face of this comparison, we present a novel segmentation technique which incorporates lin-guistic information. Results obtained in both tasks outperform all previous techniques. Index Terms: statistical machine translation, tuple segmentation, n-gram-based SMT, linguistic information