The effect of shallow segmentation on English-Tigrinya statistical machine translation
Yemane Tedla, Kazuhide Yamamoto · 2016
This paper presents initial research on English-to-Tigrinya statistical machine translation (SMT). Tigrinya is a highly inflected Semitic language spoken in Eritrea and Ethiopia. Translation involving morphologically complex languages is challenged by factors including data sparseness, word alignment and language model. We try to address these problems through morphological segmentation of Tigrinya words. As a result of segmentation, the size of the language model and its perplexity were greatly reduced. Furthermore, the increase in Tigrinya tokens decreased out-of-vocabulary ratio by 46%. We analysed phrase-based translation with unsegmented and segmented corpus to investigate the effect of segmentation on translation quality. Preliminary results demonstrate promising performance improvement from a relatively small parallel corpus.