Comparative study of factored SMT with baseline SMT for English to Kannada

K. M. Shivakumar, N. Shivaraju, Vighnesh Sreekanta, Deepa Gupta · 2016

Dravidian languages are highly agglutinative and morphologically rich in their features. Language processing for these languages requires more annotating data compared to European or Indo-European languages. In this paper we present the comparison between Statistical Machine Translation (SMT) model with linguistic and non-linguistic data models for English to Kannada languages. The experiments shows an improvement in Bleu-Score for Factored MT system against Baseline MT system for English to Kannada SMT. Kannada fonts can take ten different forms in representing a word any change of a font variant in word leads to change in meaning of the word. We model these morphological variants of Kannada lemma words, their variants and PoS as Factors in our MT System.

Read the paper · More papers on PaperTik