GIdTra: A dictionary-based MTS for translating Gujarati bigram idioms to English
Jatinderkumar R. Saini, Jatin C. Modh · 2016
Gujarati language is spoken by nearly 60 million people worldwide, making it the 26th most-spoken native language in the world. It is part of Indo-Aryan family which is a branch of Indo-European languages. Based on our literature review, we have found that there are 3 Gujarati translators available but none of them provides the accurate translation of Gujarati idioms. This is so because they follow a token-based dictionary-based approach. We have created a corpus of nearly 3000 n-gram idioms of Gujarati language and done an intensive analysis of the same. From this corpus, nearly 1800 bigram idioms have been targeted for Machine Translation using a phrase-based dictionary-based approach. Idiom is a sequence of tokens whose meaning is different from the combination of the literal meaning of the individual tokens forming the idiom. Unlike many natural language tokens, the meaning of idiom does not change with time and space. The proposed Gujarati to English Idioms Translator works well for the bigrams and being based on idioms is free from the effect of getting obsolete with time as well as free from giving unexpected output at a geographically different location throughout the world.