Development of a Complete Urdu-Hindi Transliteration System

Gurpreet Singh Lehal, Tejinder Singh Saini · International Conference on Computational Linguistics · 2012

Hindi and Urdu are variants of the same language, but while Hindi is written in the Devnagri script from left to right, Urdu is written in a script derived from a Persian modification of Arabic script written from right to left. The difference in the two scripts has created a script wedge as majority of Urdu speaking people in Pakistan cannot read Devnagri, and similarly the majority of Hindi speaking people in India cannot comprehend Urdu script. To break this script barrier, it becomes necessary to develop a high accuracy Urdu-Devnagri transliteration system. The major challenges in developing such system are handling missing diacritic marks and short vowels in Urdu, zero/multiple character mappings of Urdu in Hindi, absence of half characters in Urdu, multiple mappings of Urdu words in Hindi and word segmentation issues in Urdu including broken and merged words. Already a few Urdu-Hindi transliteration systems have developed but their accuracy is not very high and they have failed to address all the above issues. For the first time, we present a complete Urdu-Hindi transliteration system which takes care of all the above issues and has reported a transliteration accuracy of more than 97% at word level.

Read the paper · More papers on PaperTik