Word Disambiguation in Shahmukhi to Gurmukhi Transliteration
Tejinder Singh Saini, Gurpreet Singh Lehal · 2011
To write Punjabi language, Punjabi speakers use two different scripts, Perso-Arabic (referred as Shahmukhi) and Gurmukhi. Shahmukhi is used by the people of Western Punjab in Pakistan, whereas Gurmukhi is used by most people of Eastern Punjab in India. The natural written text in Shahmukhi script has missing short vowels and other diacritical marks. Additionally, the presence of ambiguous character having multiple mappings in Gurmukhi script cause ambiguity at character as well as word level while transliterating Shahmukhi text into Gurmukhi script. In this paper we focus on the word level ambiguity problem. The ambiguous Shahmukhi word tokens have many interpretations in target Gurmukhi script. We have proposed two different algorithms for Shahmukhi word disambiguation. The first algorithm formulates this problem using a state sequence representation as a Hidden Markov Model (HMM). The second approach proposes n-gram model in which the joint occurrence of words within a small window of size ± 5 is used. After evaluation we found that both approaches have more than 92% word disambiguation accuracy.