Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation

Zobia Rehman, Waqas Anwar, Usama Ijaz Bajwa · 2011

Urdu is morphologically rich language with different nature of its characters. Urdu text tokenization and sentence boundary disambiguation is difficult as compared to the language like English. Major hurdle for tokenization is improper use of space between words, where as absence of case discrimination makes the sentence boundary detection a difficult task. In this paper some issues regarding both of these language processing tasks have been identified. 1

Read the paper · More papers on PaperTik