Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation
Zobia Rehman, Waqas Anwar, Usama Ijaz Bajwa · 2011
Urdu is morphologically rich language with different nature of its characters. Urdu text tokenization and sentence boundary disambiguation is difficult as compared to the language like English. Major hurdle for tokenization is improper use of space between words, where as absence of case discrimination makes the sentence boundary detection a difficult task. In this paper some issues regarding both of these language processing tasks have been identified. 1