Character segmentation for Nastaleeq URDU OCR: A review

Aejaz Farooq Ganai, Faisal Rasheed Lone · 2016

Urdu Nastaleeq is a highly cursive, context sensitive language, written diagonally from top right to bottom left that makes it difficult to segment the partial word or a compete word into characters. Further due to stacking of characters, the segmentation at the character level is hard to perform. Some researchers have performed the ligature level segmentation and have succeeded to a great extent, but the accuracy of segmentation is still less and needs to improved. In this paper, the methodology for segmentation of Urdu nastaleeq at the character level is presented. The various challenges encountered during segmentation have been discussed in detail.

Read the paper · More papers on PaperTik