Conversion of urdu nastaliq to roman urdu using OCR

Faiza Iqbal, Aisha Latif, Nazia Kanwal, Tayyaba Altaf · International Conference on Information Systems · 2011

This paper deals with Urdu Nastaliq, which is a popular script of writing Urdu language. The complex and cursive nature of nastaliq makes optical character recognition for this script very challenging. Segmentation of urdu nastaliq is also very complex due to different levels and shapes of characters according to their position in a word. Character based segmentation technique for urdu names has been proposed in this paper. After segmenting each character, the characters are matched to their templates already saved in the database. The proposed technique handles many complexities in conversion of urdu names to roman urdu including vowel sounds produced by a single character due to diacritics and different voices of ye, ain and other such characters. At the end names written in nastaliq script are converted into roman urdu.

Read the paper · More papers on PaperTik