Multi-Script Line Identification System for Indian Languages

U. Dinesh Acharya, Rajesh Gopakumar, Prakash K. Aithal · 2010

India is a multilingual multi-script country. There are totally 18 official languages and 12 scripts in India. For Optical Character Recognition (OCR) of such a multi-lingual document, it is necessary to identify the script before feeding the text lines to the OCRs of individual scripts. In this paper, a simple and efficient technique of script identification for Kannada, Malayalam, Telugu, Tamil, Gujarati, Hindi and English text lines from a printed document is presented. The proposed system uses horizontal projection profile, Vertical projection profile and Top pitch information to distinguish the seven scripts. The knowledge base of the system is developed based on 50 different document images containing about 250 text lines of each script. The proposed system is tested on 50 different document images containing about 250 text lines of each script and an overall classification rate of 97.64% is achieved.

Read the paper · More papers on PaperTik