Language Identification of Kannada, Hindi and English Text Words Through Visual Discriminating Features
M. C. Padma, P. Vijaya · International Journal of Computational Intelligence Systems · 2008
In a multilingual country like India, a document may contain text words in more than one language.For a multilingual environment, multi lingual Optical Character Recognition (OCR) system is needed to read the multilingual documents.So, it is necessary to identify different language regions of the document before feeding the document to the OCRs of individual language.The objective of this paper is to propose visual clues based procedure to identify Kannada, Hindi and English text portions of the Indian multilingual document.