Word-wise Sinhala Tamil and English Script Identification Using Gaussian

Sukalpa Chanda, Srikanta Pal, Umapada Pal · 2008

There are many documents in Srilanka where a single document page may contain Sinhala, Tamil and English texts. For OCR development of such a document page, it is better to identify different scripts present in the page and thenfeed the identifiedportion to the respective OCR module. In this paper, a SVM based technique is proposed for word-wise identification of Sinhala, Tamil and English scripts from a single document page. Structural features, topological features and water reservoir principle basedfeatures are mainly used here for the purpose. From the experiment we obtainedencouragingresults.

Read the paper · More papers on PaperTik