Word Level Script and Language Identification for Unconstrained Handwritten Document Images
P.V. Prasanthkumar, E.D. Dileesh · 2014
Word level Script and language identification is a process of separating the script and language of each word present in a printed or handwritten multi-script document. It is an essential part of a multi-lingual Optical Character Recognizer (OCR). Most of the OCRs are solely designed for a single script. So it can't convert a document which is written in more than one script. This paper explained a system, which automatically separate the script of an unconstrained handwritten document mix with three Indian scripts and Roman script English up to word level. The process starts from extracting text-lines from the document and then separates the words from the text-line using projection profiles. 15 connected component and morphological features and 32 Gabor filter features is extracted from one word to form a feature set. Total of 15699 words are separated from 215 documents for training and 4856 words from 64 documents for testing. Three different classifiers, Support Vector Machine (SVM), Multilayer Perceptron (MLP), and K-Nearest Neighbors (KNN) classifiers are used for testing the discriminating power of the feature set. MLP classifier outperform over all others with cross validation accuracy of 81.69% across four scripts.