Text Identification in Noncursive English Handwritten Script

Sarita Ghadge, Sanika S Patankar, Jayant V Kulkarni · 2018

Handwritten script identification and converting it to printed electronic document as vast application. It helps to translate humanly readable to machine-readable text. This paper presents a technique to identify non cursive scanned, handwritten English script to digital document in notepad/MS-Office format using neural networks. Initially preprocessing of the input document image as conversion from RGB image to 8 bit gray scale, nonlinear contrast enhancement and mean normalization for correcting non uniform illumination and binarization are performed. Further the lines, words and characters of the input documents are localized in the preprocessed image and extracted from gray input image. Each extracted character of the gray input image is resized, normalized and used as input to the neural network for classification. The network consists of two layers feedforward network with sigmoid output neurons. The error back propagation algorithm is used for training. The proposed technique is tested on sample images from publically available CEDAR database as well as self created database with 1600 images for training and 620 images for testing purpose belonging to 62 classes. Overall Recognition accuracy of 95% is obtained. The performance of the proposed technique is also compared with the other existing technique.

Read the paper · More papers on PaperTik