Identifying Handwritten Text in Mixed Documents

Faisal Farooq, K. S. Sridharan, Venu Govindaraju · 2006

In this paper we present a system for classification of machine printed and handwritten text in mixed documents. The classification is performed at the word level. We propose a feature extraction algorithm for each word image based on Gabor filters followed by classification using an expectation maximization (EM) based probabilistic neural network that reduces overfitting of training data. An overall precision of 94.62% was obtained for the Arabic script using the modified neural network. The accuracies obtained using a simple backpropagation neural network and an SVM were 83.33% and 90.26% respectively

Read the paper · More papers on PaperTik