Pixel Level Handwritten and Printed Content Discrimination in Scanned Documents

Mathias Seuret, Marcus Liwicki, Rolf Ingold · 2014

Classification of the content of a scanned document as either printed or handwritten is typically tackled as a segmentation problem of pages into text lines or words. However these methods are not applicable on documents where handwritten annotations overlay printed text. In this paper we propose to treat the task as a pixel classification task, i.e., To classify individual foreground pixels into either printed or handwritten pixels. Our method uses various features of diverse nature taking the surrounding window into account. The influence of the features and their parameters are investigated and optimized on a validation set. Each foreground pixel is then classified by a multilayer perceptron using feature vectors based on a pixel neighborhood. Finally, a post-processing step corrects typical misclassifications, i.e., It removes outliers based on several heuristics. We evaluated our method on printed documents with real handwritten annotations and reached an accuracy of 96.10% on the test set. This is significantly higher than a previously published methods based on local features.

Read the paper · More papers on PaperTik