Color segmentation for historical documents using Markov random fields

Werner Pantke, Arne Haak, Volker Märgner · 2014

Binarization is often used for pixel-wise document text extraction as preprocessing step for scanned historical documents. These documents are scanned in color and high resolution today. The reduction of color to grayscale images and the subsequent binarization implies a loss of information and often results in unsatisfying processing results. In this paper, a color segmentation instead of a binarization approach is used to segment text from background in historical manuscripts. A color segmentation approach based on Markov random fields with a reduced set of required parameters is presented to segment text written in different colors from noisy page background. First tests with historical Arabic manuscripts show promising results. In case of words written in light red color, our approach shows better results than a state-of-the-art binarization approach.

Read the paper · More papers on PaperTik