Analysis and classification for complex scanned documents

M. Sezer Erkılınç, Mustafa Musa Jaber, Eli S. Saber, Peter Bauer, Dejan Depalov · SPIE Newsroom · 2011

Page-layout-classification methodologies aim to extract text and non-textual regions such as graphics, photos, or logos. These techniques have applications in digital document storage and retrieval where efficient memory consumption and quick retrieval are required.1 Such classification algorithms can also be used in the printing industry for selective or enhanced scanning and object-oriented rendering (printing different parts of a document with different resolution depending on the content).2 Additionally, these techniques can be used as an initial step for various applications. These include optical-character recognition (the electronic translation of handwritten or printed text into machine-encoded text) and graphic interpretation (classifying documents—into military, educational, and others—according to the image content).3 In the past two decades, several techniques have focused on identifying text regions in scanned documents.4, 5 In addition, comprehensive algorithms that aim to identify both text and graphic regions have been developed.6, 7 However, these systems are limited to specific documents, such as newsletters or articles, where the background region is assumed white.8, 9 This assumption not only excludes complex backgrounds and colored documents (such as book covers, advertisements, and flyers),10 but also limits practicality and feasibility when applied to non-ideal (complex) documents. We propose a page-layout-segmentation technique to extract text, image, and strong-edge or strong-line regions (actual lines in the document or transition pixels between a picture and text or a picture and background).11 The algorithm consists of four modules: pre-processing stage, text detection, photo detection, and strong-edge or strong-line detection units. We start by applying a pre-processing module that includes image scaling and enhancement, as well as color-space conversion Figure 1. Line detection results for two different documents. (a) Original image, (b) enhanced L* channel of the CIE L*a*b* space, and (c) final segmentation map where strong-edge or strong-line and text regions are colored in yellow and green, respectively.

Read the paper · More papers on PaperTik