A knowledge-based approach for textual information extraction from mixed text/graphics complex document images

Yen‐Lin Chen · 2010

A new knowledge-based technique for extracting and identifying text-lines from various real-life mixed text/graphics complex document images is presented in this paper. The proposed technique first decompose the document image into distinct object planes to separate homogeneous objects including textual regions of interest, non-text objects such as graphics and pictures, and background textures. Then a knowledge-based text extraction and identification method is performed on the resultant planes to obtain text-lines with different characteristics in each plane. This proposed system can offer high flexibility and expandability by just updating new rules for coping with more various types of real-life and future complex document images. From the experimental and comparative results, the proposed knowledge-based technique demonstrates its effectiveness and advantages on extracting text-lines with various illuminations, sizes, and font styles from various types of mixed text/graphics complex document images.

Read the paper · More papers on PaperTik