Text location in color documents

Anil Kumar Jain, Anoop Namboodiri, Keechul Jung · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2003

Many document images contain both text and non-text (images, line drawings, etc.) regions. An automatic segmentation of such an image into text and non-text regions is extremely useful in a variety of applications. Identification of text regions helps in text recognition applications, while the classification of an image into text and non-text regions helps in processing the individual regions differently in applications like page reproduction and printing. One of the main approaches to text detection is based on modeling the text as a texture. We present a method based on a combination of neural networks (texture-based) and connected component analysis to detect text in color documents with busy foreground and background. The proposed method achieves an accuracy of 96% (by area) on a test set of 40 documents.

Read the paper · More papers on PaperTik