Text and non-text region identification using texture and connected components

Ankit Vidyarthi, Namita Mittal, Ankita Kansal · 2014

Finding text area from document image i.e. an image which has text embedded with graphic is a challenging task. In the past few years, people are working on document images to extract the text from complex colored background images but results in the extraction of text with the loss of the existing graphics from the original image. However, it is a challenging problem to detect text and non-text region, because extraction of a text region from a document image has lower pixel intensity over graphics pixel intensity. In this paper, a new texture based method is proposed for extraction of Text and Non-Text area without losing the graphics from the document image using binarization and nearly connected component.

Read the paper · More papers on PaperTik