Graphics Separation System for Printed Document Images

Yanglem Loijing Khomba Khuman, Haobam Mamata Devi, T. Romen Singh, Nidhi Singh · 2020

In the development of a printed document based Optical Character Recognition (OCR) system, the removal of graphics from the documents is an essential step. The presence of graphics in the document affects the segmentation of lines, words, and characters. Due to the presence of different font types, font size and location of text and non-text areas, format analysis of printed documents such as books, newspapers is a very challenging task. In this paper, we proposed a new algorithm to distinguish graphics from printed documents using Flood-Fill Operation by detecting the edges of all objects in the image of the text. The proposed algorithm has outperformed in removing graphics from various printed documents, without depending on the shapes of graphics present in the document with the precision of 0.58, recall of 0.92, and f-measure of 0.71.

Read the paper · More papers on PaperTik