Extraction of text in images
R. Malik, Seongah Chin · 2003
In this paper we present a text segmentation technique that is useful in locating and extracting text blocks in images. The algorithm works without prior knowledge of the text orientation, size or font. It is designed to eliminate background image information and to highlight or identify the regions of the image that contain text. The algorithm uses the fact that text regions in an image may be identified by searching for several repeated instances of uniform gray intensity of approximately the same width. Combining this with the fact that the ratio of type-face stroke width to height is often fixed provides a useful technique for extracting text from images. Results of the application of this algorithm are presented.