Text Extraction from Web Images Based on Human Perception and Fuzzy Inference

Apostolos Antonacopoulos, Dìmosthenis Karatzas · ePrints Soton (University of Southampton) · 2001

There is a significant need to extract and recognise the semantically-important text contained in images on Web pages. This paper proposes a new approach to text extraction from this special class of images. The method attempts to emulate closer than before the way humans perceive colour differences in order to differentiate between text and background regions. Pixels of similar colour (as humans see it) are merged into components and a fuzzy inference mechanism (using connectivity and colour distance features) is devised to group components into larger character-like regions. 1.

Read the paper · More papers on PaperTik