Shape Code Based Word-Image Matching for Retrieval of Indian Multi-lingual Documents
Arundhati Tarafdar, Ranju Mondal, Srikanta Pal, Umapada Pal, Fumitaka Kimura · 2010
In the current scenario retrieving information from document images is a challenging problem. In this paper we propose a shape code based word-image matching (word-spotting) technique for retrieval of multilingual documents written in Indian languages. Here, each query word image to be searched is represented by a primitive shape code using (i) zonal information of extreme points (ii) vertical shape based feature (iii) crossing count (with respect to vertical bar position) (iv) loop shape and position (v) background information etc. Each candidate word (a word having similar aspect ratio and topological feature to the query word) of the document is also coded accordingly. Then, an inexact string matching technique is used to measure the similarity between the primitive codes generated from the query word image and each candidate word of the document with which the query image is to be searched. Based on the similarity score, we retrieve the document where the query image is found. Experimental results on Bangla, Devnagari and Gurumukhi scripts document image databases confirm the feasibility and efficiency of our proposed approach.