Segmentation-free Word Spotting in Historical Bangla Handwritten Binarized Document
Sugata Das, Sekhar Mandal · 2017
Content-Based Image Retrieval (CBIR) for historical handwritten documents is more challenging due to the large variety of writing style and degradation of historical manuscripts due to ageing. In this paper, we propose a segmentation-free word spotting method for historical handwritten binarized documents. The query word and the document image are converted into gray-scale images using distance transform followed by Gaussian smoothing. SIFT detector is used to locate the keypoints in both the query word and the document image. Histogram of Oriented Gradient (HOG) feature vector is used to describe each keypoint. We use an efficient search technique which calculates distance between query-word and the word (or part of a word) present in document image to spot the zone of interest in the document. The proposed method is tested on three historical handwritten Bengali data-sets and one historical English handwritten data-set. The performance is measured using standard evaluation metric which shows the efficiency of the proposed method.