Keyword Spotting in Online Chinese Handwritten Documents with Candidate Scoring Based on Semi-CRF Model
Heng Zhang, Xiangdong Zhou, Cheng‐Lin Liu · 2013
For text-query-based keyword spotting from handwritten Chinese documents, the index is usually organized as a candidate lattice to overcome the ambiguity of character segmentation. Each edge in the lattice denotes a candidate character associated with a candidate class. Character similarity (between character and class) scores are calculated on each edge, and the similarity between a query word and handwriting is obtained by combining these edge scores. In this paper, we propose a document indexing method using semi-Markov conditional random fields (semi-CRFs), which provide a principled framework for fusing the information of different contexts. For fast retrieval and to save storage space, the lattice is first purged by a forward-backward pruning approach. On the reduced lattice, we estimate the character similarity scores based on the semi-CRF model. Experimental results on a large handwriting database CASIAOLHWDB justify the effectiveness of the proposed method.