Ambiguity handling of similar categories in handwritten Chinese character recognition

Daniel Yeung, H.S. Fong · 2002

Chinese characters consist of thousands of categories, some of which have very similar structural characteristics. Ambiguity in recognition may thus arise. The authors previously (1990) proposed a technique for recognizing characters off-line based on their structural characteristics. Each input character is subject to various alternatives in stroke segmentation, and for each resulting stroke set, the strokes are matched against the category templates maintained in a knowledge base. This knowledge base is so devised to offer certain degree of tolerance to handwriting ambiguities, with respect to individual character categories. This paper aims at extending our current technique to distinguish characters within a similar category as well. Basically, Chinese characters composed of the same stroke set, but with different geometric attributes, are grouped into a similar category. A secondary knowledge base is created to store these similar groups. Our previous methodology will be modified so that once a candidate output is identified (together with its computed score of matching result), all characters belonging to the same similar category will have their matching scores recalculated and possibly new ranking information may ultimately lead to new output candidates. This may eventually improve the system's recognition performance. The complete construction of the secondary knowledge will not be possible since it is application domain specific. But a number of similar categories will be presented to demonstrate how the method works.

Read the paper · More papers on PaperTik