Visual similarity analysis of Chinese characters and its uses in Japanese OCR

Tao Hong, Stephen W. Lam, Jonathan J. Hull, Sargur N. Srihari · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1995

Traditionally, a Chinese or Japanese optical character reader (OCR) has to represent each character category individually as one or more feature prototypes, or a structural description which is a composition of manually derived components such as radicals. Here we propose a new approach in which various kinds of visual similarities between different Chinese characters are analyzed automatically at the feature level. Using this method, character categories are related to each other by training on fonts; and character images from a text page can be related to each other based on the visual similarities they share. This method provides a way to interpret character images from a text page systematically, instead of a sequence of isolated character recognitions. The use of the method for post processing in Japanese text recognition is also discussed.

Read the paper · More papers on PaperTik