Influence of Language Models and Candidate Set Size on Contextual Post-processing for Chinese Script Recognition

Yuan-xiang Li, Chew Lim Tan · 2004

In Chinese language, word is the basic syntaxmeaningful unit, however, each character also has the definite meaning itself. In this paper, we compare the perplexities of four n-gram language models (characterbased bigram, character-based trigram, word-based bigram and class-based bigram) and their influence on the performance of contextual post-processing of Chinese scripts in an offline handwritten Chinese character recognition system. We also demonstrate the influence of the candidate set size on the performance of contextual post-processing in detail, and indicate that the number of candidates should vary with each script. 1.

Read the paper · More papers on PaperTik