Lexicon adaptation with reduced character error (LARCE) - a new direction in Chinese language modeling
Yicheng Pan, Lin-shan Lee · 2007
Good language modeling relies on good predefined lexicons. For Chinese, since there are no text word boundaries and the concept of “word ” is not very well defined, constructing good lexicons is difficult. In this paper, we propose lexicon adapta-tion with reduced character error (LARCE), which learns new word tokens based on the criterion of reduced adaptation cor-pus error rate. In this approach, a multi-character string is taken as a new “word ” as long as it is helpful in reducing the er-ror rate, and minimum number of new, high-quality words can be obtained. This algorithm is based on character-based con-sensus networks. In initial experiments on Chinese broadcast news, it is shown that LARCE not only significantly outper-forms PAT-tree-based word extraction algorithms, but even out-performs manually augmented lexicons. It is believed the con-cept is equally useful for other character-based languages.