Spelling Correction in Pinyin Input

Zheng Chen · Chinese Journal of Computers · 2001

Chinese input method is one of the difficult problems of Chinese Language Processing. Because of its facility to learn and to use, Pinyin is the most popular Chinese input method. Over 97% of the users in China use Pinyin for input. Although Pinyin input method has so many advantages, it also suffers from several problems, including Pinyin to characters conversion errors, user typing errors, etc. Base on sentence based pinyin input method, we propose a new typing model to solve this problem. The system will accept correct typing, but also tolerate common typing errors. After analyzing the most popular errors made by typists, we build a typing model. The typing model is trained on real data, and learns probabilities of typing errors, including substitution errors, insertion errors and deletion errors. We also design a unified approach to Chinese statistical language modeling. This unified approach enhances trigram based statistical language modeling with automatic, maximum likelihood based methods to segment words, select the lexicon, and filter the training data. Compared to the commercial product, our system is up to 50% lower in error rate at the same memory size, and about 76% better without memory limits at all. In the Pinyin to Hanzi conversion, the probabilities of typing model are combined with the language model probabilities, to find the most probable interpretation of a sequence of Roman letters typed. Further more, spelling correction can automatically adapt to typist according to their typing skills. In a real system, skilled typist could be assigned lower LM weight, and the skill of typist can be determined by the their typing speed. Moreover, the method is applicable to any language. Compared to the baseline of system, our system gets approximate 30% error reduction on the open testset.

Read the paper · More papers on PaperTik