Chinese Spelling Errors Detection Based on CSLM
Zhaoyi Guo, Xingyuan Chen, Peng Jin, Siyuan Jing · 2015
Spelling errors are very common in various electronic documents and it leads to serious influence sometimes. To solve this problem, methods based on the n-gram language model are the most commonly used. CSLM (continuous space language model) which represents a word as a vector is different from traditional models. In this paper, we experimented with a specific CSLM, namely, the CBOW (Continuous Bag-of-Words) model, to detect spelling errors. Since spelling errors are usually considered as wrong characters rather than words in Chinese language, we trained character vectors with a large Chinese corpus, and then judged a Chinese character is right or not by its probability of the occurrence in a given context. Experimental results show that the method based on CSLM outperforms the n-gram language model.