Chinese Pinyin-to-character Conversion Based on Cascaded Reranking

Xin Li · Acta Automatica Sinica · 2014

The word n-gram language model is the most common approach for pinyin-to-character conversion. It is simple, efficient, and widely used in practice. However, in the decoding phase of the word n-gram model, the determination of a word only depends on its previous words, which lacks long distance grammatical or syntactic constraints. In this paper, we propose two reranking approaches to solve this problem. The linear reranking approach uses minimum error learning method to combine different sub-models, which includes word and character n-gram language models, part-of-speech tagging model and dependency model. The averaged perceptron reranking approach reranks the candidates generated by word n-gram model by employing features extracted from word sequence, part-of-speech tags,and dependency tree. Experimental results on Lancaster Corpus of Mandarin Chinese and People s Daily show that both reranking approaches can efficiently utilize information of syntactic structures, and outperform the word n-gram model. The perceptron reranking approach which takes the probability output of linear reranking approach as initial weight achieves the best performance.

Read the paper · More papers on PaperTik