Pinyin-to-Character Conversion Model Based on Support Vector Machines

Bingquan Liu · Zhongwen xinxi xuebao · 2007

In order to overcome the difficulty in fusing more features into n-gram,a Pinyin-to-Character conversion model based on Support Vector Machines(SVM) is proposed in this paper,providing the ability of integrating more statistical information.Meanwhile,the excellent generalization performance effectively overcomes the overfitting problem existing in the traditional model,and the soft margin strategy overcomes the noise problem to some extent in the corpus.Furthermore,rough set theory is applied to extract complicated and long distance features,which are fused into SVM model as a new kind of feature,and solve the problem that traditional models suffer from fusing long distance dependency.The experimental result showed that this SVM Pinyin-to-Character conversion model achieved 1.2% higher precision than the trigram model,which adopted absolute smoothing algorithm,moreover,the SVM model with long distance features achieved 1.6% higher accuracy.

Read the paper · More papers on PaperTik