A Chinese Text Corrector Based on Seq2Seq Model
Sunyan Gu, Fei Lang · 2017
In this paper, we build a Chinese text corrector which can correct spelling mistakes precisely in Chinese texts. Our motivation is inspired by the recently proposed seq2seq model which consider the text corrector as a sequence learning problem. To begin with, we propose a biased-decoding method to improve the bilingual evaluation understudy (BLEU) score of our model. Secondly, we adopt a more reasonable OOV token scheme, which enhances the robustness of our correction mechanism. Moreover, to test the performance of our proposed model thoroughly, we establish a corpus which includes 600,000 sentences from news data of Sogou Labs. Experiments show that our corrector model can achieve better corrector results based on the corpus.