A Hybrid Model of Combining Rule-based and Statistics-based Approaches for Automatic Detecting Errors in Chinese Text

Yangsen Zhang, Yuanda Cao · Zhongwen xinxi xuebao · 2006

Chinese text automatic proofreading is an important research subject in NLP.A hybrid model based on the combination of rules and statistics are proposed in this article.According to the distribution of Chinese single-character after word segmentation in Chinese text and the conception of non-multi-character word error,we proposed a group of rules to find errors in texts,to construct the automatic error-detection model and to implement its algorithm by combining the scattered single-character Bigram models,part-of-speech Bigram and Trigram models.Our experiment for the 30 texts that contain 578 error test points shows that the recall rate is 86.85% and accuracy rate is 69.43%,distorting rate is 30.57%.

Read the paper · More papers on PaperTik