Research on the Method of Automatic Correction of Chinese Part-of-Speech Tagging

Zheng Jia-heng · Zhongwen xinxi xuebao · 2004

The disambiguation of multi-category words is one of the difficulties in part-of-speech tagging of Chinese text, which affects the processing quality of corpora greatly. Aiming at this question, the paper describes an approach to correcting the part-of-speech tagging of multi-category words automatically. It acquires correction rules for the part-of-speech tagging of multi-category words from right-tagged corpora based on the rough sets and data mining, and then corrects the corpora based on these rules automatically. According to the results of close-test and open-test on the corpus of 500,000 Chinese characters, the accuracy of multi-category words'part-of-speech tagging can be increased by 11.32% and 5.97% respectively.

Read the paper · More papers on PaperTik