Experimental Study of Chinese POS Tagging
Liu Xiaofeng · Proceedings of the 2018 2nd International Conference on Computer Science and Artificial Intelligence · 2018
Chinese POS tagging is an important task in Chinese information processing, and many other Chinese natural language processing tasks are dependent on it. Nowadays, there exist a lot of approaches to Chinese POS tagging, and they exploit different machine learning frameworks. We conducted experiments on the People Daily corpus for 6 tagging methods from 3 families, i.e. reference method, generative model based and discriminate model based methods, and compare them in terms of tagging efficiency, effectiveness and error distribution etc. We found that the reference method is the fastest, available easily and applicable for some simple applications, the HMM based method is barely satisfactory, and the character based joint method has the best effectiveness, which increases steadily with the increase of the size of the training corpus. In addition, the distributions of these methods are alike, so no better effectiveness would be expected by the combination of them.