Experimental study of N-gram based Uyghur part of speech tagging
Nijat Najmidin · Computer Engineering and Applications Journal · 2012
There are many approaches to the problem of part-of-speech tagging,current Uyghur part-of-speech tagging is mainly based on rule based methods and does not achieve the state-of-art accuracy.A large scale of manually annotated Uyghur corpus and a number of well-conducted experiments are used to identify the efficiency of N-gram based part-of-speech tagging scheme for Uyghur texts.The N-gram language model parameters and data smoothing are analyzed,and the efficiency of Bigram and Trigram models are compared.The impacts of tag sets and size of training data on tagging accuracy are studied.The experiments show that N-gram based part-of-speech tagging for Uyghur texts has achieved good results.