Part-of-speech tagger based on maximum entropy model

Heyan Huang, Xiaofei Zhang · 2009

The maximum entropy (ME) conditional models don't force to adhere to the independence assumption such as in Hidden Markov generative models, and thus the ME-based part-of-speech (POS) tagger can depend on arbitrary, non-independent features, which are benefit to the POS tagging, without accounting for the distribution of those dependencies. Since ME models are able to flexibly utilize a wide variety of features, the sparse problem of training data is efficiently solved. Experiments show that the POS tagging error rate is reduced by 54.25% in close test and 40.56% in open test over the hidden-markov-Model-based baseline, and synchronously an accuracy of 98.01% in close test and 95.56%in open test are obtained.

Read the paper · More papers on PaperTik