Part of Speech Tagging with Discriminatively Re-ranked Hidden Markov Models

Brian Highfill · 2011

The task of part of speech tagging has been approached by various ways. Originally, constructed by way of hand-crafted rules for disambiguation, the majority of tagging is now accomplished by utilizing statistical machine learning methods. Two commonly applied statistical methods are hidden Markov models (HMM) and an extension of Markov chains combined with a maximum entropy classifier called maximum entropy Markov models (MEMM). This paper explores POS tagging by combining a standard HMM tagger separately with a maximum entropy classifier designed to re-rank the best sequence of tags produced by the HMM. Tested on the Brown corpus (Francis, 1979; Francis and Kucera, 1982), the discriminatively re-ranked tagger performed with an accuracy of 91.8%, slightly less than that of the HHM tagger alone.

Read the paper · More papers on PaperTik