Suitability of Hidden Markov Model in Part-of-Speech Tagging for Indian Languages

Bhairab Sarma, Bipul Shyam Purkayastha · 2016

Abstract Part-of-speech (POS) tagging is an important activity in Natural Language Processing (NLP) where each token is assigned with an appropriate symbol according to its lexical category. There are many applications of POS tagging in NLP; among them, two main applications are word class classification and word sense disambiguation. A number of approaches have been developed in POS tagging for different languages. However it is a challenging job to develop a universal tagger for all languages because of different grammatical rules available for different languages. Hidden Markov Model (HMM) is a statistical approaches commonly used for all languages. TNT is a popular example of POS tagger that used bi-gram and tri-gram HMM. According to HMM, a tagger tagged its tokens by estimating two kinds of probabilistic functions: observation probability and transition probability. Basically, this model estimated the probability of some hidden states based on previous known states. The category of a word could be predicted from its few previous words categories. This model is found suitable in word sense disambiguation and in information retrieval specifically for English language. Many researchers claimed up to 98% accuracy in their tagging. Unlike English, Indian languages are highly inflectional and have rich morphology. Due to their morphological richness, performance of HMM based tagger degraded. The second difficulty of Indian languages is that structurally they are free word order language. Because of these two significant problems in Indian languages, recent researcher tries to develop POS tagger applying multiple approaches. In this paper, we will discuss some pitfalls of HMM as a POS tagging approach considering few Indian languages including Hindi and try to develop alternative solution for developer. Our objective is to increase the accuracy level of probabilistic tagger specifically for Indian languages. Keywords: contextual information, HMM, observation probability, POST, transition probability

Read the paper · More papers on PaperTik