Part-of-Speech Tagging for Marathi Language Using Hidden Markov Model

Prathmesh Kondar, Sanket Jadhav, Shivprasad Landage · Technix International Journal for Engineering Research · 2026

Part-of-Speech (POS) tagging is an essential task in Natural Language Processing (NLP) that helps computers understand the grammatical roles of words. For Indian languages like Marathi, POS tagging plays an important role in improving applications such as machine translation, text classification, and speech recognition. However, Marathi is morphologically rich and has limited annotated datasets, making accurate POS tagging difficult. Existing models often perform poorly because they cannot handle word variations and complex sentence structures. This study aims to develop an efficient POS tagging system for Marathi using a Hidden Markov Model (HMM). The goal is to identify the correct grammatical category of each word in a sentence and evaluate the performance of the HMM-based approach. The HMM-based POS tagger showed good tagging accuracy, especially for common word categories such as nouns and verbs. Errors mostly occurred in handling highly inflected words and rare forms, but performance remained consistent across different sentence types. These results indicate that HMM is a reliable statistical method for Marathi POS tagging, despite the language’s complexity. Improved datasets and morphological processing can further enhance accuracy and support more advanced Marathi NLP applications.

Read the paper · More papers on PaperTik