HMM-Based POS Tagging in Hindi: A Viterbi Algorithm and Smoothing Analysis
Abhishek Ghimire, Rabin Dhakal · European Journal of Applied Science Engineering and Technology · 2025
Part-of-speech (POS) tagging plays a pivotal role in Natural Language Processing, providing valuable metadata for various downstream tasks such as syntactic parsing, information retrieval, sentiment analysis, machine translation, named entity recognition, and speech recognition. While significant progress has been made in POS tagging for English, the landscape for Hindi remains relatively underexplored. Limited availability of tagged datasets poses a unique challenge in training accurate POS taggers for Hindi. In this study, we aim to explore this gap by training a Hidden Markov Model (HMM) and evaluating its performance using the Viterbi algorithm. We specifically focus on assessing the model’s ability to handle unseen or unknown words, crucial for real-world applications, and evaluate the impact of smoothing techniques too. Our results show that the non-smoothed model achieves higher overall accuracy and significantly better performance on unknown words compared to the smoothed version. Interestingly, precision and recall across POS tags remain consistent between both models, suggesting comparable effectiveness in tagging individual categories despite differences in overall performance.