A Fine-Grained Parts-of-Speech tagging for Hindi-Devanagari Script based on Deep Learning

Aditi Bajpai, Sonal Yadav, Naresh Kumar Nagwani · 2025

Named entity recognition, information extraction, word sense disambiguation, machine translation, and many more natural language processing activities require parts-of-speech tagging as a preprocessing step. It has already achieved encouraging outcomes in English and European languages. Nevertheless, in Indian languages, it remains inadequately investigated due to the absence of supporting tools, resources, and the morphological complexity of the language. In Hindi, words often change their forms based on grammar rules like tense, gender, number, case, and person. Additionally, Hindi’s free word order structure introduces significant challenges for accurate POS tagging. The proposed study involves the development of an advanced POS tagger that categorizes adjectives by gender and number, utilizing algorithms such as Decision Tree, Random Forest, alongside the the IndicBERT model, integrating multi-feature analysis for enhanced precision and accuracy. This proposed method for POS tagging in Hindi attains a notable accuracy of 90.7%, additionally investigating the classification of adjectives based upon gender and number. Developing effective POS tagging systems for Hindi is crucial for bridging language gaps, improving comprehension, and enhancing Natural Language Processing applications within the Indian linguistic framework.

Read the paper · More papers on PaperTik