BERT for Hindi word sense disambiguation

Shailendra Kumar Patel, Rakesh Kumar, Anuj Kumar Sirohi · 2025

With rich lexical ambiguity in the Hindi language, the problem of Word Sense Disambiguation (WSD) is an important natural language processing task. This task is even more challenging when the availability of annotated datasets tailored for WSD remains limited. This paper proposes a novel approach to Hindi WSD by synergizing the BERT BASE model with the Hindi WordNet dataset. Our methodology involves adapting the pre-trained BERT BASE model to the specifics of Hindi WSD through fine-tuning with the Hindi WordNet dataset. Experimental evaluations on benchmark Hindi WSD datasets demonstrate significant improvements in the accuracy and performance metrics compared to traditional methods. Our findings underscore the efficacy of leveraging both pre-trained language models and rich lexical data source like Hindi WordNet to advance Hindi WSD. We have achieved 69.65% and 79.30% accuracy with BERT BASE and BERT BASE Multilingual, respectively, which is approximately a 3% improvement in comparison to the best-performing baseline.

Read the paper · More papers on PaperTik