Enhancing word sense disambiguation for Hindi agriculture domain: Feature engineering and machine learning approaches

Aarti Purohit, Kuldeep Kumar Yogi · Results in Engineering · 2025

• Addressed Word Sense Disambiguation (WSD) challenges in Hindi agricultural texts with frequent polysemous terms. • Created a tagged corpus using content from ICAR, DARE, and Kheti magazines. • Incorporated syntactic, semantic, and statistical features: POS tagging, context, bigrams/trigrams, and sense frequency. • Evaluated four ML models: Naïve Bayes, Decision Tree, Random Forest, and SVM. • Random Forest achieved the highest F1-score (0.8541), with SVM also performing strongly. • Feature Sets 2, 3, 4, and 10 provided the best model performance across classifiers. • Outperformed prior WSD research in other Indian languages like Bengali, Marathi, and Punjabi. • Applicable to agriculture-related NLP tools such as advisory systems, semantic search, and chatbots. • Suggested future enhancements using BERT, RNNs, and domain-specific pretraining with Hindi WordNet. The dense network of polysemous words in Hindi agricultural texts poses a challenging problem in the field of Word Sense Disambiguation (WSD). The goal of the present study is to devise a reliable method for resolving such ambiguities, thus enhancing analytical precision and elucidation in agricultural literature. A corpus was constructed from Hindi agricultural texts available in ICAR, DARE, and Kheti magazines, and it was sense-tagged. A feature set comprised of word context, part-of-speech tagging, POS tagging with bigram and trigram probabilities, as well as frequency of senses, was pivoted to classifiers in the machine learning paradigm. These traits were utilized in various models such as Naive Bayes (NB), Decision Trees (DT), Random Forest (RF), and Support Vector Machines (SVM). Random Forest and SVM appear to have an edge in accuracy and F-measure over other models in word sense identification as seen in F-measure results as well in other results. This is unlike previous works about general Hindi WSD. This is also the first sense-tagged corpus in the agricultural area as well as this is the first to assess different feature sets systematically with ML classifiers. This study makes a domain-specific contribution. This study enhances the understanding of the word sense ambiguity of the domain the study focuses on and the reported feature sets and classifiers provide a convincing argument for the inclusion of sociolinguistic, syntactic, and semantic dimensions of the disciplines with statistical measures for better sense disambiguation. The study also highlights agricultural vocabulary prominence and the lack of annotated corpora in the field as an obstacle to progress.

Read the paper · More papers on PaperTik