Word Sense Disambiguation for Setswana Using Transformer-Based Models
Gabofetswe Malema, Boago Okgetheng · 2024
This study explores the performance of various transformer-based models on Word Sense Disambiguation (WSD) for Setswana, a lowresource language.We evaluated six models: PuoBERTa, BERT-basemultilingual-cased, XLM-Roberta-base, Afro-XLMR-base, LaBSE, and MPNet.These models were tested on a dataset of over 2,000 sentences across five frequently used Setswana verbs, each annotated with multiple senses by native speakers.Using 5-fold crossvalidation, PuoBERTa, tailored for Setswana, demonstrated superior accuracy and nuanced understanding compared to general multilingual models.Specifically, PuoBERTa achieved an average accuracy of 65.5% across all verbs, significantly outperforming the other models.Our findings highlight the necessity of language-specific models in enhancing NLP tasks for low-resource languages.Future research should focus on expanding such models and fine-tuning techniques to further support linguistic diversity.