Comparative study of Moroccan Dialect Sentiment analysis: Finetuning Deep learning transformers
Kaoutar Aboukass, Houda Anoun · 2024
Research on sentiment analysis for the Moroccan dialect "Darija" is limited, with a noticeable scarcity of available resources and annotated corpora. To address this gap, we extensively explore sentiment analysis in Darija, leveraging state-of-the-art pretrained language models. Our study involves fine-tuning various models on a specific dataset covering text written in Arabic letters and Arabizi, with a focus on understanding their performance nuances. Models such as DarijaBert, DarijaBert-mix, MorrBERT, DarijaBert-arabizi, Camelbert-da, and Arabert are included. The study elaborates on the different stages of data collection and pre-processing, as well as transfer learning using pre-trained transformers and the evaluation of their performance. Our findings shed light on the effectiveness of these models, particularly those pretrained on Moroccan Darija, showcasing their performance disparities, especially in sentiment analysis tasks. Consequently, the study concludes by demonstrating the high accuracy rates achieved by DarijaBert for text written in Arabic letters and DarijaBert-mix for text in both Arabic letters and Arabizi. DarijaBert achieves an accuracy of 82.5% coupled with an F1 score of 0.80, while DarijaBert-mix achieves an accuracy of 81.92% alongside an F1 score of 0.79.