Evaluating Large Language Models for Arabic Sentiment Analysis: A Comparative Study Using Retrieval-Augmented Generation

Salma Khaled, Ensaf Hussein Mohamed, Walaa Medhat · Procedia Computer Science · 2024

Sentiment Analysis (SA) is a crucial task in natural language processing. There are numerous studies devoted to this field specially with emerging use of encoder transformers and generative Large Language Models (LLMs). Transformer based models like BERT are often used for sentiment analysis. However, this article examines the performance of the latest generative generative LLMs in Arabic Sentiment Analysis (ASA) using Retrieval-Augmented Generation (RAG) architecture. We evaluated these models using the ASAD, ArSarcasm-v2, and SemEval datasets. Our experimental studies revealed challenges due to dataset imbalances and misclassified neutral labels, which impacted the effectiveness of fine-tuning. By removing the neutral class, significant improvements in model performance were observed across all datasets, with F1-scores increasing by 17%, 18%, and 18% on ASAD, ArSarcasm-v2, and SemEval, respectively.

Read the paper · More papers on PaperTik