Are LLMs Good Zero-Shot Classifiers? Re-Evaluating the Performance of LLMs on Arabic Sentiment Analysis
Mohamed Alkaoud · 2025
Recent evaluations of Large Language Models (LLMs) in Arabic Natural Language Processing have suggested their inferior performance compared to traditional state-of-the-art (SOTA) models in zero-shot learning scenarios [1]. This paper presents a critical re-examination of this assertion, focusing specifically on Arabic sentiment analysis. We found that inconsistent class definitions in prior evaluations skewed performance comparisons between LLMs and SOTA models. While SOTA models were evaluated on a simplified three-class sentiment classification task, LLMs were tested on a more complex four-class problem. By standardizing the evaluation framework and aligning class definitions, we demonstrate that LLMs, particularly Claude Sonnet 3.5 and GPT-4o, achieve competitive performance in zero-shot settings, with F1-scores within 2.3-3.7% of the fine-tuned SOTA model. Our study also provides the first evaluation of Grok's capabilities in Arabic sentiment analysis, revealing its relatively lower performance compared to other LLMs. Additionally, we investigate the impact of prompt language (Arabic vs. English) on model performance, finding varying effects across different LLMs. These findings challenge previous conclusions about LLMs' limitations in Arabic sentiment analysis and emphasize the critical importance of standardized evaluation methodologies in comparative model assessments.