Arabic Dialect Detection Using Large Language Models: A Comparative Analysis

Hind Rahma Alowais, Ashraf Elnagar · 2025

The Arabic language's rich dialectal variation presents a significant challenge for natural language processing (NLP), especially due to code-switching, phonetic shifts, and region-specific expressions. This study investigates the performance of state-of-the-art large language models (LLMs)—GPT-4, Falcon-7B-Instruct, and LLaMA-3-8B—in classifying Arabic dialects using the MADAR Corpus, a comprehensive dataset covering city-level and regional dialects. The data was preprocessed through normalization, character standardization, and class balancing to enable consistent evaluation across five major dialect groups: Levant, Nile, Gulf, North Africa, and Iraq. In addition to zero-shot and few-shot evaluations of the general-purpose LLMs, we fine-tuned AraBERT on the balanced MADAR dataset and developed a hybrid system combining AraBERT with GPT-4 as a fallback when confidence was low. AraBERT achieved the highest accuracy (68.00%), outperforming GPT-4 (42.80%), Falcon (20%), and LLaMA (23%). The hybrid approach reached 66.13%, reflecting the potential-and complexity-of combining general and domain-specific models. These findings highlight the strengths of task-specific fine-tuning for dialect detection and reveal the limitations of multilingual LLMs when applied to linguistically diverse and underrepresented dialectal data.

Read the paper · More papers on PaperTik