Shawarma Chats: A Benchmark Exact Dialogue & Evaluation Platter in Egyptian, Maghrebi & Modern Standard Arabic—A Triple-Dialect Feast for Hungry Language Models

Kamyar Zeinalipour, Mohamed Zaky Saad, Oumaima Attafi, Marco Maggini, Marco Gori · 2025

Contentgrounded dialogue evaluation for Arabic remains underresourced, particularly across Modern Standard Arabic (MSA), Egyp tian, and Maghrebi varieties.We introduce ShawarmaChats 1 , a benchmark of 30,000 six turn conversations grounded in Wikipedia con tent, evenly split across the three dialects.To build this corpus, we prompt five frontier LLMs -GPT-4o, Gemini 2.5 Flash, Qwen-Plus, DeepSeek-Chat, and Mistral Large to generate 1,500 seed dialogues.Native Ara bic speakers evaluate these outputs to select the most effective generator and most human aligned grader.SubA dialogues undergo a two pass, rationaledriven selfrepair loop where the grader critiques and the generator revises; unresolved cases are manually corrected.We apply this pipeline to 10,000 Wikipedia para graphs to create 30,000 highquality conver sations 10,000 per dialect at modest human cost.To validate the benchmark, we LoRA finetune six open LLMs (1 B to 24 B parame ters) on ShawarmaChats and observe consistent gains in automaticgrader scores, BERTScore, BLEU and ROUGE particularly for models larger than 7 B parameters.ShawarmaChats thus establishes the first largescale, dialect aware, contentgrounded dialogue benchmark for Arabic.References

Read the paper · More papers on PaperTik