IQAD: Iraqi Arabic Dialect Dataset for Multi-Regional Dialect Classification Using Conventional and Machine Learning Approaches
Noora Aljubouri, Naderi Hassan · Journal of Techniques · 2025
The work's main contribution is creating a dataset for specifying Iraqi Arabic dialects from written texts. With the increase of Iraqi dialectal Arabic usage across social media platforms, accurate dialect identification has become an important step for such tasks as sentiment analysis, social media monitoring, and linguistic studies. We collected, annotated, and prepared normal text data: 53,146 unique text samples taken from social media, divided into three major dialects in Iraq: Middle, Western, and Southern. The lexical variability of the corpus is 78,582 unique tokens. The dataset was passed through preprocessing to clean and prepare it for classification-based tasks. To verify the quality of this dataset, we carried out experiments with two approaches for the classification: a dictionary-based methodology and a TF-IDF-based SVM classification. The SVM outperformed the dictionary-based classifier by achieving 74% accuracy and F1-score, whereas the classifier peaked at 63.6% accuracy and 63.4% F1 score. The results show the effectiveness of the dataset in supporting dialect classification tasks and its potential for use in future Iraqi Arabic NLP applications and research.