Machine Learning-Based Approach for Arabic Dialect Identification

Hamada Nayel, Ahmed M. Hassan, Mahmoud Sobhi, Ahmed El-Sawy · 2021

This paper describes our systems submitted to the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). Dialect identification is the task of automatically detecting the source variety of a given text or speech segment. There are four subtasks, two subtasks for country-level identification and the other two subtasks for province-level identification. The data in this task covers a total of 100 provinces from all 21 Arab countries and come from the Twitter domain. The proposed systems depend on five machine-learning approaches namely Complement Naive Bayes, Support Vector Machine, Decision Tree, Logistic Regression and Random Forest Classifiers. F1 macro-averaged score of Naive Bayes classifier outperformed all other classifiers for development and test data.

Read the paper · More papers on PaperTik