IUNADI at NADI 2023 shared task: Country-level Arabic Dialect Classification in Tweets for the Shared Task NADI 2023
Yash Hatekar, Muhammad Abdo · 2023
In this paper, we describe our participation in the NADI2023 shared task for the classification of Arabic dialects in tweets.For training, evaluation, and testing purposes, a primary dataset comprising tweets from 18 Arab countries is provided, along with three older datasets.The main objective is to develop a model capable of classifying tweets from these 18 countries.We outline our approach, which leverages various machine learning models.Our experiments demonstrate that large language models, particularly Arabertv2-Large, Arabertv2-Base, and CAMeLBERT-Mix DID MADAR, consistently outperform traditional methods such as SVM, XGBOOST, Multinomial Naive Bayes, AdaBoost, and Random Forests.