TF-IDF or Transformers for Arabic Dialect Identification? ITFLOWS participation in the NADI 2022 Shared Task
Fouad Shammary, Yiyi Chen, Zsolt T. Kardkovács, Mehwish Alam, Haithem Afli · 2022
This study targets the shared task of Nuanced Arabic Dialect Identification (NADI) organized with the Workshop on Arabic Natural Language Processing (WANLP).It further focuses on Subtask 1: the identification of the Arabic dialects at the country level.More specifically, it studies the impact of a traditional approach such as TF-IDF and then moves on to study the impact of advanced deep learning based methods.These methods include fully finetuning MARBERT as well as adapter based fine-tuning of MARBERT with and without performing data augmentation.The evaluation shows that the traditional approach based on TF-IDF scores the best in terms of accuracy on TEST-A dataset, while, the fine-tuned MAR-BERT with adapter on augmented data scores the second on Macro F1-score on the TEST-B dataset.This led to the proposed system being ranked second on the shared task on average.