Enhancing Arabic Text Classification: The Impact of Dataset Variety on BERT Model
Muhammad Aliah, Dmitry Berezkin, Ilya A. Kozlov · 2025
This study investigates the impact of dialectal variations on the performance of Bidirectional Encoder Representations from Transformers (BERT) models in Arabic text classification. A comparative analysis is conducted on the performance of three BERT models - Arabic BERT (AraBERT), Multidialectal Arabic BERT (MarBERT), and Multilingual BERT (MBERT) - on a dataset of Arabic tweets with varying dialects. The findings of this study demonstrate that incorporating dialect-specific data in the training set significantly improves model performance, with MarBERT and AraBERT outperforming MBERT. It is also evident that the models encounter difficulties in terms of cross-dialectal performance, underscoring the need for more diverse and representative training data. The present study underscores the significance of incorporating dialectal variations in Arabic natural language processing (NLP) research and suggests that language-specific pre-training can improve model performance. The findings of this study bear implications for the development of more robust Arabic NLP models.