Augmenting Moroccan Darija Text Data for Improved Classification: A Comparative Study of Augmentation Techniques
Rabia Rachidi, Wassima Moutaouakil, Othmane Daanouni, Mounir Omari, Bouchaib Cherradi, Hassan Silkan · 2025
The performance of learning models mostly depends on the availability and quality of training data. To address the issue of dataset adequacy, researchers have widely investigated Data Augmentation (DA) as a promising solution. This paper examines the application of data augmentation techniques for the Moroccan Darija dialect, an under-resourced language with limited linguistic resources. We investigate using four Easy Data Augmentation (EDA) methods—Random Swap, Random Deletion, Random Insertion, and Synonym Replacement—to increase the diversity of training data, ultimately improving the performance of machine learning models. These techniques were applied to a corpus of Moroccan Darija sentences to improve text classification tasks for emotion detection (Fear, Anger, and Joy). The results demonstrate that these augmentation techniques can significantly enhance data variability, leading to more robust models for the Darija dialect. By creating synthetic examples, this approach tackles the issue of data scarcity in Moroccan Darija, ultimately improving performance in Natural Language Processing (NLP) tasks for low-resource languages.