Classical Machine Learning and Transformer Models for Offensive and Abusive Language Classification on Dziri Language
Mohammed Mehdi Bouchene, Kheireddine Abainia · 2023
Offensive and abusive language identification is a challenging task, especially for low-resource languages such as Dziri, a dialect of Algerian Arabic. In this paper, we propose two approaches to enhance the performance of this task for Dziri. The first approach fine-tunes several pre-trained transformer models, including DziriBERT (a BERT-based model trained on Dziri corpus), AraBERT (a BERT-based model trained on Arabic corpus), BERT-base model (a BERT-based model trained on English corpus), and XLM-Roberta (a multilingual BERT-based model trained on 100 languages). The second approach applies a tailored preprocessing pipeline for the Dziri dialect, optimized chi-2 feature selection, and optimized SVM hyper-parameters using the Bayesian optimization. The evaluation demonstrates that our approach substantially improves the identification score compared to baseline classifiers. However, DziriBERT still outperforms our approaches by 5.83 % and 1.17 % of F -score in binary and ternary classification, respectively. The code to reproduce the results is available at: https://github.com/Bouchenemehdi24/Off-Abus-Dziri.