Sarcasm Detection in Arabic Short Text: A Novel Approach Leveraging Pre-trained Language Models and Addressing the Data Imbalance Problem
Wafa’ Q. Al-Jamal, Mostafa Z. Ali, Ahmad M. Mustafa · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
The popularity of social media and the rise of sarcasm in online conversations make sarcasm detection a crucial task in natural language processing (NLP), which still faces several challenges, including data imbalance and expression implicitness. This article presents a novel approach for sarcasm detection in Arabic short texts using deep learning techniques with a particular focus on addressing data imbalances. To ensure accurate and reliable analysis of Arabic short text data, care must be taken in considering and adapting existing techniques due to lack of standardization and changes in word forms. Thus, the proposed approach utilizes transformers to capture contextual knowledge and addresses the data imbalance problem by combining various data sampling and augmentation techniques. Using Arabic short text data, this article selects the most effective techniques for detecting sarcasm, which proposes a sampling technique that generates new samples from promising and representative words. The proposed approach was evaluated on several Arabic sarcasm detection datasets, including a new dataset we collected called ArSarcasticNews, a curated collection of Arabic short news texts for sarcasm detection. The experimental results across all datasets demonstrate promise. Specifically, the proposed model achieves F1 scores of 0.58, 0.60, 0.84, and 0.64 for the sarcastic class on ArSarcasm-v2, iSarcasmEval, IDAT data, and ArSarcasticNews, respectively. The study highlights the importance of understanding the context between sentences in identifying sarcasm in short Arabic texts and the potential of data augmentation techniques to improve model performance. Moreover, the use of the SHAP-based model provides valuable insight into the features used by the model to detect sarcasm and can aid in identifying potential biases in the model. Based on the findings, this study improves sarcasm detection and outperforms other approaches by analyzing dialectical short text and addressing data imbalances.