Enhancing Arabic Sentiment Analysis via MARBERT : Domain Adaptation With Pseudo‐Labeling and Contrastive Learning

Mohamed R. Ezzeldin, Gaber Sallam Salem Abdalla, Abdoulie Faal · Engineering Reports · 2025

ABSTRACT Cross‐domain sentiment analysis in Arabic is still challenging due to the scarcity of labeled data and the inherent complexity of identifying sarcasm, when ostensibly negative language expresses good sentiment. In our study, we address these issues with a comprehensive semi‐supervised framework that combines domain adaptation, pseudo‐labeling, and contrastive learning, with MARBERT, a pretrained Arabic language model. We adapted a model originally trained on the Large Arabic Book Reviews (LABR) dataset to the ArSarcasm dataset. This was achieved through five iterative runs of pseudo‐labeling, using a dual‐threshold confidence filter (0.85–0.75 for general samples and 0.70–0.60 for positive samples) to ensure reliable learning. Our approach yielded strong results, achieving a macro F1 score of 66.25% (±0.49) and an overall accuracy of 70.21% (±0.43) on the ArSarcasm test set. The model demonstrated particularly robust performance in classifying neutral (F1: 75.38%) and negative (F1: 70.20%) sentiments. However, detecting positive sentiment in sarcastic expressions remains challenging, as reflected in a lower F1 score of 53.18%, underscoring the complexity of this specific linguistic phenomenon. This study ultimately shows the synergistic value of integrating domain adaptation, pseudo‐labeling, and contrastive learning for semi‐supervised sentiment analysis in low‐resource, sarcasm‐heavy Arabic contexts. It also provides empirical insight into a key limitation of current transformer‐based models: accurately detecting the incongruence that defines sarcasm.

Read the paper · More papers on PaperTik