Dhati: a Fine-tuned Large Language Model for evaluating Subjectivity in Arabic Textual Data

Attia Nehar, Slimane Bellaouar, Soumia Souffi, Mounia Bouameur · 2023

Despite being a linguistically rich and morphologically complex language, Arabic remains an under-resourced language. The scarcity of large annotated datasets creates a challenge to providing accurate tools for many natural language processing (NLP) tasks such as subjectivity and sentiment analysis. The efficacy of text classification has been significantly enhanced for various languages, including English and French, due to the notable progress made in deep learning (DL) and Transformers, leading to the development of large language models. The aforementioned models have undergone pre-training using extensive datasets, followed by a process of fine-tuning for targeted downstream tasks. In this paper, we provide a tool, which we call "Dhati", for the evaluation of subjectivity in Arabic textual data by fine-tuning a large language model (XLM-RoBERTa) on the Arabic Sentiment Tweets Dataset (ASTD). Then, for comparison purposes, we provide a parallel approach, where we translate the Arabic text into English and use two existing fine-tuned models. The findings indicate that the Dhati model has superior performance compared to the parallel approach, as it achieves an accuracy rate of 82% in the Arabic subjectivity classification task using the ASTD benchmark.

Read the paper · More papers on PaperTik