Harnessing Pretrained Models for Arabic Idiomatic Expression Identification: LLMs
Salma Tace, Mossab Batal, Soumaya Ounacer, Sanaa El Filali, Mohamed Azouazi · International Journal of Computing · 2025
Researchers have increasingly focused on idiomatic expressions in recent years, particularly Arabic idiomatic expressions. These phrases, often derived from ancient stories, are characterized by deeply idiomatic and non-compositional meanings. In this study, we explore the capabilities of large language models (LLMs) to understand and identify these expressions. After collecting data on Arabic idiomatic expressions, we carried out a preprocessing phase. We conducted a comprehensive set of experiments comparing two models, ChatGPT 4 and Arabic Bidirectional Encoder Representations from Transformers (AraBERT). Using 80% of the data for training and 20% for testing, our results reveal the strong ability of LLMs to identify idiomatic expressions, with performance reaching up to 95% in terms of F1 score and accuracy. In the second part of our study, we evaluate the efficacy of the pretrained AraBERT model in detecting idiomatic expressions, comparing it to baseline models, namely Convolutional Neural Network - Long Short-Term Memory (CNN-LSTM) and Bidirectional Long Short-Term Memory (BiLSTM). The analyses show that the pretrained AraBERT model outperforms the conventional CNN-LSTM method by 14% in accuracy and F1 score, and also outperforms the BiLSTM model by 22%.