ArabNED: A Novel Dataset and Evaluation of Fine-Tuned vs. Zero-Shot Models for Arabic Named Entity Disambiguation
Jumana Nounou, Caroline Sabty · Procedia Computer Science · 2026
Named Entity Disambiguation (NED) task is an important task that resolves the ambiguity of entities to enhance other tasks like information retrieval and question answering. A lot of work has been done for English NED, however, for the Arabic language this task is still underexplored because the limited number of resources and the complexity of the language. To the best of our knowledge, we propose the first Arabic dataset for NED. The dataset is called ArabNED and includes 10,000 named entities, NER tags, and structured candidate entities. It is constructed based on the Named Entity Recognition datasets ANERCorp and AQMAR. In addition, we implemented two models for the NED task. The first one is a fine-tuned AraBERT model. The second one is a zero-shot GPT-4o Mini model. The evaluation of these two models showed that the second model of the GPT-4o Mini achieved a higher performance of F1score equal to 86.7% and 85.5% accuracy.