ARDIAL-BERT: Advancing Multidialectal Arabic Named Entity Recognition Through Continual Pretraining

Tahar Alimi, Rahma Boujelbane, Lamia Hadrich Belguith · IEEE Transactions on Artificial Intelligence · 2025

Named Entity Recognition (NER) is among the main tasks of Natural Language Processing (NLP). NER is a critical and fundamental component for several NLP applications including Information Retrieval (IR), Question-Answering (QA) and Machine Translation (MT). While several NER models for formal languages such as English and Modern Standard Arabic (MSA) have emerged, Arabic dialects remain in their infancy. We noted that the most recent researches focused on the study of a single Arabic dialect, hence the absence of a perfect multidialectal NER model. In this paper, we present the ARDIAL-BERT, the first multidialectal NER model which was built upon a continuous pretraining and then a finetuning on Arabic dialect publicly available datasets grouped by region (Levantine, Maghrebi, Egyptian and Gulf). The model was built upon the last updated version of BERT transfer transformers after several experiments on various BERT NER models.We approached our contribution on two different tasks: first we built ARDIAL-NER, an Arabic multidialect dataset extracted from existing NER datasets. ARDIAL-NER was manually annotated and contains a total of 53,539 entities of 369,372 tokens composing 21683 sentences. Second, we conducted a continual pretraining process using additional unannotated data, and then we guided a finetuning on new annotated NER datasets. The continuous learning system can be applied at different levels: updating model parameters, incorporating new language data, and training on new labels. Our results demonstrate its effectiveness to varying degrees. Our approach showed that it exhibited a greater ability by achieving superior results compared with both baselines and previous models. This demonstrates the capabilities of grouping Arabic dialects by region and the good selection of data that matches well with the baseline transformers.

Read the paper · More papers on PaperTik