Enhancing Named Entity Recognition in Arabic: Leveraging Transformer-Based Models and Open Data Sources

Faiza Belbachir, Assia Soukane · 2024

In this paper, we present our system’s design, focusing on named entity recognition (NER) in the Arabic language. Our approach employs a transformer-based language model, augmented with data from open sources. We report our experiments with NER in Arabic, where we conduct a comparative analysis of the performance between CamelBERT and two other models, namely UBC-NLP-ArBert2 and AraBERTv02. This evaluation is performed on the training and development datasets of Subtask 1 FLAT-NER for non-nested entities, proposed by the WojoodNER Shared Task 2023 [1]. Furthermore, we investigate the impact of incorporating Arabic open data sources extracted from AQMAR [2] and ANERcorp [3] through post-processing lexical filtering. Our system demonstrates promising results, achieving an F1 score of 0.9112, surpassing the organizers’ baseline performance, which yielded an F1 score of 0.8681.

Read the paper · More papers on PaperTik