Exploring Different Annotation Schemes for Single and Consecutive Named Entity Recognition in the Arabic Biomedical Domain using Transformer Models and Contextual Semantic Embeddings

Ismail Ait Talghalit, Hamza Alami, Saïd Ouatik El Alaoui · Engineering Technology & Applied Science Research · 2025

Named Entity Recognition (NER) is an important task for Natural Language Processing (NLP) in the Arabic biomedical field. However, most works on NER in the Arabic biomedical domain suffer from some limitations, such as the inability to capture the context and semantics within texts. Moreover, only a few research studies have efficiently handled biomedical consecutive entities in the Arabic language. To overcome these limitations, this study proposes an efficient method to build contextual models for biomedical NER tasks that capture context and semantics in Arabic text using transformer models and semantic embeddings. The extracted embeddings are combined with machine learning methods, including SVM, Decision Tree (DT), and AdaBoost, to identify both single and consecutive named entities accurately. Furthermore, the effect of seven annotation schemes, namely IO, IOB, IE, IOE, BI, BIES, and IOBES, was studied to determine the most suitable for Arabic biomedical NER. The experimental results showed that the BERT and AraBERT models when fine-tuned for the Arabic biomedical NER outperform well-known machine learning methods in terms of accuracy and F1 score. The findings across various annotation schemes highlight the effectiveness of the IO scheme for simple (single) entities, while IOBES and BIES annotation schemes are better suited for recognizing multi-token entities.

Read the paper · More papers on PaperTik