A Novel Dataset for Arabic Domain Specific Term Extraction and Comparative Evaluation of BERT-Based Models for Arabic Term Extraction

Abdulmohsen O. Al-Thubaity · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

Automatic term extraction from domain-specific corpora is a well-known challenge in natural language processing, with applications in machine translation, information retrieval, text classification, ontology building, and thesaurus construction. Unlike English, where various approaches have been explored, Arabic automatic term extraction has relied heavily on rule-based or statistical methods due to the lack of annotated datasets. This article introduces the first annotated dataset for Arabic automatic term extraction (AraATE) in the field of Arabic linguistics. AraATE comprises 4,148 sentences and 155,502 tokens, annotated with 4,362 single and multi-word Arabic linguistic terms. The dataset covers diverse areas of Arabic linguistics, including lexicography, semantics, pragmatics, phonetics, and semiotics. Additionally, this article presents the results of fine-tuning five BERT-based models using AraATE. The findings indicate that AraBERTv0.2-base, CAMeLBERT-MSA, and AraBERTv0.2-large exhibit comparable F1 scores (0.82, 0.81, and 0.81). However, no statistically significant difference was observed in the performance of these models. The availability of AraATE will facilitate Arabic term extraction by serving as a benchmarking dataset for different approaches. Nevertheless, the field still requires additional benchmarking datasets that cover other domains.

Read the paper · More papers on PaperTik