A Novel Central Kurdish Part-of-Speech Corpus and Deep Tagging Model Evaluation
Haneen Al-Raghefy, Halgurd S. Maghdid, Akar H. Taher · ARO-The Scientific Journal of Koya University · 2026
For many low-resource languages, including the central Kurdish language (CKL), building effective natural language processing (NLP) tools has been a challenge. This is due to the lack of annotated text. Without a large corpus that specifies how words function grammatically, it is difficult to perform basic tasks such as part-of-speech (POS) tagging, which is the building block of many language technologies. To address this issue, this study presents the first comprehensive POS-tagged corpus for CKL. This dataset consists of 108,680 words manually tagged with 86 tags. Unlike simpler tagging schemes, the 86 tags account for the complexity of Kurdish grammar and allow a single word to have multiple valid tags, reflecting the language’s natural ambiguity. Using this resource, this study benchmarks a range of deep models, including neural networks such as bidirectional long short-term memory (BiLSTM). To address the ambiguity challenge, this paper introduces a new method, adaptive tag cycling within the BiLSTM that trains the model to consider all possible tags. The most advanced model in this study, an ensemble of neural sequence taggers, achieves 92.3% accuracy with stop-words retained and 89.5% with stop-words removed on broad grammatical categories (main tags). On the full fine-grained tagset (detailed tags), the same model attains 79.0% accuracy with stop-words and 76.2% without stop-words. Therefore, this study provides two key contributions: (i) a new dataset that supports future Kurdish NLP research, and (ii) a strong performance benchmark for CKL POS tagging.