Adyan: automated annotating named entity recognition dataset for sorani kurdish language

Chovyan H Wahid, Rebwar M. Nabi · Data in Brief · 2025

This paper introduces the first high-quality, automatically annotated Sorani Kurdish Named Entity Recognition (NER) dataset, addressing the lack of annotated resources for Kurdish, a low-resource language in Natural Language Processing (NLP). The corpus was collected from publicly available Kurdish news articles published in 2024 to ensure its relevance to contemporary language use. It spans a wide range of domains, including politics, economics, sports, health, culture, interviews, and technology, providing comprehensive coverage of named entities across various contexts. To ensure the trustworthiness of the content, the news articles were carefully selected from accredited Kurdish outlets. Annotation was performed using a lexicon-based approach, leveraging a pre-defined lexicon to maintain consistency and accuracy. The dataset was preprocessed as follows: entities were labeled using seed words from the pre-defined lexicon, and the BIO (Begin, Inside, Outside) tagging scheme was applied to ensure compatibility with widely used NER models. The dataset is available in TXT format (.txt), making it readily accessible and flexible for use in a variety of research applications. The Adyan dataset can be utilized for multiple NLP tasks, including NER, sentiment analysis, machine translation, and text classification. It is publicly released to support ongoing research and development in NLP for low-resource languages.

Read the paper · More papers on PaperTik