Swahili Dholuo Topic Classification Dataset

Edwin Onkoba, Lilian Wanzare, Calvins Otieno · Data Intelligence · 2025

The Swahili Dholuo Topic Classification (SDTC) dataset is a groundbreaking resource aimed at advancing Natural Language Processing (NLP) for Swahili and Dholuo, two widely spoken but underrepresented languages in machine learning research. A critical challenge in low-resource NLP is the limited availability of annotated datasets that cover a broad range of topics. Previous studies have focused on narrow or generic topic categories, such as Business, Politics, Religion, Sports, Health and Entertainment, or specialized fields like Science/Technology, Legal, Traffic News and Foreign Affairs. However, many real-world applications require more diverse and fine-grained topic representations. To address this gap, the SDTC dataset was developed, incorporating 30,621 labeled sentences across 12 distinct topics: Social Setting, Agriculture and Food, Healthcare, Religion and Culture, Ceremony, Business and Finance, Automotive and Transport, Sports and Entertainment, Nature and Environment, Education and Technology, News and Media and History and Government. While there were some thematic categories inherently with more instances than others, class imbalance was addressed using a combination of tailored strategies. Easy Data Augmentation (EDA) operations like synonym substitution, random insertion, random swap and random deletion were used to increase the representation of underrepresented classes. The sample augmentations were also modified in the class distribution proportion to help enhance the diversity and quantity of the minority-class samples. This provided an improved balanced and effective training process for all the thematic categories. The dataset underwent rigorous data cleaning and preprocessing, followed by manual annotation by native speakers. Krippendorff’s alpha was used to validate annotation quality, achieving a high agreement score of 0.85, confirming its reliability and consistency. To demonstrate the practical applicability of the dataset, multiple models were trained and evaluated on the topic classification task. The Naive Bayes model achieved an F1-score of 0.44, while the Bi-LSTM model slightly outperformed it with an F1-score of 0.58. The best performance was recorded by the XLM-R model, a fine-tuned variant of XLM-RoBERTa, which achieved an F1-score of 0.9471, indicating its effectiveness in handling Swahili short-text classification across diverse thematic categories. By addressing challenges such as class imbalance and dialectal variations, the SDTC dataset fosters linguistic diversity in Artificial Intelligence (AI). It is publicly available on Zenodo, promoting open access and collaboration for further advancements in multilingual NLP.

Read the paper · More papers on PaperTik