Topic Modelling Swahili Using LDA and Contextualized Embeddings

Dorothy W. Kahenya, Waweru R. Mwangi, Richard Maina Rimiru · 2024

Technological innovation in today's society has led to a highly interconnected world, the fruit of which has been an unprecedented breakthrough in new technologies and a tremendous growth of information resources. The availability of online platforms with user-generated content has increased the volume of textual data, leading to a great need for better tools and methods to process the data into meaningful information. While there has been a lot of research on topic modelling in high-resource languages, there is limited research in resource-scarce languages like Swahili despite being spoken by millions world-wide. New topic modelling techniques provide an opportunity to advance NLP research in low-resource settings. This paper provides benchmark experiments for Swahili topic modelling using different transformer based pre-trained language models as embeddings and compares the results with the baseline LDA technique. We also conduct experiments to examine the impact of two dimensionality reduction techniques, PCA and UMAP, on the performance of the topic modelling framework in long Swahili texts. We evaluate the resulting topic models using the$C\_V$and Normalized pointwise Mutual information (npmi) coherence scores. The results show that BERTopic using XLM-Roberta Pretrained Language model embeddings with UMAP dimensionality reduction and HDBSCAN clustering was best suited for topic modelling Swahili, achieving a c_v coherence of 0.776 and nmpi of 0.199. This was followed by the model developed using SwahBERT embeddings with UMAP and HDBSCAN clustering, which achieved a c_v Coherence of 0.707 and npmi of 0.171. Our findings provide a benchmark NLP contribution in automatically categorising Swahili news articles using topic modelling techniques.

Read the paper · More papers on PaperTik