Comparative Analysis of Machine Learning Models for Detecting Mobile Messaging Spam In Swahili SMS
Christopher Kalolo, Jimmy Mbelwa · 2023
The prevalence of spam messages is a worldwide issue. Unsolicited text messages in Swahili via SMS have emerged as a significant concern for mobile phone users. This surge can be attributed to the increasing number of mobile phone users who favor SMS as their primary means of communication. Various techniques and strategies have been developed to counteract unwanted SMS messages in various languages. Nonetheless, languages with limited research and a lack of publicly available Swahili SMS datasets are more susceptible to spam texts. This research conducted a comparative examination of machine learning models for addressing spam in Swahili mobile messages. We employed six machine learning models and assessed their performance using various classification metrics. The comprehensive assessment favored the Support Vector Classifier (SVC) model as the top performer. The SVC model achieved an accuracy rate of 96.19%, a precision of 97.39%, a recall of 94.92%, an F-Measure of 96.14%, and an AUC-ROC score of 0.9619. Following closely was the Logistic Regression model, boasting an accuracy rate of 95.35%, precision of 96.93%, recall of 93.64%, an F-Measure of 95.26%, and an AUC-ROC score of 0.9535. The overall evaluation indicated that all machine learning models displayed promising performance. Results demonstrated that these models exhibit strong discriminatory capabilities between positive and negative classes, as evidenced by their higher AUC-ROC scores. In conclusion, we recommend further research in this domain and in other natural language processing tasks specific to the Swahili language.