Spam filtering of bi-lingual tweets using machine learning

Hammad Afzal, Kashif Mehmood · 2016 18th International Conference on Advanced Communication Technology (ICACT) · 2016

During recent years, usage of social media has increased enormously. Billions of users use Twitter, Youtube etc which has resulted in the increase in spams as well. Spammers use spam accounts and target users on online social media. Whether a user accesses this social media through smart-phone or web, he/she is prone to the spammers on social media websites. This paper analyses different classification techniques that are currently being used in spam filtering in the context of social media. The contents of tweets are unique in nature, and are different from emails due to their less content so some techniques used in emails might be effective while some might not be effective. Moreover, the conversations on social media often comprises of short-forms/slangs and incorrect spellings. Usage of social media has also become popular in local/regional languages. One such language is Urdu which is common in subcontinent Indo-Pak and is written using English alphabets. We have performed spam classification for Roman Urdu tweets, collected from five major cities of Pakistan. Some of the most commonly used algorithms and techniques for spam classification are discussed and evaluated on English and Roman Urdu tweets from Pakistan in this paper.

Read the paper · More papers on PaperTik