Study of Trend-Stuffing on Twitter through Text Classification

Danesh Irani, Steve A. R. Webb, Calton Pu, Kang Li · 2010

Twitter has become an important mechanism for users to keep up with friends as well as the latest popular topics, reaching over 20 million unique visitors monthly and generating over 1.2 billion tweets a month. To make popular topics easily accessible, Twitter lists the current most tweeted topics on its homepage as well as on most user pages. This provides a one-click shortcut to tweets related to the most popular or trending topics. Due to the increased visibility of tweets associated with a trending topic, miscreants and spammers havestarted exploitingthembypostingunrelated tweets to such topics – a practice we call trend-stuffing. We study the use of text-classification over 600 trends consisting of 1.3 million tweets and their associated web pages to identify tweets that are closely-related to a trend as well as unrelated tweets. Using Information Gain, we reduce the original set of over 12,000 features for tweets and over 500,000 features for the associated web pages by over 91% and 99%, respectively, showing that any additional features would have a low Information Gain of less than 10 −4 and 0.016 bits, respectively. We compare the use of naïve Bayes, C4.5 decision trees, and Decision Stumpsover the individual sets of features. Then, we combine classifier predictions for the tweet text and associated web page text. Although we findthat the C4.5 decision tree classifier achieves thehighest average F1-measure of 0.79 on the tweet text and 0.9 on the associated web page content, we recommend the use of the naïve Bayes classifier. The naïve Bayes classifier performs slightly worse with an average F1-measure of 0.77 on the tweet text and 0.74 on the associated web page content, but its required training time is significantly smaller. 1.

Read the paper · More papers on PaperTik