Exploiting Topical Perceptions Over Multi-Lingual Text For Hashtag Suggestion On Twitter

Amara Tariq, Asim Karim, Fernando Gómez, Hassan Foroosh · 2013

Microblogging websites, such as Twitter, provide seem-ingly endless amount of textual information on a wide variety of topics generated by a large number of users. Microblog posts, or tweets in Twitter, are often written in an informal manner using multi-lingual styles. Ignor-ing informal styles or multiple languages can hamper the usefulness of microblogging mining applications. In this paper, we present a statistical method for pro-cessing tweets according to users perceptions of top-ics and hashtags. Based on the non-classical notion of relatedness of vocabulary terms to topics in a corpus, which is quantified by discriminative term weights, our method builds a ranked list of terms related to hashtags. Subsequently, given a new tweet, our method can sug-gest a ranked list of hashtags. Our method allows en-hanced understanding and normalization of users per-ceptions for improved information retrieval applica-tions. We evaluate our method on a dataset of 14 mil-lion tweets collected over a period of 52 days. Results demonstrate that the method actually learns useful rela-tionships between vocabulary terms and topics, and that the performance is better than a Naive Bayes suggestion system.

Read the paper · More papers on PaperTik