Linguistically Informed Tweet Categorization for Online Reputation Management

Gerard Lynch, Pádraig Cunningham · 2014

Determining relevant content automatically is a challenging task for any aggregation system.In the business intelligence domain, particularly in the application area of Online Reputation Management, it may be desirable to label tweets as either customer comments which deserve rapid attention or tweets from industry experts or sources regarding the higher-level operations of a particular entity.We present an approach using a combination of linguistic and Twitter-specific features to represent tweets and examine the efficacy of these in distinguishing between tweets which have been labelled using Amazon's Mechanical Turk crowdsourcing platform.Features such as partof-speech tags and function words prove highly effective at discriminating between the two categories of tweet related to several distinct entity types, with Twitterrelated metrics such as the presence of hashtags, retweets and user mentions also adding to classification accuracy.Accuracy of 86% is reported using an SVM classifier and a mixed set of the aforementioned features on a corpus of tweets related to seven business entities.

Read the paper · More papers on PaperTik