Counting Clusters in Twitter Posts

Andrew Bates, Jugal Kumar Kalita · 2016

The Internet is full of information contained in short texts. These texts, sometimes called microblogs, can include a wealth of useful data. Twitter is a well known microblogging platform and has been mined for everything from earthquake detection to suicide prevention. One approach to nding information in Twitter posts is to cluster the tweets. The clustering associates posts that are common in some way. When using K-Means clustering, a major challenge is simply in determining a good value for the number of clusters. Our research has shown that using simple term based statistics can be used to choose the number of clusters. This approach is signi cantly faster and produces better clusters when compared to other cluster counting techniques. Additionally, the term based approach itself produces groupings of tweets that have similar quality when compared to clusters produced with K-Means.

Read the paper · More papers on PaperTik