Finding Similar Tweets and Similar Users by Applying Document Similarity to Twitter Streaming Data

Iwao Fujino, Yuko Hoshino · Institutional Repositories DataBase (IRDB) · 2013

Twitter has grown so rapidly that its users are suffering from an information ove rload.In order to help people to find interesting tweets and users from the enormous Twitter space, by applying the document similarity of Gerard Salton and Christopher Buckley to Twitter streaming data, we attempt to devise a content-based scheme that compares similarities between Twitter users by matching their tweets against each other, especially in terms of the new key-value type database environment.Considering that each tweet is a very short message, we performed data processing of individual twe et level and accumulated tweet level according to author name in our study.At the individual tweet level, finding similar tweets is functional for finding retweets of tweeted messages, which is helpful for estimating the information propagation in Twitter.Also in the meaning of security, it is also effective for finding spam tweets and dishonest copies of tweeted messages.At the accumulated tweet level, finding users who use similar words or expressions is functional for finding users who post similar c ontents or topics, which is helpful for finding friends who have similar preferences and interests.As for the concrete procedures of both levels, we first use Japanese morphological analysis to pick up terms from Twitter data.Then we calculate tfidf to provide a weight parameter for each term.Finally we calculate the document similarity from weight parameter vector between any two documents, which shows how much the tweets or the authors are similar to each other.As a confirmation work, we build a computer system to search tweets by keywords query and to show user similarity between any two users.Distribution graphs of similarity in both tweet level and author level are also achieved.

Read the paper · More papers on PaperTik