Generating stochastic data to simulate a twitter user

Jason Li, Abdolreza Abhari · Communications and Networking Symposium · 2017

Twitter is a popular social network that carries information in short messages. A user's tweets can contain information that is similar to another user's tweets. In this research, we aim to provide stochastic tweets that can be used for testing recommender systems with large data. For this reason, we used term frequency and inverse document frequency (tf-idf) to analyze users' aggregated tweets. The empirical results show Weibull distribution fits the model of tf-idf of the words in users' tweets. Then Weibull distribution is used to generate stochastic data for users' tweets. A simulation of a recommender system was also conducted to test classification of users based on stochastic tweets. The recommender system uses collaborative filtering to find similarity between users. The simulation used k-means clustering to verify the similarity of the stochastic data versus real data.

Read the paper · More papers on PaperTik