Classification of Private Tweets Using Tweet Content
Qiaozhi Wang, Jaisneet Bhandal, Shu Rong Huang, Bo Luo · 2017
Online social networks (OSNs) like Twitter provide an open platform for users to easily convey their thoughts and ideas from personal experiences to breaking news. With the increasing popularity of Twitter and the explosion of tweets, we have observed large amounts of potentially sensitive/private messages being published to OSNs inadvertently or voluntarily. The owners of these messages may become vulnerable to online stalkers or adversaries, and they often regret posting such messages. Therefore, identifying tweets that reveal private/sensitive information is critical for both the users and the service providers. However, the definition of sensitive information is subjective and different from person to person. To develop a privacy protection mechanism that is customizable to fit the needs of diverse audiences, it is essential to accurately and automatically classify potentially sensitive tweets. In this paper, we make the first attempt to classify private tweets into 14 categories, such as alcohol & drugs, family information, etc. We model tweet semantic with term distribution features as well as users' topic-preferences based on personal tweet history. Experiments show that our method can boost classification accuracy compared with the well-known Bag-of-Words and tf-idf methods.