TRANSFER LEARNING TO IDENTIFY PRIVACY LEAKS IN TWEETS
Medrano, Saul Medrano · Digital Repository at the University of Maryland (University of Maryland College Park) · 2016
Privacy concerns have arisen in social networks, because a lot of sensitive information has been disclosed by users intentionally or unintentionally, allowing different organizations such as the government, advertising and credit companies to exploit that information. In this dissertation, I focus on identifying privacy leaks in content of tweets mainly on two sensitive topics: pregnancy and drunkenness. I built classifiers to detect whether a tweet contains private information in two domains. In the feature selection process, I performed a sentiment analysis to evaluate whether the polarity and the sensibility of a tweet can help us identify whether a tweet is private or not. In addition, due to the increasing amount of data and the cost of labeling, I applied transfer learning algorithms to reuse labeled instances from one domain in the training process in the other domain, such that we do not need to label a lot of instances in each domain.