Privacy-Encoding Models for Preserving Utility of Machine Learning Algorithms in Social Media
Sara Salim, Nour Moustafa, Benjamin Turnbull · 2020
Social media has become a vital platform in our daily life, where users can interact with their friends and other people throughout the world. The vast data generated by these platforms is unique in its variety and sensitivity, and although it potentially has significant utility, but also the potential for misuse. Although social media providers apply some existing privacy techniques, such as encryption and anonymization, the techniques cannot achieve a solid level of data privacy while maintaining the highest level of data utility. This paper proposes new Privacy-Encoding (PE) models that contain two-levels of data privacy: 1) data perturbation-based encoding techniques, and 2) data normalization-based scaling techniques. The data perturbation-based encoding techniques involve label encoder and one-hot encoder ones, while data normalization-based scaling techniques include min-max and z-score normalization ones. The aim of the two-levels is to transform original data into perturbed data, along with balancing the high level of data utility using machine learning algorithms. To evaluate the data utility, the proposed models are applied on the adult dataset as well as a simulated social media dataset and the accuracy of the results is compared with several machine learning algorithms. The experiment results reveal that the models could achieve high privacy and utility levels in terms of variance, accuracy and f-measure metrics.