Data Forensics On Social Media
Mannat Amit Doultani, M. Vijayalakshmi · 2019
Authorship Attribution (AA), is a process to identify an author based on input text data given to the system based on its characteristics is a problem with a long history. In this project, we study the problem of authorship attribution for forensic purposes and present machine learning techniques and stylometric features of the author tweets. For this purpose micro-blogging site Twitter is taken for experimentation purpose. On this site people share their ideas, likes, dislikes, interest, opinion, thoughts in the form of short messages called tweets. More than thousand tweets are posted every second and the possibility of sensitive, illicit text sharing cannot be ignored. This system downloads live twitter tweets, and takes text file as the input. The text file contains tweets of random author. Our system finds that that tweet downloaded belongs to which author. For classification of the author some important features are used. Important features include calculation of smiley, calculation of stop words, calculation of punctuations, and calculation of similarity words. Basically this system is divided into two stage process, where in the first stage, stylometric information is extracted from the collected dataset and in the second stage classification algorithm is trained to predict authors of unseen text. The effort is to find out which combination of features help in accurate prediction of the author thus maximizing the accuracy of predictions with optimum amount of data.