Twitter Data Preprocessing for Spam Detection
Myungsook Klassen · 2013
Abstract—Detecting Twitter spammer accounts using various classification machines learning algorithms was explored from an aspect of data preprocessing techniques. Data normalization, discretization and transformation were methods used for preprocessing in our study. Additionally, attribute reduction was performed by computing correlation coefficients among attributes and by other attribute selection methods to obtain high classification rates with classifiers,such as Support Vector Machine, Neural Networks, J4.8, and Random Forests. When top 24 attributes were selected and used for these classifiers, the overall classification rates obtained were very close in range 84.30 % and 89%. There was no unique subset of attributes which performed the best, and there were various different sets of attributes playing important roles. Keywords-data preprocessing;spam detection; social network; classification. I.