Deep-learnt features for Twitter spam detection

Xinbo Ban, Chao Chen, Shigang Liu, Yu Wang, Jun J. Zhang · 2018

Twitter spam has become one of the most critical problems in recent years. Despite the efforts of researchers and security companies, the growing number of spam is not stopping. Machine learning is a very popular technology in network security and is also used for spam detection. An important step of applying machine learning for Twitter spam detection is feature engineering. Existing works mainly use URL based features, meta-data based features and social relation based features to detect spam tweets. All of the above mentioned works require human effort to extract features. More recently, deep learning has shed its light on automated feature engineering in extracting features from text. In this paper, we propose a new feature engineering mechanism based on a deep neural network trained using Bi-LSTM. We name the extracted features “deep-learnt features”. We compare our feature set with word2vec features and statistical features in the experimental evaluation. The results show that machine learning models trained using deep-learnt features can detect Twitter spam more accurately than models trained using word2vec features and statistcal features.

Read the paper · More papers on PaperTik