An N-gram-based Approach for Detecting Social Media Spambots
Tianyu Wang, li-Chiou Chen, Yegin Genc · Journal of the Association for Information Systems · 2018
Ubiquitous nature of social media has transformed the businesses and society.With the rise of AI, artificial accounts, known as bots, have become more pervasive in the society accomplishing complicated tasks.However, they are accompanied with detrimental effects.They can be easily used for spreading fake news or ill-intended information.Detecting such artificial accounts before they cause any harm is an important however complicated task.To detect these spambots at their early stage, we propose a machine learning method that uses content features including n-grams (n many consecutive words) and information entropy.Our method builds up an n-gram dictionary from the content of spam tweets.This dictionary is then used as a benchmark for comparing the similarity of later tweets with the keywords of previous spam tweets.Our proposed n-gram based features have a better performance than the entropy-based feature alone.However, the best performance is achieved when n-gram features and entropy-based features combined.In addition, by using only the first 5% of the data for building n-gram benchmarks, we achieved 85% accuracy in detecting the source authenticity in the remaining data.Our methodology provides insights into the early detection of spambots as well as distinguishing the differences between machine-generated and humangenerated information.