Bot Detection in Social Media Using GraphSage and BERT
Abhishek J. Deshmukh, Melody Moh, Teng-Sheng Moh · 2024
The rise of digital platforms has resulted in an influx of widespread misinformation and automated bot activities, all of which continues to post a significant threat to information integrity and societal discourse. Misinformation and disinformation are often disguised as credible information spread through bots that have been designed to manipulate societal discourse. Hence, detecting and combating this issue requires advanced detection strategies when compared with traditional Machine Learning (ML)-based approaches. This paper introduces a novel graph-based detection system to tackle the increasing challenge of identifying bot users on social media. The strength of GraphSage (Graph Sample and Aggregation) for network-based pattern recognition has been integrated with the strength of BERT (Bidirectional Encoder Representations from Transformers) for deep contextual analysis together to create a robust bot detection system. This system concatenates BERT embeddings with Graph-Sage embeddings into a comprehensive feature vector, it therefore captures a rich blend of textual and network characteristics. SVM (Support Vector Machine) is then used to process the embeddings and classify the accounts. The system has achieved an impressive accuracy of 98.68% on the Cresci-15 dataset, outperforming all the compared models. On the Twibot-22 dataset, the system has achieved an accuracy of 74.62%, placing it in the middle tier of the state-of-art models compared with. These results highlight the efficacy and scalability of the proposed system across large-scale datasets and complex network graphs. Concatenating the embeddings to generate a more diverse feature set may be widely applicable to other areas of bot detection and misinformation classification.