A Novel Methodology Design to Predict Hate Speech on Social Media Using Machine Learning Principles
Anitha G. S., G. Karthik, Rajkumar Chadge, M. G. Dinesh, T.D. Subha, Ala’a Al Sherideh · 2024
Hate speech detection on social media platforms has become increasingly critical in maintaining a safe and inclusive online environment. This study proposes a novel methodology combining Bidirectional Encoder Representations from Transformers (BERT) with Convolutional Neural Networks (CNN) to improve the accuracy and robustness of hate speech detection models. The BERT-CNN hybrid model leverages BERT's contextual understanding of language and CNN's ability to capture local patterns in text data. The methodology involves comprehensive data collection from Twitter, including both labeled and unlabeled datasets, followed by rigorous data preprocessing, including text cleaning, tokenization, and feature extraction using Transformer-based embeddings. The proposed model was evaluated against nine existing models, including LR, SVM, RF, NB, DT, CNN, LSTM, BiLSTM, and BERT. The results demonstrate that the BERT-CNN hybrid model outperforms all other models, achieving an accuracy of 92.34 %, indicating its effectiveness in both precision and recall metrics. This study's findings highlight the potential of hybrid models in advancing hate speech detection technology on social media platforms.