A Comparative Exploration in Text Classification for Hate Speech and Offensive Language Detection Using BERT-Based and GloVe Embeddings
S. Santhiya, Uma S, N. Abinaya, P. Jayadharshini, Somala Priyanka, M. N. Dharshini · 2024
On social media platforms, hate speech and provocative language have significantly increased over the past few years, generating major concerns about their effects on both individuals and society. The proposed work analyzes four models, namely BERT Tokenizer(Bidirectional Encoder Representations from Transformers), BERT-CNN(Convolutional Neural Network), and BERT-RNN(Recurrent Neural Network), and a 1D CNN model with GloVe embeddings, to identify hate speech and offensive language. The aim of the proposed research is to examine the accuracy and other evaluation criteria of these models' performance in identifying hate speech and offensive language. A labeled Kaggle dataset with instances of offensive and hateful words is used for training and testing. To rectify the issue of class imbalance, oversampling techniques are utilized to equalize the class distribution within the dataset. The BERT Tokenizer model utilizes the BERT tokenizer for text tokenization and encoding to classify hate speech and offensive language. BERTCNN model incorporates a CNN layer, while the BERT-RNN model employs an LSTM layer. 1D CNN with GloVe embeddings model uses a 1D CNN layer with pre-trained GloVe word embeddings for classification. The models are trained, validated, and tested using appropriate train-validation-test splits. The 1D CNN with GloVe embeddings model, which obtains the 96.91% as highest accuracy of the four models, is demonstrated to be effective by the experimental findings. An in-depth analysis of recall, precision, and F1-score for every classification enables an improved understanding of the model's performance.