A Hybrid Deep Learning Techniques Using BERT and CNN for Toxic Comments Classification
Adelia Jessica, Migel Sastrawan Sugiarto, Jerry Jerry, Said Achmad, Rhio Sutoyo · 2024
Cyberbullying is a pervasive issue across all forms of media, affecting various demographics and platforms indiscriminately. From social media networks to online forums and comment sections on news sites, the harmful behavior of cyberbullying manifests in many forms, including harassment, threats, and demeaning comments. This ubiquity underscores the need for effective detection mechanisms. This paper explores the identification of toxic traits in online comments using advanced hybrid models, specifically BERT-CNN and BERT-LSTM. The research methodology involved constructing and configuring multiple layers within these models to optimize their ability to detect harmful content. For the BERT-CNN model, BERT's powerful language understanding capabilities are combined with CNN's strength in feature extraction through convolutional layers, capturing spatial hierarchies of features. In the BERT-LSTM model, BERT is integrated with LSTM layers to leverage their ability to learn long-term dependencies and sequential patterns in text. Various configurations of these models were tested, including adj ustments in the batch sizes and the length of the word tokens, to enhance performance in identifying toxic language. The best model which is BERT-CNN using the configuration 256 as the unit and 32 as the batch size achieved 94.48% accuracy in detecting toxic comments. The result from the model indicates that by harnessing BERT's contextual embeddings and the respective benefits of CNN's and LSTM's processing capabilities, it is possible to significantly reduce the frequency of cyberbullying through effective detection, ultimately fostering safer online environments across different media platforms.