Deciphering Disruptive Discourse: Leveraging BERT and CNN for Detrimental Content Detection

Zethindra Mekala, CH. Vinay, M Tharun, M. Kanipriya · 2024

Identifying harmful communication in social media is of greatest significance. The present paper introduces a new method that employs BERT Base embeddings and Convolutional Neural Networks to detect harmful communication. Specifically, this research extends previous studies and addresses the challenge of harmful content identification across multiple contexts. The distinguishing factor of this approach is that BERT embeddings are used in all relevant contexts. Thus, the obtained information is transferred to CNN and used to classify data. According to the reviewed literature, the results of this investigation are relatively surprising since previous studies have mainly focused on one specific context, and even studies regarding two contexts have used different architectures like RNNs and LSTMs. Thus, the present approach advances theoretical and practical knowledge considerably. To simplify the study, we opted for benchmark datasets for each kind of harmful content. Specifically, a Mendeley dataset for hate speech identification was employed, popular Kaggle datasets for emotional distress, and an IEEE publication dataset for cyberbullying were utilized. Findings prove the methodology’s successful application in identifying harmful content. When BERT Base embeddings were paired with CNN, our classification of different kinds of harmful content was 88 percent accurate.

Read the paper · More papers on PaperTik