Cyberbullying Detection in Code-Mixed Languages: Dataset and Techniques
Krishanu Maity, Sriparna Saha, Pushpak Bhattacharyya · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022
The advent of the Internet is a boon to society. However, many of its banes cannot be undermined, and cyberbullying is one of them. In this work, we have created a benchmark corpus for cyberbullying detection in code-mixed languages. In India, most communications on different social media platforms are based on Hindi and English languages, and language switching is a common practice in digital communication. To investigate how code-mixed data can be handled effectively, BERT language model and VecMap based bilingual embedding along with a two-channel convolutional neural network model, namely BERT+VecMap-CNN, have been used. The input to one channel is the BERT language model and that to the other is the bilingual word embedding based on VecMap. As a baseline, we used standard machine learning models, as well as deep neural network models like CNN and LSTM. Our proposed model outperforms the baselines with overall accuracy and F1-measure values of 81.12%, and 81.03%, respectively. Furthermore, a different benchmark code-mixed dataset has been considered to show the robustness of our proposed model.1