Detection of Cyberbullying on Social Media Code Mixed Data
Kavisha Mathur, Krishna Nikhil Mehta, Keerthana Shivakumar, Uma D · 2022
Cyberbullying or cyber-harassment is a form of bullying or harassment using online means such as the internet. The majority of previous research on this topic has been conducted in the English language. Code mixing, which is the practice of blending two or more languages or language varieties in communication, is gaining attention in research because it is extensively used in social media. With the emergence of code-mixed languages in all social media and everyday conversation, we aimed to understand how to identify and prevent cyberbullying in order to promote cultural sensitivity and awareness among online communities. In our study, we used tweets to detect cyberbullying comments and analysed them using various machine learning models such as Naive Bayes, Logistic Regression, SVM, Decision Tree, KNN and Random Forest. We then constructed a hybrid model that predicts outputs based on the outputs of the base model. This hybrid model was applied to code-mixed social media comments in Bengali-English, Tamil-English, and Kannada-English languages and resulted in an accuracy of 88.06%.