Comparision of Varied Embedding and Machine Learning Classifiers for Fine Grained Offensive Content Identification
Sneha Chinivar, M. S. Roopa, J S Arunalatha, K R Venugopal · 2022
Online offensive behaviour is becoming a serious threat with the increased access to technology for all. Specifically, social media is being used more to exhibit this derogatory behaviour. Researchers tried to address these hostile practices using numerous techniques, but most previous works have ignored the multi-class categorization of online bullying content. Our work aims to efficiently identify the online offensive content and further classify the recognized bullying content into multi-class categories viz age, gender, religion, ethnicity, and others to understand what qualities or features most online offenders target. Used a balanced dataset and experimented with numerous embedding models to generate the most fitting vectors. These vectors are further given as input to baseline machine-learning algorithms for the appropriate classification of identified offensive content into fine-grained categories. We have compared the results obtained from the various combinations of word embedding techniques and Machine Learning classifiers.