Machine Learning-based Multilabel Toxic Comment Classification
Nitin Kumar Singh, Satish Chand · 2022 International Conference on Computing, Communication, and Intelligent Systems (ICCCIS) · 2022
The emergence of social media marked the beginning of a revolution not just in the realm of digitalization but also in that of communication. In spite of the fact that social media platforms make it possible for anyone located anywhere in the world to express their viewpoints and interact with a large audience, social media has also evolved into a venue for cruel behaviour, offensive language, cyberbullying, personal assaults, and the use of profane language. To tackle this challenge, we are detecting the toxicity level in the Jigsaw dataset by Google, which consists of six different classes, including toxic, severe_toxic, obscene, threat, insult, and identity_hate. This is a multilabel classification in which one comment can fall under more than one class. We study the impact of Multinomial NB, Logistic Regression, and Support Vector Machine with TF-IDF on identifying toxicity in text. These models were trained using the training data and after training were tested on the test data provided in the dataset. Experimental results show that Logistic Regression trumps the other models in terms of accuracy and hamming loss.