Toxic Comments Classification using Machine Learning and Word Embedding Techniques
Sneh Prabha, Aditya Yadav, Ashok Yadav, Akash Kumar, Manu Tyagi · 2025
Toxic comments are comments which are disrespectful, unreasonable and infuriating that make the reading uncomfortable. Sometimes these comments are inappropriate and not good for the public. These comments make daily content consumers uncomfortable and sometime even cause emotional as well as emotional harm. There are some people whose purpose for coming online is to create trouble for others and make their work and daily life difficult by making inappropriate comments. This kind of behavior can be described as anti-social and harmful to society. Many steps have already been taken to catch these troublemakers and punish them accordingly by blocking users who are found guilty. This cannot be done personally by checking one by one, to achieve this purpose we need to make a model. We have applied three Machine Learning Models logistic regression, Naive bayes, SVM and three Word Embedding Techniques, Word2Vec, FastText and Glove which can classify comments, and the content creator take appropriate action against users who makes these comments by reporting and even by deleting their account if needed. Out of all six techniques SVM has outperformed for classification of toxic and non-toxic comments and normal text classification in general. SVM gives the best results with 93.61% accuracy and F1 score of 93%. Proposed technique has improvised accuracy score as compared to existing techniques.