Comparative Analysis of Vectorization Techniques and Machine Learning Models for Hate Speech Detection

Sushil Dalavi, Tanvesh Nivelkar, Sarvesh Patil, Aadesh Sawant, Amit Aylani · 2023

It is now more important than ever to identify hate speech in digital communication in order to preserve a welcoming and safe online community. Utilizing a three-class classification dataset, this research study gives a thorough analysis of hate speech detection for English textual data. We methodically investigated a variety of vectorization and embedding methods, such as Tfidf Vectorizer, Count Vectorizer, Word2Vec, and GloVe, along with a wide range of machine learning models, including Random Forest, AdaBoost, and Logistic Regression. In order to determine the best method for detecting hate speech in textual content, we rigorously evaluated and conducted extensive experiments to evaluate the effectiveness of different combinations.

Read the paper · More papers on PaperTik