Hate Speech Detection using Machine Learning Techniques
Anchal Rawat, Santosh Kumar, Surender Singh Samant · 2024
The widespread transmission of dangerous online information is one notable concern raised by the rise in internet usage among people from a variety of cultural and educational backgrounds. The main difficulty is identifying hate speech and derogatory language in the context of automatically detecting harmful text material. This research endeavor presents a meticulous and exhaustive comparative examination of Machine Learning (ML) algorithms tailored for the identification of hate speech. Its primary objective is to reveal an optimal algorithmic amalgamation characterized by simplicity, ease of implementation, efficiency, and the capacity to deliver robust detection performance. In this study, we conducted a comparative analysis of three ML techniques— Decision Tree (DT), Gradient Boosting (GB) and Random Forest (RF), and—to categorize tweets on Twitter into two distinct categories: those containing hate speech and those devoid of by applying the term frequency-inverse document frequency (TF-IDF) technique. The RF model yielded the most promising results, a commendable accuracy rate of $82 \%$ and a recall rate of $84 \%$ were achieved.