Ensemble Text Classification with TF-IDF Vectorization for Hate Speech Detection in Social Media
R. Sathishkumar, Karthikeyan T, Praveen kumar P, S M Shamsundar · 2023
The development of artificial intelligence (AI) has changed how hate speech is detected. In hate speech identification using machine learning, a number of methods are used to automatically find text that uses vocabulary that is considered to be derogatory, discriminatory, or motivated by hatred. Supervised learning techniques like neural networks, decision trees, and SVMs need a labelled dataset comprising samples of hate speech and non-hate speech. This project investigates the use of AI and machine learning techniques to automatically detect material that uses offensive, intolerant, or hostile words. A voting classifier and TF-IDF representations are combined to improve classification accuracy. The ensemble of classifiers, powered by AI approaches, shows impressive accuracy in identifying hate speech by training five different classifiers (Random Forest, Bagging, Support Vector Machine, AdaBoost, and Gradient Boosting) on a labelled dataset of tweets. The TF-IDF representation prioritises textual terms, whereas the ensemble method uses classifier diversity to capture distinctive patterns. Results from experiments show the strategy's effectiveness, with precision 0.95, recall 0.96, f1-score 0.95 and accuracy 0.97 for detecting hate speech. By successfully utilising AI's capacity to fight hate speech, this research helps the development of a diverse and secure online environment. The suggested approach works well for automatically identifying hate speech, making the internet a safer and more welcoming place for all users.