Automatic Hate Speech Detection using Ensemble Method and Natural Language Processing Techniques
N Sahana, R Prerana, Sing H Niharika, S Rakshitha, K J Bhanushree · 2023
The use of social media has exploded in recent years, and sharing information has numerous benefits for society. Hateful content has increased as a result of the increased usage of social media. Hate speech can exist in any sort of content designed to slander, dishonor, or incite hatred toward certain affecting various communities, organizations and individuals. It is critical to distinguish between hate and offensive texts to detect hate and offensive speech in any given text. Since no previous work has been done utilizing both English and Hinglish (Hindi- English code combined) data sets for multi class prediction, an ensemble model has been proposed to categorize any given input sentence into one of the three categories: hatred, offensive, or neither. Data collection, data pre-processing, feature extraction, and text categorization are significant processes needed for the proposed approach to detect hate speech. The data set collected for this model is a publicly available Twitter data set in English and Hindi-English code mix language to which data preprocessing is done. Extraction of n-grams as features is done using the term frequency-inverse document frequency (TFIDF) extraction method. Several single classification methods, such as Decision tree, SVC, Logistic Regression, Random Forest, and Naive Bayes are considered, evaluated and compared, combinations of various single classification models are done to form a better ensemble model with a combination of Random Forest and Support Vector Classifier. When tested against the testing dataset, the proposed ensemble model achieved an accuracy of 90.7.