Improving the hate speech analysis through dimensionality reduction approach

Neha Rai, Pooja Meena, Chetan Agrawal · 2020

One of the key challenges during automatic hate speech analysis on the social media is the separation of hate speech from other parts of the offensive language. Normally hate speech analysis or in case of lexical dictionary approach the classifier may produce low precision because they classify the message as hate speech by recognizing some of the well known specific terms. This is the reason traditional algorithms many times fail to recognize the classes properly. This work is based on crowd-sourced hate speech lexicon which can collect tweets containing hate speech keywords as well as dimensionality reduction is also applied to perform classification and this shows how the implemented approach performs better than existing algorithms. This work has classified the tweets into three categories first one is hate speech, second one is offensive language, and the last one is neutral. We train a multi-class classifier to recognize these various classifications. This work has implemented the classification algorithm i.e. logistic regression through dimensionality reduction approach with an accuracy of 83% which is better than other existing algorithms like Naïve Bayes with 71.33% and SVM with 80.56%.

Read the paper · More papers on PaperTik