Handling Imbalance Issue in Hate Speech Classification using Sampling-based Methods

Heng Rathpisey, Teguh Bharata Adji · 2019

In a text classification problem, imbalance nature in dataset oftentimes has been disregarded even though it might significantly have an impact on the results of the classification model performance. This issue also occurred in the hate speech detection where most collected datasets are highly unbalanced. Among state-of-the-art methods that deal with classifying disparity data, the sampling-based technique is the most effective approach in classifying an imbalanced data. In this paper, four resampling methods include Random Oversampling (ROS), Synthetic Minority Technique (SMOTE), Adaptive Synthetic (ADASYN) and Random Undersampling (RUS) are used as an answer to the inequality of class distribution in a hate speech dataset. With three basic machine learning classifiers i.e. Support Vector Machine, Logistic Regression and Naïve Bayes, the evaluation results show that the oversampling approach improves the accuracy and the overall performance of three classifiers. Among all resampling techniques and machine learning algorithms, Logistic Regression enforced by ROS performed the best with an overall accuracy of 91 percent and F1-Score of 0.95.

Read the paper · More papers on PaperTik