Handling Imbalanced Dataset on Hate Speech Detection in Indonesian Online News Comments
Almira Diva Sanya, Lya Hulliyyatus Suadaa · 2022
In the current technological era, people can get information quickly, up to date, and in large quantities by accessing online news. A comment column is usually provided as a feature in the online news to express criticism, suggestions, and opinions on certain news. This convenience is a form of freedom of expression for everyone. However, it increases the opportunity for users to express their hate in the comment column. The rise of posts containing hate speech is a problem that needs to be addressed by automatic hate speech detection. Naturally, there is an imbalance between the number of comments that include hate speech and those that do not. Therefore, in this study, the detection of Indonesian hate speech posts in the comments was carried out under several imbalanced conditions. Models used in this study are the SVM model, SVM with various resampling methods, fine-tuned IndoBERT, and fine-tuned mBERT. Based on the results, an imbalanced state can reduce the performance of SVM and the fine-tuned models. Even under an extreme imbalanced condition, the fine-tuned IndoBERT can capture the sentence contexts better than SVM. Then, applying oversampling strategies to SVM can improve the performance of SVM under imbalanced settings. The combination of SVM and SMOTE models perform the best in handling imbalanced problem.