Detection of Hate Speech by Employing Support Vector Machine with Word2Vec Model

Nina Sevani, Iwan Aang Soenandi, Adianto, Jeremy Wijaya · 2021 7th International Conference on Electrical, Electronics and Information Engineering (ICEEIE) · 2021

Social media can be seen as one prominent assimilation of technology into human life interaction. Its presence is now likely inseparable to us and its usage, whether individually or communally, evidently has been impactful by means of news spread-both positive and negative ones. As a highlight, Indonesia records more than 3,640 hate speech cases from 2018 to this day. This issue has been the main drive of our research. We aim to produce a model for hate speech detection posted on a social media platform. The data was obtained from github-a hosting provider, consisting of tweets. Word2vec was employed as the method for feature extraction while support vector machine (SVM) with RBF as kernel function was used for data classification. The model was built and tested with a 70:30 ratio of data training and testing, in which we achieved the highest accuracy level of 85% with the settings of gamma =0.1 and C-value =10. The accuracy dropped to 69.7% when the model was tested with different datasets. With the development of hate speech detection models, we are optimistic towards a better society where social media users are less prone to negatively-intended information spread.

Read the paper · More papers on PaperTik