Hate Speech Classification in Indonesian Tweets Using TF-IDF and Data Augmentation
I Putu Widiarta Nandana Githa, Alleycia Syananda, Regina Faustine, Ivan Sebastian Edbert, Derwin Suhartono · 2024
Currently suicide is a public health concern as it can affect people of all ages. The number of people attempting suicide is increasing as the World Health Organization (WHO) notes that over 700.000 people in the world die because of suicide every year. The factors that cause suicide are various, one of which is cyberbullying. One form of cyberbullying is hate speech, which has far-reaching impacts on society as it may cause a serious emotional health issues and fatal consequences such as suicide. So, the detection of hate speech as early as possible may be essential to prevent adverse effects. In this study, we proposed a model that can classify hate speech in Indonesian tweets using three machine learning algorithms: Support Vector Machine, Random Forest, and ensemble method combined with data augmentation to handle imbalanced class. The results show that the model will achieve higher accuracy and become more accurate when we perform data augmentation before training the model. The ensemble method obtained 11 percent higher evaluation score when using a balanced dataset. With an 11 percent increase in accuracy, we can create a more accurate model for classifying hate speech and prevent the negative impact of hate speech.