Indonesian Hate Speech and Abusive Tweets Classification with Deep Learning Pre-trained Language Models

Bagus Tri Yulianto Darmawan, Bassamtiano Renaufalgi Irnawan, Yoshimi Suzuki · 2023

Hate speech and abusive language on social media can spread massively and escalate into conflict between two or more individuals or parties. Previous studies about multi-label classification to identify whether the tweet contains hate speech, abusive language, or not have been conducted with machine learning techniques. They applied machine learning with a combination of feature extraction and achieved a good result in recognizing the tweet with hate speech content. In this research, we apply two approaches: first, we use vanilla pre-trained models of IndoBERT, IndoBERTweet, and Indonesian RoBERTa; second, we combine the three previous models with CNN. Our experiment shows that our proposed method has better results compared to previous research with the best performance achieved by IndoBERTweet+CNN with 93.9% Accuracy.

Read the paper · More papers on PaperTik