IndoBERT-based Indonesian Cyberbullying Detection with Multi-stage Labeling

Yohanes Deny Novandian, Ardytha Luthfiarta, Dhiaka Shabrina Assyifa, Johanes Setiawan, Lailatul Cahyaningrum, Noval Althoff, Mufida Rahayu, Adhitya Nugraha, Rismiyati · 2024

The problem of cyberbullying on social media in Indonesia needs to be addressed to protect users from irresponsible behavior. An efficient and effective text classification model is required to detect indications of bullying. This study collected a dataset from Twitter by searching for keywords related to body shaming. The dataset was labeled by using Large Language Model (LLM) and Vader Lexicon approaches for more accurate labeling. Due to the imbalanced dataset labeling, SMOTE was applied for oversampling the data to get a balanced dataset. The detection model was conducted by using IndoBERT. IndoBERT resulted in 96.7% accuracy with fine-tuning. The proposed model was then examined on real data based on a survey related to bullying sentences conducted with 20 students from Dian Nuswantoro University. The results showed that our models accurately predicted that these sentences contained elements of bullying.

Read the paper · More papers on PaperTik