Exploring the Performance of BERT Models for Multi-Label Hate Speech Detection on Indonesian Twitter
Muhammad Razi Mahardika, I Putu Janardana Wijaya, Arvin Rayhandi Prayoga, Henry Lucky, Irene Anindaputri Iswanto · 2023
Due to its widespread distribution and the anonymity it offers users, hate speech on social media platforms, particularly Twitter, is a major problem. Because of Twitter's echo chamber algorithm and viral nature, hate speech can spread quickly and lead to societal unrest. The use of BERT-based models for hate speech identification in the Indonesian language has been studied in previous research. This study compares how well several pre-trained BERT models, IndoBERT and mBERT, perform in identifying hate speech on Twitter. The methodology includes gathering datasets, preparing the data, and utilizing the proper metrics to assess the models. IndoBERT successfully outperforms mBERT model according to the average of each evaluation metric shown. By carrying out this study, we hope to advance knowledge of techniques for identifying hate speech on Indonesian social media platforms and enhance the use of BERT models for such analysis.