Experiments on IndoBERT Implementation for Detecting Multi-Label Hate Speech with Data Resampling through Synonym Replacement Method

Michael Adriel Darmawan, Nathanael William Boentoro, Kevin Christian Surya, Rhio Sutoyo · 2023

The freedom to share information online leads to increased online hate speech. Addressing this issue is crucial to creating better online communities. A study in 2019 used machine learning approaches to detect multi-label hate speech in the Indonesian language, achieving an accuracy of 66.12%. Recently, several researchers have shown improved performance when using the deep learning method. Looking at the opportunity, this work is developing an Indonesian hate speech model using IndoBERT. Various word preprocessing techniques and dataset stratification were explored for optimal results. In addition, this work also tried to solve the imbalance dataset problem in the multi-label dataset through the synonym replacement method. As a result, the IndoBERT model of this research achieved an improved accuracy of 88.23%. It was also implemented into a web-based application to show the potential usage.

Read the paper · More papers on PaperTik