IndoBERT-Based Ensemble Learning for Multi-Level Multi-Label Hate Speech Detection in Indonesian Social Media
Imam Fadhkur Rokhim, Riyanarto Sarno, Abdullah Faqih Septiyanto, Agus Tri Haryono, Shoffi Izza Sabilla · 2024
Hate speech on social media platforms has become a pressing issue, with harmful content often leading to social tensions and emotional harm. In Indonesia, the complex linguistic and cultural context of online discourse presents additional challenges for effective hate speech detection. This study addresses these challenges by presenting an ensemble learning approach for hate speech detection in Indonesian social media. Leveraging IndoBERT for language understanding and combining it with Bi-LSTM and Bi-GRU models for sequence processing, we developed a robust multi-model architecture that effectively captures linguistic patterns and contextual nuances unique to Indonesian. The proposed ensemble framework was tested on a comprehensive dataset with multiple hate speech labels, including categories such as Religion, Race, Gender, and Severity. Experimental results demonstrate that the ensemble model achieved an accuracy of 86% and an Fl-score of 63%, significantly outperforming individual models across most categories. This approach highlights the potential of ensemble learning for automated content moderation in Indonesian social media, providing a promising solution for managing diverse forms of online hate speech.