Exploring the Impact of Lexicon-based Knowledge Transfer for Hate Speech Detection in Indonesia Code-Mixed Languages
Endang Wahyu Pamungkas, Dian Purworini, Diah Priyawati, Rona Rizkhy Bunga Chasana · 2023
In this study, our objective is to examine the influence of external knowledge from a lexicon on knowledge transfer for mitigating the language shift issue in the detection of hate speech in code-mixed languages. To accomplish this, we constructed a lexicon based on findings from previous studies. Subsequently, we implemented several machine learning models with various scenarios to assess the impact of the lexicon. The experimental results demonstrate that incorporating lexicon features leads to improved performance in detecting hate speech within code-mixed languages. Particularly, utilizing a lexicon that encompasses both implicit and explicit lexicons yields the most favorable outcomes in this investigation. This research provides valuable insights into understanding the detection of hate speech in code-mixed Indonesian languages and contributes to the advancement of more robust systems aimed at fostering a safer and more inclusive online environment. By leveraging the lexicon and exploring the interplay between implicit and explicit elements in hate speech, this study enhances our understanding of addressing hate speech challenges in Indonesia code-mixed languages and paves the way for developing more effective detection mechanisms.