Examining the Generalizability of English Cyberbullying Detection Models on Malay Informal Text Using Direct Translation
Shu Xian Chew, Jasy Liew Suet Yan, Wan Abdul Rahman Wan Ibrisam Fikry, Noor Farizah Ibrahim · International Journal of Asian Language Processing · 2022
As cyberbullying on social media platforms becomes more rampant in Malaysia, there is a need for automatic cyberbullying detection models that can handle text in the local Malay language. Although Malay is widely used in Malaysia, it remains a low resource language particularly the availability of high-quality Malay cyberbullying corpora required to train machine learning models to effectively identify cyberbullying in the local context. Our study explores the possibility of borrowing from an existing cyberbullying corpus (Formspring with 13,153 posts) in a resource rich language to train an English cyberbullying detection model, and then evaluate the performance of the model on a test set containing 3663 Malay WhatsApp messages translated to English. By using direct translation, we reveal that the model performance greatly relies on the quality and accuracy of the English-to-Malay translation, a problem that is exacerbated by many informal Malay expressions and slangs in WhatsApp messages.