Improving Sinhala Hate Speech Detection Using Deep Learning

Kavishka Gamage, Viraj Welgama, Ruvan Weerasinghe · 2022

Automatic Hate Speech Detection is a fine-grained sentiment analysis task that has been the focus of many researchers around the world. This has been a difficult task due to challenges such as the usage of native languages and distinct vocabularies, as well as the distortion of words. However, based on the findings of previous studies on Sinhala hate speech identification, this has proven to be more difficult for low-resource languages like Sinhala. The effectiveness of pretrained embedding for Sinhala hate speech detection has not been investigated. We investigated several embeddings as well as frequency-based features, including bag of words, n-grams, and TF-IDF to address this shortcoming. We present results from several machine learning experiments, including deep learning experiments and transfer learning experiments on state-of-the-art cross-lingual transformers. With an f1-score of 0.764 and a recall value of 0.788 in our study, the XLMR model outperformed other baseline algorithms and deep learning models.

Read the paper · More papers on PaperTik