Back Translation-EDA and Transformer for Hate Speech Classification in Indonesian

Anita Desiani, Muhammad Adrezo, Endang Sri Kresnawati, Ermatita Ermatita, Muhammad Akbar, Muhammad Said Hasibuan · 2023

Hate speech is speech or writing carried out by individuals or groups with the aim of pitting parties against each other for various purposes. The impact of hate speech is very dangerous. Machine learning can help with early detection of hate speech, especially on social media. Unfortunately, data providing hate speech in Indonesian is still limited. This study combines augmentation and classification techniques. The augmentation techniques used are EDA and back translation. By combining these two techniques, hate speech data in Indonesia can be produced with many variations. In classification, this research applies machine learning using the transformer method with the Leaky ReLU function in the hidden layer. The purpose of using the Leaky ReLU function is to avoid neuron death because it has a negative value. The performance results in this research using the proposed method obtained an accuracy of 85.21%, precision of 83.09%, recall of 88.78%, and f1-score of 85.84%. These results show the effectiveness of the proposed method approach with good performance for classifying hate speech in Indonesian on social media such as Twitter.

Read the paper · More papers on PaperTik