Classification of Abusive Thai Language Content in Social Media Using Deep Learning
Ruangsung Wanasukapunt, Suphakant Phimoltares · 2021
This paper presents binomial and multinomial models for Thai language abusive speech classification in social media. While previous similar research focused on using traditional machine learning models for binomial classification, we showed that deep learning models have better performance. Our binomial and multinomial models achieved F1 scores of 0.8510 and 0.9067, respectively. These scores were significantly better than the machine learning models' respective best F1 scores of 0.7452 and 0.8090. While the bidirectional LSTM performed well, the DistilBERT had higher accuracy and recall. Moreover, the recall was especially higher for the “figurative” class where certain words were more likely to have different meanings depending on context.