Comparative Analysis of CNN and LSTM Performance for Hate Speech Detection on Twitter

Artisa Bunga Syahputri, Yuliant Sibaroni · 2023

The internet's current development has become one factor that gives social media users opportunities to leave comments and posts containing hate speech. Detecting hate speech on social media, particularly on Twitter, has recently become a widely researched topic. Research that has been conducted usually applies a standard machine learning approach. The deep learning approach has become popular because it provides better and more effective results. However, it's still rare to be applied to detect hate speech in Indonesian language texts. This research shows the results of a performance comparison from the deep learning approach using CNN, LSTM, and CNN+LSTM architecture models for detecting hate speech in tweets using the Indonesian language. The dataset used is divided into a general dataset which is the entire dataset and a specific topic dataset that deals with the topic of government, which was taken from the general dataset. The research shows better results when the CNN architecture model is implemented on Indonesian language tweet data compared to the results obtained from the LSTM architecture model and the combination of CNN+LSTM with accuracy and F1-score reaching 81%. Furthermore, the implementation of deep learning models in detecting hate speech performs better than previous research using the same dataset but applying machine learning models with feature extraction. This research also shows that specific data discussing a particular topic significantly impact the model's performance. Thus the version of the model becomes better when applied to data with a general topic and a more extensive vocabulary.

Read the paper · More papers on PaperTik