Systematic Literature Review: Toxic Comment Classification

Felix Museng, Adelia Jessica, Nicole Wijaya, Anderies Anderies, Irene Anindaputri Iswanto · 2022

Over the last decade, deep learning models have surpassed machine learning models in text classification. However, with the continuity of the digital age, many are exposed to the dangers of the internet. One of the dangers would be cyberbullying. In an attempt to decrease cyberbullying, much toxic text detection and classification research has been done. In this paper, we aim to understand the effectiveness of deep learning models compared to machine learning models along with the most common models used by researchers in the last 5 years. We will also be providing insight on the most common data sets utilized by researchers to detect toxic comments. To achieve this, we have compiled the datasets of research papers and analyze the algorithm used. The findings indicate that Long Term Short Memory is the most routinely mentioned deep learning model with 8 out of26 research papers. LSTM has also repeatedly yielded high accuracy results with above 79% for around 9000 data which could be adjusted depending on the pre-processing method used. There have been attempts to combine more than one deep learning algorithms, however these hybrid models might not result in a better accuracy than an original model. Furthermore, the most frequent sources of datasets came from Kaggle and Wikipedia datasets and a total of 13 researchers that used Wikipedia's talk page edits as their dataset.

Read the paper · More papers on PaperTik